1 파일 구조 & 수치 출처
scene_bench_paper/
├── iclr2027_conference.tex # 메인 (스타일 파일 불변, 익명 저자, AI use/Ethics/Repro 성명)
├── macros.tex # \bench, \flashm/\sonnetm, \todo/\ph, 전략명
├── sections/{introduction,related,benchmark,protocol,experiments,analysis,conclusion}.tex
└── iclr2027_conference.bib # 확실 9건 + TODO 스텁 12건
| 논문 표 | 수치 출처 (로그 실측) | 핵심 값 |
| Tab 3 Tier-1 격자 | 260819/exp-s3_budget_sweep.md (find n=353) + 260821/exp-v1_full_eval.md (restore 11/resume 7) | vis flash 0.690 vs sonnet 0.477; textlog restore/resume 1.000 포화(A2 발동) |
| Tab 4 층별 분해 | S3 층화 표 | textlog 옮긴 0.80 / 본 0.17 — vis 정반대 0.61/0.77 |
| Tab 5 온라인+Δ재생 | 260824/exp-online_full_eval.md 최종(n=353) | frames 0.436 ≈ vis 0.433 > textlog 0.385; 텍스트 온라인 이득 +0.09~+0.13 |
| Tab 6 cap 스터디 | 260822/exp-tagmem_first_sweep.md 260824 업데이트 | vis −0.15 vs tagmem −0.03, crossover@8 (0.55≥0.54) |
주장 수위 교정: 대표본에서 tagmem "최악-케이스 최강"이 소멸(vis와 접전) → 논문은 cap 면역·역할 분담(text=사건/image=수동 커버리지)·읽기 효율만 주장, Limitations에 "부분 우위" 명시. 덱 stale 수치는 전혀 사용하지 않음.