Egocentric 비디오에서 시간적으로 먼 이벤트에 대한 질문에 답하려면 어떤 memory 구조가 적합한가? 프레임 단위 VQA의 한계를 극복할 수 있는 구조적 memory representation은?
VLM으로 프레임 쌍에서 structured triplet 추출 → 시간축으로 연결된 graph 구축. Graph는 3종 노드(Object, Action, Keyframe)와 3종 엣지(spatial, participates, temporal)로 구성.
| Decision | Choice | Alternative | Rationale |
|---|---|---|---|
| VLM model | Gemini 2.5 Flash | GPT-4V, Claude | Cost ($0.085/min), structured output quality, speed. 가격 대비 성능 최적. |
| Graph structure | Object+Action+Keyframe | Object+Relation only | Action을 first-class node로 두면 "무엇을 했나" 질문에 직접 답 가능. Relation만으로는 temporal reasoning이 약해짐. |
| Entity resolution | VLM dedup + Camera pos | Embedding similarity, object detector | Camera position이 가장 신뢰도 높았음. 실험으로 검증됨. |
| Scene nodes | Removed (scene-free) | Include scene hierarchy | Scene 노드가 query를 복잡하게 만들고 정보량 대비 비용 높음. 제거 후 query 단순화 확인. |
| Sampling rate | 2s intervals | 1s, 0.5s, adaptive | 2s가 cost/quality 균형점. 주방 action은 보통 2s 이상 지속. 1s는 중복 triplet이 급증. |