Projects
Graph-VLM
Egocentric VQA — graph memory
15
reports
3
weeks
exp (1)
analysis (5)
plan (2)
research (2)
progress (4)
eng (1)
260420 — 260426
3 reports · 14 tables
3
04-21
progress
Query Plan V3: 3-Bug Fix 구현 (multi-seed / plan-union / keyword-subgraph)
Query Plan 3-Bug Fix 구현(V1→V3):
\1
multi-seed, plan union+
\1
금지, keyword-subgraph fallback. Context 평균 37K→9K(8배↓), fallback 5/10→0/10, hardfilt10 정확도 10%→20%.
4 tables
04-21
eng
Hard QA 필터링 파이프라인 + Debug 인프라 구축
Description-proof Hard QA 자동 선별 파이프라인(
\1
) 구축 + debug 인프라 정비. Letsur native video 불가 확정, 2-stage 필터로 20후보→10 hard QA 선별(60% 하드율).
4 tables
04-21
analysis
Hard 벤치마크 3-Way 비교: hardfilt10 / hardest10 / L34hard10
3개 Hard 10-Q 벤치마크 3-way 비교(Graph V3 vs Describe vs Frame-300): hardfilt10 20%/10%/50%, hardest10 20%/30%/40%, L34hard10 40%/70%/60%. Clean QA 기준 Graph V3 40% > Describe 20% (+20%p).
6 tables
260413 — 260419
7 reports · 17 tables
7
04-17
progress
Graph-VLM QA 벤치마크 구축 및 전체 영상 평가
560개 QA 벤치마크 구축 및 27.5분 전체 영상 평가: Graph 67.9% vs Frame Baseline 70.0% (Norm1 기준), L4 장기기억에서 Graph 75% > Baseline 66.7%
3 tables
04-17
analysis
Wrong Answer 분석: Graph-VLM vs Frame Baseline — 실패 패턴 해부
Wrong Answer 심층 분석 (Graph vs Frame Baseline): Graph 77개 고유 정답(EM3/T2/T3 강점), Baseline 89개 고유 정답(EM1/L1 시각 디테일·EM2/L2 부정확인 약점) — 5가지 실패 패턴 분류
2 tables
04-17
analysis
3-Way 비교: Graph-VLM vs Frame Baseline vs Describe-Once
3-Way 비교 (Graph vs Frame vs Describe-Once): Describe-Once 80.7%로 압도, Graph 67.9% vs Frame 70.0% — 그래프가 같은 description의 lossy compression임을 확인, L4/EM4에서만 동률
3 tables
04-13
progress
Graph_VLM — Weekly Progress
260406-260412 · long-horizon VQA via spatio-temporal graph
1 table
04-13
analysis
Graph_VLM — Graph Quality v3→v4 Analysis
260406-260412 · −47% object nodes via 4 prompt rules
2 tables
04-13
plan
Graph_VLM — QA Pipeline & Roadmap
260406-260412 · 10-step roadmap update
3 tables
04-13
research
Graph_VLM — Research Notes
260406-260412 · 8 benchmarks · Everyday Episodic Memory positioning
3 tables
260406 — 260412
5 reports · 8 tables
5
04-08
research
Research Note — Temporal Scene Graph Memory for Egocentric VQA
Egocentric 비디오에서 시간적으로 먼 이벤트에 대한 질문에 답하려면 어떤 memory 구조가 적합한가? 프레임 단위 VQA의 한계를 극복할 수 있는 구조적 memory representation은?
1 table
04-08
exp
Graph VLM Experiments
Graph VLM Experiments
2 tables
04-08
analysis
Graph VLM Deep Analysis — Camera-Aware Resolution, Prompt Optimization, Action Segmentation
Camera parameter 기반 entity resolution에서 FoV overlap은 주방 환경에서 구별력이 없고 position distance가 핵심 지표임을 확인. 프롬프트 v3→v4 개선만으로 ideal graph와의 gap을 71% 해소(74→39 노드, 목표 ~25). Action segmentation 과분할 문제도 그룹핑으로 해결했으나 하드코딩 한계 잔존.
3 tables
04-08
plan
Graph VLM - Project Plan
Graph VLM
1 table
04-08
progress
Graph VLM Progress Report — 2026.04.08
Egocentric 비디오(EPIC-KITCHENS)에서 장기 활동을 이해하기 위해 시공간 그래프를 자동 구축하는 시스템. 기존 VQA는 프레임 단위 처리라 "5분 전에 칼을 어디 뒀는가?" 같은 장기 기억 질문에 답할 수 없다.
1 table