Graph VLM Experiments

VLM API 기반 Scene Graph 파이프라인 실험 기록 -- EPIC-KITCHENS P01_01
4 Experiments
~110 API Calls
$0.12 Total Cost
47% Duplication Reduction

Overview

이 프로젝트는 GPU 학습 대신 Gemini VLM API를 사용하여 비디오에서 Scene Graph를 구축한다. EPIC-KITCHENS P01_01 영상에 대해 반복적으로 파이프라인을 개선하며, VLM description, entity resolution, prompt optimization 등의 실험을 수행했다. SLURM 작업 없이 API 호출만으로 진행되므로, 비용과 응답 품질이 핵심 지표다.

Experiment Timeline

Mar 24
Exp 1: VLM Description Test
10s clip, Strategy B (direct JSON) 검증
Apr 1
Exp 2: 1-Minute Full Pipeline
60s clip 스케일링, entity duplication 발견
Apr 1
Exp 3: Camera-Aware Entity Resolution
EPIC-Fields 카메라 파라미터 활용 시도
Apr 6
Exp 4: Prompt Optimization v4
프롬프트 규칙으로 47% 중복 감소 달성

Experiment Details

Exp 1: VLM Description Test
API Mar 24
Hypothesis
Gemini 2.5 Flash가 별도 파싱 없이 직접 structured JSON triplet을 출력할 수 있다.
Setup: P01_01 4:30~4:40 (10s), 2s interval, 5 frame pairs, max_tokens=8192
5/5 Phase 1 Success
3/5 Phase 2 Success
~15 API Calls
$0.02 Cost
Result
Phase 1 전수 성공, Phase 2는 5건 중 3건 성공. 2건 실패는 thinking token이 max_tokens를 소진하여 JSON 출력이 잘린 것이 원인.
Insight
Strategy B (direct JSON) 방식이 동작함을 확인. max_tokens=8192 필수 (4096은 truncation 발생). Phase 2에는 thinking 기능이 없는 Gemini 2.0 Flash가 더 적합할 수 있음.
Exp 2: 1-Minute Clip Full Pipeline
API Apr 1
Hypothesis
파이프라인이 10s에서 60s로 스케일링 시 품질 저하 없이 동작한다.
Setup: P01_01 300~360s (60s), 2s interval, 30 frame pairs
14 Actions Detected
74 Object Nodes
479 Edges
$0.085 Cost
Result -- Issues Found
1분 스케일에서 파이프라인은 동작하나, entity resolution이 병목.

Duplication Analysis

Detected / Unique
74 / ~25
Insight
파이프라인은 1분 스케일에서 기능적으로 동작하나, entity resolution이 핵심 병목. 동일 객체를 다른 이름으로 반복 생성하는 문제를 해결해야 함.
Exp 3: Camera-Aware Entity Resolution
CAMERA Apr 1
Hypothesis
EPIC-Fields 카메라 파라미터로 객체를 구별할 수 있다.
Setup: camera_loader.py (quaternion to viewing direction, position distance) + entity_resolver.py 확장
0.2m Same-Area Dist
3.36m Diff-Area Dist 1
9.15m Diff-Area Dist 2
0.999 FoV Similarity
Result
카메라 위치 거리(position distance)는 같은 영역(0.2m)과 다른 영역(3.36m, 9.15m)을 명확히 분리. 반면 FoV(시야각) 유사도는 주방 환경에서 0.999로 거의 동일하여 무용.
Insight
카메라 위치(position)가 핵심 신호이며, 시야 방향(viewing direction)은 쓸모없음. 단, 정지 객체(stationary objects)에만 유효하다는 한계.
Exp 4: Prompt Optimization v4
PROMPT Apr 6
Hypothesis
프롬프트 규칙만으로 object duplication을 30% 이상 줄일 수 있다.
Rules added: (1) Egocentric person 제외 (2) Generic ID 금지 (3) Noise exclusion (4) Base name 통일
47% Duplication Reduction
39 Object Nodes (from 74)
6 → 1 Washing Segments
$0.00 Additional Cost
Result
가설(30%) 초과 달성: 47% 감소 (74 → 39). Washing segmentation도 6 → 1로 해소. 추가 비용 없이 프롬프트만으로 달성한 개선.

Before / After

Object nodes
74 → 39
Wash segments
6 → 1
Insight
비용 없는 개선이지만 VLM compliance가 ~80-90%로 완벽하지 않음. Post-processing (entity resolution) 여전히 필요.

Cost Analysis

Per-Experiment Breakdown

ExperimentAPI CallsCostTime
VLM Test (10s)~15~$0.02~30s
1-min Full Pipeline~45~$0.085~2min
Entity Resolution+3-5+$0.01+10s
Total (v4 run)~50~$0.10~2.5min

Scaling Projection

Video LengthEst. API CallsEst. CostEst. Time
1 min50$0.102.5 min
10 min~500~$1.00~25 min
1 hour~3,000~$6.00~2.5 hours

Key Takeaways