VGGRPO Technical Analysis

Deep technical validation of LGM connector, implementation audit, depth calibration, and learning dynamics
4 Analyses
2 Validated
1 Critical
1 Pending

Analysis 1

LGM Connector Design Validation Validated

VGGRPO 논문은 VAE latent에서 4D geometry로의 연결에 "lightweight connector"를 사용한다. 단일 Conv3d가 충분한지 검증.

Method

Conv3d (k=5x5x5, s=1x2x2, p=2x2x2) 단독 connector를 Any4D decoder에 연결하여 end-to-end forward pass 수행. Gradient flow 존재 여부 확인.

VAE Latent
[B, 16, T, H, W]
Conv3d
2.05M params
Any4D Input
[B, 1024, T, H/2, W/2]
4D Geometry

Results

2.05M Parameters
<1% Pipeline Share
5.6GB Peak GPU
Pass Grad Flow

Takeaway

단일 Conv3d가 VAE ↔ Any4D 연결에 충분함을 확인. 다만 5 step은 의미 있는 geometry 생성에 부족하며, 논문 수준 재현에는 25K+ step이 필요.

Analysis 2

Implementation Audit: Paper vs Code 1 Fix Found

재구현 시 논문과 코드의 불일치가 성능 차이의 주요 원인. 학습 전 사전 감사 수행.

Audit Results

Component Paper Our Code (Before) Status
Stitching Layer Encoder 부분 교체 전체 encoder 교체 Fixed → 부분 교체로 수정
DPT Input Dims [1024, 768, 768, 768] 일치 OK
Conv3d Spec k=5, s=(1,2,2) 일치 OK
Info Sharing Alternating attention 일치 OK

Takeaway

1건 불일치 사전 발견 및 수정 완료 (stitching layer: 전체 encoder 교체 → 부분 교체). 나머지 구조는 논문과 일치 확인.

Analysis 3

Depth Label Quality Problem Critical

LGM 학습에서 GT depth 사용 시 Any4D 출력과 DROID GT 간 좌표계 일치 여부가 학습 품질의 upper bound를 결정.

Problem

Any4D output : relative depth (scale/shift ambiguous) DROID GT : metric depth (absolute scale) Conversion required: Procrustes alignment or scale-shift fitting

Current Status

Unresolved Pointcloud alignment experiment pending
30% — DROID pose format identified, alignment experiment not yet run

Implications

Depth calibration이 해결되지 않으면 GT mode 학습이 무의미. Pseudo-label mode만 사용 가능하지만, 그 경우 reward quality가 떨어짐. 이 문제가 전체 파이프라인 성능의 ceiling을 결정.

Analysis 4

LGM Learning Dynamics Bottleneck

190 steps (Phase A + Phase B)로는 학습이 거의 불가능함을 확인. 충분한 step 학습이 전제 조건.

Evidence

190 Current Steps
25K Min Target
50K Paper Steps
0.4% of paper target (190 / 50K)

Job Status

Implication

LGM 학습이 전체 파이프라인의 bottleneck. 충분한 step 학습 완료 후에야 reward quality 판단이 가능. 현재 결과로는 LGM connector 자체의 capacity 문제인지, 단순히 학습 부족인지 구분 불가.