VGGRPO 논문은 VAE latent에서 4D geometry로의 연결에 "lightweight connector"를 사용한다. 단일 Conv3d가 충분한지 검증.
Conv3d (k=5x5x5, s=1x2x2, p=2x2x2) 단독 connector를 Any4D decoder에 연결하여 end-to-end forward pass 수행. Gradient flow 존재 여부 확인.
재구현 시 논문과 코드의 불일치가 성능 차이의 주요 원인. 학습 전 사전 감사 수행.
| Component | Paper | Our Code (Before) | Status |
|---|---|---|---|
| Stitching Layer | Encoder 부분 교체 | 전체 encoder 교체 | Fixed → 부분 교체로 수정 |
| DPT Input Dims | [1024, 768, 768, 768] | 일치 | OK |
| Conv3d Spec | k=5, s=(1,2,2) | 일치 | OK |
| Info Sharing | Alternating attention | 일치 | OK |
LGM 학습에서 GT depth 사용 시 Any4D 출력과 DROID GT 간 좌표계 일치 여부가 학습 품질의 upper bound를 결정.
Rotation.from_euler('xyz', rot)) — axis-angle이 아님190 steps (Phase A + Phase B)로는 학습이 거의 불가능함을 확인. 충분한 step 학습이 전제 조건.