Dual-View 3D Consistency via Decoder Stitching
로봇 데이터에서 Ego view (고정 카메라)와 Wrist view (이동 카메라)의 dual-view video generation에 VIST3A 스타일 3D decoder를 stitching한다. 양쪽 view를 explicit 3D Gaussian space로 올리고, differentiable rendering으로 3D consistency loss를 부여한다. 기존 GGA(Geometry-Guided Attention)를 제거하고 3D representation 기반 일관성으로 대체.
dual-view generation
3D consistency
gaussian splatting
decoder stitching
GGA 제거
Text-to-3D by Stitching a Multi-view Reconstruction Network to a Video Generator
ICLR 2026 Oral. Wan Video Diffusion의 VAE latent를 AnySplat 3D reconstruction에 stitching하여 text-to-3D를 실현. Video diffusion의 생성 능력과 feed-forward 3D reconstruction의 구조적 일관성을 결합.
3-Phase Training Pipeline
Phase 1
CKA로 stitching layer 탐색
+ Conv3D fitting
→
Phase 2
Stitching + AnySplat
LoRA 학습
→
Phase 3
DRF로 Wan Transformer
LoRA 학습
Robotics Dual-View Stitching Pipeline
Input: Text + Ego video (conditioning)
Output: Wrist view video (generation target)
Architecture Flow
Input: Ego + Wrist
↓
Wan Transformer (LoRA)
↓
Denoised Latent → Split ego / wrist
↓
Stitching Conv3D (16→1024 ch, k5x3x3, s1x2x2)
↓
AnySplat Decoder → 3D Gaussians
↓
gsplat Rendering
↓
Cross-view rendering loss + Diffusion SFT loss
VIST3A
— ICLR 2026 Oral
Text-to-3D by stitching a multi-view reconstruction network to a video generator. 본 연구의 base architecture.
AnySplat
Feed-forward 3D Gaussian prediction. Stitching 대상인 3D decoder로 사용.
Wan 2.2
Video diffusion model (flow matching). Backbone video generator로 사용.
gsplat
Differentiable Gaussian splatting renderer. Cross-view rendering loss 계산에 사용.
DROID Dataset
Robot manipulation dual-view data. 학습 데이터셋으로 활용.