Research

EgoX v2: VIST3A 3D Decoder Stitching

로봇 Dual-View Video Generation을 위한 3D Decoder Stitching 연구 배경 정리
5 Related Works
3 Open Questions
4 Key Decisions

Core Idea

Dual-View 3D Consistency via Decoder Stitching
로봇 데이터에서 Ego view (고정 카메라)와 Wrist view (이동 카메라)의 dual-view video generation에 VIST3A 스타일 3D decoder를 stitching한다. 양쪽 view를 explicit 3D Gaussian space로 올리고, differentiable rendering으로 3D consistency loss를 부여한다. 기존 GGA(Geometry-Guided Attention)를 제거하고 3D representation 기반 일관성으로 대체.
dual-view generation 3D consistency gaussian splatting decoder stitching GGA 제거

VIST3A Architecture (Base Paper)

Text-to-3D by Stitching a Multi-view Reconstruction Network to a Video Generator
ICLR 2026 Oral. Wan Video Diffusion의 VAE latent를 AnySplat 3D reconstruction에 stitching하여 text-to-3D를 실현. Video diffusion의 생성 능력과 feed-forward 3D reconstruction의 구조적 일관성을 결합.

3-Phase Training Pipeline

Phase 1
CKA로 stitching layer 탐색
+ Conv3D fitting
Phase 2
Stitching + AnySplat
LoRA 학습
Phase 3
DRF로 Wan Transformer
LoRA 학습

Our Adaptation for Robotics

Robotics Dual-View Stitching Pipeline
Input: Text + Ego video (conditioning)
Output: Wrist view video (generation target)

Architecture Flow

Input: Ego + Wrist Wan Transformer (LoRA) Denoised Latent Split ego / wrist Stitching Conv3D (16→1024 ch, k5x3x3, s1x2x2) AnySplat Decoder 3D Gaussians gsplat Rendering Cross-view rendering loss + Diffusion SFT loss

Key Technical Decisions

enc_2 Stitching Layer
5x3x3 Conv3D Kernel
16→1024 Channel Expansion
1x2x2 Conv3D Stride

Decisions

Related Work

VIST3A — ICLR 2026 Oral
Text-to-3D by stitching a multi-view reconstruction network to a video generator. 본 연구의 base architecture.
AnySplat
Feed-forward 3D Gaussian prediction. Stitching 대상인 3D decoder로 사용.
Wan 2.2
Video diffusion model (flow matching). Backbone video generator로 사용.
gsplat
Differentiable Gaussian splatting renderer. Cross-view rendering loss 계산에 사용.
DROID Dataset
Robot manipulation dual-view data. 학습 데이터셋으로 활용.

Open Questions