단일 Exocentric 비디오에서 Egocentric + Wrist 뷰 비디오를 동시 생성하는 diffusion 모델. Wan 2.2를 base로 VIST3A 스타일 3D decoder를 stitching하여 multi-view consistent video generation.
Flow matching 기반 Phase 분리 학습 — Phase 1: high noise region, Phase 2: low noise region으로 나누어 coarse-to-fine 학습.
max_norm=1.0) + Phase 2 warmup schedule (500 steps)