Experiments

VGGRPO — SLURM Details + Hypothesis per Experiment
15 Jobs
6 Experiments
~60h GPU Time
40% Preemption Rate

SLURM Summary

5 Completed
5 Preempted
~6 Cancelled
1 Running
GPU time breakdown (effective ~35h / total ~60h)
Effective (~35h) Lost to preemption/cancel (~25h)
Job IDDurationStateExperiment
#2814213mCOMPLETEDExp 1
#2814823mCOMPLETEDExp 1
#281603h 9mCOMPLETEDExp 2
#298073mCANCELLEDExp 3
#298081h 56mCANCELLEDExp 3
#298431h 50mPREEMPTEDExp 3
#3003715mCOMPLETEDExp 4
#305614h 20mCOMPLETEDExp 4
#3109312h 26mPREEMPTEDExp 4
#3017624mCANCELLEDExp 5
#3074545mCANCELLEDExp 5
#307587mCANCELLEDExp 5
#307599h 43mCANCELLEDExp 5
#3177018h 30mCANCELLEDExp 5
#360316h 46mPREEMPTEDExp 5
#3645710h 4mPREEMPTEDExp 5
#365357h 9mPREEMPTEDExp 5
#3739325h+RUNNINGExp 6

Experiments

Exp 1: Phase 1 Validation Apr 1
Hypothesis
Wan2.1-1.3B + Any4D can produce valid outputs on DROID data
#28142 13m COMPLETED #28148 23m COMPLETED
Result
Wan generation 33f 480x832 OK. Any4D depth/pose/flow OK on DROID wrist video.
Insight
Pipeline works E2E. DROID pose is Euler XYZ, not axis-angle.
Exp 2: LGM Smoke Test Apr 1
Hypothesis
Conv3d connector can bridge VAE latent to Any4D
#28160 3h9m COMPLETED
Result
Pseudo-label loss 3.54 → 1.81, GT loss 1.23 → 1.06 (5 steps each)
Insight
Gradient flows through. But 5 steps insufficient for meaningful geometry.
Exp 3: LGM Short Training Apr 3
Hypothesis
190 steps produces usable geometry
#29807 3m CANCELLED #29808 1h56m CANCELLED #29843 1h50m PREEMPTED
Result
Loss decreases but geometry is noisy point cloud, not useful.
Insight
Minimum 25K steps needed (paper uses 50K). Submitted longer jobs.
Exp 4: LGM Medium Training Apr 3–4
Hypothesis
Longer training converges to usable geometry quality
#30037 15m COMPLETED #30561 4h20m COMPLETED #31093 12h26m PREEMPTED
Result
Gradual improvement but still not converged.
Insight
Preemption is major obstacle for long training. Auto-checkpoint every 1000 steps needed.
Exp 5: Extended Training Apr 4–6
Hypothesis
25K+ steps with checkpointing can survive preemptions
#31770 18h30m CANCELLED #36031 6h46m PREEMPTED #36457 10h4m PREEMPTED #36535 7h9m PREEMPTED
Result
Frequent preemptions prevent convergence.
Insight
Need own-QOS for critical long jobs, or implement robust auto-resume.
Exp 6: Current Long Run Apr 7
Hypothesis
25K+ step LGM training with auto-checkpoint reaches convergence
#37393 25h+ RUNNING
Result
Ongoing — longest uninterrupted run so far.
Insight
Since Apr 7 04:47. Monitoring for preemption.

Training Configuration

Model LGM: Conv3d connector + Any4D decoder (frozen) Optimizer AdamW Data DROID processed (156 episodes dev, 3162 full) GPU B200 x 1 (LGM is lightweight)

GPU Usage

~60h+ Total GPU
~35h Effective
~40% Preemption Rate
~25h Time Lost
Preemption rate higher than other projects. Primary bottleneck for convergence. Mitigation: own-QOS reservation or auto-resume with checkpoint-on-SIGTERM.