What we ran
THE DATA
gr1_arms_waist.PlateToCardboardBox
- 10,055 teleop episodes, ~2M frames
- GR-1 humanoid, RoboCasa tabletop sim
- Pick target object from plate, place into box. 2 objects on plate
THE SCORING
Skyu scores every episode- Each episode compared against what a lightly trained policy expects
- Single A10 GPU
- Runs on-prem; no data leaves building
THE OUTPUT
Ranked manifest + HTML report- Severity 0–10 per episode with reasoning
- Deterministic: rerun gives the same ranking, bit for bit
- Full manifest, eval logs available
What flagged demos look like
The test: train with and without them
| ARM | TRAINING SET | EPISODES |
|---|---|---|
| A · Control | 929 random + 71 flagged | 1000 |
| B · Flagged removed | without 71 flagged | 929 |
| C · Random removed | without 71 random episodes | 929 |
Same recipe for every arm
- GR00T N1.5 fine-tune, diffusion transformer tuned, per NVIDIA's recipe
- 10000 steps, batch 32 (about 1.6 passes over the data), 2 seeds per arm: 6 models, ~78 min each on one A100
- Evaluated in NVIDIA's RoboCasa GR-1 tabletop sim: 50 fixed scenes, 720 max steps, same scenes for every model
Why only 1000 episodes
The 71 flagged episodes are 0.7% of the full set. Removing 0.7% from a fine-tuned foundation model would show nothing either way. At 7% of the training set the effect is measurable if it exists. The claim is therefore scoped: in a 1000-episode training set, what do these 71 cost?
Removed flagged episodes: +4 points. Removed random ones: 0
What this does and does not show
The evaluation is sound
NVIDIA's released checkpoint scores 42% (21/50) on our setup; NVIDIA reports 44% for this task. Every number in this post is on NVIDIA's own footing: NVIDIA's model, NVIDIA's recipe, NVIDIA's simulator.
Direction is right, size is unproven
+4 for flagged removal, 0 for random, over 100 rollouts per arm. But arm A's two seeds differ by 6 points, so this could still be luck. Five seeds per arm settles it: nine more fine-tunes, about 90 minutes on 8×A100, one day end to end.
Why the effect is small here
Sim teleop data is clean by construction: 0.7% anomalous. We had to concentrate the flagged episodes to 7% to measure anything. Real-robot teleop (resets, operator errors, dropouts) flags more, and each flag costs more. This benchmark is the hard case for the tool.
