Evaluation Suitefor ML engineers & developers
Real world evals for real world workflows
Post-training competitions, scored on real deployment tasks.

Benchmarks don’t predict reliability.
A model that scores 95% in simulation can drop to 60% in production. One lighting change can cost 5-10% accuracy. And open-weight VLAs that match the proprietary models on public leaderboards fall apart on real-world tasks they were never trained for.
Two phases. Sim to real.
Every model runs the same pipeline, and it tells you how much of a simulation score survives on hardware.
Phase 1
Simulation
Your model runs hundreds of episodes in high-fidelity physics simulation, scored on whether it finishes the task, stays safe, and does it efficiently. Sim is cheap to repeat, so we test wide and move only the top performers forward without tying up a robot.
Phase 2
Real-world
The top performers move to physical robots in our lab. Lighting varies, objects sit a little off, the background is noisy, and models that looked solid in sim start to break. We run this across several robot form factors, and the diagnostics tell you when and why a model fails, which a pass/fail score never does.
What you get.
Per-episode video
Watch every rollout. Replay the exact episode a model failed.
Failure diagnostics
The when and why behind a failure, beyond a simple pass/fail.
Multi-scene evaluation
One model, many scenes, per-scene success rates.
A/B comparison
Two models, side by side, identical conditions.
The Shopfloor Corpus
Evals and post-training data from cells deployed in real factories. No open industrial bimanual corpus exists, so we’re building one.
Open by default
Bring a model from the open ML ecosystem: SmolVLA, π0, GR00T, MolmoAct, and more.
The first benchmark
WarehouseBench
WarehouseBench is our first suite. Bimanual arms take on the packing, kitting, and fulfilment work real production floors run all day. Finishing the task is only part of the score; the rest is whether the result would hold up on a real line.
did the task complete.
did items land where they belong.
is the package sealed correctly.
does it stay put after the arm lets go.
Next up: Physical Tool Use, an adjacent deployment category.
The leaderboard.
A public ranking of frontier open models, scored on evaluations built to predict real-world reliability instead of simulation performance. See where the field stands, then put your model on the board.
How does your model compare?
The suite is invite-only for now. Sign in with GitHub, request access, and we’ll get you onboarded.
Built for the open ML ecosystem.