Evaluation Suitefor ML engineers & developers

Real world evals for real world workflows

Post-training competitions, scored on real deployment tasks.

Evaluation Suite platform shown on a laptop

Benchmarks don’t predict reliability.

95% → 60%simulation score to production success rate
5-10%accuracy lost to a single lighting change
0-shotwhere open VLAs that top leaderboards collapse on real-world tasks

A model that scores 95% in simulation can drop to 60% in production. One lighting change can cost 5-10% accuracy. And open-weight VLAs that match the proprietary models on public leaderboards fall apart on real-world tasks they were never trained for.

Two phases. Sim to real.

Every model runs the same pipeline, and it tells you how much of a simulation score survives on hardware.

Phase 1

Simulation

Your model runs hundreds of episodes in high-fidelity physics simulation, scored on whether it finishes the task, stays safe, and does it efficiently. Sim is cheap to repeat, so we test wide and move only the top performers forward without tying up a robot.

Phase 2

Real-world

The top performers move to physical robots in our lab. Lighting varies, objects sit a little off, the background is noisy, and models that looked solid in sim start to break. We run this across several robot form factors, and the diagnostics tell you when and why a model fails, which a pass/fail score never does.

What you get.

01

Per-episode video

Watch every rollout. Replay the exact episode a model failed.

02

Failure diagnostics

The when and why behind a failure, beyond a simple pass/fail.

03

Multi-scene evaluation

One model, many scenes, per-scene success rates.

04

A/B comparison

Two models, side by side, identical conditions.

05

The Shopfloor Corpus

Evals and post-training data from cells deployed in real factories. No open industrial bimanual corpus exists, so we’re building one.

06

Open by default

Bring a model from the open ML ecosystem: SmolVLA, π0, GR00T, MolmoAct, and more.

The first benchmark

WarehouseBench

WarehouseBench is our first suite. Bimanual arms take on the packing, kitting, and fulfilment work real production floors run all day. Finishing the task is only part of the score; the rest is whether the result would hold up on a real line.

01Success rate

did the task complete.

02Containment accuracy

did items land where they belong.

03Closure quality

is the package sealed correctly.

04Post-action stability

does it stay put after the arm lets go.

Next up: Physical Tool Use, an adjacent deployment category.

The leaderboard.

A public ranking of frontier open models, scored on evaluations built to predict real-world reliability instead of simulation performance. See where the field stands, then put your model on the board.

RankModelReal-world score
Submissions are open. Be among the first models on the board.

How does your model compare?

The suite is invite-only for now. Sign in with GitHub, request access, and we’ll get you onboarded.

Built for the open ML ecosystem.