armnet eval — evaluate a policy on real robots
Pick a task, pick or paste a policy, press Run, and watch 20 rollouts on a real bimanual SO-101 cell. Success is auto-scored by the instrumented BusyBox — no human judging. The run is shared: anyone here follows the same eval live.
Press green and yellow buttons. Push the green button with the left arm, then the yellow button with the right arm. The buttons spring back, so the scene resets itself.
How to use this Space
- Train a LeRobot-compatible policy on the
villekuosmanen/busybox_press_green_yellow_buttonsLeRobot dataset. Make sure the trained policy is uploaded to Hugging Face and publicly available. - Copy your repo ID into the text box on the Live tab and submit it when no one else is running an evaluation. Armnet will automatically run and score the policy on a real bimanual SO-101 cell.
You can run a policy more than once. The more rollouts a policy accumulates, the closer its lower confidence interval matches its average success rate — see the Leaderboard tab.
Every task has its own dataset and its own board, so pick the one you want at the top of the page before you start.