armnet eval — evaluate a policy on real robots

armnet.dev

Pick a task, pick or paste a policy, press Run, and watch 20 rollouts on a real bimanual SO-101 cell. Success is auto-scored by the instrumented BusyBox — no human judging. The run is shared: anyone here follows the same eval live.

Task

Press green and yellow buttons. Push the green button with the left arm, then the yellow button with the right arm. The buttons spring back, so the scene resets itself.

How to use this Space

  1. Train a LeRobot-compatible policy on the villekuosmanen/busybox_press_green_yellow_buttons LeRobot dataset. Make sure the trained policy is uploaded to Hugging Face and publicly available.
  2. Copy your repo ID into the text box on the Live tab and submit it when no one else is running an evaluation. Armnet will automatically run and score the policy on a real bimanual SO-101 cell.

You can run a policy more than once. The more rollouts a policy accumulates, the closer its lower confidence interval matches its average success rate — see the Leaderboard tab.

Every task has its own dataset and its own board, so pick the one you want at the top of the page before you start.