Call for Open-Source Robotics Benchmarks
We are funding academic teams to design open benchmarks for bimanual manipulation and to run real-world evaluations of today’s public models against them.
Build the benchmark. Keep the robots.
All team requirements
Design the benchmark
Build a benchmark for bimanual arms with at least 10 tasks.
Run real-world evaluations
Collect real-world evaluations of YAM arms on your benchmark across at least 3 publicly accessible models — VLAs such as MolmoAct2 and Isaac 0.5, and LLMs such as Gemini 3.7 Flash.
Technical support is provided through Inspect Robots, the open-source physical AI evaluation harness.
What each team receives
- A pair of YAM arms (ABC Research Kit), including 3× Intel RealSense D405 cameras.
- $10,000 cash
- $1,000 in evaluation bill of materials
- $1,000 in compute credits
- Co-authorship on a mega-author paper
Teams can keep all equipment, including the YAM arms and cameras, after the program.

What happens next
Invitation emails will be sent to everyone who registers their expression of interest.
Words from the experts
Researchers, engineers, and forecasters on Robocurve.
“We've all seen the demo where a robot does a backflip or jumps rope, and then you try to use it for anything real and, in the best case, it doesn't break your own robot. That gap is exactly why independent, reproducible benchmarks matter, and why I'm excited Robocurve is building them in the open.”
Liane GalantiPhD Student in Computer Science at Princeton“Timelines to robot automation is an important input into our models of takeoff, yet it is one that I (and in my experience, many other people) have a lot of uncertainty about. This project appears well-positioned to clarify trends in robotics capabilities.”
Gabe WuAlignment Researcher at OpenAI“Measuring robotics capabilities over time seems like a very important input for forecasting AI takeoff speeds. Currently, there are basically no widely-used high-quality robotics benchmarks, and additional work could make a big difference in helping us understand the automation of manual labor.”
Nikola JurkovicMember of Technical Staff at METR“It is difficult to honestly evaluate progress in robotics. This bounty program to support large-scale community-driven benchmarks takes a step in the right direction.”
Neehar PeriPostdoctoral Research Fellow at Caltech“There is a need for good robotics benchmarks to capture capability improvements that may emerge in the coming years. This project could mark a strong contribution to this space.”
Julian JacobsResearch Scientist and Economist at Google DeepMind“Physical AI may be one of the most transformative technologies in human history, and yet we cannot say with any confidence what robots can do today, how fast it is changing, or when it will start to matter for the economy. Robocurve's benchmarking initiative is therefore a much needed project that will change the way we think about robotics progress.”
Sebastian SartorPhD Student in Mechanical Engineering at MIT“Supply chain automation timelines are a crucial input to ASI timelines in hardware-dependent AI takeoff scenarios, as well as forecasting progress in AI military technologies. This kind of benchmarking and forecasting may impact MATS' field-building priorities.”
Ryan KiddCEO & Co-Founder at MATS Research