Robocurve
Expression of Interest

Call for Open-Source Robotics Benchmarks

We are funding academic teams to design open benchmarks for bimanual manipulation and to run real-world evaluations of today’s public models against them.

Award pool
$500,000In prizes
Hardware
A pair of YAM armsYours to keep after the program
Output
Co-authored paperOne mega-author publication
Participation

Build the benchmark. Keep the robots.

All team requirements

  1. Design the benchmark

    Build a benchmark for bimanual arms with at least 10 tasks.

  2. Run real-world evaluations

    Collect real-world evaluations of YAM arms on your benchmark across at least 3 publicly accessible models — VLAs such as MolmoAct2 and Isaac 0.5, and LLMs such as Gemini 3.7 Flash.

    Technical support is provided through Inspect Robots, the open-source physical AI evaluation harness.

What each team receives

  • A pair of YAM arms (ABC Research Kit), including 3× Intel RealSense D405 cameras.
  • $10,000 cash
  • $1,000 in evaluation bill of materials
  • $1,000 in compute credits
  • Co-authorship on a mega-author paper

Teams can keep all equipment, including the YAM arms and cameras, after the program.

A pair of YAM robot arms on a tabletop

What happens next

Invitation emails will be sent to everyone who registers their expression of interest.

Get started
Testimonials

Words from the experts

Researchers, engineers, and forecasters on Robocurve.

We've all seen the demo where a robot does a backflip or jumps rope, and then you try to use it for anything real and, in the best case, it doesn't break your own robot. That gap is exactly why independent, reproducible benchmarks matter, and why I'm excited Robocurve is building them in the open.
Liane GalantiPhD Student in Computer Science at Princeton
Timelines to robot automation is an important input into our models of takeoff, yet it is one that I (and in my experience, many other people) have a lot of uncertainty about. This project appears well-positioned to clarify trends in robotics capabilities.
Gabe WuAlignment Researcher at OpenAI
Measuring robotics capabilities over time seems like a very important input for forecasting AI takeoff speeds. Currently, there are basically no widely-used high-quality robotics benchmarks, and additional work could make a big difference in helping us understand the automation of manual labor.
Nikola JurkovicMember of Technical Staff at METR
It is difficult to honestly evaluate progress in robotics. This bounty program to support large-scale community-driven benchmarks takes a step in the right direction.
Neehar PeriPostdoctoral Research Fellow at Caltech
There is a need for good robotics benchmarks to capture capability improvements that may emerge in the coming years. This project could mark a strong contribution to this space.
Julian JacobsResearch Scientist and Economist at Google DeepMind
Physical AI may be one of the most transformative technologies in human history, and yet we cannot say with any confidence what robots can do today, how fast it is changing, or when it will start to matter for the economy. Robocurve's benchmarking initiative is therefore a much needed project that will change the way we think about robotics progress.
Sebastian SartorPhD Student in Mechanical Engineering at MIT
Supply chain automation timelines are a crucial input to ASI timelines in hardware-dependent AI takeoff scenarios, as well as forecasting progress in AI military technologies. This kind of benchmarking and forecasting may impact MATS' field-building priorities.
Ryan KiddCEO & Co-Founder at MATS Research
Register

Register your interest

Register your interest

You’ll receive an invitation email at the address above.