Real-world evaluations of physical AI
We measure and report frontier robotics capabilities to the public.
Pull one block out of the tower without bringing it down and placing the block on top of the tower.
Independent analysis of robotics capabilities
Robocurve raises $10M to independently evaluate frontier AI in the physical world
Seed funding to independently evaluate frontier AI in the physical world.

A technology this big must be measured in the open
Source: METR, Task-Completion Time Horizons of Frontier AI Models
Robotics capabilities are progressing rapidly
Frontier labs are racing to develop general-purpose robots in the next two years.
No one knows where the frontier is
The Internet is filled with cherry-picked demo videos. No continuous, standardized evaluations for robotics exist.
We are an independent evaluator for robotics
We measure robotics capabilities in the real world, independent of any agenda.
We build open-source tools to help society understand robotics progress
Open-source evaluation framework
Inspect Robots is released under the open-source MIT license. Free to use, modify, distribute, and sell.
Inspect Robots
Run any model on any embodiment on any benchmark, with full trace logs and live Rerun visualization. If you know Inspect AI, this is that for robotics.
- Real-world first, with simulation support
- Runs VLAs, WAMs, LLMs, and coding agents to control robots
- First-class integrations with ROS, Isaac Lab, Cap-X, and XPolicyLab
- Open-source license (MIT)
Robocurve is a Public Benefit Corporation helping society understand the frontier of physical AI.
Our mission
Robocurve is a Public Benefit Corporation bound by law to serve the public good. General-purpose robots may arrive within years, with profound implications for the economy and the labor market.
We build open-source tools and independent benchmarks to measure robotics capabilities. We report progress to the public rigorously and neutrally.
Words from the experts
Researchers, engineers, and forecasters on Robocurve.
“Supply chain automation timelines are a crucial input to ASI timelines in hardware-dependent AI takeoff scenarios, as well as forecasting progress in AI military technologies. This kind of benchmarking and forecasting may impact MATS' field-building priorities.”
Ryan KiddCEO & Co-Founder at MATS Research“Timelines to robot automation is an important input into our models of takeoff, yet it is one that I (and in my experience, many other people) have a lot of uncertainty about. This project appears well-positioned to clarify trends in robotics capabilities.”
Gabe WuAlignment Researcher at OpenAI“There is a need for good robotics benchmarks to capture capability improvements that may emerge in the coming years. This project could mark a strong contribution to this space.”
Julian JacobsResearch Scientist and Economist at Google DeepMind“We've all seen the demo where a robot does a backflip or jumps rope, and then you try to use it for anything real and, in the best case, it doesn't break your own robot. That gap is exactly why independent, reproducible benchmarks matter, and why I'm excited Robocurve is building them in the open.”
Liane GalantiPhD Student in Computer Science at Princeton“Measuring robotics capabilities over time seems like a very important input for forecasting AI takeoff speeds. Currently, there are basically no widely-used high-quality robotics benchmarks, and additional work could make a big difference in helping us understand the automation of manual labor.”
Nikola JurkovicMember of Technical Staff at METR“Physical AI may be one of the most transformative technologies in human history, and yet we cannot say with any confidence what robots can do today, how fast it is changing, or when it will start to matter for the economy. Robocurve's benchmarking initiative is therefore a much needed project that will change the way we think about robotics progress.”
Sebastian SartorPhD Student in Mechanical Engineering at MIT