Robocurve
Backed byY Combinator

Real-world evaluations of physical AI

We measure and report frontier robotics capabilities to the public.

Opus 5 and the Tower of Hanoi


Test‑time scaling of Opus 5 on robot arms


$5 $10 $20 $50 $100 API cost per run (log) 0 20 40 60 80 100 score Opus (low) Opus (medium) Opus (high) Sol (high) Opus (low) Opus (medium) Opus (high) Sol (high)

LLM inference speed over time


Jan '24 Jul '24 Jan '25 Jul '25 Jan '26 Jul '26 20 50 100 200 500 1,000 2,000 ×2.3/yr Claude 3 Haiku Celeris-1 ×7.8/yr o1 Mercury 2 ×3.4/yr GPT-5 Gemini 3.5 Flash-Lite ×2.1/mo Claude Fable 5 Gemini 3.7 Flash Intelligence level GPT-3.5 Turbo o1 GPT-5 Fable 5

Opus 5 on block stacking


NVIDIA Inception Program
Claudefor Open Source
Opus 5 pulling and placing a Jenga block

Pull one block out of the tower without bringing it down and placing the block on top of the tower.

With support from experts at:
MIT
Stanford University
Harvard University
Princeton University
Caltech
MATS
Foresight Institute
University of Washington

Independent analysis of robotics capabilities

Robocurve raises $10M to independently evaluate frontier AI in the physical world

Seed funding to independently evaluate frontier AI in the physical world.

Why it matters

A technology this big must be measured in the open

02 h4 h6 h8 h10 h12 h14 h16 h18 h20192020202120222023202420252026Task length at 50% success (human hours)Model release dateGPT-2 (2019-02-14): 0.1 minGPT-3 (davinci-002) (2020-05-28): 0.1 minGPT-3.5 Turbo Instruct (2022-03-15): 0.6 minGPT-4 (0314) (2023-03-14): 4.0 minGPT-4 (1106) (2023-11-06): 4.0 minClaude 3 Opus (2024-03-04): 4.0 minGPT-4 Turbo (2024-04-09): 3.7 minGPT-4o (2024-05-13): 7.0 minClaude 3.5 Sonnet (Jun 2024) (2024-06-20): 11.4 mino1-preview (2024-09-12): 20.3 minClaude 3.5 Sonnet (Oct 2024) (2024-10-22): 20.5 mino1 (2024-12-05): 38.8 minClaude 3.7 Sonnet (2025-02-24): 1.0 ho3 (2025-04-16): 2.0 hClaude Opus 4 (2025-05-22): 1.7 hClaude Opus 4.1 (2025-08-05): 1.7 hGPT-5 (2025-08-07): 3.4 hGemini 3 Pro (2025-11-18): 3.7 hGPT-5.1-Codex-Max (2025-11-19): 3.7 hClaude Opus 4.5 (2025-11-24): 4.9 hGPT-5.2 (2025-12-11): 5.9 hClaude Opus 4.6 (2026-02-05): 12.0 hGPT-5.3-Codex (2026-02-05): 5.8 hGemini 3.1 Pro (2026-02-19): 6.4 hGPT-5.4 (2026-03-05): 5.7 hClaude Mythos Preview (early) (2026-04-07): 17.4 h

Source: METR, Task-Completion Time Horizons of Frontier AI Models

The stakes

Robotics capabilities are progressing rapidly

Frontier labs are racing to develop general-purpose robots in the next two years.

The problem

No one knows where the frontier is

The Internet is filled with cherry-picked demo videos. No continuous, standardized evaluations for robotics exist.

Our answer

We are an independent evaluator for robotics

We measure robotics capabilities in the real world, independent of any agenda.

Our work

We build open-source tools to help society understand robotics progress

Open-source evaluation framework

Inspect Robots is released under the open-source MIT license. Free to use, modify, distribute, and sell.

Contribute

Inspect Robots

Run any model on any embodiment on any benchmark, with full trace logs and live Rerun visualization. If you know Inspect AI, this is that for robotics.

  • Real-world first, with simulation support
  • Runs VLAs, WAMs, LLMs, and coding agents to control robots
  • First-class integrations with ROS, Isaac Lab, Cap-X, and XPolicyLab
  • Open-source license (MIT)
About us

Robocurve is a Public Benefit Corporation helping society understand the frontier of physical AI.

Our mission

Robocurve is a Public Benefit Corporation bound by law to serve the public good. General-purpose robots may arrive within years, with profound implications for the economy and the labor market.

We build open-source tools and independent benchmarks to measure robotics capabilities. We report progress to the public rigorously and neutrally.

Testimonials

Words from the experts

Researchers, engineers, and forecasters on Robocurve.

“Supply chain automation timelines are a crucial input to ASI timelines in hardware-dependent AI takeoff scenarios, as well as forecasting progress in AI military technologies. This kind of benchmarking and forecasting may impact MATS' field-building priorities.”
Ryan KiddCEO & Co-Founder at MATS Research
“Timelines to robot automation is an important input into our models of takeoff, yet it is one that I (and in my experience, many other people) have a lot of uncertainty about. This project appears well-positioned to clarify trends in robotics capabilities.”
Gabe WuAlignment Researcher at OpenAI
“There is a need for good robotics benchmarks to capture capability improvements that may emerge in the coming years. This project could mark a strong contribution to this space.”
Julian JacobsResearch Scientist and Economist at Google DeepMind
“We've all seen the demo where a robot does a backflip or jumps rope, and then you try to use it for anything real and, in the best case, it doesn't break your own robot. That gap is exactly why independent, reproducible benchmarks matter, and why I'm excited Robocurve is building them in the open.”
Liane GalantiPhD Student in Computer Science at Princeton
“Measuring robotics capabilities over time seems like a very important input for forecasting AI takeoff speeds. Currently, there are basically no widely-used high-quality robotics benchmarks, and additional work could make a big difference in helping us understand the automation of manual labor.”
Nikola JurkovicMember of Technical Staff at METR
“Physical AI may be one of the most transformative technologies in human history, and yet we cannot say with any confidence what robots can do today, how fast it is changing, or when it will start to matter for the economy. Robocurve's benchmarking initiative is therefore a much needed project that will change the way we think about robotics progress.”
Sebastian SartorPhD Student in Mechanical Engineering at MIT
Contact

We work with labs of all sizes

You’re interested in: