Robocurve

Robocurve raises $10M to independently evaluate frontier AI in the physical world

Robocurve announces $10M in seed funding for independent evaluation of frontier AI in the physical world, beside a robot arm.

At Robocurve, we evaluate how well frontier AI models control real robots as an independent third-party auditor.

Our research shows a further generalization of Richard Sutton’s bitter lesson, where large language models (LLMs) could outperform state-of-the-art robotics vision-language-action models (VLAs) on simple tasks. Some simple tasks that earlier models struggled to complete a month ago can now be completed easily by GPT-6 Astra.

We are not alone in witnessing the rapid acceleration of robotics progress and the emergent capabilities of LLMs in robotics. MIT professor and computer vision pioneer Phillip Isola calls this the beginning of robot-use agents, following the advent of computer-use agents merely a couple of years ago. Frontier labs also take the robotics capabilities of frontier AI models seriously, with Anthropic studying the capabilities of Claude in robotics since 2025.

Latency is still a big bottleneck for LLMs to control robots in real time, but current trends show that the output token speed of frontier LLMs is improving by 2.1x/month, while specialized hardware for faster inference continues to advance through efforts by Cerebras, Groq, and Lamb Labs. If these trends continue, inference could become fast enough for real-time robot control by Fable-class LLMs between late 2026 and 2029.

Such direct extrapolation of latency trends might not hold due to external limitations, and it is not yet clear whether the robotic competency of LLMs generalizes to embodiments with many degrees of freedom, such as hands or whole-body control of humanoids. However, the possibility of intelligent general-purpose robots arriving before the end of the decade is not to be dismissed lightly.

The widespread deployment of general-purpose robots could have profound implications for the social contract and the labor market, and deserves a whole-of-society approach to help families prepare for the changes ahead, even if the precise timing of the arrival of such technology still has large uncertainties. When the clouds are heavy, we don’t need to know exactly when it will rain to carry an umbrella.

Our goal at Robocurve is to fulfill that responsibility of societal preparedness. We are incorporated as a Public Benefit Corporation with a legal duty to independently evaluate the robotics capabilities of frontier AI systems, and report these capabilities to the public as a neutral third party.

Our work follows three principles:

  • Independence and neutrality. We set our own research agenda. Frontier labs do not direct our evaluations or control our methodology or public results.
  • Public benefit. We support the safe development and deployment of AI by helping society understand the capabilities and risks of frontier AI in the physical world.
  • Support academic research. We fund academic teams and provide robot hardware to build open-source benchmarks.

We will work proactively with governments, policymakers, and civil society to inform the public of the pace of AI progress in the physical world.

In the past three months since our incorporation, our research has been viewed 6M+ times and our open-source evaluation harness (Inspect Robots) has been downloaded 97k+ times. Researchers from 200+ institutions, including 19 of the world’s top 20 universities, have signed up to build benchmarks with us.

In pursuit of this public benefit mission, Robocurve has raised a $10M seed round led by Initialized Capital, with participation from Notable Capital, Decasonic, Y Combinator, Halcyon Futures, and many others. The funding will let us expand our research team, test a wider range of robots and tasks, and support academia in building open benchmarks.

The field of robotics benchmarking is still nascent. If you build robots or frontier AI, reach out to help the world understand what your systems can do. If you are interested in building robotics benchmarks, apply to our open-source benchmarking program, where we award $500k in total funding and free YAM arms to academic groups.