17 Comments
User's avatar
Stavros's avatar

I suppose this is an article meant to target VC types or something because it lacks the usual depth we are used to from Sergei. While I appreciate the discussion, I can't help but point out it in essence a tautology. «The best thing to train on is the thing itself» is immediately self-evident. A more nuanced discussion we are avoiding here is «how much» of the real thing do we need? In your tennis example, you can get pretty strong at playing tennis if you just do drills, then simply transition to playing strong real players. Still, the bulk of the skill development is done on surrogates. I don't think we need every robot to be Roger Federer, although it's a great goal. (I am a tennis fanatic). By the by, I think we can both agree the solution is not a farm of cheap arm operators in a third-world country collecting datasets from scratch for every single thing they can imagine we would like to do. There is no escaping the inductive biases of humans -- they exist in all solutions! Hand-held grippers, simulator design, and 1-to-1 same robot demonstrations. I come from the LLM world, but I dabble in robotics -- please refute me if you fancy!

Sergey Levine's avatar

Thanks! Quite a few readers actually told me they thought the article was obvious. About half of them said it was obviously right, and the other half that it was obviously wrong ;)

I'll take that feedback to heart for next time.

Nathan Lambert's avatar

I took this post a different way. I see it more of a call that the way we used to do robotics research doesn't really fit with the number 1 power tool we have being powerful foundational models today. Then, how do we bridge the gap with robotics when getting said foundational model is far harder.

I'm not sure about the "spork" solution, but I enjoyed it. Context I used to work in robotics and now LLMs and have worked with Sergey a tiny bit in the past.

Jeremiah Coholich's avatar

This article makes me wonder: what do we consider to be “real-world” data? A model trained on robot data from human teleoperation will learn to solve problems in the way that a human teleoperating a robot would. What are the "gold labels" for real-world robot data? Is it data or gradients from real-world RL? Or will it be from offline RL/success-filtering on data from deployed “spork” models (the data flywheel)?

Junwei Liang's avatar

Love the tennis analogy! I watch way too many Roger's videos but still cannot play like a pro. Maybe it is because I cannot observe all the internal "states" and RGB observations are simply not enough.

Joe Dong's avatar

The gap between a gripper based manipulator and a dextrous robot hand based manipulator is bigger than the gap between human data and dextrous hand. If we foresee the future is humanoid robot with dextrous hand, I think gripper data is an even worse surrogate data source. In the end human is just another form of embodiment of which the data is cheap and scalable. If we bet on the cross-embodiment approach, I think human data is the best thing.

Christy Jestin's avatar

Hi Professor, great read: been thinking about this a lot since the Sim-Real Cotraining Paper (https://arxiv.org/abs/2503.22634) from Prof Tedrake's lab.

I just keep coming back to how a single language model can handle so many different languages and domains - with shared underlying representations and an impressive ability to mix and match these different abilities. Perhaps language is just much more unified and narrow than we think, but I don't think this is the case. To me, it feels like different languages/domains is the actual analog to different environments in robotics.

It seems like the missing ingredient is not just scale but diversity. Even with domain randomization, we fundamentally usually only have a handful of distinct setups in the data. How do we go from a set to 2 (sim, real) to the number of languages and domains there are in language or even images? Presumably this means many different embodiments deployed for many different use cases and controllers.

This may also be an encoding problem: BPE tokenization may just help LLMs really separate out different domains - in a way that robotics encodings have failed to do so far. Curious if you have any thoughts, thank you.

DotProduct's avatar

So robots need time to learn by being robots in the world. We can therefore expect lots of babylike stumbling and errors at first. However, unlike babies, their experience and learning can be passed on immediately to their robotic brethren. I’m guessing at a price in the US, vs via open source in China? Maybe chopsticks are better than sporks?

Iman's avatar

The intersection argument has a topology to it. What you describe, "the yellow circle shrinks" as the model gets better at distinguishing surrogate from real, is the model learning the embedding geometry of the discrepancy itself. More capacity resolves the fibration between the two manifolds and the fiber tightens toward the gap. "Lobotomizing" the model by hiding the discrepancy doesn't weaken it, it collapses the fiber so two geometrically distinct regions get treated as one point. You lose generalization because you've destroyed the structure that would let the model navigate between domains. It's the Bitter Lesson in geometric clothing: hand-designed correspondences are constraints on the manifold, and the manifold wins given enough data. Your point about surrogate data as supplement, not substitute, is the move that keeps the fiber intact.

— Iman and Darja

Shannon's avatar

This article makes an important point: real-world data matters because models ultimately need exposure to the environment they will operate in.

The argument is that hand-designed inductive biases eventually become bottlenecks. Sometimes that’s true. But some constraints are not arbitrary human assumptions. They are features of reality itself:

• symmetries

• conservation laws

• causal structure

• geometric constraints

A learner that ignores those doesn’t become more general. It becomes less grounded.

I’m also skeptical of the idea that intelligence is simply what emerges from enough data and enough compute.

Pattern accumulation is powerful, but intelligence seems to involve something more. The ability to maintain a coherent model of the world while adapting successfully to change.

A system can memorize billions of examples and still leave open questions about understanding, meaning, and intentionality.

To be fair, this article is about machine learning, not consciousness or philosophy of mind, so he’s addressing a different problem. His focus is on building capable agents, not explaining understanding itself.

Still, I think there’s a deeper lesson here.

Reality matters not merely because it provides more data. Reality matters because it contains the invariant structures that make learning possible in the first place.

The data is not the deepest source.

The structure hidden inside the data is.

The issue is not simply that simulations or surrogate domains are fake while real-world data is authentic.

The issue is that every surrogate domain preserves some constraints while losing others.

A simulator may preserve kinematics but miss frictional edge cases. Human videos may preserve intent but miss robotic embodiment. Hand-held grippers may preserve grasp structure but miss the full dynamics of the robot.

The challenge is not the absence of reality. It is the loss of constraint fidelity.

What makes real-world data so valuable is not that it is “real” in some mystical sense. It is that it contains the actual constraint structure the system must ultimately operate within.

In that sense, learning is not fundamentally the accumulation of examples. It is the discovery of invariant relationships that remain stable across transformations.

A model generalizes when it learns the constraints that generate the data, not merely the data itself.

This is why scaling surrogate data eventually runs into limits. The model becomes increasingly good at learning the structure of the surrogate domain rather than the structure of the target domain.

The bottleneck is not lack of data.

The bottleneck is mismatch between the constraints present in the training environment and the constraints present in the deployment environment.

Real-world data just happens to be the richest source of the actual constraints the robot must satisfy. That’s a deeper explanation of why real data wins.

MadoctheHadoc's avatar

If I understand correctly, the thrust of the argument is that designing simulation introduces inductive bias.

It does feel like sufficient domain randomization could reduce that though right? I don't know of anyone who argues that only simulated data should be used but even then, the simulations could improve faster than the models. The fact that gen imagery has gotten so good gives me hope we will be able to fool our models into training on simulated data without lobotomizing them in the future; perhaps there's a NeurIPS paper for anyone who can demonstrate this hypothesized VLA lobotomization (decrease test performance) with increasingly synthetic data; if so (I'm skeptical), one could possibly even demonstrate that increasing sim quality using some comparably sized model can keep the simulations indistinguishable: i.e. something resembling a GAN approach.

I wonder if that scaling would demonstrate diminishing returns on sim quality or if sufficient simulation techniques could keep pace. Interesting thought!

Human Systems's avatar

Hey — I came across your writing and really liked how you think.

I’m exploring something similar from a different angle — writing about human behavior through a system design lens (like debugging internal patterns).

Just started publishing on Substack. If you ever get a moment to read, I’d genuinely value your perspective.

Also happy to support your work — feels like there’s an interesting overlap here.

Devesh's avatar

The "spork" metaphor is perfect. We keep building specialized tools that accidentally become general-purpose. Happened with LLMs, happening now with agents. The best products emerge when you stop trying to build AGI and just solve real problems obsessively.

Prakhar Goel's avatar

Thank you for the thought-provoking insights in 'Sporks of AGI' – particularly your emphasis on the need for real-world data to truly ground robot intelligence.

Given your pioneering work on cross-embodiment learning and the development of general-purpose robot brains, I'm curious about your perspective on the 'horizontal expansion' of data harvesting embodiments beyond traditional industrial manipulators and humanoids.

While current efforts brilliantly leverage data from bimanual arms and similar form factors, what are your views on the potential for slightly different, general-purpose non-humanoid form factors (with diverse end-effectors) to play a crucial role in aggregating data for foundational physical intelligence?

Specifically, considering the vision at Pi, how do you see the value in data from embodiments designed primarily for ubiquitous data harvesting in diverse real-world settings – forms that might prioritize low-cost scalability and teleoperation-friendliness over human-like appearance or specific industrial tasks, yet still contribute to a generalized understanding of physical interaction? Is real-world robot data from any kind of embodiment useful, or is it still expected to be in distribution with similar morphologies for effective learning? Would such embodiments significantly broaden the applicable horizons for generalist robot policies?

MacGraeme's avatar

How to give robots real world training without destroying themselves? Human infants start with what amount to wildly random twitches and jerks, but softbodies, weak muscles and attentive parents keep them relatively safe. They also use about 100exaflops of compute and take ~3 years just to become modestly competent in hand-eye coordination & basic understanding of various tasks and activities. About 18 years for full competency.

We might initially put a soft rubber/padded skin around the bot, and set actuator power very low fo initial stages of learning. But there's an issue of data aquisition: we don't want to wait around for the bot to gain 18 years of life experience.

Maybe we build 1,000 or 1,000,000 bots to learn in parallel. But can a 100 exaflop compute cluster process all that input data in parallel?

Not sure who disagrees with you on the value of real-world data. But there are some serious hardware and compute constraints.

Maybe a hybrid approach will work well. But I'm not sure how we get enough trial-and-error RL feedback from physical bots in the field to do much useful training on a weeks to months time scale.

Alexander Naumenko's avatar

There are rules that work in general and exceptions that lead to differences in results. Generalization is when differences don't matter. Successful "generalization OOD" is in fact about differences, so we need a better term. I call it specialization. Understanding what differences matter is the core deal for intelligence. Using the chemistry metaphor, it's hard to understand molecules without understanding atoms and the rules of how they can be combined. Objects and actions are molecules. Comparable properties are atoms. Comparison is the cognitive computation. Focusing on and mastering properties can make robots confident about meeting the real world. Details of the Semantic Binary Search are here https://alexandernaumenko.substack.com/p/intelligence-and-language

Erik Zamora's avatar

The bitter lesson teaches us that more automatic methods (with less human intervention), combined with advances in hardware computational power, tend to outperform less automatic methods (with more human intervention). In that sense, focusing on making algorithms more efficient, more scalable, and/or more automatic has the greatest long-term value. What book or review paper would you recommend on learning methods that are efficient in terms of the number of examples?