Robot Startups Turn to Human Workers as Simulation Hits Manipulation Limits
Companies are recording everyday work, teleoperating machines and capturing motion with wearables. Simulation is scaling quickly, but its toughest transfer problems leave real-world examples in demand.
Listen to this story
The audio brief
Story brief
3 key pointsAi2’s MolmoBot is a notable simulation-first attempt to scale robot training: its open suite uses 1.7 million expert trajectories across 11,000-plus objects and 94,000 procedurally generated environments, then transfers to real robots without fine-tuning on tested tasks. The result strengthens the case for synthetic data in routine manipulation, but does not resolve contact-heavy work, deformable objects, or fluid...
- 01
The largest openly available robot-task dataset cited contains only 3,500 hours of demonstrations.
- 02
Ai2 says MolmoBot transfers to static and mobile manipulation tasks across two robot platforms without fine-tuning.
- 03
In a hammering experiment, 25 human demonstrations reduced simulated training from roughly 50 hours to six.
Robot makers are trying to manufacture the ingredient that language models found on the internet: a huge supply of examples. The emerging data race runs through camera-equipped workers, people remotely operating robots, motion-capture gear and ever-larger simulated worlds—because the largest openly available robot-task dataset cited has only 3,500 hours of demonstrations.
Two routes to a usable robot corpus
The first route begins with people rather than robots. Data-collection companies are recording human actions, paying operators to control robots directly, and using gloves or exoskeletons to capture movements in forms intended to be useful for robot training. Shift, for example, offered free New York City apartment cleaning while workers wore cameras mounted beneath baseball caps; the company planned to record the work for potential sale as robotics data.
The second route attempts to replace much of that expensive collection with virtual experience. In March, Ai2 released MolmoBot, an open manipulation-model suite trained entirely on simulation data, along with a dataset of 1.7 million expert trajectories covering more than 11,000 objects and more than 94,000 procedurally generated environments. Ai2 says its best model transferred without fine-tuning to real-world static and mobile manipulation tasks.
Why walking is not the same problem as working
Simulation has a clear advantage when the task and its feedback can be specified. Figure said it trained a digital twin of its Figure 02 robot for walking and obtained years of simulated demonstrations in a few hours. Large-scale variation can also make learned behavior more robust: a Skild model trained for the equivalent of 1,000 simulated years controlled a quadruped after engineers sawed its legs in half.
Manipulation is harder because a robot acting randomly may never stumble into a successful sequence, leaving reinforcement learning with no reward to reinforce. Developers can create intermediate rewards—such as touching a hammer, lifting it, then bringing it to a nail—but that task-by-task work is labor-intensive. Imperfect simulation of details such as friction can also teach behavior that does not hold up outside the virtual environment.
What human demonstrations can add
- A cited hammering experiment could not learn from a sparse success-or-failure reward alone; shaped rewards took about 50 hours and still produced awkward technique.
- Adding 25 human demonstrations cut the simulated training time to about six hours and produced better hammering technique.
A promising result, with a boundary
Ai2 frames MolmoBot as evidence that sufficiently diverse simulation can close much of the sim-to-real gap. Its released system covers pick-and-place work, articulated-object manipulation and door opening across two robot platforms. But the cited research did not address harder contact-rich work, deformable objects, or jobs requiring accurate fluid or granular dynamics. Those boundaries leave room for the physical-data strategies now being pursued alongside simulation.
The long-term prize is a deployment flywheel: robots that do useful work could generate new training data while operating, helping their makers improve faster. Yet that loop has a prerequisite. A robot must first be capable enough to perform useful work at scale, and getting to that point still appears to require a scalable source of training examples outside commercial deployment.
Editorial analysis
Our Read
The strategic contest is not simply over who can collect the most footage. It is over whether simulated diversity can remove the need for costly physical demonstrations, or whether the companies that reach useful deployments first will accumulate a hard-to-copy stream of real operating data. Ai2’s open MolmoBot release makes the simulation case testable. The more decisive evidence will be whether sim-trained systems handle the contact-heavy and deformable-object work that remains outside the cited research scope, and whether robot deployments generate enough useful data to begin the proposed flywheel.
Citation desk / original work
Cite this
Citation desk / original work
Cite this
The strategic contest is not simply over who can collect the most footage.
/posts/robot-startups-turn-to-human-workers-as-simulation-hits-manipulation-limits#finding-1
Sources
- allenai.orgMolmoBot: Training robot manipulation entirely in simulation | Ai2
- scale.comScale AI and Universal Robots: Enabling Physical AI for Real-World Deployment