Businesspublished

Robot Startups Turn to Human Workers as Simulation Hits Manipulation Limits

Companies are recording everyday work, teleoperating machines and capturing motion with wearables. Simulation is scaling quickly, but its toughest transfer problems leave real-world examples in demand.

By 3 min read
Robot Startups Turn to Human Workers as Simulation Hits Manipulation Limits
Robot Startups Turn to Human Workers as Simulation Hits Manipulation Limits

Listen to this story

The audio brief

About 1:44
0:001:44
Read transcript
Twenty-five human demonstrations cut simulated hammer-training from roughly fifty hours to six—and produced better technique. That result captures why robot companies are turning back to people, even as simulation scales rapidly. The basic problem is data. The largest openly available robot-task dataset cited contains only about 3,500 hours of demonstrations. One response is physical collection: workers record everyday tasks, remotely operate robots, or wear gloves and exoskeletons that capture human motion. Scale AI and Universal Robots, for example, are collecting production-robot data with both visual information and force feedback. The other response is to create experience in software. Ai2’s MolmoBot used 1.7 million expert trajectories across more than 11,000 objects and 94,000 procedurally generated environments. Ai2 says the system transferred to static and mobile manipulation tasks on two robot platforms without fine-tuning. That is meaningful progress for routine work, but manipulation remains a different challenge from walking. In simulation, a robot can train rapidly when the task and rewards are easy to specify. Picking up a hammer, making reliable contact, and driving a nail is harder: random attempts may never succeed, and simulated friction may not match reality. Human examples can supply the missing guidance. The unresolved boundary is contact-heavy work involving deformable objects, fluids, or granular materials. The long-term goal is a deployment flywheel, where working robots generate more data. But the immediate constraint is clear: robots may need real-world examples before they can produce enough of their own.

Story brief

3 key points

Ai2’s MolmoBot is a notable simulation-first attempt to scale robot training: its open suite uses 1.7 million expert trajectories across 11,000-plus objects and 94,000 procedurally generated environments, then transfers to real robots without fine-tuning on tested tasks. The result strengthens the case for synthetic data in routine manipulation, but does not resolve contact-heavy work, deformable objects, or fluid...

  1. 01

    The largest openly available robot-task dataset cited contains only 3,500 hours of demonstrations.

  2. 02

    Ai2 says MolmoBot transfers to static and mobile manipulation tasks across two robot platforms without fine-tuning.

  3. 03

    In a hammering experiment, 25 human demonstrations reduced simulated training from roughly 50 hours to six.

Robot makers are trying to manufacture the ingredient that language models found on the internet: a huge supply of examples. The emerging data race runs through camera-equipped workers, people remotely operating robots, motion-capture gear and ever-larger simulated worlds—because the largest openly available robot-task dataset cited has only 3,500 hours of demonstrations.

Two routes to a usable robot corpus

The first route begins with people rather than robots. Data-collection companies are recording human actions, paying operators to control robots directly, and using gloves or exoskeletons to capture movements in forms intended to be useful for robot training. Shift, for example, offered free New York City apartment cleaning while workers wore cameras mounted beneath baseball caps; the company planned to record the work for potential sale as robotics data.

The second route attempts to replace much of that expensive collection with virtual experience. In March, Ai2 released MolmoBot, an open manipulation-model suite trained entirely on simulation data, along with a dataset of 1.7 million expert trajectories covering more than 11,000 objects and more than 94,000 procedurally generated environments. Ai2 says its best model transferred without fine-tuning to real-world static and mobile manipulation tasks.

Why walking is not the same problem as working

Simulation has a clear advantage when the task and its feedback can be specified. Figure said it trained a digital twin of its Figure 02 robot for walking and obtained years of simulated demonstrations in a few hours. Large-scale variation can also make learned behavior more robust: a Skild model trained for the equivalent of 1,000 simulated years controlled a quadruped after engineers sawed its legs in half.

Manipulation is harder because a robot acting randomly may never stumble into a successful sequence, leaving reinforcement learning with no reward to reinforce. Developers can create intermediate rewards—such as touching a hammer, lifting it, then bringing it to a nail—but that task-by-task work is labor-intensive. Imperfect simulation of details such as friction can also teach behavior that does not hold up outside the virtual environment.

What human demonstrations can add

  • A cited hammering experiment could not learn from a sparse success-or-failure reward alone; shaped rewards took about 50 hours and still produced awkward technique.
  • Adding 25 human demonstrations cut the simulated training time to about six hours and produced better hammering technique.

A promising result, with a boundary

Ai2 frames MolmoBot as evidence that sufficiently diverse simulation can close much of the sim-to-real gap. Its released system covers pick-and-place work, articulated-object manipulation and door opening across two robot platforms. But the cited research did not address harder contact-rich work, deformable objects, or jobs requiring accurate fluid or granular dynamics. Those boundaries leave room for the physical-data strategies now being pursued alongside simulation.

The long-term prize is a deployment flywheel: robots that do useful work could generate new training data while operating, helping their makers improve faster. Yet that loop has a prerequisite. A robot must first be capable enough to perform useful work at scale, and getting to that point still appears to require a scalable source of training examples outside commercial deployment.

Editorial analysis

Our Read

The strategic contest is not simply over who can collect the most footage. It is over whether simulated diversity can remove the need for costly physical demonstrations, or whether the companies that reach useful deployments first will accumulate a hard-to-copy stream of real operating data. Ai2’s open MolmoBot release makes the simulation case testable. The more decisive evidence will be whether sim-trained systems handle the contact-heavy and deformable-object work that remains outside the cited research scope, and whether robot deployments generate enough useful data to begin the proposed flywheel.

Citation desk / original work

Cite this

Permanent attributionView citation
Finding 01

The strategic contest is not simply over who can collect the most footage.

/posts/robot-startups-turn-to-human-workers-as-simulation-hits-manipulation-limits#finding-1

Sources

  1. allenai.orgMolmoBot: Training robot manipulation entirely in simulation | Ai2
  2. scale.comScale AI and Universal Robots: Enabling Physical AI for Real-World Deployment