Simple AI Opens 2,000 Hours of Robot Training Data Collected Without Robots

The release makes a sizable human-capture corpus available for commercial use. Its early parity results are promising, but the robot-free system used more than 10 times as many demonstrations as teleoperation in key comparisons.

By 3 min read
Simple AI Opens 2,000 Hours of Robot Training Data Collected Without Robots
Simple AI Opens 2,000 Hours of Robot Training Data Collected Without Robots

Listen to this story

The audio brief

About 1:44
0:001:44
Read transcript
Simple AI has released two thousand hours of robot-training data captured without using a target robot or a teleoperation rig. Called HiFi-UMI-2K, the dataset contains more than 482,000 demonstrations across over 110 scenes, and is available for commercial use under a C-C B-Y four point oh license, with attribution. The idea is to separate data collection from a particular robot. People perform two-handed tasks using portable, sensor-equipped grippers. Head-mounted stereo cameras and inertial sensors capture their movements, while cameras on each hand record the objects and actions. Simple AI then reconstructs those demonstrations as robot trajectories, and filters them through simulation. The company reports roughly 98 percent pass rates for both reconstruction and simulation-replay validation. The early results are promising, but the comparison needs context. In tests across four bimanual tabletop task groups and three policy architectures, robot-free training came within minus 2.5, plus 3.1, and minus 0.6 percentage points of teleoperation-trained policies. A precision-insertion task reached 85 percent success. But the robot-free policies used about 3,200 demonstrations per task, versus roughly 300 for teleoperation—more than ten times as many. The evaluation robot also shared its gripper and wrist-camera setup with the capture system, so transfer to different hands, sensors, robot bodies, or contact-heavy tasks remains unproven. Simple AI says its broader, 20,000-hour corpus improves offline action prediction and real-robot success. The key question is whether cheaper, higher-volume human capture makes up for lower data efficiency.

Story brief

3 key points

Simple AI’s CC BY 4.0 release gives researchers a large, robot-independent source of bimanual demonstrations: 482,000-plus episodes across 110 scenes, with trajectories, video, language labels and validation metadata. In tests, policies trained on the data approached teleoperation results, including 85% success on precision insertion, but used over 10 times more demonstrations and benefited from shared gripper and...

  1. 01

    HiFi-UMI-2K reports 2,000 hours, 482,000-plus episodes and approximately 98% reconstruction and simulation-replay validation pass rates.

  2. 02

    Robot-free policies differed from teleoperation-trained results by -2.5, +3.1 and -0.6 percentage points across tested task groups.

  3. 03

    The evaluation does not establish transfer to different grippers, sensor layouts, robot bodies or broader contact-heavy tasks.

Simple AI has released HiFi-UMI-2K, a 2,000-hour dataset designed to train robot arms from human demonstrations captured without a target robot or teleoperation rig. The experiment is not whether people can show a task, but whether those recordings are precise enough to replace part of the costly robot-operated data pipeline.

The dataset contains more than 482,000 episodes from over 110 scenes. It pairs synchronized multi-view video with two-handed trajectories, gripper states, language annotations, subtask boundaries and quality-control metadata. Simple AI released it under CC BY 4.0, allowing commercial use with attribution.

Capturing the action before a robot enters the picture

HiFi-UMI records people doing two-handed manipulation with portable, sensor-equipped grippers. The system then reconstructs those demonstrations as trajectories intended for real robot arms. That separates collection from the particular arm and control setup normally used in teleoperation.

What the capture system is built to preserve

  • Head-mounted stereo cameras and inertial sensors record the operator’s view and movement.
  • Each hand module uses two non-parallel fisheye cameras, while a shared hardware trigger aligns camera and inertial-sensor data.
  • Simple AI reports roughly 3-millimeter local end-effector accuracy, synchronization below 40 microseconds and about 200 degrees of visual coverage around each hand.

After capture, the company reconstructs trajectories and tries to replay them in simulation, rejecting failures before they reach training. Simple AI reports about 98% pass rates for both trajectory reconstruction and simulation replay validation.

Near parity, but not equal data efficiency

Across three policy architectures and four bimanual tabletop task groups, policies trained only on HiFi-UMI data differed from teleoperation-post-trained policies by -2.5, +3.1 and -0.6 percentage points in aggregate success rate. On the strongest robot-free policy for a precision-insertion task, success reached 85%.

Those results show that high-fidelity human capture can work in the tested setup. They do not establish that it delivers the same performance from the same number of examples: the robot-free side had far more demonstrations. The teleoperation baseline was also collected in the evaluation scene, while the human demonstrations came from other locations.

A hardware boundary remains

The evaluation robot had different arm kinematics from the capture setup, but shared its gripper and wrist-camera configuration. The tests therefore do not yet answer how well the approach transfers to sharply different hands, sensor layouts or robot bodies, or to a wider range of contact-heavy tasks.

The public release is only a slice of the operation Simple AI describes. The company says it has processed a broader corpus of more than 20,000 hours and 4.32 million episodes across over 480 scenes. It also reports that pre-training on 4,000 hours from that corpus reduced offline action-prediction error by 41% across 10 unseen tasks and lifted real-robot success by 18.1 percentage points in its evaluated setup.

HiFi-UMI-2K gives researchers a commercial-use dataset to test that proposition themselves. Its central open question is economic as much as technical: whether capturing many more human demonstrations is cheap and fast enough to outweigh the lower per-trajectory efficiency suggested by the comparison.

Loading discussion...