Figure Tests Helix 2.5 in 30 Unseen Homes, Reports 56% Zero-Shot Success

The humanoid test is a notable attempt to measure whether robot skills transfer between real homes. But its reported 56% end-to-end success rate also shows how far dependable household automation remains from settled.

By 3 min read
Figure Tests Helix 2.5 in 30 Unseen Homes, Reports 56% Zero-Shot Success
Figure Tests Helix 2.5 in 30 Unseen Homes, Reports 56% Zero-Shot Success

Listen to this story

The audio brief

About 1:38
0:001:38
Read transcript
Figure says its Helix 2.5 humanoid robot completed 56 percent of full-task trials across 30 Bay Area homes it had never seen before. The homes were excluded from training, and the robots collected no new data, received no fine-tuning, and made no environment-specific adjustments during evaluation. The test covered three defined behaviors: tidying a living room, folding towels, and making a bed. A trial counted only if the robot finished the entire assignment—putting all the toys in a basket, folding and placing all the towels, or positioning both pillows and the comforter. There was no partial credit. So the result tests more than a fixed arm movement: the robot has to perceive unfamiliar surroundings, walk to objects, position its body, and coordinate two-handed manipulation. Figure says the key difference was Index pretraining, using a dataset of human-behavior video. An otherwise matched policy without that pretraining completed 9 percent of the blind trials. Figure also says Helix 2.5 matched an earlier Helix 02 behavior with half as much task-specific adaptation data. That points to a possible way to reduce robot-specific data requirements, but it is not evidence of open-ended household help. The 56 percent success rate still leaves many complete failures, and the evaluation was run by the company. Figure’s separate eightfold scaling study predicted robot-action loss within 0.54 percent of observed variation—not household success. The next constraint is whether the transfer holds across more tasks, homes, and independent evaluations.

Story brief

3 key points

Figure reports that its Helix 2.5 humanoid policy completed 56% of full-task trials across 30 Bay Area homes absent from training, versus 9% for an otherwise matched policy without Index pretraining. The evaluation covered tidying, towel folding, and bed-making with unfamiliar layouts and objects, but not open-ended household assistance. The result suggests broad human-behavior pretraining could reduce task-specific...

  1. 01

    Helix 2.5 was announced September 17 and tested without fine-tuning, collected data, or adaptation in the evaluation homes.

  2. 02

    The comparison held architecture, optimization, downstream data, and conditions constant; Index pretraining was the stated difference.

  3. 03

    Trials required complete outcomes with no partial credit, combining locomotion, perception, and two-handed manipulation.

A household robot cannot be useful only in the room where it was trained. Figure says its new Helix 2.5 policy cleared an early version of that hurdle: the company evaluated humanoid robots in 30 Bay Area homes excluded from training, asking them to tidy living rooms, fold towels and make beds without adapting to those homes first.

Figure announced Helix 2.5 on September 17. Its central claim is not that a robot learned a new household chore, but that a single pretrained base model could carry three specified behaviors into new layouts, with new furniture and unfamiliar task objects. Figure calls that setup zero-shot because it collected no data, performed no fine-tuning and made no adaptation in the test homes or on the objects handled there.

A test of transfer, not open-ended help

The distinction matters. The tasks were defined through training data gathered elsewhere; “zero-shot” describes the homes and objects used for evaluation, not an ability to perform any household request. Each task required a complete outcome under a fixed rubric: all toys into a basket, all towels folded and placed in a basket, or both pillows and the comforter positioned on the bed. Figure gave no partial credit.

That makes this a whole-body control problem rather than a fixed-arm demo. A robot may need to walk to find an item, shift its stance to reach it, then coordinate perception, locomotion and two-handed manipulation. Figure says Helix 2.5 can also step back, reposition or move around a bed when an earlier approach leaves it poorly placed.

The pretraining claim carries the announcement

Figure attributes most of Helix 2.5’s reported generalization to Index, its dataset of human-behavior video. In a comparison holding architecture, optimization, downstream task data and evaluation fixed, the company says a policy trained from random initialization completed 9% of blind zero-shot trials. The Index-pretrained policy completed 56%.

What Figure held constant in its comparison

  • The task-specific training data did not include evaluation homes or their objects.
  • The two policies used the same architecture, optimization settings and downstream data.
  • A successful trial meant finishing the assigned task end to end.

Figure also says Helix 2.5 matched the success rate of a representative earlier Helix 02 behavior with half as much task-specific adaptation data. That comparison points to a potential economic benefit: less robot-specific data may be needed to specify a behavior. It does not mean the model has solved general household assistance, and the reported 56% result leaves substantial room for failed full-task attempts.

A scaling signal, with a narrower target

The company paired the home evaluation with a second result about training scale. It trained four models on nested Index subsets spanning an eightfold data range, while holding model size and downstream training fixed. Figure says held-out robot-action prediction loss fell predictably as data increased and that smaller runs forecast the largest run’s loss with an error equal to 0.54% of variation across that range.

That metric is about predicting recorded robot actions, not directly proving higher household completion rates. Still, Figure’s release offers a concrete proposition for humanoid development: pretrain broadly on human experience, then adapt behaviors with smaller task datasets and test them across new physical spaces. The next meaningful test is whether that recipe holds across more tasks, homes and independent evaluations.

Sources

  1. figure.aiHelix 2.5: Zero-Shot 30-Home Generalization

Loading discussion...

YOUR READING SPACE

Notifications