Figure Tests Helix 2.5 in 30 Unseen Homes, Reports 56% Zero-Shot Success
The humanoid test is a notable attempt to measure whether robot skills transfer between real homes. But its reported 56% end-to-end success rate also shows how far dependable household automation remains from settled.
Listen to this story
The audio brief
Story brief
3 key pointsFigure reports that its Helix 2.5 humanoid policy completed 56% of full-task trials across 30 Bay Area homes absent from training, versus 9% for an otherwise matched policy without Index pretraining. The evaluation covered tidying, towel folding, and bed-making with unfamiliar layouts and objects, but not open-ended household assistance. The result suggests broad human-behavior pretraining could reduce task-specific...
- 01
Helix 2.5 was announced September 17 and tested without fine-tuning, collected data, or adaptation in the evaluation homes.
- 02
The comparison held architecture, optimization, downstream data, and conditions constant; Index pretraining was the stated difference.
- 03
Trials required complete outcomes with no partial credit, combining locomotion, perception, and two-handed manipulation.
A household robot cannot be useful only in the room where it was trained. Figure says its new Helix 2.5 policy cleared an early version of that hurdle: the company evaluated humanoid robots in 30 Bay Area homes excluded from training, asking them to tidy living rooms, fold towels and make beds without adapting to those homes first.
Figure announced Helix 2.5 on September 17. Its central claim is not that a robot learned a new household chore, but that a single pretrained base model could carry three specified behaviors into new layouts, with new furniture and unfamiliar task objects. Figure calls that setup zero-shot because it collected no data, performed no fine-tuning and made no adaptation in the test homes or on the objects handled there.
A test of transfer, not open-ended help
The distinction matters. The tasks were defined through training data gathered elsewhere; “zero-shot” describes the homes and objects used for evaluation, not an ability to perform any household request. Each task required a complete outcome under a fixed rubric: all toys into a basket, all towels folded and placed in a basket, or both pillows and the comforter positioned on the bed. Figure gave no partial credit.
That makes this a whole-body control problem rather than a fixed-arm demo. A robot may need to walk to find an item, shift its stance to reach it, then coordinate perception, locomotion and two-handed manipulation. Figure says Helix 2.5 can also step back, reposition or move around a bed when an earlier approach leaves it poorly placed.
The pretraining claim carries the announcement
Figure attributes most of Helix 2.5’s reported generalization to Index, its dataset of human-behavior video. In a comparison holding architecture, optimization, downstream task data and evaluation fixed, the company says a policy trained from random initialization completed 9% of blind zero-shot trials. The Index-pretrained policy completed 56%.
What Figure held constant in its comparison
- The task-specific training data did not include evaluation homes or their objects.
- The two policies used the same architecture, optimization settings and downstream data.
- A successful trial meant finishing the assigned task end to end.
Figure also says Helix 2.5 matched the success rate of a representative earlier Helix 02 behavior with half as much task-specific adaptation data. That comparison points to a potential economic benefit: less robot-specific data may be needed to specify a behavior. It does not mean the model has solved general household assistance, and the reported 56% result leaves substantial room for failed full-task attempts.
A scaling signal, with a narrower target
The company paired the home evaluation with a second result about training scale. It trained four models on nested Index subsets spanning an eightfold data range, while holding model size and downstream training fixed. Figure says held-out robot-action prediction loss fell predictably as data increased and that smaller runs forecast the largest run’s loss with an error equal to 0.54% of variation across that range.
That metric is about predicting recorded robot actions, not directly proving higher household completion rates. Still, Figure’s release offers a concrete proposition for humanoid development: pretrain broadly on human experience, then adapt behaviors with smaller task datasets and test them across new physical spaces. The next meaningful test is whether that recipe holds across more tasks, homes and independent evaluations.
Sources
- figure.aiHelix 2.5: Zero-Shot 30-Home Generalization
Loading discussion...
Reader comments
Newest comments first. Replies stay oldest first.