Screenshot of Figure 03 doing household chores in a Figure livestream.
John Koetsier
Four weeks ago Figure told the world it would spend a billion dollars over 12 months paying strangers to film themselves folding laundry, making beds and picking things up off the floor.
Just one month later, that money has already bought a 6X increase in humanoid robot performance: nine percent to 56%.
That’s the zero-shot success rate of Figure’s humanoid doing household chores in 30 rented Bay Area homes it had never been trained in, with and without pretraining on Index, the crowdsourced dataset those creators are filling. Without Index data: 9% success rate. With Index data: 56%. Same robot, same task data, same architecture, same optimizer, same hyperparameters, same blind evaluation.
Figure changed exactly one variable and success went up more than 6X.
That jump just might be the most important number in humanoid robotics this year, especially for robots intended to work in our homes. The number that got most of the attention was the 56% success rate, or its corollary: the 44% failure rate.
That’s fair: a 44% failure rate doesn’t get you an A in any school I’ve attended. But the success/failure rate tells you where humanoids are today: interesting but not amazing. The 6X boost tells you which direction they’re moving and why. It also tells you that Figure’s investment is already paying off.
And the four week time span tells you that change is happening quickly. (Note: Index was running for some time before Figure announced it, so the training data that enabled this improvement is for more than four weeks.)
In case you didn’t hear the details: Figure rented 30 homes, sent in its Figure 03 humanoid with no data collected in any of them, and graded a variety of whole-body behaviors — tidying 13 to 15 scattered toys into a basket, folding towels and placing them in a basket, and making a bed with both pillows and both comforter corners at the top and the comforter smoothed.
None of the toys, towels or bedding had appeared in training data. The robots used each home’s own beds, couches and folding surfaces … each of which, of course, was slightly different. Each task ran a single checkpoint across all 30 houses. Out of 420 trials, 237 succeeded, with success meaning the entire task finished and any safety intervention scored as a failure.
This is strong evidence that video of a human body doing a task transfers to a humanoid robot doing similar tasks in similar settings. (Note that’s not even the same tasks in the same settings.)
Compare that against previous baselines from just a few months ago.
In April I covered Stanford HAI’s 2026 AI Index, which found humanoids fully and safely completing only about 12% of real household tasks, against 89.4% in simulation. On BEHAVIOR-1K, the 1,000-task benchmark built from chores real people said they wanted done, the best teams managed 25% success rates at merely acceptable quality.
The gap between what these models can do in a controlled setting and what they can handle in the real world is still wide,” that report said.
Five months later a robot is finishing whole chores in strangers’ houses 56% of the time. Note: the measurements aren’t directly equivalent, and three tasks is not a thousand. But the quick five-month difference is still striking.
It sounds a little unhinged, but this might just make Index the most valuable thing Figure owns.
The numbers on it are moving fast enough to be slightly absurd. When I wrote about the platform in August, it had 264,000 downloads across 108 countries, 44,000 weekly active contributors, 16 million uploaded videos and $15 million already paid out. Figure’s Helix 2.5 post put uploads at roughly 35 minutes of new human experience every second. A day later, AI director Corey Lynch posted 55 minutes per second and 115,000 weekly users, with the note “Index is going vertical.” CEO Brett Adcock’s version of the same week was 53 minutes per second and 2.4 million video uploads.
Behind all of it sits $3.5 billion in committed compute through a deal with Nscale for up to 100,000 Nvidia Vera Rubin GPUs. Arguably, Figure is now a data acquisition company that also plans to manufacture humanoid robots.
Strategically, that might set up Figure for a physical AI platform model down the road: sell its own robots, sure, but also make its physical AI platform available to other robot makers and become the Microsoft/Windows of the new robot era. Those are big shoes to fill and there are a lot of unknowns before anything that significant happens, but the potential is there.
There’s a second result in the paper that may end up mattering even more.
Figure trained four models on nested subsets of Index spanning an 8X data range and found that downstream robot-action prediction loss fell smoothly enough with each doubling that it could forecast the largest run’s loss to four decimal places before training started. That’s essentially a human-to-robot training data transfer scaling law for physical AI.
If it holds, Figure can price humanoid performance improvements before buying them.
That’s a big if, and some caveats apply. New failures might be challenging to predict, like towels that slip out of robot hands, pillows that catch on the headboard, toys that roll under the couch (I personally want to see humanoids try to deal with that one). Figure says as much in its own conclusion: the point “is not that general humanoid robotics is solved.”
Other players are working on similar things, of course. On the 14th, Reward AI came out of stealth with OM-1, a policy it says was learned entirely from humans wearing an instrumented seven-degree-of-freedom hand, no teleoperation and no on-robot data at all. On the 17th, Figure published the 30 homes project. On the 18th, XPENG detailed XPACE, which trains its IRON humanoid on human videos plus simulated mistakes. Sunday Robotics has been shipping $200 capture gloves for its Memo robot, and X Square built TwinDEX, a wearable three-finger rig it claims delivers 5.3X the collection throughput of teleoperating a robot.
Clearly, human data works. The big question is how you get it, and how much of it you can get, at what price.
There are at least four answers in the market right now, and I’ve written about all of them. Apptronik built a 90,000 square foot humanoid data factory. Shift will clean your house free in exchange for the footage. Figure pays a hundred thousand strangers to film themselves. And Antioch raised $32 million from Greylock on the argument that most of this belongs in simulation, because almost nobody except Figure can afford a billion dollars of real-world collection.
“You still need to collect some real-world data, but we help you be a lot more sample efficient,” cofounder Harry Mellsop said when I covered that round.
The industry split here is fidelity versus breadth. Sunday, Reward AI and X Square build purpose-made capture hardware because they’re betting that signal quality is what makes a demonstration transfer. Figure built a consumer app because it knows that breadth is what makes a dataset big, and it is betting that enough data volume overcomes any data quality issues.
One side is optimizing signal per demonstration, the other demonstrations per dollar, and Helix 2.5 is a very serious datapoint for the breadth side.
Of course, there are differing viewpoints on the quality of Figure’s robot work so far.
Tony Zhao, cofounder and CEO of rival Sunday Robotics, responded within hours: “Love my friends at figure, but failing half the time is not ‘doing real useful work.’ See the 237/420 success rate below taken from the blog. Doing useful work = generalization + reliability.”
Sunday’s counter-number: 778 successful folds out of 785 attempts, 99.1%, across nine garment types in unseen environments.
But those two percentages describe different universes. Sunday counts based on an individual garment, while Figure counts an entire task chain where one dropped towel out of four fails the trial. Figure’s towel-only rate was 87 of 140, or 62%, and its per-task spread ran from 67% on bed making down to 40% on toy tidying, the task with the most individual pickups and therefore the most chances to break the chain.
Essentially, whoever defines the yardstick wins the narrative, and right now everybody defines their own. A neutral home-task benchmark that Figure, Sunday and 1X all agreed to run would be a huge step forward.
Another important metric that would be useful, especially with a shared way of measuring it: throughput. How long tasks take is important.
Figure doesn’t state how long the chores took, but you can infer the ceiling from the grading rubric in the appendix: one minute per toy, three minutes per towel, one minute per pillow and per side of the comforter. A robot that tidied 15 toys could have spent close to 15 minutes doing it and still passed. That’s the outer bound rather than the average, though, and plenty of the four hours of raw footage Lynch posted moves faster.
Reward AI, notably, led its OM-1 announcement with the claim that its policy executes at human speed. That’s impressive, and it’s the metric that the average customer is certainly going to look at first.
The last caveat is the sample. Thirty short-term rentals staged with toys is a far harder test than a lab and a far easier one than thirty lived-in homes, a point several replies under Lynch’s thread made … one asking how it handles a house “occupied by two toddlers,” another noting the homes all look alike. Rentals are decluttered by definition.
No pets, no laundry piles, no toddler pulling toys back out of the basket.
Ultimately, though, Figure has shown that paying 115,000 people a week to film their chores moves a robot from 9% success to 56% success in houses it has never seen, in just a few short weeks.
The big question: can Figure get to 95% success in a few more months?
That’s not high enough in an industrial setting. But it’s certainly good enough in a home, where no-one dies if you have to re-attempt a chore a few times per day.

Leave a comment