Figure Index: The Global Economy of Robot Training Data

Figure's new app, Index, pays ordinary people to film themselves doing chores — folding laundry, washing dishes, mopping floors — and turns those first-person videos into robot training fuel. Launched globally on August 26 after four months of testing, Index already covers 108 countries, has collected 16 million clips, and has paid creators $15 million; Figure has committed more than $1 billion to data and compute over the next 12 months. The consumer-app story is almost beside the point. The structural signal is that the humanoid race has entered its data phase: after bodies and models, the next battlefield is real human behavior — and it has just become a purchasable commodity.

Why "data" became the bottleneck

LLMs got their training fuel almost free from the internet — GPT-4 alone trained on more than 13 trillion tokens of text. Embodied robots have no such internet. Industry estimates put embodied-AI data demand at more than 1,000× what LLMs need; the world's 8.1 billion people produce roughly 100 billion hours of physical-interaction data every day, and almost none of it is ever recorded. That asymmetry is why Figure — a company valued around $39 billion — treats Index not as a side project but as core strategy. Its Helix model is an end-to-end VLA (vision-language-action) system: a 7B-parameter vision-language model that understands tasks, plus an 80M-parameter visuo-motor network that handles real-time control. Such a model demands data that is massive, diverse, and physically grounded — exactly what no vendor could supply.

How Index works

Sign up, record first-person footage of yourself doing household or workplace tasks, upload it, and get paid based on data quality and quantity. You can also hire gig workers to do chores while you film. What Figure buys is not labels or annotations — it is raw, native video of humans solving physical problems in real environments: how fingers pinch a folded shirt, how the wrist rotates while mopping, how you judge a dish clean. Simulation cannot synthesize that granularity, and laboratories cannot capture it at scale.

The platform stats show what "diverse data" means operationally: every 1,000 hours of Index video contains, on average, 373 unique tasks, 1,146 manipulated objects, and 116 unique environments. The back end ingests 30 minutes of video upload per second, running 24/7. The machine isn't a labeling pipeline — it is a data mine with a payment rail attached.

Five routes to robot data

China's leading embodied players are running the same race with different equipment:

  • Figure — crowdsourcing platform + economic incentive + global reach. The only model with simultaneous mass upload.
  • Zhiyuan (AgiBot) — a self-built 2,000 m² collection site with 850 TB of real data (AgiBot World, 217 tasks, 3,000+ objects), open-sourced in part; by some counts NVIDIA's GR00T N1 trains on more than 80% real data from Zhiyuan's open datasets.
  • Unitree — a hardware-plus-data-pipeline model; its UnifoLM-WBT teleoperation dataset (340 hours, 1.89M trajectories) relies on community contribution to keep growing.
  • Zizhang ("independent variable" in Chinese) — 100% real-scene collection, coined "milk data" (noisy home reality) versus "sugar-water data" (clean lab data); its QUANXTA Zero body-free collection scheme claims to cut data cost for simple tasks by about 60%.
  • UBTech and others — centralized collection centers and closed internal loops, or outsourced third-party labeling.

Every route except Figure's has a ceiling: heavy-asset sites don't scale cheaply, open communities grow slowly, outsourcing makes quality hard to control. None of them has several million people uploading from 108 countries at once. Few-shot routes (GEN-1.5 learns a task from a few seconds of demo) save data but cannot substitute for exponentially scaled real-world diversity.

The flywheel and the moat

Index is a classic network-effects loop: more uploads → more task and environment diversity → stronger models → more capable robots → a more valuable platform → more uploaders paid to join. Once spinning, that flywheel is hard to chase. Three frictions slow copycats, especially in China: privacy and compliance (filming inside private homes trips portrait-rights and data-security regulation far stricter than the user-consent plus anonymization path Figure can take in the West), incentive economics (Chinese labor is cheaper, but so is the motivational power of a small payout — a local Index would need heavy subsidies to make people's time worth selling), and data diversity (a single-country dataset cannot replicate 108 countries' worth of climates, building types, and everyday routines). China's counterweights are real — 1.4 billion people, the world's most complete manufacturing chain, lower operating costs — but the data gap is currently widening, not closing. The data pipeline and the 10,000-unit wall in robot mass production are two ends of the same bottleneck: scale data first, then scale manufacturing.

Standards: ponds vs. a sea

The deeper problem is interoperability. Every player today uses its own collection formats, its own annotation conventions, its own scene definitions — datasets that cannot talk to each other or train together. Scale, under those conditions, produces isolated ponds rather than a sea. Whoever defines the common rules for embodied data earns what amounts to the interface-definition power of the entire industry. The UBS industrial analyst reading of the market is blunt: the "EV moment" for humanoids has not arrived — most shipments still go to research institutions and data-collection centers, and the real commercial inflection point is still ahead. Today's data race is precisely how the chips are being stacked for that moment.

What to watch next

Three signals worth tracking: first, whether Index's flywheel holds — watch the per-1,000-hours diversity metrics (373 tasks, 1,146 objects, 116 environments) over the promised $1 billion year. Second, whether a "Chinese Index" emerges; if 1.4 billion people's household reality gets wired into one platform, the compliance-and-incentive design will be the real product. Third, standards: Zhiyuan already open-sourced part of AgiBot World, and NVIDIA's GR00T N1 feeds on it — that is the template for data that "speaks" across companies. When embodied data becomes tradable and interoperable, the race stops being about who has the biggest dataset and starts being about who owns the pipes.

A related thread: this data-economy push is happening in parallel with the shared-brain paradigm shift in embodied AI — cross-embodiment models that let rival robots run on one brain. Fewer, more general brains plus bigger, more diverse data is a compounding combination, and it is exactly the pair that decides which robots actually leave the lab.

FAQ

Q: Does Figure really pay people to film their chores?
A: Yes. In four months of testing, Index collected 16 million clips across 108 countries and paid creators $15 million, with payments tiered by data quality and quantity — and Figure has pledged over $1 billion for the next 12 months of data and compute.

Q: Why can't robots just learn from simulation instead of paying humans?
A: Simulation struggles with the messy, fine-grained physics of real life — the "milk data" of noisy homes versus the "sugar-water data" of clean labs. Embodied models need real human behavior at a scale estimated at 1,000× an LLM's data appetite, which is why direct capture beats synthetic generation for now (generated-world-model training is the other route, see the Veeda virtual robot training ground).

Q: Is this just crowdsourced data labeling?
A: No. Labeling annotates existing content; Index captures never-before-recorded human physical behavior — first-person video of real tasks in real environments — which becomes raw training material for VLA systems like Helix (7B vision-language model + 80M visuo-motor network).

Leave a Comment

Scroll to top