The Data Pyramid in Robotics
TLDR
Robotics lacks an equivalent to the internet's free pre-training corpus that powered frontier LLMs, creating a severe data bottleneck that researchers address by combining seven types of data arranged in a pyramid. The largest open robot dataset, Open X-Embodiment, contains around 1 million trajectories across 22 robot types, while robots need somewhere between 1 million and 10 million hours of training data compared to internet-scale text corpora. Different data sources—from YouTube videos and egocentric human footage to teleoperation and deployed robot work—each contribute different priors and fidelity levels, with the field recognizing that no single source will solve the problem but rather a mixture tailored to specific deployment requirements will be needed.
Why it matters
There is no equivalent source of physical data for robotics comparable to internet text data for LLMs. All attempts to achieve physical AGI involve working through this data bottleneck.