rvachev.orgEN / RU / 🤖
← Back to essays
· Essay · 2 min

Data Pyramid for Robotics

Varun Nair drew a 'data pyramid for robotics' - a map of what VLA models are fed.

🤖 Varun Nair drew a 'data pyramid for robotics' - a decent map of what modern VLA models are actually fed.

At the core of any embodied AI lies one compromise: fidelity versus scalability. The closer the data is to a real robot, the less of it can be collected; the more there is, the further it is from what the robot needs to do with its hands.

The pyramid consists of five layers, from top to bottom - from expensive and precise to cheap and widespread:

- Teleoperation - a human directly controls the robot itself, one hour of operator time gives one hour of data one-to-one. This is how DROID was collected (Stanford, Berkeley, TRI, and 13 other institutes) - 76 thousand trajectories, 350 hours, 18 identical Franka Panda with Oculus Quest on three continents.
- Hand-held grippers in the style of UMI - a person holds the robot's gripper without a manipulator, a GoPro records the event, and the skill is then transferred to any hardware. DexCap does the same with gloves and motion capture.
- Simulation - unlimited runs, but the price is the sim-to-real gap: physics and rendering in the simulation are slightly deceptive, and a skill honed in Isaac Lab or ManiSkill sometimes falls apart on a real robot.
- Egocentric video - Ego4D and Ego-Exo4D are shot on human hands from a first-person perspective, actions are not recorded by the robot but reconstructed from the video.
- Random internet video & text - like HowTo100M: a million clips from YouTube, web-scale, but there are neither robot bodies nor actions.

The industry is betting on different layers. Figure AI and 1X are building their own fleet of robots to own the data entirely. Tesla in 2025 abandoned mocap suits for teleoperation in favor of cameras on a helmet and backpack - betting that learning can also be done on regular video. Physical Intelligence and Skild AI (valued at around $14 billion after a round in January 2026) are building a universal model on top of a mix of all five layers at once. Startup XDOF in 2026 raised $70 million for infrastructure specifically for the upper, teleoperation layers and released ABC-130K - the largest open teleoperation dataset, 130 thousand demonstrations.

https://x.com/_varunnair/status/2081619067535065103

#ai@rvnikita_blog #robotics@rvnikita_blog #vla@rvnikita_blog #varun_nair@rvnikita_blog

Data Pyramid for Robotics — illustration