truelabelRequest dataEarnRequest

Primary-source evidence matrix

Robot Foundation Model Data Evidence Matrix: What Egocentric Video Can and Cannot Prove

Most content says "egocentric video helps robots." That sentence is true and nearly useless, because it hides which claim is being made and which source backs it. This matrix normalizes the primary literature into twelve rows. Each row names the source, the reported unit or modality, exactly what the source supports, exactly what it does not support, and what a buyer should specify because of it. The one-line synthesis across all twelve rows: egocentric human video is strong, well-evidenced pretraining and affordance data, and it is consistently insufficient on its own — every deployable-policy result pairs it with action labels, sensorized human demonstrations, or robot-native data, and validates on held-out physical evaluation. Read the rows and you'll see the same shape repeat: a real gain in a specific setup, plus a specific limit. We built this before writing anything narrative, on purpose. If a claim isn't in a row with a source and a limit, it doesn't belong in the cluster.

Updated 2026-07-1913 min read
By Truelabel Team
Reviewed by Truelabel Team ·
evidence for egocentric data in robot learning

How to read the matrix

Each claim has a stable ID (`EGO-VLA-0NN`) so it can be cited and re-checked. Numbers are reported exactly as the source states them, with their unit and denominator preserved — we don't normalize "≈21K hours" into a round number, and we don't compare figures across papers unless the source itself does. All rows were checked on 2026-07-19. Source tiers: paper and official project page / model card are the only tiers used for numeric and performance claims here. The Jim Fan / Sequoia talk that seeded this topic is a Tier-3 discovery source and appears in no row. NVIDIA's WAM technical blog and the RLDS/LeRobot/MCAP format docs are official context and are referenced in the methodology, not as evidence rows.

The evidence matrix (scan view)

IDSource URLModality / reported unitClaim supportedLimitationBuyer/research implication
EGO-VLA-001RT-2 — https://proceedings.mlr.press/v229/zitkovich23a.htmlVLA; ~6,000 robot eval trialsWeb vision-language knowledge transfers to robot control; emergent semanticsThat web/text replaces physical robot dataBudget for robot action data even behind a strong VLM
EGO-VLA-002OpenVLA — https://arxiv.org/abs/2406.09246 · card — https://huggingface.co/openvla/openvla-7b7B params; 970k robot episodes (Open X-Embodiment); discrete action tokensOpen VLAs depend on robot episodes; diversity can beat scale (+16.5% abs. vs 55B RT-2-X in its eval)Zero-shot beyond represented embodiments/domains; not a diffusion headMatch training embodiments/domains to your target robot
EGO-VLA-003DreamZero — https://dreamzero0.github.io/ · paper — https://arxiv.org/abs/2602.15922World/action model; >2×, 7Hz, >42%, 30-min adaptJoint future-video+action prediction can improve generalization in reported settingsBroad/production consensus; single-team lab resultTreat WAM gains as setting-specific; require your own eval
EGO-VLA-004Deep Visual Foresight — https://arxiv.org/abs/1610.00696Action-conditioned video prediction + MPC; unlabeled robot interactionPredictive models can plan manipulation in controlled settingsThat passive video alone trains deployable policiesAction-conditioning (or a proxy) is what makes prediction useful for control
EGO-VLA-005Ego4D — https://ego4d-data.org/ · paper — https://arxiv.org/abs/2110.070583,670 h; 923 participants; 74 locations; 9 countriesLarge-scale first-person hand-object activity sourceRobot action supervision or commercial-use clearanceUse as evidence/benchmark; plan labels and rights separately
EGO-VLA-006Ego-Exo4D — https://ego-exo4d-data.org/ · paper — https://arxiv.org/abs/2311.182591,286.3 h; 740 wearers; 13 cities; synced ego+exoPaired ego/exo covers first-person blind spots for skilled activityThat every task needs both viewsPair views when body pose / scene geometry matter
EGO-VLA-007EgoScale — https://research.nvidia.com/labs/gear/egoscale/ · paper — https://arxiv.org/abs/2602.1671020,854 h action-labeled ego; R²=0.9983; +54% on 22-DoF handEgo pretraining + aligned robot mid-training improves dexterous manipulation; hours scale predictablyEgo video alone sufficiency; all robots/tasks/productionBudget for aligned robot data + action labels, not just hours
EGO-VLA-008HRP — https://arxiv.org/abs/2407.18911Hand/object/contact affordance labels; >15%; 5 tasks; 3,000+ trials; 3 morphologiesHuman-video affordance labels improve robot pretraining in reported tasksGeneralization beyond source contextSpecify affordance labels (contact, grasp points), not just object boxes
EGO-VLA-009UMI — https://umi-gripper.github.io/ · paper — https://arxiv.org/abs/2402.10329Sensorized handheld gripper + wrist camera demosSensorized human demos bridge human scale and robot-compatible learningPassive ego video sufficiency (its own stated embodiment gap)Sensorization/retargeting is the bridge; spec the interface
EGO-VLA-010DROID — https://droid-dataset.github.io/ · paper — https://arxiv.org/abs/2403.1294576,000 trajectories; 350 h; 564 scenes; 86 tasksRobot-native demos give embodiment-aligned trajectories + real-world grounding/evalThat teleop is always sufficient or easily scalableKeep a robot-native slice for alignment and evaluation
EGO-VLA-011BridgeData V2 — https://rail-berkeley.github.io/bridgedata/ · paper — https://proceedings.mlr.press/v229/walke23a.html~60,000 trajectories; 24 environments; language/goal conditioningReal robot data with language/goal conditioning grounds policy learningUniversal task or hardware coverageDon't assume public robot data covers your environment
EGO-VLA-012EgoNCE++ — https://arxiv.org/abs/2405.17719Egocentric video-language benchmark/limitation studyCurrent ego VLMs struggle with fine-grained hand-object dynamicsA blanket dismissal of egocentric dataTest fine-grained interaction; more ego video ≠ better action understanding

The evidence matrix (detail cards)

Each card carries the full field set: ID, source, source type, canonical URL, checked date, a close excerpt of the reported evidence, supported claim, limitation, and buyer implication.

EGO-VLA-001 — RT-2Source / type: RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control · paper (PMLR/CoRL) • URL: https://proceedings.mlr.press/v229/zitkovich23a.html · checked 2026-07-19 • Reported evidence (close excerpt): Co-fine-tunes a vision-language model on web vision-language tasks and robot trajectories, expressing actions as tokens; reports on the order of 6,000 robot evaluation trials and emergent semantic generalization to concepts absent from the robot data. • Supports: VLAs can carry web-scale semantic knowledge into robot control. • Does not support: Physical mastery; that internet data removes the need for physical demonstrations. • Buyer implication: A strong VLM backbone still needs paired robot action data for the target behavior.

EGO-VLA-002 — OpenVLASource / type: OpenVLA: An Open-Source Vision-Language-Action Model · paper + official model card • URL: https://arxiv.org/abs/2406.09246 · card https://huggingface.co/openvla/openvla-7b · checked 2026-07-19 • Reported evidence (close excerpt): 7B-parameter model trained on 970,000 robot episodes from Open X-Embodiment, predicting discretized action tokens; reports outperforming the 55B RT-2-X by 16.5% absolute task success in its evaluation; the model card states zero-shot use is limited to embodiments/domains in the training mix. • Supports: Modern open VLAs depend on robot-episode supervision; data diversity can outweigh parameter count. • Does not support: Zero-shot generality beyond represented embodiments. Also corrects a common error: OpenVLA does not use a diffusion action head. • Buyer implication: Specify the embodiments and domains you need represented; "open VLA" is not "works on any robot."

EGO-VLA-003 — DreamZero (world/action model)Source / type: DreamZero · official project page + paper • URL: https://dreamzero0.github.io/ · https://arxiv.org/abs/2602.15922 · checked 2026-07-19 • Reported evidence (close excerpt): A world action model on a video-diffusion backbone that jointly predicts future world states and robot actions; reports >2× improvement over VLA baselines on new tasks/environments, 7Hz closed-loop control after optimization, >42% improvement from 10–20 minutes of video-only demonstrations, and adaptation to a new YAM robot with 30 minutes of play data. • Supports: Joint future-video-and-action prediction can improve generalization in the reported experiments. • Does not support: Broad industry consensus or production readiness; these are recent, single-team, lab-reported results. • Buyer implication: If you're betting on WAM-style methods, plan your own held-out evaluation; don't generalize the numbers.

EGO-VLA-004 — Deep Visual ForesightSource / type: Deep Visual Foresight for Planning Robot Motion · paper (Google Research / ICRA) • URL: https://arxiv.org/abs/1610.00696 · checked 2026-07-19 • Reported evidence (close excerpt): Combines action-conditioned video prediction with model-predictive control using unlabeled robot interaction; performs nonprehensile pushing of novel objects without calibrated cameras, instrumented setups, depth, or explicit 3D object models. • Supports: Action-conditioned prediction can support real manipulation planning in controlled settings — and it predates current WAM branding. • Does not support: That passive, action-free video alone yields deployable control. • Buyer implication: The action-conditioning is the load-bearing part; raw video without an action signal (or proxy) is weaker for control.

EGO-VLA-005 — Ego4DSource / type: Ego4D: Around the World in 3,000 Hours of Egocentric Video · dataset + paper • URL: https://ego4d-data.org/ · https://arxiv.org/abs/2110.07058 · checked 2026-07-19 • Reported evidence (close excerpt): 3,670 hours of first-person daily activity from 923 participants across 74 locations and 9 countries, with modalities including audio, gaze, 3D scans, stereo, synchronized cameras, and narrations; benchmarks include hand-object interaction and forecasting. • Supports: A large-scale source of first-person hand-object activity for research. • Does not support: Robot action labels or commercial-use clearance by itself. • Buyer implication: Use Ego4D to scope terminology and benchmarks; plan action labels and a rights review separately.

EGO-VLA-006 — Ego-Exo4DSource / type: Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives · dataset + paper • URL: https://ego-exo4d-data.org/ · https://arxiv.org/abs/2311.18259 · checked 2026-07-19 • Reported evidence (close excerpt): 1,286.3 hours of skilled activity from 740 camera wearers across 13 cities, with synchronized first-person Aria glasses and 4–5 exocentric GoPros plus gaze, IMU, point clouds, audio, camera pose, and expert commentary. • Supports: Paired ego/exo capture helps cover first-person blind spots (body pose, scene geometry) for skilled activity. • Does not support: That every robotics task requires both views. • Buyer implication: Add an exocentric view when body pose or workspace layout matters to the task; otherwise it's overhead.

EGO-VLA-007 — EgoScaleSource / type: EgoScale · official NVIDIA GEAR project + paper • URL: https://research.nvidia.com/labs/gear/egoscale/ · https://arxiv.org/abs/2602.16710 · checked 2026-07-19 • Reported evidence (close excerpt): Pretrains a VLA on 20,854 hours of action-labeled egocentric video using wrist motion and retargeted hand actions, followed by aligned human-robot mid-training; reports a log-linear scaling law between data hours and validation loss (R²=0.9983) and a 54% average success-rate improvement over no-pretraining on a 22-DoF dexterous hand. • Supports: Egocentric pretraining plus aligned robot data can improve dexterous manipulation in the reported setup; ego data hours scale predictably against loss. • Does not support: That egocentric video alone is enough; that the recipe generalizes to all robots, tasks, or production. • Buyer implication: The scaling win rides on aligned robot mid-training and action labels — budget those, not just hours of raw video.

EGO-VLA-008 — HRPSource / type: HRP: Human Affordances for Robotic Pre-training · paper (RSS) • URL: https://arxiv.org/abs/2407.18911 · checked 2026-07-19 • Reported evidence (close excerpt): Extracts hand, object, and contact affordance labels from human video to pretrain robot visual representations; reports over 15% improvement over baselines across 5 real-world tasks, 3,000+ robot trials, and 3 robot morphologies. • Supports: Human-video affordance labels (where/how to interact) can improve robot pretraining in the reported tasks. • Does not support: Generalization of the reported gains outside that context. • Buyer implication: Ask for affordance-level labels — contact events, grasp points — not just object bounding boxes.

EGO-VLA-009 — UMISource / type: Universal Manipulation Interface (UMI) · project + paper • URL: https://umi-gripper.github.io/ · https://arxiv.org/abs/2402.10329 · checked 2026-07-19 • Reported evidence (close excerpt): Uses a handheld gripper with a wrist camera and a carefully designed policy/action interface to collect human demonstrations that transfer to robot policies; the paper states that pure robot teleoperation is costly and that unstructured human video carries a large embodiment gap. • Supports: Sensorized human demonstrations are a practical bridge between scalable human behavior and robot-compatible policy learning. • Does not support: That passive egocentric video is sufficient — the authors say the opposite. • Buyer implication: If you want action-relevant human data, sensorize or retarget it; specify the capture interface up front.

EGO-VLA-010 — DROIDSource / type: DROID: A Large-Scale In-the-Wild Robot Manipulation Dataset · dataset + paper • URL: https://droid-dataset.github.io/ · https://arxiv.org/abs/2403.12945 · checked 2026-07-19 • Reported evidence (close excerpt): 76,000 demonstration trajectories (350 hours) collected across 564 scenes and 86 tasks, with synchronized RGB, calibration, depth, and language. • Supports: Robot-native demonstrations provide embodiment-aligned trajectories and real-world grounding for training and evaluation. • Does not support: That teleoperation is always sufficient or that it scales as easily as human video. • Buyer implication: Keep a robot-native slice in the plan specifically for embodiment alignment and evaluation.

EGO-VLA-011 — BridgeData V2Source / type: BridgeData V2: A Dataset for Robot Learning at Scale · dataset + paper (PMLR/CoRL, Walke et al.) • URL: https://rail-berkeley.github.io/bridgedata/ · https://proceedings.mlr.press/v229/walke23a.html · checked 2026-07-19 • Reported evidence (close excerpt): Roughly 60,000 real manipulation trajectories across 24 environments, with language and goal conditioning. • Supports: Real robot manipulation data with language/goal conditioning grounds policy learning. • Does not support: Universal task or hardware coverage. • Buyer implication: Public robot data is a starting point, not domain coverage — check environment and embodiment fit before relying on it.

EGO-VLA-012 — EgoNCE++Source / type: EgoNCE++ · paper (egocentric video-language limitation study) • URL: https://arxiv.org/abs/2405.17719 · checked 2026-07-19 • Reported evidence (close excerpt): Argues current egocentric video-language benchmarks overemphasize closed-set visual concepts, that models are biased toward objects rather than temporal dynamics, and that they show diminished performance on fine-grained hand-object interactions. • Supports: Current egocentric video-language models struggle with fine-grained hand-object understanding. • Does not support: A blanket dismissal of egocentric data — it's a bounded limitation, not a verdict. • Buyer implication: Don't assume more egocentric video fixes action understanding; test fine-grained interaction explicitly.

Methodology

Inclusion rule. A source enters the matrix only if it's a primary paper, an official project page, or an official model card, and only if it makes a claim about data for VLA or world-model training that we can state with a preserved unit and a limitation. That's why widely cited market numbers, vendor blog stats, and talk claims are absent — they don't clear the bar.

Number discipline. Figures are reported exactly as the source states them (hours, episodes, trajectories, participants, R², absolute percentage points, Hz) with the denominator kept. We do not convert approximate figures to exact ones, and we do not compare numbers across papers unless a source makes the comparison itself. The one cross-paper comparison shown — OpenVLA vs RT-2-X at 16.5% absolute — is reported by OpenVLA's own paper, not assembled by us.

Source tiers used. Tier 1 (papers, official project pages, model cards) for all numeric/performance rows. Tier 2 (official technical blogs, format docs) as context only: NVIDIA's WAM technical blog for the VLA/WAM/hybrid framing, and the RLDS[1], LeRobot[2], and MCAP[3] docs for delivery-format grounding. Tier 3 (the Jim Fan / Sequoia talk, podcasts, interviews) for discovery only — cited in no row.

Checked date. All rows were verified against their cited sources on 2026-07-19. Recent 2026 preprints (DreamZero, EgoScale) are labeled as recent and lab-reported; independent replication is still pending and the rows say so.

What this matrix does not prove

• It does not prove egocentric video alone trains a deployable robot. Every row that shows a manipulation gain also shows a dependency on action labels, sensorized demos, aligned robot data, or evaluation. • It does not rank the datasets. Ego4D, DROID, and BridgeData V2 answer different questions; "best" is task-dependent. • It does not establish that public datasets are commercially usable. Licensing and intended-use review are separate steps that live outside this table. • It does not endorse any TrueLabel performance claim. TrueLabel's contribution here is the normalization discipline — naming source, unit, limit, and implication — not a proprietary result. • It does not treat forecasts as evidence. Timeline predictions from talks and keynotes are excluded by design.

Use these to move from category-level context into specific task, dataset, format, and comparison detail.

External references and source context

  1. RLDS GitHub repository

    Primary or official source cited by the authored page

    GitHub
  2. LeRobot GitHub repository

    Primary or official source cited by the authored page

    GitHub
  3. MCAP

    Primary or official source cited by the authored page

    Foxglove
  4. Physical AI data-spec generator

    Authored internal route

    truelabel.ai
  5. RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

    EGO-VLA-001: RT-2 connects vision-language pretraining with executable robot action outputs and reports semantic generalization. Limitation: The source does not show that web or language data replaces physical robot trajectories.

    Proceedings of Machine Learning Research
  6. OpenVLA: An Open-Source Vision-Language-Action Model

    EGO-VLA-002: OpenVLA is trained on robot episodes and maps visual observations plus language to robot actions. Limitation: Do not infer universal zero-shot control outside represented embodiments and domains or describe the base model as diffusion-headed.

    arXiv
  7. World Action Models are Zero-shot Policies

    EGO-VLA-003: DreamZero reports a world-action-model recipe that jointly predicts video/world states and actions. Limitation: Its performance and adaptation results are recent lab reports in specific robots, tasks, environments, and baselines.

    arXiv
  8. Deep Visual Foresight for Planning Robot Motion

    EGO-VLA-004: Action-conditioned video prediction can support robot manipulation planning in a controlled research setting. Limitation: The work does not establish that passive human egocentric video alone trains deployable control policies.

    arXiv
  9. Ego4D: Around the World in 3,000 Hours of Egocentric Video

    EGO-VLA-005: Ego4D establishes egocentric daily-life video as a large research substrate for activity, forecasting, and hand-object study. Limitation: Dataset scale is neither robot action supervision nor automatic commercial-use clearance.

    arXiv
  10. Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives

    EGO-VLA-006: Ego-Exo4D provides paired actor and external viewpoints for studying skilled activity and first-person blind spots. Limitation: The source does not prove that every robotics task needs both viewpoints.

    arXiv
  11. EgoScale: Scaling Dexterous Manipulation with Diverse Egocentric Human Data

    EGO-VLA-007: EgoScale separately reports log-linear scaling between egocentric data hours and validation loss, plus a task-success improvement after egocentric pretraining and aligned human-robot training in its dexterous-hand evaluation. Limitation: The reported dexterous-hand results do not prove that egocentric video alone is sufficient or that the result generalizes to all robots.

    arXiv
  12. HRP: Human Affordances for Robotic Pre-Training

    EGO-VLA-008: HRP shows a concrete route from human-video affordances to robot representation pretraining in its reported tasks. Limitation: Reported improvements cannot be generalized beyond the paper's tasks, cameras, robot morphologies, and evaluation protocol.

    arXiv
  13. Universal Manipulation Interface: In-The-Wild Robot Teaching Without In-The-Wild Robots

    EGO-VLA-009: UMI demonstrates sensorized human collection designed to produce robot-relevant observations and action representations. Limitation: The carefully designed hardware and policy interface are evidence against treating unstructured passive video as equivalent supervision.

    arXiv
  14. DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset

    EGO-VLA-010: DROID supplies robot-native demonstrations for real-world embodiment grounding and evaluation research. Limitation: The dataset does not prove teleoperation is sufficient for every task or easy to scale to arbitrary coverage.

    arXiv
  15. BridgeData V2: A Dataset for Robot Learning at Scale

    EGO-VLA-011: BridgeData V2 is primary evidence for grounding goal- and language-conditioned robot learning in real trajectories. Limitation: Its tasks, environments, and hardware do not establish universal deployment coverage.

    Proceedings of Machine Learning Research
  16. Do Egocentric Video-Language Models Truly Understand Hand-Object Interactions?

    EGO-VLA-012: EgoNCE++ provides limitation evidence that current egocentric video-language models can miss fine-grained interaction dynamics. Limitation: It is a targeted benchmark finding, not a blanket rejection of egocentric video pretraining.

    arXiv

FAQ

What evidence supports using egocentric data for robot foundation models?

Twelve primary sources, normalized above. In short: Ego4D and Ego-Exo4D establish first-person data at scale; HRP and EgoScale show human video improving robot pretraining in specific setups; RT-2 and OpenVLA anchor the VLA side; DreamZero and Deep Visual Foresight anchor action-conditioned prediction; UMI, DROID, and BridgeData V2 supply the sensorized-human and robot-native layers; EgoNCE++ marks the limits. Each row cites its source and states what it does not prove.

Does egocentric video prove a robot policy will transfer?

No. It supports pretraining and affordance learning. Transfer to a deployable policy is demonstrated only with action alignment (retargeting, sensorized demos, or robot data) and held-out physical evaluation — see rows EGO-VLA-007, EGO-VLA-009, and EGO-VLA-010.

Why isn't the Jim Fan / Sequoia video in the matrix?

Because a talk is a discovery source, not evidence. It pointed us toward world models, sensorized human data, and egocentric pretraining; the technical claims are then re-supported by the papers and project pages in the rows. Any statement that only traces back to the talk was dropped.

How current are these numbers?

All rows were checked on 2026-07-19. Older, well-replicated work (Deep Visual Foresight, Ego4D, DROID) is stable. The 2026 preprints (DreamZero, EgoScale) are flagged as recent and lab-reported, and their rows explicitly limit generalization.

What should I request in a dataset because of this matrix?

Specify the layer you need and its labels: viewpoint and sync, task language, hand/wrist pose, contact and object-state labels, embodiment-aligned actions where control is the goal, rights/consent artifacts, and an evaluation split. Draft it with the data-spec generator.

Want the matrix turned into a data request?

Use the physical AI data-spec generator to translate these rows into capture, annotation, and rights requirements, or post a spec to scope a rights-cleared sample packet with QA evidence. TrueLabel normalizes robotics-data evidence and delivers in RLDS, LeRobot, MCAP, and custom schemas with per-trajectory provenance — no model-gain or delivery-time promises.

Post a spec