truelabelRequest dataEarnRequest

World-action models

World-action models: what data each stage actually needs

A world-action model (WAM) is not one fixed data recipe. In NVIDIA’s two-stage framing, video / world-model pretraining and robot-action fine-tuning need different inputs; the survey also distinguishes passive, controllable, inverse-dynamics, joint, and latent formulations. This guide maps what each primary source can prove, what it cannot, and which quantities remain unknown.

Updated 2026-07-227 min read
By Truelabel Team
Reviewed by Truelabel Team ·
world action model

Comparison

World-action models: what data each stage actually needs comparison table
ClaimEntitySource-reported versionNormalized fieldNormalized valueUnitPrimary source URLSource typeSource IDExact locatorChecked dateRetrieval hashConfidenceReview statusWAM stageParadigmModality / reported unitSupported claimLimitationBuyer implication
WAM-STAGE-001nvidia-wam-stage-framingNVIDIA Technical Blog, 2026-07-15stage boundaryvideo/world-model pretraining; robot-action fine-tuningstagehttps://developer.nvidia.com/blog/pretrained-to-imagine-fine-tuned-to-act-the-rise-of-world-action-models/vendordocs-nvidia-world-action-modelsSections ‘Pretrained to Imagine’ and ‘Fine-Tuned to Act’2026-07-22unknown — no immutable publisher snapshot is exposedmediumhumanpretraining → robot-action fine-tuningtwo-stage hybridvideo for pretraining; robot trajectories for fine-tuning; quantity not reportedThe source separates video/world-model pretraining from later robot-action fine-tuning.source-reported, not independently validated; this vendor framing is not proof that every WAM uses paired actions during pretraining.Budget passive video and embodiment-aligned robot actions as different acquisition stages.
WAM-DISCOURSE-002awesome-wam-research-areaAwesome-WAM repository, checked 2026-07-22research-area statusnamed curated areadiscourse signalhttps://github.com/OpenMOSS/Awesome-WAMprojectproject-github-com-openmoss-awesome-wamRepository README taxonomy and paper list2026-07-22unknown — repository commit was not exposed by the reviewed pagemediumhumanfield taxonomymultiple WAM formulationsnot applicable; catalog, not a training reportThe repository curates work under the World Action Model name.source-reported, not independently validated; a curated list does not prove performance, dataset scale, or one settled definition.Treat WAM as an evolving family name and require a stage-specific data specification from each model team.
WAM-PARADIGM-003world-model-survey-taxonomyarXiv:2605.00080world-model paradigmpassive; controllable; inverse-dynamics; joint; latentparadigmhttps://arxiv.org/pdf/2605.00080paperpaper-arxiv-org-abs-2605-00080Survey taxonomy sections2026-07-22unknown — no retrieval snapshot is committedmediumhumanarchitecture selection before data specificationpassive / controllable / inverse-dynamics / joint / latentvaries by paradigm; the survey does not report one universal data unitThe survey distinguishes multiple world-model formulations rather than one universal action-label requirement.source-reported, not independently validated; this preprint taxonomy may change and does not by itself validate a robot policy.Ask which formulation is being trained before specifying action labels, controls, or evaluation data.
WAM-JOINT-004wa-rl-joint-optimizationCVPR 2026 Workshops proceedingsonline optimization targetworld model and actor jointly optimizedmodel componenthttps://openaccess.thecvf.com/content/CVPR2026W/GigaBrainChallenge/supplemental/Qian_WA-RL_World-Action_Model_CVPRW_2026_supplemental.pdfpaperpaper-cvf-wa-rl-2026Method and supplemental experiments2026-07-22unknown — no retrieval snapshot is committedhighhumanonline robot optimizationjoint world model + actorexpert trajectories and online environment interaction; universal quantity not reportedWA-RL is a concrete formulation that jointly optimizes a world model and actor through online interaction.source-reported, not independently validated; the workshop result is setting-specific and does not establish a generally deployable recipe.Plan an online-interaction and held-out evaluation budget if the selected formulation updates both components after demonstrations.

How this matrix was built

Each row ties one durable claim to one primary source and tags the WAM stage and paradigm it applies to. TrueLabel analysis is limited to stage/paradigm attribution and cross-tabulation; source-reported text remains separate. We do not assert a quantity the source does not state, and the HTML matrix, CSV, and JSON are deterministic projections of one record set.

WAM-STAGE-001: pretraining → robot-action fine-tuning

The source separates video/world-model pretraining from later robot-action fine-tuning. The evidence is the primary source.

FieldValue
ClaimWAM-STAGE-001
Entitynvidia-wam-stage-framing
Source-reported versionNVIDIA Technical Blog, 2026-07-15
Normalized fieldstage boundary
Normalized valuevideo/world-model pretraining; robot-action fine-tuning
Unitstage
Primary source URLhttps://developer.nvidia.com/blog/pretrained-to-imagine-fine-tuned-to-act-the-rise-of-world-action-models/
Source typevendor
Source IDdocs-nvidia-world-action-models
Exact locatorSections ‘Pretrained to Imagine’ and ‘Fine-Tuned to Act’
Checked date2026-07-22
Retrieval hashunknown — no immutable publisher snapshot is exposed
Confidencemedium
Review statushuman
WAM stagepretraining → robot-action fine-tuning
Paradigmtwo-stage hybrid
Modality / reported unitvideo for pretraining; robot trajectories for fine-tuning; quantity not reported
Supported claimThe source separates video/world-model pretraining from later robot-action fine-tuning.
Limitationsource-reported, not independently validated; this vendor framing is not proof that every WAM uses paired actions during pretraining.
Buyer implicationBudget passive video and embodiment-aligned robot actions as different acquisition stages.
WAM-STAGE-001 evidence detail

WAM-DISCOURSE-002: field taxonomy

The repository curates work under the World Action Model name. The evidence is the primary source.

FieldValue
ClaimWAM-DISCOURSE-002
Entityawesome-wam-research-area
Source-reported versionAwesome-WAM repository, checked 2026-07-22
Normalized fieldresearch-area status
Normalized valuenamed curated area
Unitdiscourse signal
Primary source URLhttps://github.com/OpenMOSS/Awesome-WAM
Source typeproject
Source IDproject-github-com-openmoss-awesome-wam
Exact locatorRepository README taxonomy and paper list
Checked date2026-07-22
Retrieval hashunknown — repository commit was not exposed by the reviewed page
Confidencemedium
Review statushuman
WAM stagefield taxonomy
Paradigmmultiple WAM formulations
Modality / reported unitnot applicable; catalog, not a training report
Supported claimThe repository curates work under the World Action Model name.
Limitationsource-reported, not independently validated; a curated list does not prove performance, dataset scale, or one settled definition.
Buyer implicationTreat WAM as an evolving family name and require a stage-specific data specification from each model team.
WAM-DISCOURSE-002 evidence detail

WAM-PARADIGM-003: architecture selection before data specification

The survey distinguishes multiple world-model formulations rather than one universal action-label requirement. The evidence is the primary source.

FieldValue
ClaimWAM-PARADIGM-003
Entityworld-model-survey-taxonomy
Source-reported versionarXiv:2605.00080
Normalized fieldworld-model paradigm
Normalized valuepassive; controllable; inverse-dynamics; joint; latent
Unitparadigm
Primary source URLhttps://arxiv.org/pdf/2605.00080
Source typepaper
Source IDpaper-arxiv-org-abs-2605-00080
Exact locatorSurvey taxonomy sections
Checked date2026-07-22
Retrieval hashunknown — no retrieval snapshot is committed
Confidencemedium
Review statushuman
WAM stagearchitecture selection before data specification
Paradigmpassive / controllable / inverse-dynamics / joint / latent
Modality / reported unitvaries by paradigm; the survey does not report one universal data unit
Supported claimThe survey distinguishes multiple world-model formulations rather than one universal action-label requirement.
Limitationsource-reported, not independently validated; this preprint taxonomy may change and does not by itself validate a robot policy.
Buyer implicationAsk which formulation is being trained before specifying action labels, controls, or evaluation data.
WAM-PARADIGM-003 evidence detail

WAM-JOINT-004: online robot optimization

WA-RL is a concrete formulation that jointly optimizes a world model and actor through online interaction. The evidence is the primary source.

FieldValue
ClaimWAM-JOINT-004
Entitywa-rl-joint-optimization
Source-reported versionCVPR 2026 Workshops proceedings
Normalized fieldonline optimization target
Normalized valueworld model and actor jointly optimized
Unitmodel component
Primary source URLhttps://openaccess.thecvf.com/content/CVPR2026W/GigaBrainChallenge/supplemental/Qian_WA-RL_World-Action_Model_CVPRW_2026_supplemental.pdf
Source typepaper
Source IDpaper-cvf-wa-rl-2026
Exact locatorMethod and supplemental experiments
Checked date2026-07-22
Retrieval hashunknown — no retrieval snapshot is committed
Confidencehigh
Review statushuman
WAM stageonline robot optimization
Paradigmjoint world model + actor
Modality / reported unitexpert trajectories and online environment interaction; universal quantity not reported
Supported claimWA-RL is a concrete formulation that jointly optimizes a world model and actor through online interaction.
Limitationsource-reported, not independently validated; the workshop result is setting-specific and does not establish a generally deployable recipe.
Buyer implicationPlan an online-interaction and held-out evaluation budget if the selected formulation updates both components after demonstrations.
WAM-JOINT-004 evidence detail

Limitations

WAM terminology may change, so this page anchors on the durable separation between video/world-model learning and action/control alignment. It shows that the cited claims are checkable; it does not settle the architecture debate, prove independent replication, or resolve the pre-existing generic world-model query overlap.

Turn the stage boundary into a data brief

First identify the paradigm and training stage. Then specify the observed modality, action representation, embodiment, synchronization, provenance, and held-out physical evaluation separately. Use the linked canonical guides for neighboring procurement and data-stack questions instead of treating this page as their owner.

  1. 01

    Name the formulation

    Record whether the system is passive, controllable, inverse-dynamics, joint, latent, or another explicitly documented formulation.

  2. 02

    Separate pretraining from alignment

    Do not infer robot action labels from a passive-video pretraining source. Record each stage's actual inputs independently.

  3. 03

    Reserve held-out physical evaluation

    A plausible predicted video is not evidence that a robot completed the task. Keep target-embodiment evaluation outside the training mixture.

Use these to move from category-level context into specific task, dataset, format, and comparison detail.

FAQ

Does every world action model pretrain on paired video and robot actions?

No. The cited stage framing separates video or world-model pretraining from robot-action fine-tuning, and the survey describes multiple formulations. Check the selected model's exact stage and paradigm before specifying labels.

Can this matrix prove that one WAM will work on my robot?

No. It records what each primary source supports and where that evidence stops. Target-embodiment evaluation and domain-specific data remain required.

Turn the stage boundary into a sample brief

TrueLabel is a physical AI data marketplace: post a specification, then matched suppliers return sample packets for review with rights, consent, and per-trajectory provenance artifacts attached.

Source paired video→action data for your WAM pipeline