World-action models
World-action models: what data each stage actually needs
A world-action model (WAM) is not one fixed data recipe. In NVIDIA’s two-stage framing, video / world-model pretraining and robot-action fine-tuning need different inputs; the survey also distinguishes passive, controllable, inverse-dynamics, joint, and latent formulations. This guide maps what each primary source can prove, what it cannot, and which quantities remain unknown.
Comparison
| Claim | Entity | Source-reported version | Normalized field | Normalized value | Unit | Primary source URL | Source type | Source ID | Exact locator | Checked date | Retrieval hash | Confidence | Review status | WAM stage | Paradigm | Modality / reported unit | Supported claim | Limitation | Buyer implication |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| WAM-STAGE-001 | nvidia-wam-stage-framing | NVIDIA Technical Blog, 2026-07-15 | stage boundary | video/world-model pretraining; robot-action fine-tuning | stage | https://developer.nvidia.com/blog/pretrained-to-imagine-fine-tuned-to-act-the-rise-of-world-action-models/ | vendor | docs-nvidia-world-action-models | Sections ‘Pretrained to Imagine’ and ‘Fine-Tuned to Act’ | 2026-07-22 | unknown — no immutable publisher snapshot is exposed | medium | human | pretraining → robot-action fine-tuning | two-stage hybrid | video for pretraining; robot trajectories for fine-tuning; quantity not reported | The source separates video/world-model pretraining from later robot-action fine-tuning. | source-reported, not independently validated; this vendor framing is not proof that every WAM uses paired actions during pretraining. | Budget passive video and embodiment-aligned robot actions as different acquisition stages. |
| WAM-DISCOURSE-002 | awesome-wam-research-area | Awesome-WAM repository, checked 2026-07-22 | research-area status | named curated area | discourse signal | https://github.com/OpenMOSS/Awesome-WAM | project | project-github-com-openmoss-awesome-wam | Repository README taxonomy and paper list | 2026-07-22 | unknown — repository commit was not exposed by the reviewed page | medium | human | field taxonomy | multiple WAM formulations | not applicable; catalog, not a training report | The repository curates work under the World Action Model name. | source-reported, not independently validated; a curated list does not prove performance, dataset scale, or one settled definition. | Treat WAM as an evolving family name and require a stage-specific data specification from each model team. |
| WAM-PARADIGM-003 | world-model-survey-taxonomy | arXiv:2605.00080 | world-model paradigm | passive; controllable; inverse-dynamics; joint; latent | paradigm | https://arxiv.org/pdf/2605.00080 | paper | paper-arxiv-org-abs-2605-00080 | Survey taxonomy sections | 2026-07-22 | unknown — no retrieval snapshot is committed | medium | human | architecture selection before data specification | passive / controllable / inverse-dynamics / joint / latent | varies by paradigm; the survey does not report one universal data unit | The survey distinguishes multiple world-model formulations rather than one universal action-label requirement. | source-reported, not independently validated; this preprint taxonomy may change and does not by itself validate a robot policy. | Ask which formulation is being trained before specifying action labels, controls, or evaluation data. |
| WAM-JOINT-004 | wa-rl-joint-optimization | CVPR 2026 Workshops proceedings | online optimization target | world model and actor jointly optimized | model component | https://openaccess.thecvf.com/content/CVPR2026W/GigaBrainChallenge/supplemental/Qian_WA-RL_World-Action_Model_CVPRW_2026_supplemental.pdf | paper | paper-cvf-wa-rl-2026 | Method and supplemental experiments | 2026-07-22 | unknown — no retrieval snapshot is committed | high | human | online robot optimization | joint world model + actor | expert trajectories and online environment interaction; universal quantity not reported | WA-RL is a concrete formulation that jointly optimizes a world model and actor through online interaction. | source-reported, not independently validated; the workshop result is setting-specific and does not establish a generally deployable recipe. | Plan an online-interaction and held-out evaluation budget if the selected formulation updates both components after demonstrations. |
How this matrix was built
Each row ties one durable claim to one primary source and tags the WAM stage and paradigm it applies to. TrueLabel analysis is limited to stage/paradigm attribution and cross-tabulation; source-reported text remains separate. We do not assert a quantity the source does not state, and the HTML matrix, CSV, and JSON are deterministic projections of one record set.
WAM-STAGE-001: pretraining → robot-action fine-tuning
The source separates video/world-model pretraining from later robot-action fine-tuning. The evidence is the primary source.
| Field | Value |
|---|---|
| Claim | WAM-STAGE-001 |
| Entity | nvidia-wam-stage-framing |
| Source-reported version | NVIDIA Technical Blog, 2026-07-15 |
| Normalized field | stage boundary |
| Normalized value | video/world-model pretraining; robot-action fine-tuning |
| Unit | stage |
| Primary source URL | https://developer.nvidia.com/blog/pretrained-to-imagine-fine-tuned-to-act-the-rise-of-world-action-models/ |
| Source type | vendor |
| Source ID | docs-nvidia-world-action-models |
| Exact locator | Sections ‘Pretrained to Imagine’ and ‘Fine-Tuned to Act’ |
| Checked date | 2026-07-22 |
| Retrieval hash | unknown — no immutable publisher snapshot is exposed |
| Confidence | medium |
| Review status | human |
| WAM stage | pretraining → robot-action fine-tuning |
| Paradigm | two-stage hybrid |
| Modality / reported unit | video for pretraining; robot trajectories for fine-tuning; quantity not reported |
| Supported claim | The source separates video/world-model pretraining from later robot-action fine-tuning. |
| Limitation | source-reported, not independently validated; this vendor framing is not proof that every WAM uses paired actions during pretraining. |
| Buyer implication | Budget passive video and embodiment-aligned robot actions as different acquisition stages. |
WAM-DISCOURSE-002: field taxonomy
The repository curates work under the World Action Model name. The evidence is the primary source.
| Field | Value |
|---|---|
| Claim | WAM-DISCOURSE-002 |
| Entity | awesome-wam-research-area |
| Source-reported version | Awesome-WAM repository, checked 2026-07-22 |
| Normalized field | research-area status |
| Normalized value | named curated area |
| Unit | discourse signal |
| Primary source URL | https://github.com/OpenMOSS/Awesome-WAM |
| Source type | project |
| Source ID | project-github-com-openmoss-awesome-wam |
| Exact locator | Repository README taxonomy and paper list |
| Checked date | 2026-07-22 |
| Retrieval hash | unknown — repository commit was not exposed by the reviewed page |
| Confidence | medium |
| Review status | human |
| WAM stage | field taxonomy |
| Paradigm | multiple WAM formulations |
| Modality / reported unit | not applicable; catalog, not a training report |
| Supported claim | The repository curates work under the World Action Model name. |
| Limitation | source-reported, not independently validated; a curated list does not prove performance, dataset scale, or one settled definition. |
| Buyer implication | Treat WAM as an evolving family name and require a stage-specific data specification from each model team. |
WAM-PARADIGM-003: architecture selection before data specification
The survey distinguishes multiple world-model formulations rather than one universal action-label requirement. The evidence is the primary source.
| Field | Value |
|---|---|
| Claim | WAM-PARADIGM-003 |
| Entity | world-model-survey-taxonomy |
| Source-reported version | arXiv:2605.00080 |
| Normalized field | world-model paradigm |
| Normalized value | passive; controllable; inverse-dynamics; joint; latent |
| Unit | paradigm |
| Primary source URL | https://arxiv.org/pdf/2605.00080 |
| Source type | paper |
| Source ID | paper-arxiv-org-abs-2605-00080 |
| Exact locator | Survey taxonomy sections |
| Checked date | 2026-07-22 |
| Retrieval hash | unknown — no retrieval snapshot is committed |
| Confidence | medium |
| Review status | human |
| WAM stage | architecture selection before data specification |
| Paradigm | passive / controllable / inverse-dynamics / joint / latent |
| Modality / reported unit | varies by paradigm; the survey does not report one universal data unit |
| Supported claim | The survey distinguishes multiple world-model formulations rather than one universal action-label requirement. |
| Limitation | source-reported, not independently validated; this preprint taxonomy may change and does not by itself validate a robot policy. |
| Buyer implication | Ask which formulation is being trained before specifying action labels, controls, or evaluation data. |
WAM-JOINT-004: online robot optimization
WA-RL is a concrete formulation that jointly optimizes a world model and actor through online interaction. The evidence is the primary source.
| Field | Value |
|---|---|
| Claim | WAM-JOINT-004 |
| Entity | wa-rl-joint-optimization |
| Source-reported version | CVPR 2026 Workshops proceedings |
| Normalized field | online optimization target |
| Normalized value | world model and actor jointly optimized |
| Unit | model component |
| Primary source URL | https://openaccess.thecvf.com/content/CVPR2026W/GigaBrainChallenge/supplemental/Qian_WA-RL_World-Action_Model_CVPRW_2026_supplemental.pdf |
| Source type | paper |
| Source ID | paper-cvf-wa-rl-2026 |
| Exact locator | Method and supplemental experiments |
| Checked date | 2026-07-22 |
| Retrieval hash | unknown — no retrieval snapshot is committed |
| Confidence | high |
| Review status | human |
| WAM stage | online robot optimization |
| Paradigm | joint world model + actor |
| Modality / reported unit | expert trajectories and online environment interaction; universal quantity not reported |
| Supported claim | WA-RL is a concrete formulation that jointly optimizes a world model and actor through online interaction. |
| Limitation | source-reported, not independently validated; the workshop result is setting-specific and does not establish a generally deployable recipe. |
| Buyer implication | Plan an online-interaction and held-out evaluation budget if the selected formulation updates both components after demonstrations. |
Limitations
WAM terminology may change, so this page anchors on the durable separation between video/world-model learning and action/control alignment. It shows that the cited claims are checkable; it does not settle the architecture debate, prove independent replication, or resolve the pre-existing generic world-model query overlap.
Turn the stage boundary into a data brief
First identify the paradigm and training stage. Then specify the observed modality, action representation, embodiment, synchronization, provenance, and held-out physical evaluation separately. Use the linked canonical guides for neighboring procurement and data-stack questions instead of treating this page as their owner.
- 01
Name the formulation
Record whether the system is passive, controllable, inverse-dynamics, joint, latent, or another explicitly documented formulation.
- 02
Separate pretraining from alignment
Do not infer robot action labels from a passive-video pretraining source. Record each stage's actual inputs independently.
- 03
Reserve held-out physical evaluation
A plausible predicted video is not evidence that a robot completed the task. Keep target-embodiment evaluation outside the training mixture.
Related pages
Use these to move from category-level context into specific task, dataset, format, and comparison detail.
FAQ
Does every world action model pretrain on paired video and robot actions?
No. The cited stage framing separates video or world-model pretraining from robot-action fine-tuning, and the survey describes multiple formulations. Check the selected model's exact stage and paradigm before specifying labels.
Can this matrix prove that one WAM will work on my robot?
No. It records what each primary source supports and where that evidence stops. Target-embodiment evaluation and domain-specific data remain required.
Turn the stage boundary into a sample brief
TrueLabel is a physical AI data marketplace: post a specification, then matched suppliers return sample packets for review with rights, consent, and per-trajectory provenance artifacts attached.
Source paired video→action data for your WAM pipeline