ModelsArchitectures & capability
NVIDIA’s Cosmos 3 positions omnimodal world models as policy-verification infrastructure for physical AI
Ming-Yu Liu’s interview frames Cosmos 3 as a model family linking language, video, audio and action, with near-term value in synthetic data, post-training and policy checkpoint triage—while leaving real-world validation and safety assurance unresolved.
Cosmos 3 is presented as an omnimodal physical-AI model family spanning text, image, video, audio and action for understanding, generation, simulation and action-related tasks. [9] [11] [17]
The described architecture separates autoregressive reasoning from diffusion-based generation, implying different serving, latency and evaluation considerations for text versus continuous modalities such as video, audio and actions. [1] [2] [8]
The reviewed evidence shows NVIDIA is treating world models as a broader development stack for robots, autonomous systems and physical-AI workflows, not just as video generators.
Read the full assessment
Official materials and the interview support claims about multimodal scope, model tiers and post-training recipes. The implication for practitioners is that Cosmos 3 may be most useful as infrastructure for synthetic data, debugging and policy selection. The business risk is mistaking simulated ranking or vendor benchmarks for deployment-grade safety evidence.
Executive brief
The interview positions NVIDIA Cosmos 3 as an “omnimodal” world-model family for physical AI: one model family intended to connect language, image, video, audio and action for physical-world reasoning, video/world generation, forward dynamics, inverse dynamics and robot policy generation. NVIDIA’s official Cosmos 3 page and GitHub release describe the same broad scope: a shared model architecture for understanding, generation, simulation and action across those modalities. Cosmos 3 — Cosmos Lab For practitioners, the most actionable claim is not that a neural simulator is ready to replace field testing.
Read the full section
The interview positions NVIDIA Cosmos 3 as an “omnimodal” world-model family for physical AI: one model family intended to connect language, image, video, audio and action for physical-world reasoning, video/world generation, forward dynamics, inverse dynamics and robot policy generation. NVIDIA’s official Cosmos 3 page and GitHub release describe the same broad scope: a shared model architecture for understanding, generation, simulation and action across those modalities. Cosmos 3 — Cosmos Lab
For practitioners, the most actionable claim is not that a neural simulator is ready to replace field testing. Liu’s practical claim is narrower: early value may come from policy verification and checkpoint triage, where the simulator only needs to preserve the ranking of policy candidates well enough to decide which policies deserve scarce real-world evaluation. In the transcript, this is explicitly framed as a way to reduce real-world test burden, not eliminate it (YouTube transcript, ~12:53–15:16).
For researchers and developers, the technical novelty discussed is a dual-tower Mixture-of-Transformers architecture: an autoregressive reasoner/VLM-like tower and a diffusion generator tower for continuous modalities. NVIDIA’s model card describes Cosmos3 as a MoT transformer with an autoregressive transformer for discrete token generation and a diffusion transformer for continuous multimodal generation. README.md · nvidia/Cosmos3-Edge at main
Evidence remains uneven. NVIDIA has released code, checkpoints, datasets and evaluation assets under OpenMDW-1.1, which improves inspectability and reproducibility relative to closed models. Cosmos 3: Omnimodal World Models for Physical AI But most headline performance statements are vendor-reported. Independent context and benchmark-methodology work, including RoboArena’s distributed double-blind robot evaluation design were found in the reviewed sources, but not broad independent reproduction of Cosmos 3’s strongest claims across robotics, driving, generation and safety-critical deployment. RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies
What changed and event timeline
Baseline
NVIDIA’s earlier Cosmos World Foundation Model platform framed physical AI as requiring both a “digital twin” of the AI policy and a world model of the environment. It also acknowledged that the world-foundation-model problem was not solved.
More detail
That paper described pretraining world foundation models and post-training them for camera control, robot manipulation and autonomous driving.
The Cosmos 3 technical report appeared as arXiv:2606.02800 and states that code, model checkpoints, curated synthetic datasets and evaluation benchmarks are available under the Linux Foundation’s OpenMDW-1.1 license.
Launch
NVIDIA announced Cosmos 3 at GTC Taipei on May 31, 2026, calling it an open physical-AI foundation model with MoT architecture for reasoning, world simulation and action prediction. The GitHub release was posted June 1, 2026.
The Hugging Face Cosmos3-Edge card lists Edge release on July 20, 2026, alongside model-version information for Edge and the earlier Nano/Super releases.
The Cosmos3-Edge model card records an update to the Edge generator checkpoint, runtime defaults, usage examples and benchmark results; this matters for reproducibility because results may depend on snapshot date.
Machine Learning Street Talk published the Liu interview
The episode’s key new value is interpretive: Liu explains how NVIDIA wants developers to think about Cosmos 3 as a foundation for policy verification, action modelling and eventually closed-loop neural simulation. The episode states that it is a paid NVIDIA partnership ().
Capabilities and access
Known model sizes and intended deployment tiers: Access paths are documented through Hugging Face model cards, the NVIDIA/cosmos GitHub repo, NVIDIA docs, and NVIDIA NIM/vLLM/TensorRT-LLM paths. The GitHub README warns that Nano/Super use Cosmos3OmniForConditionalGeneration, while Edge uses a separate integration and should not be loaded with the Omni class.
Read the full section
Known model sizes and intended deployment tiers:
- Cosmos3-Super — 64B: data-center tier; inputs include text, image, video and action; outputs include text, image, video, sound and action; suited, per NVIDIA docs, for high-quality synthetic-data generation and teacher-model use. Model Matrix — Cosmos
- Cosmos3-Nano — 16B: workstation/data-center tier; positioned as a speed/quality balance for post-training workflows. Model Matrix — Cosmos
- Cosmos3-Edge — 4B: edge/on-device tier; intended for Jetson AGX Orin, Thor or RTX PRO 6000-class deployment; inputs/outputs include text, image, video and action. Model Matrix — Cosmos
Access paths are documented through Hugging Face model cards, the NVIDIA/cosmos GitHub repo, NVIDIA docs, and NVIDIA NIM/vLLM/TensorRT-LLM paths. The GitHub README warns that Nano/Super use Cosmos3OmniForConditionalGeneration, while Edge uses a separate integration and should not be loaded with the Omni class. GitHub - NVIDIA/cosmos: NVIDIA Cosmos is an open platform of world models, datasets, and tools that enables developers to build Physical AI for robots, autonomous vehicles, smart infrastructure, and more. · GitHub
Technical analysis for researchers and developers
The interview describes a staged construction: start with a language model or vision-language model; connect video/image encoders; then initialize a generator from VLM weights. Liu contrasts an autoregressive reasoning tower with a bidirectional diffusion generator tower for video, audio and action (~02:46–04:57).
Read the full section
Architecture
The interview describes a staged construction: start with a language model or vision-language model; connect video/image encoders; then initialize a generator from VLM weights. Liu contrasts an autoregressive reasoning tower with a bidirectional diffusion generator tower for video, audio and action (~02:46–04:57). NVIDIA’s model card corroborates the high-level architecture: an autoregressive transformer for discrete token generation and a diffusion transformer for continuous multimodal generation, with non-text modalities synthesized through iterative denoising. README.md · nvidia/Cosmos3-Edge at main
The important implementation implication is that text reasoning and continuous world/action generation are not treated as the same decoding problem. Text can be generated next-token style; images, video, audio and action trajectories are generated through a diffusion-style process. That affects serving, latency, batching, evaluation and failure modes.
Action modelling
The interview frames a world model as a toolset rather than a single agreed definition. Liu highlights three robotics-relevant functions: forward dynamics — predict future observations given current state and action; inverse dynamics — infer action from visual transition; and policy — choose actions to achieve a task (~05:02–06:15). NVIDIA’s technical materials similarly describe action post-training for forward dynamics, inverse dynamics and policy generation. Develop Physical AI Reasoning, World, and Action Models with NVIDIA Cosmos 3 | NVIDIA Technical Blog
Time alignment across modalities
A key engineering issue is mismatched sampling rates: video frames, audio samples and robot actions operate at different frequencies. Liu says Cosmos uses a temporal-position-embedding scheme to normalize signals onto a common time axis so tokens can identify co-temporal and relatively distant events (~06:15–07:28). This claim in the interview and related secondary excerpts was found in the reviewed sources, but not enough independently extracted technical-report text to validate exact implementation details beyond NVIDIA’s architecture description was found in the reviewed sources.
Policy verification, not certified simulation
Liu’s most cautious and useful claim is that neural simulators may first be useful for policy verification, especially ranking many checkpoints before expensive real-world testing. He explicitly notes the simulator need not produce exact success rates if it preserves ordering between policy A and policy B (~12:53–15:16). This is plausible as a development accelerator, but it requires evidence of rank correlation against real-world tests for each target domain.
Claims and evidence
- Cosmos 3 unifies language, image, video, audio and action for reasoning/generation/action.
- Cosmos 3 uses MoT with AR reasoner and diffusion generator.
- Open checkpoints/code/data/eval assets are available under OpenMDW-1.1. — Vendor/open-source documentation; directly checkable through repo/HF.
Read the full section
| Material claim | Evidence status |
| Cosmos 3 unifies language, image, video, audio and action for reasoning/generation/action. | Vendor-documented by NVIDIA site, GitHub and model cards; interview consistent. Cosmos 3 — Cosmos Lab |
| Cosmos 3 uses MoT with AR reasoner and diffusion generator. | Vendor-documented in model card and NVIDIA technical blog; interview consistent. README.md · nvidia/Cosmos3-Edge at main |
| Open checkpoints/code/data/eval assets are available under OpenMDW-1.1. | Vendor/open-source documentation; directly checkable through repo/HF. Cosmos 3: Omnimodal World Models for Physical AI |
| Cosmos 3 leads multiple physical-AI benchmarks. | Vendor-reported; NVIDIA cites public leaderboards, but no broad independent reproduction was found in the reviewed sources. NVIDIA Launches Cosmos 3, the Open Frontier Foundation Model for Physical AI | NVIDIA Newsroom |
| Policy verification can reduce real-world testing burden if simulator rankings correlate with real world. | Interview argument; conceptually sound but deployment-specific; requires empirical validation per domain. Transcript ~12:53 |
| Edge can support on-device robot workloads. | Vendor-documented model size and deployment tier; real product suitability remains workload-dependent. Model Matrix — Cosmos |
Context and prior work
Cosmos 3 sits at the intersection of three research lines: The DROID dataset is especially relevant because the interview says Cosmos was post-trained on DROID for pick-and-place policy work. DROID is an in-the-wild robot manipulation dataset with teleoperated demonstrations across many scenes and tasks.
Read the full section
Cosmos 3 sits at the intersection of three research lines:
- Vision-language-action models. RT-2 framed robot control as a VLA problem: map observations and language into actions while transferring knowledge from web-scale VLM pretraining. RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
- Open generalist robot policies. Octo showed an open transformer-based generalist policy trained on diverse robot trajectories and designed for fine-tuning across sensors and action spaces. Octo: An Open-Source Generalist Robot Policy
- World models / interactive simulation. Google DeepMind’s Genie 3 is another major world-model reference point, focused on real-time interactive world generation from text descriptions rather than open robot-action policy models. Genie 3 — Google DeepMind
The DROID dataset is especially relevant because the interview says Cosmos was post-trained on DROID for pick-and-place policy work. DROID is an in-the-wild robot manipulation dataset with teleoperated demonstrations across many scenes and tasks. nvidia/Cosmos3-DROID · Datasets at Hugging Face
Limitations, safety and contested findings
NVIDIA’s own model cards caution that Cosmos3 outputs should not be treated as physically accurate simulation, reliable ground-truth reasoning or safety-certified decision-making. They state that robotics control, autonomous systems, scientific simulation and safety-critical planning require additional validation, external constraints, system-level safety analysis and domain-specific guardrails.
Read the full section
NVIDIA’s own model cards caution that Cosmos3 outputs should not be treated as physically accurate simulation, reliable ground-truth reasoning or safety-certified decision-making. They state that robotics control, autonomous systems, scientific simulation and safety-critical planning require additional validation, external constraints, system-level safety analysis and domain-specific guardrails. README.md · nvidia/Cosmos3-Super-Text2Image at adf30429aaaf93b7453aa45944cc0f4021f2a92d
The interview also acknowledges several hard limits:
- Manipulation is harder than navigation because contact introduces occlusion, deformation and complex interaction dynamics (~19:46–20:43).
- Simulator exploitation is possible. Liu acknowledges that policies may hack neural simulators or exploit artifacts, analogizing to adversarial examples in deep learning (~15:52–16:36).
- Ambiguous tasks need system-level planning. Liu frames “System 2” planning, harnesses, tools, memory and post-execution checks as necessary layers above the model (~11:06–12:52).
Independent practitioner analysis makes the same governance point: a world model may reason, generate scenarios, infer dynamics or propose actions, but product teams must specify authority, operating envelope, fallback, human role and evidence gates before giving it actuation authority. World Models and Robot Policies: Make the Product Decision Before the Model Decision
Business and practitioner implications
For business leaders, Cosmos 3 is best understood as a platform bet: NVIDIA is trying to make physical-AI development run through open models, CUDA/NIM/vLLM/TensorRT-LLM deployment paths, Jetson/RTX/DGX hardware and synthetic-data workflows. Axios independently framed Cosmos 3 as part of NVIDIA’s move beyond chips into models and software for robotics and autonomous systems.
Read the full section
For business leaders, Cosmos 3 is best understood as a platform bet: NVIDIA is trying to make physical-AI development run through open models, CUDA/NIM/vLLM/TensorRT-LLM deployment paths, Jetson/RTX/DGX hardware and synthetic-data workflows. Axios independently framed Cosmos 3 as part of NVIDIA’s move beyond chips into models and software for robotics and autonomous systems. Nvidia's Cosmos 3 open AI world model helps robots, autonomous vehicles
Near-term useful applications are likely:
- synthetic data generation for perception and scenario coverage;
- offline policy checkpoint triage;
- post-training starting points for robot embodiments;
- debugging and visualization of policy rollouts;
- edge experimentation with smaller models where latency and connectivity matter.
Riskier near-term applications are:
- direct safety-critical control without independent runtime constraints;
- assuming generated videos are calibrated physics simulations;
- training policies only on generated worlds without field validation;
- using vendor benchmark rank as a proxy for product readiness.
Sources
- Primary commentary/interview transcript: Machine Learning Street Talk / Spotify episode; companion YouTube URL: https://www.youtube.com/watch?v=L6tLBApQN-g
- NVIDIA Cosmos 3 project page. Cosmos 3 — Cosmos Lab
- NVIDIA/cosmos GitHub repo and release notes. GitHub - NVIDIA/cosmos: NVIDIA Cosmos is an open platform of world models, datasets, and tools that enables developers to build Physical AI for robots, autonomous vehicles, smart infrastructure, and more. · GitHub
Read the full section
- Primary commentary/interview transcript: Machine Learning Street Talk / Spotify episode; companion YouTube URL: https://www.youtube.com/watch?v=L6tLBApQN-g
- NVIDIA Cosmos 3 project page. Cosmos 3 — Cosmos Lab
- NVIDIA/cosmos GitHub repo and release notes. GitHub - NVIDIA/cosmos: NVIDIA Cosmos is an open platform of world models, datasets, and tools that enables developers to build Physical AI for robots, autonomous vehicles, smart infrastructure, and more. · GitHub
- Cosmos 3 arXiv record / technical report metadata. Cosmos 3: Omnimodal World Models for Physical AI
- Hugging Face Cosmos3 model cards, including Edge and safety limitations. README.md · nvidia/Cosmos3-Edge at main
- NVIDIA Cosmos 3 developer blog and model matrix. Develop Physical AI Reasoning, World, and Action Models with NVIDIA Cosmos 3 | NVIDIA Technical Blog
- Axios independent reporting on Cosmos 3 launch. Nvidia's Cosmos 3 open AI world model helps robots, autonomous vehicles
- RoboArena paper for independent robot-policy evaluation methodology. RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies
- RT-2, Octo, DROID and Genie 3 context sources. RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
The source trail.
Sources (19)
How Physical AI Learns Across Language, Video and Action — Ming-Yu Liu
Transcript retrieved via companion_video_youtube_manual_captions; language en-GB. Timestamped text, not direct audiovisual review. Source: https://www.youtube.com/watch?v=L6tLBApQN-g. Automatic captions/transcription may contain errors.
podcasters.spotify.comTranscript: How Physical AI Learns Across Language, Video and Action — Ming-Yu Liu
companion_video_youtube_manual_captions
www.youtube.comNew interview with Ming-Yu Liu, who leads Cosmos research at @NVIDIAAI. Cosmos 3 takes video in, generates it out, and outputs robot actions, so developers can test a pol
Related coverage; assess separately
x.com