Cosmos3
Netw0rkNoob/VulnClawCosmos3Jun 1, 2026by mingyuliutw
AI Summary
Cosmos 3 is NVIDIA's next-generation family of open omnimodal world foundation models designed for Physical AI. By unifying language, images, video, audio, and actions, it enables developers to build agents that can perceive, reason, simulate, and act within the physical world.
Key Highlights
- Unified omnimodal architecture supporting text, image, video, audio, and action modalities.
- Integrated Reasoner + Generator architecture for world understanding and generation.
- Flexible generation capabilities for Text-to-Image, Image-to-Video, and Video-to-World tasks.
- Native support for robot action generation and policy learning via Cosmos3-Policy models.
- Full open release of models, code, datasets, and evaluation tooling.
New Features
- Cosmos3-Nano (16B) model optimized for efficient deployment.
- Cosmos3-Super (64B) model for advanced reasoning and simulation.
- Cosmos3-Super-Text2Image and Cosmos3-Super-Image2Video specialized generation models.
- Cosmos3-Nano-Policy-DROID for robot manipulation and control policies.
- Access to the Cosmos Cookbook and inference tooling.
Full Release Notes
# Cosmos 3 is Here 🚀 Today, we're excited to release Cosmos 3 — NVIDIA's next-generation family of open omnimodal world foundation models for Physical AI. Cosmos 3 unifies language, images, video, audio, and actions within a single architecture, enabling developers to build agents that can understand, reason about, simulate, and act in the physical world. From world generation and simulation to robotics and embodied AI, Cosmos 3 serves as a general-purpose foundation model for Physical AI. What's new: - 🌍 Unified omnimodal world model supporting text, image, video, audio, and action modalities - 🧠 Integrated Reasoner + Generator architecture for world understanding and world generation - 🎬 Flexible generation across Text-to-Image, Image-to-Video, Video-to-World, and multimodal simulation tasks - 🤖 Native support for robot action generation and policy learning through Cosmos3-Policy models - 📈 State-of-the-art open model performance across world understanding, generation, and robotics benchmarks - 🔓 Open release of models, code, datasets, evaluation benchmarks, and inference tooling for the Physical AI community :book: [Read the Paper](https://research.nvidia.com/labs/cosmos-lab/cosmos3/technical-report.pdf) | :point_right: [Download the Models](https://huggingface.co/collections/nvidia/cosmos3) | 🧑🍳 [Explore the Cosmos Cookbook](https://github.com/NVIDIA/cosmos/tree/main/cookbooks/cosmos3) The Cosmos 3 release includes: - Cosmos3-Nano (16B) – Compact omnimodal world foundation model optimized for efficient deployment and development. - Cosmos3-Super (64B) – High-capacity world model for advanced reasoning, generation, simulation, and Physical AI applications. - Cosmos3-Super-Text2Image – State-of-the-art text-to-image generation model built on Cosmos 3. - Cosmos3-Super-Image2Video – High-fidelity image-to-video generation model with strong temporal consistency and controllability. - Cosmos3-Nano-Policy-DROID – Open robot foundation model for learning manipulation and control policies directly from demonstrations. Cosmos 3 represents a major step toward general-purpose world models that can perceive, reason, simulate, and act—bringing us closer to a future where Physical AI can learn from both the real world and generated worlds at scale.