aiexpert
Home / News / Brief
Research · Aug 06, 2026, 03:31 PM · 4 sources

NVIDIA Cosmos 3 ranks #1 on open-weights benchmarks for world models, robotics policy

NVIDIA released Cosmos 3, an open-source world foundation model for physical AI built on a mixture-of-transformers architecture that combines vision reasoning, world generation, and action prediction in a single system. The model is available in three sizes under the Linux Foundation's OpenMDW-1.1 license: Cosmos 3 Super (64B parameters) for high-fidelity world modeling, Cosmos 3 Nano (16B) for efficient reasoning and deployment, and Cosmos 3 Edge (4B, announced but not yet released) for on-device robotics and edge inference. The model was trained on 20 trillion tokens of multimodal data, including nearly 1 billion images, 400 million videos, ambient audio, and robot action trajectories.

Cosmos 3 ranks #1 on Artificial Analysis for open-weights text-to-image and image-to-video generation, leads PAI-Bench for world generation and Physics-IQ for image-to-video, and ranks #1 on RoboLab for robot policy. Cosmos 3 Super is also the highest-ranked open model on VANTAGE-Bench for vision understanding across fixed-camera warehouse, transportation, and smart-space footage. Unlike prior video-only foundation models, Cosmos 3 natively outputs robot action signals—joint angles, gripper positions, trajectory points—enabling direct policy training and synthetic data generation for embodied AI systems.

NVIDIA launched the NVIDIA Cosmos Coalition, a global collaboration with robotics and AI leaders including Agile Robots, Black Forest Labs, Generalist, LTX, Runway, and Skild AI to contribute models, research, and evaluation methods. Developers can post-train Cosmos 3 on proprietary data and hardware, and deployment partners including Baseten, CoreWeave, Microsoft Azure, and Deep Infra offer containerized inference. For architects, Cosmos 3 offers a single unified model replacing separate pipelines for world reasoning, simulation, and policy training—reducing development cycles from months to days.

Sources

Everything this brief rests on
  1. 01 Primary source blogs.nvidia.com
  2. 02 nvidianews.nvidia.com nvidianews.nvidia.com “NVIDIA Cosmos 3 is built on a breakthrough mixture-of-transformers architecture that combines vision reasoning, world generation and action prediction in a single system”
  3. 03 axios.com axios.com “NVIDIA trained Cosmos 3 on 20 trillion tokens of multimodal data, including nearly a billion images, 400 million real and synthetic videos”
  4. 04 marktechpost.com marktechpost.com “Generated video then becomes more than a prediction. It represents how the world should change in response to an action”