Details
- NVIDIA AI announces new 4-step Cosmos 3 Super models for both image and video generation.
- The updated Super models generate images and video up to 25x faster than the original Cosmos 3 Super variants, targeting lower latency for developers.
- Cosmos 3 Super is a 64B-parameter multimodal omnimodel designed for advanced world simulation, physical reasoning, and synthetic data generation.
- Independent benchmarking platform Artificial Analysis ranks Cosmos 3 Super among the top open-weight models, including leading positions on text-to-image and image-to-video leaderboards.
- The tweet highlights current standings: #1 for image-to-video (no audio) and #2 for text-to-image among open-weight models on Artificial Analysis.
- The new 4-step pipelines are accessible via Hugging Face, where Cosmos 3 Super checkpoints are hosted for text, image, and video generation workflows.
- Earlier Cosmos 3 releases positioned the Super family as state-of-the-art for physical AI, robotics training, and physics-accurate video generation, and the 4-step upgrade focuses on inference speed rather than changing the open-weight nature or core capabilities.
- Faster generation is particularly relevant for interactive applications, real-time content tools, and iterative robotics simulation where turnaround time is crucial.
- The announcement reinforces NVIDIA’s strategy of pairing high-capacity open models with its GPU ecosystem, encouraging developers to run or fine-tune Cosmos 3 Super on Hopper and Blackwell datacenter GPUs.
Impact
By pushing Cosmos 3 Super to generate images and video dramatically faster while retaining top open-weight benchmark rankings, NVIDIA strengthens its position in open multimodal foundation models and physical AI tooling. The speed gains make Cosmos 3 more practical for real-time or near-real-time workflows, narrowing the experiential gap with proprietary frontier video and image generators and reinforcing Hugging Face as a key distribution channel for high-end open models.