AI

NVIDIA Blackwell Ultra sets DeepSeek-V3 671B pre-training performance world record

Tuesday, July 21, 2026Read Original

Details

  • NVIDIA AI reports that the NVIDIA Blackwell Ultra platform has set a world record for pre-training the DeepSeek-V3 671B Mixture-of-Experts language model.
  • The system achieved 1,648 TFLOPs of training throughput per GPU on DeepSeek-V3 671B, using the GB300 NVL72 configuration and the latest NVIDIA NeMo 26.06 software stack.
  • NVIDIA states this represents roughly a 1.3x per-GPU throughput uplift over earlier Blackwell Ultra results on the same workload, and around 3x the delivered performance of the previous-generation GB200 NVL72 system at equivalent scale.
  • The gains are attributed to extreme full-stack co-design, including Blackwell Ultra GPU architecture, NVLink/NVLink Switch interconnects, Spectrum-X networking, and optimizations in software layers such as Megatron-Core, Transformer Engine, NCCL, and CuTe DSL.
  • This DeepSeek-V3 run aligns with the latest MLPerf Training v6.0 results, where Blackwell Ultra-based GB300 NVL72 systems posted the fastest time-to-train across all submitted workloads and demonstrated significant speedups over prior-generation platforms.
  • DeepSeek-V3 itself is a 671B-parameter open Mixture-of-Experts model with 37B parameters active per token, designed for high efficiency and competitive performance among frontier open-source LLMs.
  • NVIDIA emphasizes that the record training performance was achieved without changing the underlying silicon, underscoring how ongoing software and network-stack tuning can compound hardware advances over short timeframes.
  • The announcement highlights Blackwell Ultra as a leading platform for large-scale frontier model training, particularly for massive MoE workloads that stress memory capacity, interconnect bandwidth, and end-to-end training efficiency.

Impact

By demonstrating world-record DeepSeek-V3 671B pre-training throughput on Blackwell Ultra, NVIDIA reinforces its lead in frontier model training performance and scale, putting competitive pressure on other hyperscalers and chip vendors focused on ultra-large LLMs. The roughly multi-x uplift over prior-generation systems shows how coordinated hardware, interconnect, and software co-design can materially cut time-to-train and GPU-hour costs for trillion-token-scale runs. This is likely to accelerate adoption of Blackwell Ultra for commercial and open-source MoE models, and strengthens NVIDIA’s position in MLPerf benchmarks that increasingly shape procurement decisions for AI infrastructure.

Rift Dispatch