AI

NVIDIA Puts Groq 3 LPX Into Full Production, Nebius Named First AI Cloud Adopter

Monday, August 24, 2026Read Original

Details

  • NVIDIA announced that Groq 3 LPX, its interactive AI inference accelerator, is now in full production as an extension of the Vera Rubin platform.
  • The launch centers on NVIDIA Vera Rubin NVL72 systems, Nebius’s Nebius Token Factory inference cloud, and Groq as an early-purpose-built inference cloud adopter.
  • Groq 3 LPX boosts AI inference by delivering ultrafast token generation, reaching about 3,400 output tokens per second on Gemma 4 31B with a 100,000-token context for agentic workloads.
  • Compared with prior Vera Rubin-only configurations, LPX offloads low-latency token generation while NVL72 handles large-context processing, enabling up to 4x faster responsiveness versus NVIDIA’s cited nearest alternative platform.
  • Nebius plans to integrate Groq 3 LPX into Nebius Token Factory via the existing API, giving developers rack-scale access to generation-optimized inference within the same cloud platform used for open-weight models.

Impact

The move strengthens NVIDIA’s position in low-latency inference for agentic AI, where long-context reasoning and rapid token generation are increasingly critical for coding agents and tool-using systems. Making LPX available via AI clouds like Nebius lowers deployment friction for enterprises and is likely to steer R&D and spending toward architectures that split context processing and token generation over the next 12–24 months.

Rift Dispatch