Details
- NVIDIA announced that Groq 3 LPX, its interactive AI inference accelerator, is now in full production as an extension of the Vera Rubin platform.
- The launch centers on NVIDIA Vera Rubin NVL72 systems, Nebius’s Nebius Token Factory inference cloud, and Groq as an early-purpose-built inference cloud adopter.
- Groq 3 LPX boosts AI inference by delivering ultrafast token generation, reaching about 3,400 output tokens per second on Gemma 4 31B with a 100,000-token context for agentic workloads.
- Compared with prior Vera Rubin-only configurations, LPX offloads low-latency token generation while NVL72 handles large-context processing, enabling up to 4x faster responsiveness versus NVIDIA’s cited nearest alternative platform.
- Nebius plans to integrate Groq 3 LPX into Nebius Token Factory via the existing API, giving developers rack-scale access to generation-optimized inference within the same cloud platform used for open-weight models.
Impact
The move strengthens NVIDIA’s position in low-latency inference for agentic AI, where long-context reasoning and rapid token generation are increasingly critical for coding agents and tool-using systems. Making LPX available via AI clouds like Nebius lowers deployment friction for enterprises and is likely to steer R&D and spending toward architectures that split context processing and token generation over the next 12–24 months.