AI Briefing
KO

NVIDIA Enters Full Production with Groq 3 LPX AI Inference Accelerator... Achieves Record Token Generation Speeds Combined with Vera Rubin

·2026.08.25 09:00

Key point

NVIDIA has begun mass production of the Groq 3 LPX, significantly boosting inference speeds on the Vera Rubin platform.

1 / 2

Details

At Hot Chips 2026, NVIDIA officially announced the start of full production (mass production) for the Groq 3 LPX AI inference accelerator chip. Following the announcement of mass production for the Vera CPU and Vera Rubin servers, this demonstrates that NVIDIA's AI roadmap is progressing smoothly.

The Groq 3 LPX serves as an expansion component for the Vera Rubin NVL72 platform, supporting the ultra-fast token generation required for agentic AI workloads. The architecture involves NVIDIA Rubin GPUs handling large-scale context processing, while the LPX accelerates latency-sensitive decoding workloads.

In actual benchmarks, the system achieved a record of generating 3,400 tokens per second (3,400 TPS) while running the Gemma 4 31B open model. This is the fastest performance recorded for this model with a 100,000-token context window.

This combination reduces response times for agentic AI tasks such as coding by 4x compared to previous standards, enabling tasks that previously took hours to be completed in minutes. As a result, AI factories can deliver faster, more predictable inference and smoother agent interactions.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.