AI Briefing
KO

Groq 3 LPX Enters Mass Production; NVIDIA Expands Vera Rubin Agent Inference

·2026.08.25 00:00

Key point

NVIDIA has begun mass production of Groq 3 LPX, significantly enhancing agent inference performance based on Vera Rubin.

Details

NVIDIA has officially announced the mass production of Groq 3 LPX, supporting rapid token generation for agent systems via the Vera Rubin NVL72 platform. In Artificial Analysis benchmarks, when running the open-source agent model Gemma 4 31B, it generated 3,400 output tokens per second in a long-context environment of 100,000 tokens, demonstrating performance four times faster than competing platforms.

NVIDIA is maximizing AI factory efficiency by designing compute, networking, and inference acceleration as an integrated system through 'Extreme Codesign'. Spectrum-X Ethernet efficiently handles large-scale data flows, while Groq 3 LPX provides token generation speeds optimized for agent workloads with its ultra-low latency inference architecture.

Key partners have already adopted the Vera Rubin platform.

  • Nebius: The first AI cloud service provider to adopt Groq 3 LPX, offering developers an ultra-fast token generation environment.
  • CoreWeave: Deployed Spectrum-X Multiplane in production to connect Vera Rubin racks, building a high-bandwidth AI network.
  • SpaceXAI: Plans to apply NVIDIA Vera CPU to next-generation agent AI infrastructure to accelerate CPU-intensive tasks such as orchestration and simulation.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.