
NVIDIA begins full production of Groq 3 LPX accelerator
- NVIDIA (NASDAQ:NVDA) announced that its Groq 3 LPX AI inference accelerator has officially entered full production.
- The company stated the product delivers 3,400 output tokens per second, driving up to 4x faster responsiveness for latency-sensitive workloads.
- The company stated the new hardware targets agentic AI workloads and integrates directly into its Vera Rubin AI factory architecture.
NVIDIA (NASDAQ:NVDA) placed its Groq 3 LPX AI inference accelerator into full production to process agentic workloads at 3,400 tokens per second.
This launch extends the Vera Rubin NVL72 platform to handle complex, multi-step inference sequences more efficiently than previous systems.
The company stated this performance represents the fastest recorded benchmarking result for the Gemma 4 31B model within a 100,000-token context.
NVIDIA reported that the accelerator provides up to 4x faster responsiveness for latency-sensitive workloads compared to the nearest alternative platform.
Following the announcement, NVIDIA's share price was down at $209.11.
The company stated that Nebius will become the first AI cloud to adopt the hardware through its production inference platform.
The accelerator integrates into the broader Vera Rubin AI factory architecture alongside BlueField-4 DPUs, Vera CPU racks, and Spectrum-6 Ethernet.



_1447169_640x360_2508609603994-640x360.webp&w=3840&q=75)
_1447326_640x360_2508632131665-640x360.webp&w=3840&q=75)