Nvidia has begun full production of the Groq 3 LPX, an inference chip that extends its Vera Rubin platform, and is telling customers it delivers the fastest token generation ever recorded on an open model. The claim rests entirely on Nvidia’s own benchmark demonstration. A closer look at how that number was produced complicates the pitch considerably.
The Groq 3 LPX racks attach to Vera Rubin NVL72 systems, splitting inference work so that the Rubin side crunches through prefill and the LPX side pushes out decode, the step that governs how fast an agent produces its next token. Nvidia unveiled the milestone at Hot Chips 2026, framing it as the payoff of the roughly $20 billion Groq license it signed in late December, a deal that also brought Groq founder Jonathan Ross and president Sunny Madra into Nvidia’s ranks.
In a demonstration run on the open model Gemma 4 31B with a 100,000 token context window, Nvidia’s rack hit 3,400 tokens per second. That figure came from Artificial Analysis, which measured it across 50 back to back requests, and it holds the record for that model. Nvidia then compared it to Cerebras’ chip, which the same benchmark clocked at 882 tokens per second, and used the gap to advertise a fourfold speed advantage for agentic tasks such as coding.
The Decoder, reporting on 25 August 2026 and drawing on analysis from The Register, points to what that comparison leaves out. Each Groq LPU holds only 500 MB of onboard memory, a fraction of the 288 GB packed into a single Rubin GPU. That forces large models to be spread across multiple accelerators connected by Ethernet, and a single rack can hold up to 256 of the LPUs.
Gemma 4 31B is close to a best case for this architecture: it is dense enough to fit inside a single rack. The Register notes that a larger mixture of experts model would not fare the same way. Serving DeepSeek V3 on this architecture would take 1,342 of the accelerators, spread across a little more than five full racks.
The Cerebras comparison also omits chip counts entirely. Cerebras can run the same model on just one or two chips, while Nvidia’s side of the comparison uses a minimum of 64. Nvidia’s benchmark also excludes Cerebras’ newest CS-4 generation, so the fourfold claim compares a brand new Nvidia product against an older Cerebras chip, not the one Cerebras currently sells.
A speed benchmark that leaves out how much hardware produced the result is not a like for like comparison. Anyone budgeting for inference capacity needs tokens per second per chip, not tokens per second per rack, and Nvidia has not published that number.
Among the cloud vendors lining up, Nebius has moved fastest, folding the accelerator into its Token Factory service so customers can test the claim on real workloads rather than Nvidia’s demo. Teams evaluating inference infrastructure this quarter should ask both vendors for chip count normalized throughput before treating the fourfold figure as a purchasing input.
Wccftech reported Nvidia’s Groq 3 LPX production announcement on 24 August 2026, and The Decoder reported the independent analysis on 25 August 2026.