Nvidia has moved its Groq 3 LPX inference accelerator into full production, targeting fast token generation for agentic AI systems. The chip, built on the Vera Rubin platform, is designed to deliver ultrafast token generation, with performance claims that could significantly reduce coding tasks to minutes instead of hours. According to a benchmark from Artificial Analysis, the Groq 3 LPX achieved 3,400 tokens per second on the Gemma 4 31B model with a 100,000-token context window, the highest figure ever recorded for this model. Nvidia asserts that this makes the accelerator four times faster than the next best option, the Cerebras chip, which reached 882 tokens per second.

The benchmark results, however, have drawn scrutiny from experts who argue the comparison is stacked in Nvidia's favor. Groq's SRAM-heavy dataflow architecture requires at least 64 chips to reach the 3,400 tokens per second mark, while Cerebras can achieve similar performance with just one or two accelerators. Additionally, the comparison does not account for Cerebras' latest CS-4 generation, which may offer improved performance. The Register notes that Groq's architecture has a limitation in memory, with each LPU having only 500 MB of memory, compared to the 288 GB available on a Rubin GPU. This necessitates splitting models across multiple accelerators, with a single rack holding up to 256 LPUs.

The context of the benchmark also highlights that Gemma 4 31B is a best-case scenario, as it fits entirely within a single rack. Scaling this setup to larger models like DeepSeek V3, which would require 1,342 accelerators or over five racks, remains an open question. The Cerebras comparison also overlooks the number of chips required, with Cerebras needing just one or two accelerators for the model, while Nvidia needs at least 64.

Source: thedecoder