Nvidia Puts Groq 3 LPX Into Production for Faster AI-Agent Responses
The first Nebius deployment will test Nvidia’s argument that specialized token generation can make long, tool-using AI sessions more responsive while leaving GPUs at the center of the system.
Listen to this story
The audio brief
Story brief
3 key pointsNvidia’s Groq-derived Groq 3 LPX has reached full production as a rack-scale inference system aimed at speeding token-by-token decoding for latency-sensitive AI agents. Nebius will be its first announced customer, integrating 256-chip racks with Vera CPUs and Rubin GPUs in its Token Factory platform later this year. Artificial Analysis measured 3,400 tokens per second on Gemma 4 31B with a 100,000-token context. The...
- 01
Each LPX rack combines 256 Groq 3 chips and 500MB of on-chip SRAM to reduce memory bottlenecks during decoding.
- 02
Nvidia frames LPX as a premium low-latency layer, not a replacement for flexible GPUs used across training and inference.
- 03
Nebius plans to deploy LPX with Vera and Rubin hardware in Token Factory, making rollout the first commercial proof point.
Nvidia has put its Groq 3 LPX inference chip into full production, turning technology acquired from Groq into a rack-scale product for AI systems where delayed output can make an agent feel slow. Nebius is slated to be the first announced deployment, with racks expected online later this year.
A dedicated engine for the response phase
LPX is designed to accelerate token generation, also called the decode phase: the part of model serving that produces an answer token by token. Nvidia positions the chips for low-latency inference rather than as replacements for GPUs, which remain flexible enough for training and inference.
That division addresses a practical constraint in agent workloads. Systems that reason through long context, call tools, or write code can spend substantial time generating their next tokens. Nvidia’s pitch is that a dedicated processor can handle that response-sensitive stage while Vera CPUs and Rubin GPUs handle other parts of the stack.
Nebius becomes the commercial proving ground
Nvidia plans to deploy the LPX racks at Nebius alongside Vera central processors and Rubin graphics processors. Nebius has committed to use the hardware in its Token Factory production inference platform, making the cloud provider the first announced customer.
The hardware is built at rack scale. Nvidia packages 256 Groq 3 chips into an LPX rack, and the Groq architecture includes 500 megabytes of SRAM, a fast memory placed on the chip itself to reduce memory-related bottlenecks.
What Nvidia is selling to cloud providers
- A specialized decode layer for latency-sensitive model serving, rather than a general-purpose GPU substitute.
- A premium service opportunity for customers that require especially latency-sensitive service agreements, according to Nvidia.
- A combined configuration in which Groq racks operate beside Vera CPUs and Rubin GPUs at Nebius.
A crowded push toward low-latency inference
Nvidia is entering an active market for faster model output. AMD has said it will integrate rack-scale systems with Cerebras chips for low-latency inference, while OpenAI’s Ultrafast mode promises 750 tokens per second and is powered by Cerebras.
The LPX product also closes the loop on Nvidia’s $20 billion December purchase of Groq assets, described as the company’s largest acquisition. Nvidia’s commercial claim now depends on the Nebius rollout and on whether the split between GPUs and dedicated decode hardware delivers the responsiveness customers will pay for.
Editorial analysis
Our Read
Nvidia’s LPX launch is a systems-design bet, not simply a faster-chip claim. The company is carving out token generation as a distinct service tier, while retaining Vera Rubin GPUs for the broader workload. That division is strategically useful if agentic applications make responsiveness valuable enough for cloud providers to sell premium service. The key event to watch is Nebius bringing the racks online later this year: it will be the first disclosed production deployment of the combined Vera, Rubin, and Groq configuration. The 3,400-token-per-second result is notable, but it comes from one named model and context-window setup.
Sources
- cnbc.comNvidia says Groq racks will be online this year following $20 billion purchase
- siliconangle.comNvidia's dedicated inference accelerator Groq 3 LPX enters full production to supercharge AI agents - SiliconANGLE