Productspublished

Nvidia Puts Groq 3 LPX Into Production for Faster AI-Agent Responses

The first Nebius deployment will test Nvidia’s argument that specialized token generation can make long, tool-using AI sessions more responsive while leaving GPUs at the center of the system.

By 3 min read
Nvidia Puts Groq 3 LPX Into Production for Faster AI-Agent Responses

Listen to this story

The audio brief

About 1:38
0:001:38
Read transcript
Nvidia has put its Groq 3 LPX inference system into full production, with Nebius set to become its first announced customer later this year. The bet is straightforward: for AI agents that spend a long time generating responses, specialized hardware could make each next token arrive faster, while GPUs remain the flexible workhorses for training and much of inference. An LPX rack combines 256 Groq 3 chips with 500 megabytes of on-chip SRAM, memory designed to reduce bottlenecks during token generation, also called the decode phase. Nvidia says that makes LPX a premium, low-latency layer rather than a replacement for GPUs. At Nebius, the racks will be deployed inside Token Factory alongside Vera central processors and Rubin graphics processors. The headline result comes from Artificial Analysis, which measured 3,400 tokens per second on Gemma 4 31B with a 100,000-token context window. That kind of performance matters for agents handling long context, calling tools, or writing code, where delays between generated tokens can make the whole interaction feel sluggish. But the commercial question is still open. Nvidia is asking cloud providers to pay for dedicated decode capacity alongside conventional GPUs, and the market is crowded: AMD is working with Cerebras, while OpenAI’s Ultrafast mode, powered by Cerebras, promises 750 tokens per second. The key test is whether Nebius’s rollout proves that this split delivers responsiveness customers will actually pay for—completing the commercial test of Nvidia’s $20 billion Groq acquisition.

Story brief

3 key points

Nvidia’s Groq-derived Groq 3 LPX has reached full production as a rack-scale inference system aimed at speeding token-by-token decoding for latency-sensitive AI agents. Nebius will be its first announced customer, integrating 256-chip racks with Vera CPUs and Rubin GPUs in its Token Factory platform later this year. Artificial Analysis measured 3,400 tokens per second on Gemma 4 31B with a 100,000-token context. The...

  1. 01

    Each LPX rack combines 256 Groq 3 chips and 500MB of on-chip SRAM to reduce memory bottlenecks during decoding.

  2. 02

    Nvidia frames LPX as a premium low-latency layer, not a replacement for flexible GPUs used across training and inference.

  3. 03

    Nebius plans to deploy LPX with Vera and Rubin hardware in Token Factory, making rollout the first commercial proof point.

Nvidia has put its Groq 3 LPX inference chip into full production, turning technology acquired from Groq into a rack-scale product for AI systems where delayed output can make an agent feel slow. Nebius is slated to be the first announced deployment, with racks expected online later this year.

A dedicated engine for the response phase

LPX is designed to accelerate token generation, also called the decode phase: the part of model serving that produces an answer token by token. Nvidia positions the chips for low-latency inference rather than as replacements for GPUs, which remain flexible enough for training and inference.

That division addresses a practical constraint in agent workloads. Systems that reason through long context, call tools, or write code can spend substantial time generating their next tokens. Nvidia’s pitch is that a dedicated processor can handle that response-sensitive stage while Vera CPUs and Rubin GPUs handle other parts of the stack.

Nebius becomes the commercial proving ground

Nvidia plans to deploy the LPX racks at Nebius alongside Vera central processors and Rubin graphics processors. Nebius has committed to use the hardware in its Token Factory production inference platform, making the cloud provider the first announced customer.

The hardware is built at rack scale. Nvidia packages 256 Groq 3 chips into an LPX rack, and the Groq architecture includes 500 megabytes of SRAM, a fast memory placed on the chip itself to reduce memory-related bottlenecks.

What Nvidia is selling to cloud providers

  • A specialized decode layer for latency-sensitive model serving, rather than a general-purpose GPU substitute.
  • A premium service opportunity for customers that require especially latency-sensitive service agreements, according to Nvidia.
  • A combined configuration in which Groq racks operate beside Vera CPUs and Rubin GPUs at Nebius.

A crowded push toward low-latency inference

Nvidia is entering an active market for faster model output. AMD has said it will integrate rack-scale systems with Cerebras chips for low-latency inference, while OpenAI’s Ultrafast mode promises 750 tokens per second and is powered by Cerebras.

The LPX product also closes the loop on Nvidia’s $20 billion December purchase of Groq assets, described as the company’s largest acquisition. Nvidia’s commercial claim now depends on the Nebius rollout and on whether the split between GPUs and dedicated decode hardware delivers the responsiveness customers will pay for.

Editorial analysis

Our Read

Nvidia’s LPX launch is a systems-design bet, not simply a faster-chip claim. The company is carving out token generation as a distinct service tier, while retaining Vera Rubin GPUs for the broader workload. That division is strategically useful if agentic applications make responsiveness valuable enough for cloud providers to sell premium service. The key event to watch is Nebius bringing the racks online later this year: it will be the first disclosed production deployment of the combined Vera, Rubin, and Groq configuration. The 3,400-token-per-second result is notable, but it comes from one named model and context-window setup.

Sources

  1. cnbc.comNvidia says Groq racks will be online this year following $20 billion purchase
  2. siliconangle.comNvidia's dedicated inference accelerator Groq 3 LPX enters full production to supercharge AI agents - SiliconANGLE