Google and Nvidia Say AI Agents Are Raising Demand for CPUs
GPUs still generate model output, but agents also need processors to run tools, execute code and move results through a growing working context.
Loading page…
GPUs still generate model output, but agents also need processors to run tools, execute code and move results through a growing working context.
Listen to this story
At the AI Infra Summit, Google’s Dave Patterson and Nvidia’s Ian Buck argued that autonomous agents are changing infrastructure economics: GPUs generate model outputs, but CPUs, memory systems, software execution, and interconnects determine how quickly agents can act on them. Nvidia is positioning Vera for code execution and tool calls, while NVLink Fusion connects third-party processors to Nvidia systems.
Agent coding loops require CPUs to execute generated code, run tests, and return results before the model can decide its next action.
Buck cited context growth from roughly 1,000 tokens in 2023 benchmarks to 142,000 in 2026; the figures were his assessment, not an industrywide measurement.
Nvidia positions Vera for agent-side code execution and tool running, complementing rather than replacing GPUs.
AI agents are pushing CPU performance back into the center of AI infrastructure. At the AI Infra Summit, Google Fellow Dave Patterson said autonomous agents are increasing demand for CPUs alongside GPUs, while Nvidia’s Ian Buck presented Vera as a CPU for running the software and tools those agents call.
The distinction is visible in a coding task. A GPU can generate code, but a CPU can execute it, run tests and return the result for the model’s next decision. Buck said slow tool execution can delay an agent even when its AI processors are ready.
That differs from a conventional chatbot exchange, where a person may stop to read an answer before making another request. An agent can proceed immediately to another calculation or tool call. Patterson and Buck described that repeated loop as a reason CPU capacity matters alongside the accelerators that run the model.
The CPU step is not the only constraint Buck identified. He said model inputs had grown from roughly 1,000 tokens in 2023 benchmarks to about 142,000 tokens by 2026. That increase raises memory demands as an agent retains code, test results and other information needed for subsequent actions; Buck also said each context length needs its own software optimization.
The figures were Buck’s assessment at the summit, not an industrywide measurement. But they illustrate Nvidia’s case that agent workloads place demands on the supporting system as well as on the GPU: more information must be stored, retrieved and passed back into the model’s next step.
Nvidia is positioning Vera for the code execution and tool-running work surrounding the model. Its NVLink Fusion offering takes a different role: it enables processors designed by other companies to connect to Nvidia computing systems through a connecting chip and supporting designs. The pairing presents Nvidia as both a supplier of a CPU for agent tasks and a provider of the links that can join outside processors to its systems.
The summit remarks do not diminish GPUs’ role in generating answers. They instead describe an agent as a chain of dependent tasks, where fast model output alone cannot eliminate delays in executing tools or handling the information returned from them.
Loading discussion...
Join the conversation
Share which step you expect to slow an agent first.
Be the first to share a perspective or an experience.
Reader comments
Newest comments first. Replies stay oldest first.