NVIDIA Adds PAIR to Spread Local AI Jobs Across Nearby PCs
The free beta is designed to keep separate agent tasks from queuing behind one GPU, while its utility depends on having compatible machines with spare local capacity.
Listen to this story
The audio brief
Story brief
3 key pointsNVIDIA’s open-source PAIR beta turns nearby computers into a shared inference layer for local agents, routing independent subagent calls to whichever compatible machine has capacity. It runs across Windows, macOS, and Linux, supports RTX 20-series and newer GPUs plus Apple M4 silicon, and integrates with Ollama and LM Studio. The practical payoff is keeping a primary PC responsive while work is distributed locally,...
- 01
PAIR adapts as devices join or leave a local network, assigning work without requiring a fixed cluster.
- 02
NVIDIA reports up to 1.9× llama.cpp throughput on an RTX 5090; other gains vary by backend and hardware.
- 03
Hermes Agent adds one-click Windows setup that detects the GPU, selects a model, and configures llama.cpp.
A local AI agent can now send separate pieces of work to different nearby computers instead of making them wait behind one graphics processor. NVIDIA has released PAIR, a free open-source beta that discovers compatible PCs on a local network and routes independent inference requests to systems with available capacity.
The tool is aimed at agent workflows that divide a larger task into smaller jobs that can run at the same time. PAIR assigns those independent jobs among available machines, rather than putting every request in line for a single GPU. That makes the tool a local traffic director for work that can be separated.
From subagent requests to available computers
PAIR works with Ollama and LM Studio and adapts when devices join or leave the network. NVIDIA illustrates the process with Hermes Agent dividing an inbox-sorting task among subagents; PAIR then distributes the resulting requests across available PCs. The company says this can leave a main system free for gaming, creative work or other tasks while local agent jobs run elsewhere.
The beta spans Windows, Mac and Linux
- PAIR is available for Windows, macOS and Linux through graphical and terminal interfaces.
- It supports GeForce RTX 20 Series GPUs and newer, Turing-or-newer RTX PRO workstation GPUs, DGX Spark, and Apple M4-or-newer silicon.
NVIDIA paired the router with efforts to reduce local-model setup. Hermes Agent, OpenClaw and Perplexity Portable Computer are adding simplified NVIDIA GPU setup, with Windows availability differing by application. NVIDIA says Hermes offers one-click setup on Windows that detects the GPU, selects a model and configuration, and runs it through integrated llama.cpp; Linux support is planned.
Perplexity Portable Computer is available on Linux systems with NVIDIA RTX GPUs carrying at least 24GB of VRAM, and Windows support is coming soon. It can seek permission to send parts of a task to more than 15 cloud models when additional research or reasoning is needed.
Speed claims remain tied to named backends and hardware
NVIDIA says new llama.cpp optimizations deliver up to 1.9 times higher throughput on a GeForce RTX 5090. It also reports vLLM performance of 1.2 times on an RTX PRO 6000 Blackwell Workstation Edition and up to 1.4 times on two DGX Spark clusters. The improvements are available through the llama.cpp and vLLM backends, and through LM Studio and Ollama.
The software update arrives before NVIDIA RTX Spark Windows PCs, scheduled for October 2026, including newly announced Acer and Lenovo designs. NVIDIA describes the platform as an RTX Blackwell GPU delivering one petaflop, with up to 128GB of unified memory and a 20-core Grace CPU. Acer showed its compact SFF RTX Spark design at IFA and said device availability would be announced later.
Sources
- blogs.nvidia.comSparks Fly: NVIDIA Accelerates Local AI at IFA 2026
- news.acer.comAcer Showcases Design Powered by NVIDIA RTX Spark™ at IFA 2026