Productspublished

NVIDIA Adds BlueField-4 Scale-In for 800 Gb/s AI-Factory Infrastructure Offload

The new layer targets a growing bottleneck around GPU clusters: moving, securing and governing data without consuming the host resources assigned to AI workloads. Its performance case remains NVIDIA-supplied.

By 3 min read
NVIDIA Adds BlueField-4 Scale-In for 800 Gb/s AI-Factory Infrastructure Offload

Listen to this story

The audio brief

About 1:42
0:001:42
Read transcript
NVIDIA is adding a new Scale-In layer to its AI-factory networking architecture, built around BlueField-4 DPUs, DOCA software, and Spectrum-X Ethernet. The goal is to keep the work surrounding GPU clusters—moving data, enforcing security, provisioning systems, and collecting telemetry—off the host CPUs running AI workloads. BlueField-4 is rated by NVIDIA for up to 800 gigabits per second. It combines a 64-core Grace CPU, PCIe Gen6 connectivity, inline accelerators, and an 800-gigabit network interface. NVIDIA says that means twice BlueField-3’s network bandwidth and four times its memory bandwidth. In practice, Scale-In targets north-south traffic: connections between an AI factory and users, enterprise data, applications, services, and external storage. Its software layer, DOCA, can run containerized microservices directly on the DPU for networking, security, storage, telemetry, provisioning, and updates. Spectrum-X is intended to manage congestion and isolate traffic across shared infrastructure, including multi-tenant environments. NVIDIA describes Scale-In as its fifth networking pillar, alongside Scale-Up, Scale-Out, Scale-Across, and Context Memory for shared KV-cache storage. The company also claims up to 1.45 times the storage-path throughput of off-the-shelf Ethernet. But those performance figures are NVIDIA-supplied, with no independent validation or production deployment results in the post. The key constraint is whether this co-designed Vera Rubin architecture delivers the promised offload in real AI factories.

Story brief

3 key points

NVIDIA is adding a “Scale-In” layer to its AI-factory networking architecture, using BlueField-4 DPUs, DOCA software and Spectrum-X Ethernet to process storage, security, provisioning and telemetry away from host CPUs. BlueField-4 is rated for up to 800 Gb/s, with twice BlueField-3’s network bandwidth and four times its memory bandwidth. The strategy targets increasingly multi-tenant, agent-driven systems, but the...

  1. 01

    BlueField-4 combines a 64-core Grace CPU, PCIe Gen6 connectivity, inline accelerators and an 800 Gb/s interface.

  2. 02

    DOCA microservices can run on the DPU for networking, security, storage, telemetry, provisioning and updates.

  3. 03

    Scale-In addresses north-south traffic between AI factories and users, enterprise data, applications and external storage.

NVIDIA says its new Scale-In layer is designed to prevent infrastructure around AI compute from becoming a bottleneck. It combines BlueField-4 DPUs, DOCA software and Spectrum-X Ethernet to handle access, security, storage and telemetry outside host CPUs. BlueField-4 supports up to 800 Gb/s, according to the company.

Scale-In is NVIDIA’s fifth AI networking pillar. It sits alongside Scale-Up, which links GPUs; Scale-Out, which connects servers; Scale-Across, which connects distributed AI factories; and Context Memory, which NVIDIA positions for shared KV-cache storage and reusable inference state.

The distinction is architectural. North-south networks carry traffic between an AI factory and its users, applications, enterprise data, services and external storage. NVIDIA’s argument is that those services need their own processing domain as AI systems add more agents, tenants and continuously accessed data sources.

BlueField-4 combines a 64-core Grace CPU, inline acceleration engines, PCIe Gen6 host connectivity and an 800 Gb/s network interface. NVIDIA says this design allows software to set policy and provisioning while hardware handles networking, storage and security locally, without sending the work back to the host CPU.

Diagram showing Scale-Up, Scale-Out, Scale-Across and Scale-In as separate AI factory infrastructure domains.
NVIDIA positions Scale-In as the layer that connects, secures, provisions and observes infrastructure surrounding AI compute. Source: developer.nvidia.com.

DOCA is the software bridge in the design. NVIDIA says its containerized microservices can run directly on BlueField-4, while its libraries and SDKs expose accelerated capabilities for networking, security, storage, telemetry, provisioning and updates.

The operational targets

  • Virtual private clouds: NVIDIA identifies isolated tenant and application traffic on shared AI infrastructure as a Scale-In use case.
  • Security: NVIDIA says BlueField-4 can enforce network, file and object access controls outside the host operating system.
  • Storage: BlueField-4 is designed to accelerate storage protocols, virtualization and data movement between AI compute and external data systems.

NVIDIA says Spectrum-X Ethernet manages congestion and traffic isolation across the Scale-In path. The company also says its BlueField-4 and Spectrum-X storage path can deliver up to 1.45 times the throughput of off-the-shelf Ethernet. These are company-supplied architecture and performance claims; the post provides no independent validation or production deployment results.

NVIDIA says Scale-In is co-designed with Vera Rubin infrastructure so processor performance, memory, PCIe, networking, acceleration and software can scale together. The proposition is broader than a faster network card: AI factories will need the systems feeding and governing GPUs to keep pace with the GPUs themselves.

Editorial analysis

Our Read

NVIDIA’s Scale-In proposal extends its AI-factory design beyond connecting GPUs to controlling the traffic and services around them. That is strategically important because agentic workloads increase pressure not only on model serving, but also on data access, tenant controls and observability. The next meaningful evidence will be deployment results that show whether BlueField-4’s offload model improves end-to-end AI operations under shared, multi-tenant load—not simply component bandwidth. NVIDIA’s recent Vera Rubin performance claims make that distinction more important: system-level gains depend on the infrastructure around the accelerator as well as the accelerator itself.

Sources

  1. developer.nvidia.comNVIDIA BlueField-4 Powers New Scale-In Network Infrastructure for Agentic AI Factories | NVIDIA Technical Blog