Oracle and Alluxio Publish AI Storage Gains, With GPU Utilization Only Emulated
The compute-side cache sharply reduced latency in company tests. But the accelerator figures came from emulation, and MLCommons has not verified the result.
Loading page…
The compute-side cache sharply reduced latency in company tests. But the accelerator figures came from emulation, and MLCommons has not verified the result.
Listen to this story
Oracle and Alluxio’s Oct. 6 results make a practical case for caching repeated AI data reads near compute: operators can keep object storage as the authoritative copy while serving cached data from local NVMe. The reported storage gains came from physical equipment, but the 90–98% H100 utilization figures used emulation, and Oracle disclosed that MLCommons did not review the result. Whether the approach pays off depends on how much active data fits in cache, network capacity, and whether teams share GPU-node resources or deploy dedicated cache nodes.
On physical WARP tests, Alluxio Enterprise 3.6 cut average latency for 10 KiB objects from 24.1 ms to 0.45 ms at 16 concurrent connections.
For 1 GiB objects in an approximately 500 GB dataset, reported average throughput rose from 6.5 to 33.3 GiB/s; peak throughput was 40.9 GiB/s.
A separate six-node test reported 240,800 IOPS with Alluxio versus 11,800 at baseline for small-object reads.
Oracle and Alluxio’s October 6 benchmark publication makes a case for keeping AI accelerators busy by moving frequently used data closer to compute. Their tests showed much faster storage access with Alluxio’s distributed cache. But the headline accelerator-utilization figures came from emulated H100s—not physical H100 GPUs—and the MLPerf result remains unverified by MLCommons.
The reference architecture places Alluxio beside Oracle Cloud Infrastructure’s GPU compute. When requested data is missing from the cache, Alluxio retrieves it from object storage and keeps a copy on local NVMe drives across participating workers. Later reads can use those cached copies rather than return to remote storage.
Persistent object storage remains the authoritative home for the dataset. Customers do not have to bulk-migrate or reformat source data solely to add the cache. The design supports OCI Object Storage and other supported backends, including Amazon S3, Google Cloud Storage, Microsoft Azure Blob Storage and on-premises object storage.
The physical WARP test environment used Alluxio Enterprise 3.6 on OCI BM.DenseIO.E5.128 instances. Client workloads ran as Kubernetes jobs connected to Alluxio workers, with OCI Object Storage behind them. The tested network used 100 GbE over TCP/IP, without RDMA, a different method of moving data between machines.
For 10 KiB objects at 16 concurrent connections, the companies reported roughly 53 times lower average latency with Alluxio.
For 1 GiB objects in an approximately 500 GB dataset, reported average throughput rose roughly fivefold; peak throughput reached 40.9 GiB/s.
A separate small-object comparison used 48 concurrent connections and a six-node cluster. The companies reported 240,800 input/output operations per second, or IOPS, with Alluxio, against 11,800 at baseline—roughly a twentyfold increase.
The reported 90–98% accelerator utilization came from MLPerf Storage v2.0’s standard accelerator-emulation methodology. Storage activity, network traffic and caching ran on physical equipment, but H100 utilization was evaluated through emulation. Oracle’s MLPerf disclosure says the result did not undergo MLCommons review and may use methods or workload implementations inconsistent with verified-result specifications.
The publication also presents a 60%-to-95% GPU-utilization comparison. Those numbers are illustrative model inputs, not measured GPU results. The authors argue that reducing storage waits can improve effective infrastructure costs for appropriately configured, data-access-constrained workloads.
Deployment brings a resource choice. Alluxio can share GPU instances and their local storage, avoiding separate caching nodes while using CPU, memory and storage alongside the AI workload. Alternatively, dedicated Dense I/O nodes let operators scale the cache independently. The architecture’s benefits depend on cache capacity, network bandwidth and how much frequently requested data fits locally.
Loading discussion...
Join the conversation
Explain what evidence would make the case convincing.
Be the first to share a perspective or an experience.
Reader comments
Newest comments first. Replies stay oldest first.