Aug 27, 2026InfrastructureBenchmarkInfrastructureDatabricks Says Async Save Cut 20B-Parameter PyTorch Checkpoints to 9 Seconds on 32 H100sThe AI Runtime components pair distributed recovery with local-NVMe data caching. The remaining operational test is whether training jobs save their input position as well as their model state.3 min
Aug 27, 2026InfrastructureBenchmarkInfrastructureHeidi Cuts ASR GPUs From 16 to 4 With CUDA MPS, Holding Sub-Second LatencyThe production design turns spare capacity in small speech-inference requests into concurrency, but it depends on workload-specific scheduling and protections against MPS failure modes.3 min
Aug 25, 2026InfrastructureBenchmarkInfrastructureOpenAI’s Jalapeño Claims Up to 1.9x Efficiency Gain, but Leaves Nvidia Training IntactThe Broadcom co-developed inference processor is headed for deployment by year-end. Its published gains are meaningful, but they exclude model training and Nvidia’s Vera Rubin platform.3 min
Aug 25, 2026InfrastructureBenchmarkInfrastructureOpenAI’s Jalapeño Claims 1.5–1.9x More AI Work Per Watt, Faces 2027 Scale TestThe custom chip gives OpenAI an early efficiency claim against available Nvidia systems, but production qualification, broader workloads and a small initial rollout will determine whether it becomes a meaningful serving platform.4 min
Aug 24, 2026InfrastructureBenchmarkInfrastructureNVIDIA Says Vera Rubin Delivers 30x More Agentic Throughput per Megawatt, Pending ReviewThe preliminary comparison shifts attention from fixed prompt benchmarks toward the power required to keep long, tool-using AI sessions responsive.3 min
Aug 18, 2026InfrastructureBenchmarkInfrastructureLMCache Reworks the Cache Plumbing That Can Stall Long-Running AgentsThe changes address a specific failure mode in long, concurrent agent sessions: reusable model state can fill a shared pool before any request makes progress. The reported gains are promising, but they come from targeted validations rather than a broad production benchmark.3 min
Aug 18, 2026InfrastructureBenchmarkInfrastructureNvidia Dynamo’s Agentic Speedups Target the Work Around the ModelThe reported gains come from changing how serving systems route, retain and stream state under concurrency. The next challenge is whether these AgentX-specific improvements hold across production mixes and unfinished scheduling work.4 min
Aug 18, 2026InfrastructureBenchmarkInfrastructureAgentX Replays Claude Code Sessions to Test AI Serving SystemsThe replay dataset is designed to measure infrastructure behind agentic coding workloads, not whether a model writes better code. Its value depends on whether synthetic traces preserve the traffic patterns operators need to serve.3 min