Weekly Digest: Capacity, controls, and agent systems
AWS adds cross-Region GPT-5.6 inference, while agent systems expose fresh operational limits.
By Saeed Ezzati8 min read

The audio edition
Listen to this newsletter
0:004:04
Read transcript
AWS is adding cross-Region inference for OpenAI’s GPT-5.6 Sol, Terra, and Luna on Amazon Bedrock, across more than 25 AWS Regions. The important choice is not just model access; it is where processing can happen. A US geographic profile routes requests among predefined US destinations. A global profile draws on available capacity across supported commercial Regions, which may improve throughput under load but can move data across Regions. Workloads with geographic processing requirements should use the matching geographic profile, or call one Region directly. Applications send an inference-profile ID from a source Region, and Bedrock records both locations: CloudTrail shows the source, with an inferenceRegion field identifying where processing occurred. Cross-Region access also requires IAM permissions for the profile and foundation model in every eligible Region. The three models accept text and image inputs, return text, and offer a one-million-token context window, reasoning mode, server-side tool calling, and prompt caching. Teams can use Bedrock’s OpenAI-compatible Responses or Chat Completions APIs, or its Converse interfaces. One policy needs its own review: AWS says GPT-5.6 content flagged by automated abuse classifiers may be retained for up to 30 days for offline abuse detection. For teams already using OpenAI formats, the code change may be narrow. The harder decision is whether extra capacity fits the workload’s data-location and retention rules. That same capacity-versus-constraint tradeoff appears underneath the model, in the serving stack. SemiAnalysis says its AgentX benchmark helped partners produce more than 50 upstream pull requests across eight inference-software layers. The target is long-lived agents, where growing attention state, or KV cache, must be retained and moved—not discarded after one response. The work spans routing, tokenization, schedulers, engines, kernels, cache managers, and transfer infrastructure, including vLLM, SGLang, TensorRT-LLM, ROCm AITER, and others. Session-aware routing can keep an agent near its cached state, while cache transfer and offload manage bursts from subagents. Some changes remain open proposals, and the benchmark is not a broad performance result. The practical signal is that agent reliability increasingly depends on state management as much as raw accelerator speed. The infrastructure story widens further when capacity has to be financed. Nvidia says partnerships with Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs, and KKR are intended to mobilize more than 500 billion dollars for AI infrastructure. The proposed platforms cover GPUs, networking, cooling, electricity, and real estate, including Nvidia’s DSX AI factories. This is not Nvidia writing a 500-billion-dollar check; asset managers and private-equity firms would provide most funding, while Nvidia could offer strategic capital, credit support, or financing partnerships. Critics argue the eventual risk could move through private credit and insurers if borrowers cannot service data-center obligations. That is a scenario, not evidence of realized losses. Deal terms, borrower performance, and Nvidia’s final risk retention remain the key unknowns. And at the application layer, fresh data becomes another operational boundary. Databricks says its Feature Store can move a Kafka event into an online feature store in 200 milliseconds at the 99th percentile, using Spark Real-Time Mode, Lakebase, and Model Serving. The figure measures feature availability, not full model-decision time. Continuous processing updates rolling aggregates as events arrive, while Lakebase serves the latest values to inference. Databricks says the design keeps exactly-once guarantees, but recovery may replay up to five minutes of Kafka data. Across these stories, the useful question for next week is not simply how much AI capacity is being added. It is where state, data, and financial risk sit when that capacity is under pressure.


