Prime Intellect Launches AI Model Serving Built for Long-Running Agents
Prime Inference brings an internally used platform to public customers. Its engineering targets crowded, long-context conversations, but the published speed gains come from company-run GLM-5.3 tests.
Prime Intellect launched Prime Inference on October 2, offering shared serverless endpoints and reserved GPU capacity for open models. Its infrastructure separates client endpoints from the model fleet and supports failover, positioning the service for workloads where agent sessions run for a long time and reuse conversation history. The launch makes the service publicly accessible, but its performance evidence is company-run and tailored to GLM-5.3; incomplete published pricing makes it difficult for buyers to compare costs.
01
Prime says it has supported customer deployments in production since January; nearly one trillion tokens per day refers to internal pre-release workloads, not public customer traffic.
02
In its GLM-5.3 tests, Prime reported about 50% more cache capacity at the same memory budget using NVFP4 compression, and nearly 40% lower 90th-percentile gaps between generated tokens.
03
Prime’s scheduling configuration cut median queue wait from 550 to 110 milliseconds by processing fewer prompt tokens per step, but added overhead for long prompts without cached history.
Prime Intellect released Prime Inference on October 2, 2026, promising reliable access to open AI models through shared, on-demand endpoints or reserved computing capacity. The platform has a production history, not just a launch-day demonstration. But its detailed performance case rests on company-run GLM-5.3 tests tuned for long-running agents—where preserving conversation history and keeping existing sessions moving are central to the design.
The service runs on Prime’s GPUs across multiple data centers. Serverless endpoints handle variable demand; reserved capacity serves sustained workloads. Existing tools using OpenAI’s software interface can connect with a Prime endpoint and API key, rather than requiring a different client interface.
Prime says it has served customer deployments in production since January. Before the public release, its internal workloads processed nearly one trillion tokens daily, covering reinforcement-learning runs, synthetic data generation, evaluations and coding agents. That figure describes internal pre-release use, not a disclosed volume of public customer traffic.
Its first public model deployment, GLM-5.3, went live on OpenRouter on September 22. For reliability, Prime separates the public API from the model fleet, allowing capacity to move without changing the client endpoint. Automatic failover routes traffic to healthy deployments, backed by cluster monitoring and a round-the-clock on-call team.
Prime’s test workload mixes returning agent sessions with new requests carrying long, uncached prompts. A typical turn adds about 6,000 tokens to a 140,000-token prompt, reusing most of the conversation. Its benchmark uses SemiAnalysis’s AgentX session replays, with additional long prompts that have no cached history.
Processing a long prompt can interrupt answer generation when both jobs share GPUs. Prime assigns prompt processing, called prefill, and token generation, called decode, to separate GPU groups. It reports nearly 40% lower 90th-percentile gaps between generated tokens in its tests—a measure of delays toward the slower end, rather than the average response.
The platform also routes requests toward workers holding reusable prompt history. Mooncake supplies another cache tier in host memory, letting workers retrieve that history instead of recomputing it. The serving stack combines Mooncake with NVIDIA Dynamo, vLLM and FlashInfer.
Scheduling involves a tradeoff. Halving the prompt tokens processed per step cut median queue wait from 550 to 110 milliseconds in Prime’s configuration. Smaller steps added overhead for long, uncached prompts, but suited its workload because most turns reused existing history. Those traffic characteristics are essential context for the gains.
Developers can point an OpenAI-compatible client at https://api.pinference.ai/api/v1, or test GLM-5.3 through Prime’s command-line tool. The launch also presents serving as part of a training loop: deployed models generate experience that can feed later training, extending the company’s existing post-training infrastructure.
prime inference chat 'z-ai/glm-5.3' "Write a haiku about KV caches."
Pricing remains a practical gap. MarkTechPost’s October 2 article noted that per-model prices were not fully published in the documentation. It also identified batch inference and one-click dedicated deployments as roadmap items, not available features. The launch provides a concrete serving architecture and public access; buyers still need the pricing detail to judge its economics.
Sources
primeintellect.aiPrime Inference: Fast, Reliable Serving for Frontier Open Models
marktechpost.comPrime Intellect Launches Prime Inference: Serverless and Reserved Serving for Frontier Open Models
Reader comments
Newest comments first. Replies stay oldest first.