Baseten Says AI-Built Inference Software Beat Its Baselines by Up to 90%
Two experimental servers improved language and image-model performance in company-run tests. The work depended on narrow targets, accuracy checks and substantial computing resources.
Baseten reports that coding agents built two specialized inference servers that outperformed their comparison systems: VibeQwen reached up to 90% higher generation speed on repetitive, structured text, while Sammie handled image segmentation at 50% higher throughput than Meta’s reference. The experiments suggest targeted agent-driven optimization could make bespoke serving software economical with limited human effort, but the gains are benchmark-specific. Neither server handles production traffic, and Sammie’s result does not prove that reusing the knowledge base caused its improvement.
01
VibeQwen ran Qwen-3.6-35B-A3B in NVFP4 on one NVIDIA B200; with 32 concurrent requests, Baseten reported 71% more output throughput than tuned vLLM 0.25.1.
02
On one H100, Sammie processed 91 images per second, according to Baseten.
03
VibeQwen took about a week and 200 B200 hours to develop, consuming 1.7 billion tokens; Sammie took a couple of days and about 200 million tokens.
AI agents can optimize the software that runs AI models, not just write applications around them. In results published October 2, Baseten says two agent-built servers beat their respective baselines, including up to 90% higher text-generation speed in one test. Neither server handles production traffic. The experiments, conducted in late July, paired largely autonomous coding with human steering and checks on output accuracy.
The work used MetaInfer, a toolkit that gives coding agents a knowledge base of requirements and constraints for building inference engines—the software that executes a model. Its approach requires neither additional model training nor a specialized agent harness. Agents test candidate changes, check their results and add what they learn to the knowledge base. Baseten extended that approach beyond the engine to the surrounding serving system.
Give the agent a narrow target—and firm boundaries
Baseten’s premise is specialization. General-purpose engines such as vLLM support many models, accelerators and workloads. A custom engine can instead target one particular combination. But speed alone is a dangerous objective: Baseten warns that an agent can satisfy a performance goal by producing useless output unless accuracy-preserving constraints are part of the task.
For its language-model experiment, Baseten gave Claude Code, running Fable 5, a goal: beat vLLM by 20% across performance metrics without losing accuracy against the chosen lower-precision model. It also gave the agent hardware access and tools to deploy candidates. Unlike the original MetaInfer experiment, Baseten allowed references to existing open-source engines.
How Baseten made the task concrete
Reuse existing low-level GPU routines from vLLM and TensorRT-LLM where available, rather than write every component from scratch.
Measure deployed services with multiple replicas behind a load balancer, using AIPerf to generate traffic and record performance.
Pause for human approval when a proposed change alters accuracy, with fixed checks governing the optimization process.
Those boundaries were not a requirement for identical outputs. Baseten ultimately permitted small numerical differences from the NVFP4 lower-precision reference, provided overall accuracy against the full-precision BF16 baseline was at least as good. The engineer also occasionally redirected the agent when it focused too heavily on a particular traffic pattern.
The largest gain came on structured text
The resulting engine, VibeQwen, ran Qwen-3.6-35B-A3B in NVFP4 on a single NVIDIA B200. Baseten compared it with a tuned vLLM 0.25.1 deployment. Its headline single-stream result used repetitive, structured text—a material condition on the 90% gain, rather than a promise of that improvement for every request.
Baseten says VibeQwen beat vLLM across all tested traffic patterns. The development effort lasted roughly a week and consumed about 200 B200 hours and 1.7 billion tokens, overwhelmingly cached input. The agent reached parity within the first few days; Baseten then let it continue searching for further gains.
A second server probes whether the learning transfers
Baseten reused the expanded knowledge base for Sammie, a server for SAM 3.1 image segmentation. Given an image and a short text prompt, it returns masks identifying matching objects. On one H100, Sammie processed 91 images per second, 50% more than Meta’s reference server, according to Baseten. Development took a couple of days and roughly 200 million tokens.
The faster second project does not establish that accumulated knowledge caused the improvement. Baseten calls the evidence for knowledge-base reuse suggestive, not definitive: the projects used different model architectures and accelerators, and there was no control test. That caveat concerns the value of reuse, distinct from Sammie’s reported throughput comparison.
Baseten puts the experiments’ costs in the hundreds of dollars for Sammie and thousands for VibeQwen, with only hours of human labor. Its business case is that even small efficiency gains can justify such work at large inference budgets. For now, those are development costs and experimental speed gains—not demonstrated savings from operating the servers in production.
Editorial illustration for Baseten Says AI-Built Inference Software Beat Its Baselines by Up to 90%.
Reader comments
Newest comments first. Replies stay oldest first.