Mercor Says Queue Redesign Increased AI Evaluation Throughput 6–20 Times
The Studio scheduler delays jobs before they occupy paid containers, then adjusts concurrency to provider demand. Gateway limits still require manual configuration.
Mercor’s scheduler redesign addresses resource contention by deciding which evaluation jobs receive compute before containers are allocated, using model-gateway queue depth to adjust concurrency every 10 seconds. The company says six high-volume lanes now exceed their previous throughput caps by 6–20×, while its platform handles over 2.5 million workloads weekly across 241 models. The gains concern lane capacity, not end-to-end evaluation speed; manually configured gateway limits and workloads outside the scheduler remain open work.
01
At peak load, 90th-percentile evaluation trajectories spent 82% of their time waiting for containers; grading spent 75% waiting.
02
Admission rules allocate slots by priority tier, rotate among projects, and start each project’s oldest waiting job first.
03
Mercor says the feedback loop settles on a concurrency level within minutes, replacing a setting it previously adjusted by hand.
Mercor says its six busiest AI evaluation lanes now deliver 6–20 times their previous throughput caps, without adding model-provider capacity. In a September 30 engineering post, the company described a redesigned Studio scheduler that holds jobs back before they consume computing resources. The change targets a costly mismatch: starting more jobs does not help when the model provider is already backed up.
The work runs on Studio, Mercor’s platform where experts create tasks for model testing and reinforcement-learning training data. Each evaluation gives a model a task, tools and an isolated computing environment, then records its steps and tool calls. Mercor reports running more than 2.5 million evaluation workloads a week across 241 models.
Growing evaluation volume showed up as longer completion times, rather than a distinct new error. Mercor initially suspected provider rate limits. Its investigation instead found problems with both how it divided computing capacity and how it used the provider capacity available to it.
Popular models had been assigned separate queues with capacity carved out of a global limit. That protected smaller workloads from large batches, but adding more dedicated lanes lowered each lane’s ceiling. Mercor needed capacity to follow concentrated demand while still sharing it across models when traffic spread out.
Another queue already existed at Mercor’s model gateway, the layer that sends requests to providers. It prevented providers from being overloaded, but requests reached it only after their jobs had started. A job kept its container—the isolated environment allocated to it—throughout that wait. The gateway could delay a request, but could not stop the job from occupying resources.
The redesign adds a queue before container allocation. Jobs wait there without incurring container costs, rather than all starting with their batch. Every ten seconds, the scheduler decides how many jobs each model should run and which waiting jobs receive those slots.
Its feedback signal is the depth of each model’s gateway queue. When that queue is shallow and jobs are waiting, the scheduler raises the number allowed to run together. When the queue grows beyond its target, it lowers that limit; between those conditions, it holds steady.
Counting jobs alone cannot predict demand: one job might make a single model request or branch into dozens, and change behavior midway through. Mercor says the feedback loop settles within a few minutes on a concurrency level it previously set by hand.
Choosing which jobs start also changed. Under the old first-come system, a 20-job customer delivery could sit behind all 5,000 jobs of an earlier evaluation batch. The new admission rules divide available slots in three steps:
Priority tiers take 80% of the slots remaining to them, leaving lower-priority work room to keep moving.
Within a tier, projects take turns one job at a time. Batch size alone does not buy priority.
Within each project, the oldest waiting job goes first.
Priority passes from an account to its projects and batches. Mercor says raising an account’s priority is the control it uses most often. That preserves an explicit way to favor urgent work without letting one large batch automatically crowd out smaller projects.
The reported gains compare individual lanes with their old throughput caps; they are not a claim that every evaluation finishes that many times faster. Mercor’s next steps remain unfinished: gateway limits are still configured manually, and other model-heavy workloads do not yet use the scheduler. The company wants the gateway to adjust its own limits and the scheduler to become the default path for jobs dominated by model waits.
Sources
mercor.comHow we 20x LLM eval throughput | Mercor Engineering
Reader comments
Newest comments first. Replies stay oldest first.