Baseten Says Its Speech System Is 5× Faster Than OpenAI’s in Coval Benchmark

Baseten also reported the lowest word error rate in the early-access test. The result highlights how infrastructure choices can shape a voice agent’s responsiveness.

By 2 min read
Baseten Says Its Speech System Is 5× Faster Than OpenAI’s in Coval Benchmark
Baseten Says Its Speech System Is 5× Faster Than OpenAI’s in Coval Benchmark

Listen to this story

The audio brief

About 1:40
0:001:40
Read transcript
Baseten says its speech-to-text system reached the quality-and-latency frontier in Coval’s early-access benchmark, with the lowest word error rate and latency roughly five times lower than OpenAI’s tested deployment. In practical terms, that means a voice agent could turn spoken audio into text with less delay, while also making fewer transcription mistakes—at least in this specific test configuration. The system Coval evaluated was Qwen3 ASR 1.7B Streaming, running through Baseten’s standard application programming interfaces and dedicated inference infrastructure. Coval is designed to measure speed and accuracy together, rather than treating transcription quality as the only score. Its Pareto frontier identifies deployments where no alternative improves both measures at once. Baseten says no other provider in the current benchmark was both faster and lower in word error rate. The important qualification is that this is a deployment comparison, not a claim that one model is universally superior. Coval fixes the test region and accounts for model versions, normalization pipelines, and latency measurement methods. That improves reproducibility, but real-world performance can change when users and services are far apart geographically. And speech recognition is only the first handoff. Orchestration, networking, autoscaling, language-model processing, and text-to-speech all affect how quickly a user hears a reply. The result gives teams a useful baseline; the open question is how much of Baseten’s advantage survives in a production voice stack with different placement and traffic patterns.

Story brief

3 key points

Baseten’s early-access Coval result suggests speech-to-text performance is becoming an infrastructure race, not just a model race. Its Qwen3 ASR 1.7B Streaming deployment reportedly combined the benchmark’s lowest word error rate with latency roughly five times lower than OpenAI’s tested deployment. The comparison is configuration-specific: Coval fixes the test region and evaluates providers’ models, APIs,...

  1. 01

    Coval evaluates speed and transcription quality together, identifying deployments on a Pareto frontier rather than ranking model accuracy alone.

  2. 02

    Baseten’s tested configuration used Qwen3 ASR 1.7B Streaming and standard APIs backed by dedicated inference infrastructure.

  3. 03

    The benchmark’s fixed-region setup improves reproducibility but may not reflect latency when users and services are geographically separated.

Baseten says its speech-to-text deployment reached the quality-latency frontier in Coval’s early-access voice AI benchmark, with the lowest word error rate and latency about five times lower than OpenAI’s deployment. The result puts deployment infrastructure into the contest over whether a voice agent feels responsive in live conversation.

The early-access benchmark measures speech-to-text speed and quality together. Word error rate, or WER, represents transcription quality, while latency represents responsiveness.

A result tied to one deployed configuration

For the speech test, Coval evaluated Qwen3 ASR 1.7B Streaming. Baseten says no other provider in the current benchmark was both faster and lower in WER, putting its configuration on a Pareto frontier: results where no option improves both measures at once.

The test measures more than a model label

Coval’s published methodology identifies the dataset, provider model versions, normalization pipelines and latency measurement methods behind its leaderboard. That makes this a comparison between deployments, not simply a claim that one speech model is inherently better than another.

Baseten’s explanation is that voice systems add delay at every handoff. A typical pipeline sends audio to speech-to-text, forwards a transcript to a language model, then sends its reply to text-to-speech. Orchestration, networking, autoscaling and the physical placement of those services can add more time before a user hears a response.

A reproducible baseline, then a production test

Coval runs the benchmark from a fixed region to support fair, reproducible comparisons. But Baseten notes that workload placement can materially affect production performance, especially when users and services are separated by network distance.

Baseten says its endpoints enter Coval’s test as standard APIs while relying on dedicated inference infrastructure intended to improve consistency as well as speed. The benchmark offers teams a reproducible baseline, but deployment architecture determines how that baseline translates in production.

Sources

  1. baseten.coBaseten leads Coval’s voice AI benchmark

Loading discussion...