Baseten Says Its Speech System Is 5× Faster Than OpenAI’s in Coval Benchmark
Baseten also reported the lowest word error rate in the early-access test. The result highlights how infrastructure choices can shape a voice agent’s responsiveness.
Listen to this story
The audio brief
Story brief
3 key pointsBaseten’s early-access Coval result suggests speech-to-text performance is becoming an infrastructure race, not just a model race. Its Qwen3 ASR 1.7B Streaming deployment reportedly combined the benchmark’s lowest word error rate with latency roughly five times lower than OpenAI’s tested deployment. The comparison is configuration-specific: Coval fixes the test region and evaluates providers’ models, APIs,...
- 01
Coval evaluates speed and transcription quality together, identifying deployments on a Pareto frontier rather than ranking model accuracy alone.
- 02
Baseten’s tested configuration used Qwen3 ASR 1.7B Streaming and standard APIs backed by dedicated inference infrastructure.
- 03
The benchmark’s fixed-region setup improves reproducibility but may not reflect latency when users and services are geographically separated.
Baseten says its speech-to-text deployment reached the quality-latency frontier in Coval’s early-access voice AI benchmark, with the lowest word error rate and latency about five times lower than OpenAI’s deployment. The result puts deployment infrastructure into the contest over whether a voice agent feels responsive in live conversation.
The early-access benchmark measures speech-to-text speed and quality together. Word error rate, or WER, represents transcription quality, while latency represents responsiveness.
A result tied to one deployed configuration
For the speech test, Coval evaluated Qwen3 ASR 1.7B Streaming. Baseten says no other provider in the current benchmark was both faster and lower in WER, putting its configuration on a Pareto frontier: results where no option improves both measures at once.
The test measures more than a model label
Coval’s published methodology identifies the dataset, provider model versions, normalization pipelines and latency measurement methods behind its leaderboard. That makes this a comparison between deployments, not simply a claim that one speech model is inherently better than another.
Baseten’s explanation is that voice systems add delay at every handoff. A typical pipeline sends audio to speech-to-text, forwards a transcript to a language model, then sends its reply to text-to-speech. Orchestration, networking, autoscaling and the physical placement of those services can add more time before a user hears a response.
A reproducible baseline, then a production test
Coval runs the benchmark from a fixed region to support fair, reproducible comparisons. But Baseten notes that workload placement can materially affect production performance, especially when users and services are separated by network distance.
Baseten says its endpoints enter Coval’s test as standard APIs while relying on dedicated inference infrastructure intended to improve consistency as well as speed. The benchmark offers teams a reproducible baseline, but deployment architecture determines how that baseline translates in production.
Sources
- baseten.coBaseten leads Coval’s voice AI benchmark
Loading discussion...
Reader comments
Newest comments first. Replies stay oldest first.