Compare speech-to-speech models on live calls
Publishes benchmarks for nine realtime voice models tested as complete phone agents across 82 scenarios, with three runs per scenario.
Loading page…
Audio and video / product dossier
Benchmarks realtime voice models on live phone calls, with public transcripts and reliability metrics.
Product brief
Cekura Bench publishes voice AI benchmarks you can verify. Our new speech-to-speech benchmark tests 9 realtime voice models, including GPT Realtime 2.1, Gemini Live, Grok and Phonic, as complete phone agents on live calls: 82 scenarios, three runs each. Models are ranked on reliability, data accuracy, stalled calls, response time and cost, and every call transcript is public. Cekura Bench also covers voice agent benchmarks and STT benchmarks, with TTS benchmarks coming soon.
Why we selected it
Offers concrete evidence for voice-agent model selection: nine realtime models tested across 82 live-call scenarios, with three runs each and public transcripts. Reliability, accuracy, latency, and cost measurements are,
Product preview saved with our daily selection
Capability scan
Only capabilities supported by the product information we collected are listed here.
Publishes benchmarks for nine realtime voice models tested as complete phone agents across 82 scenarios, with three runs per scenario.
Ranks models on reliability, data accuracy, stalled calls, response time, and cost.
Makes every benchmark call transcript public so readers can inspect the conversations behind the results.
Covers voice-agent benchmarks and speech-to-text benchmarks alongside its speech-to-speech evaluation.
Best-fit use cases
FAQ
It tests nine realtime voice models as complete phone agents on live calls, using 82 scenarios and three runs per scenario.
The named models include GPT Realtime 2.1, Gemini Live, Grok, and Phonic. The supplied description does not list all nine models.
Models are ranked on reliability, data accuracy, stalled calls, response time, and cost.
Yes. Every call transcript is public.
Cekura Bench includes speech-to-text benchmarks. Text-to-speech benchmarks are described as coming soon, rather than currently available.