OpenAI Releases GPT-Live-1 API for Voice Apps That Talk and Listen Together

The developer release targets the pauses and interruptions that make automated calls feel rigid, while allowing builders to select backend models for different speed, cost and reasoning needs.

By 2 min read
OpenAI Releases GPT-Live-1 API for Voice Apps That Talk and Listen Together
OpenAI Releases GPT-Live-1 API for Voice Apps That Talk and Listen Together

Listen to this story

The audio brief

About 1:29
0:001:29
Read transcript
OpenAI has released GPT-Live-1, an API for voice apps that can listen and speak at the same time. The goal is to make automated calls feel less like a sequence of handoffs, where the system waits for a person to finish, processes the turn, and only then responds. GPT-Live-1 supports full-duplex interaction, and OpenAI prices it at five cents per minute. Developers can also choose backend models based on the reasoning depth, speed, and cost a particular voice task requires, instead of using one model for every exchange. The API includes twelve voices across accents, dialects, and languages, along with automatic speech-recognition transcripts and response text. OpenAI says its own testing showed an 80.1 percent score, compared with 45.4 percent for GPT-Realtime-2.1. It also reports turn-taking latency of 0.8 seconds, versus 1.4 seconds, and tool-calling accuracy of 87 percent, versus 60 percent. Those are company-run comparisons, though. In a banking-support benchmark, GPT-Live-1 passed 32 percent of tasks, up from 12.4 percent for the previous model—progress, but still a majority of tasks missed. Yelp is already using it for phone reservations, with chief technology officer Alex Levy reporting improved call handling. The key question now is whether that smoother interaction translates into reliable customer service beyond benchmarks and this early deployment.

Story brief

3 key points

OpenAI’s GPT-Live-1 gives developers a $0.05-per-minute API for full-duplex voice applications, with 12 voices, transcripts, and configurable backend models. OpenAI reports higher scores, better tool calling, and lower turn-taking latency than GPT-Realtime-2.1, but those are company-run comparisons. Its banking-support result was only a 32% pass rate, making deployment evidence important. Yelp is already using the...

  1. 01

    GPT-Live-1 supports simultaneous listening and speaking, reducing rigid wait-for-the-user turn-taking in automated calls.

  2. 02

    OpenAI reports 0.8-second turn-taking latency versus 1.4 seconds for GPT-Realtime-2.1.

  3. 03

    The API includes 12 voices plus automatic speech-recognition transcripts and response text.

Automated voice apps can now respond without treating every spoken exchange as a rigid handoff. OpenAI has released GPT-Live-1 as a developer API, offering a model that can listen and speak at the same time for live voice interactions.

The release is aimed at the pauses common in automated calls, where a system waits for a person to finish before responding. GPT-Live-1 supports full-duplex interaction, and OpenAI lists the API price at $0.05 per minute.

Voice interaction with a configurable backend

Developers can pair GPT-Live-1 with different backend models based on the reasoning depth, speed and cost their application needs. That gives builders a way to make different model choices for different voice tasks rather than treating every conversation alike.

Included in the release

  • Twelve new voices spanning accents, dialects and languages.
  • Automatic speech-recognition transcripts and response text by default.
  • Backend-model choices based on reasoning depth, speed and cost.

OpenAI’s case rests on faster exchanges

OpenAI says its testing shows the model is more responsive and capable in voice workflows than GPT-Realtime-2.1. Those comparisons are company benchmark results, not independent evaluations of deployed voice agents.

A tougher test in banking support

The company also reports a 32% pass rate for GPT-Live-1 in a banking voice-support benchmark, up from 12.4% for the previous model. The result suggests progress, but it also shows a system that did not pass most tasks in that reported test.

Phone reservations offer an early real-world use

Yelp is using GPT-Live-1 for phone-based reservations and has reported improved call handling, according to CTO Alex Levy. Reservations may be a useful early test for the model because callers can change or clarify a request while speaking.

For developers, the immediate proposition is a more fluid voice interface and a choice of backend models. Whether that produces better customer-service outcomes will depend on how those systems perform in real deployments, beyond OpenAI’s benchmarks and Yelp’s early report.

Editorial analysis

Our Read

GPT-Live-1’s practical test is whether a more natural-sounding conversation improves the tasks behind it. OpenAI’s benchmarks show large gains in interactivity, latency and tool-calling accuracy, while Yelp says it has improved phone-based reservation handling. But the API also lets developers pair the voice model with backends chosen for different reasoning depth, speed and cost. That flexibility makes the voice experience only one part of the eventual service. The next evidence to watch is whether customer deployments publish outcome measures beyond early call-handling reports and company benchmarks.

Sources

  1. the-decoder.comOpenAI's GPT-Live-1 API lets developers build apps that talk and listen at the same time

Loading discussion...