Modelspublished

Phonely Opens Alma, a Voice Model Trained on 10 Million Calls, to Outside Builders

The new external offering turns Phonely’s own call traffic into a model-development pitch, combining voice-specific training with company-reported speed and pricing advantages. Whether those results transfer to other companies’ call flows remains the practical test.

By 2 min read
Phonely Opens Alma, a Voice Model Trained on 10 Million Calls, to Outside Builders
Phonely Opens Alma, a Voice Model Trained on 10 Million Calls, to Outside Builders

Listen to this story

The audio brief

About 1:34
0:001:34
Read transcript
Phonely is taking Alma, its voice-agent model trained on more than ten million phone conversations, outside the company for the first time. The pitch is that phone calls behave differently from text chats: people interrupt, talk over background voices, change direction mid-sentence, and generate imperfect transcripts. Phonely says Alma was trained around those disruptions, rather than adapted from a general-purpose model with a long prompt. The commercial argument is speed and cost. Phonely reports that Alma starts producing an answer in under 185 milliseconds, with roughly 200-millisecond latency at the 99th percentile. Its comparison points are about 500 milliseconds and 2,000 milliseconds for GPT-4.1. Phonely also lists Alma at 55 cents per million tokens, versus three dollars and fifty cents for GPT-4.1 and five dollars and sixty-three cents for GPT-5.4. Those are company-reported comparisons, not independent benchmarks. Builders can keep their existing speech-recognition and text-to-speech providers, while enterprise customers can get dedicated rate limits and infrastructure options. Phonely says Alma already handles millions of its own calls each month, creating a feedback loop from live traffic into later model improvements. That makes the launch more than a model release: it is a test of whether an internal system trained on Phonely’s calls transfers reliably to other callers, workflows, and peak loads. That external performance is the key thing to watch.

Story brief

3 key points

Phonely is turning Alma from an internal voice-agent model into an outside-facing offering for builders, with enterprise deployments, dedicated rate limits, and compatibility with existing transcription and text-to-speech providers. The company says Alma was trained on 10 million-plus real calls and tuned for interruptions, background speech, and transcription errors. Phonely reports sub-185-millisecond first-token...

  1. 01

    Phonely says Alma already handles millions of calls monthly, creating a live-traffic feedback loop for model improvements.

  2. 02

    Reported latency is under 185ms for first token and about 200ms at p99, versus roughly 500ms and 2,000ms for GPT-4.1.

  3. 03

    Reported blended pricing is $0.55 per million tokens, compared with $3.50 for GPT-4.1 and $5.63 for GPT-5.4.

Phonely has launched Alma, a language model built for voice agents and trained on more than 10 million phone conversations. The company is now offering its model beyond Phonely’s own platform, arguing that live calls require a system built around the disruptions of speech rather than text-first interactions.

Training for the interruptions

Alma’s training set is the central distinction in Phonely’s pitch. Phonely says it used real conversations containing interruptions, background voices and transcription errors, and built Alma for those conditions. The company presents that data as an alternative to adapting a general-purpose model for a phone interaction.

Phonely says text-trained models remain the common foundation for voice agents, but can require extensive prompting to sound natural and may struggle when callers interrupt or change their mind mid-sentence. Those are the behaviors Alma is intended to address, rather than established performance results for the new model.

The speed-and-cost proposition

Phonely’s commercial case rests on delay and token cost. It reports that Alma produces its first token in under 185 milliseconds, versus roughly 500 milliseconds for GPT-4.1, and lists a blended price of $0.55 per million tokens. Phonely cites $3.50 for GPT-4.1 and $5.63 for GPT-5.4.

Call traffic as a feedback loop

The deeper distinction is not simply a faster inference setting. Phonely says it operates Alma across live conversations, identifies where exchanges break down and uses that performance feedback to adjust later responses without waiting for a new training cycle. In its account, call traffic is both the workload and an input to later model behavior.

Phonely says Alma already answers millions of its calls each month and is now available to other voice-application builders. It also says the model works with existing transcription and text-to-speech providers, allowing teams to retain those parts of their voice stack.

What Phonely is offering builders

  • Access to Alma for teams building voice agents.
  • Compatibility with transcription and text-to-speech providers, according to Phonely.
  • Custom deployments with dedicated rate limits and infrastructure-tailored options for enterprises.

An external test for an internal model

The launch turns Phonely’s internal model into an external option for teams seeking voice specialization without replacing their speech-recognition or speech-generation vendors. Its latency and price comparisons are Phonely-reported figures. The consequential question is how Alma performs across customers’ own callers, workflows and peak-volume conditions.

Sources

  1. siliconangle.comPhonely launches Alma, a voice AI model trained on over 10 million phone conversations - SiliconANGLE
  2. finance.yahoo.comPhonely Launches Alma, a Voice LLM That’s 63% Faster and 84% Cheaper Than OpenAI