Phonely Opens Alma, a Voice Model Trained on 10 Million Calls, to Outside Builders
The new external offering turns Phonely’s own call traffic into a model-development pitch, combining voice-specific training with company-reported speed and pricing advantages. Whether those results transfer to other companies’ call flows remains the practical test.
Listen to this story
The audio brief
Story brief
3 key pointsPhonely is turning Alma from an internal voice-agent model into an outside-facing offering for builders, with enterprise deployments, dedicated rate limits, and compatibility with existing transcription and text-to-speech providers. The company says Alma was trained on 10 million-plus real calls and tuned for interruptions, background speech, and transcription errors. Phonely reports sub-185-millisecond first-token...
- 01
Phonely says Alma already handles millions of calls monthly, creating a live-traffic feedback loop for model improvements.
- 02
Reported latency is under 185ms for first token and about 200ms at p99, versus roughly 500ms and 2,000ms for GPT-4.1.
- 03
Reported blended pricing is $0.55 per million tokens, compared with $3.50 for GPT-4.1 and $5.63 for GPT-5.4.
Phonely has launched Alma, a language model built for voice agents and trained on more than 10 million phone conversations. The company is now offering its model beyond Phonely’s own platform, arguing that live calls require a system built around the disruptions of speech rather than text-first interactions.
Training for the interruptions
Alma’s training set is the central distinction in Phonely’s pitch. Phonely says it used real conversations containing interruptions, background voices and transcription errors, and built Alma for those conditions. The company presents that data as an alternative to adapting a general-purpose model for a phone interaction.
Phonely says text-trained models remain the common foundation for voice agents, but can require extensive prompting to sound natural and may struggle when callers interrupt or change their mind mid-sentence. Those are the behaviors Alma is intended to address, rather than established performance results for the new model.
The speed-and-cost proposition
Phonely’s commercial case rests on delay and token cost. It reports that Alma produces its first token in under 185 milliseconds, versus roughly 500 milliseconds for GPT-4.1, and lists a blended price of $0.55 per million tokens. Phonely cites $3.50 for GPT-4.1 and $5.63 for GPT-5.4.
Call traffic as a feedback loop
The deeper distinction is not simply a faster inference setting. Phonely says it operates Alma across live conversations, identifies where exchanges break down and uses that performance feedback to adjust later responses without waiting for a new training cycle. In its account, call traffic is both the workload and an input to later model behavior.
Phonely says Alma already answers millions of its calls each month and is now available to other voice-application builders. It also says the model works with existing transcription and text-to-speech providers, allowing teams to retain those parts of their voice stack.
What Phonely is offering builders
- Access to Alma for teams building voice agents.
- Compatibility with transcription and text-to-speech providers, according to Phonely.
- Custom deployments with dedicated rate limits and infrastructure-tailored options for enterprises.
An external test for an internal model
The launch turns Phonely’s internal model into an external option for teams seeking voice specialization without replacing their speech-recognition or speech-generation vendors. Its latency and price comparisons are Phonely-reported figures. The consequential question is how Alma performs across customers’ own callers, workflows and peak-volume conditions.
Sources
- siliconangle.comPhonely launches Alma, a voice AI model trained on over 10 million phone conversations - SiliconANGLE
- finance.yahoo.comPhonely Launches Alma, a Voice LLM That’s 63% Faster and 84% Cheaper Than OpenAI