SpaceXAI Releases Grok Voice Transcribe 2.0 at Its Existing API Prices

The new batch and live-transcription model is available now, while SpaceXAI’s strongest accuracy results come from its own testing and version 1.0 remains the current default.

By 2 min read
SpaceXAI Releases Grok Voice Transcribe 2.0 at Its Existing API Prices
SpaceXAI Releases Grok Voice Transcribe 2.0 at Its Existing API Prices

Listen to this story

The audio brief

About 1:31
0:001:31
Read transcript
SpaceXAI has released Grok Voice Transcribe 2.0, an upgraded speech-to-text model that keeps the existing API prices and is available now for both recorded audio and live streams. The company says it cuts word error rate from 20.6 percent to 6.8 percent on its internal multilingual test, where lower is better. But that headline result is company-reported, so the practical question is whether the improvement holds on the audio your product actually receives. Version 1.0 remains the default for now. SpaceXAI says version 2.0 will take over soon, with version 1.0 set for deprecation in the following weeks. Existing integrations should not need code changes, and customers can pin the older version during the transition. The API includes speaker labels through diarization, word-level timestamps, confidence scores, support for up to eight audio channels, key-term biasing, formatting for numbers and email addresses, and filler-word removal. SpaceXAI also says the model handles accents, poor phone connections, overlapping speakers, and language switching across dozens of languages. Its other major comparison comes from Artificial Analysis, where the company says 2.0 ranks first among 32 streaming models. Atlassian’s Loom is already using it for every video, but that does not establish performance across other workloads. Pricing stays at ten cents per batch audio hour and twenty cents per streaming hour. The next constraint is customer-side testing before the default changes.

Story brief

3 key points

SpaceXAI’s Grok Voice Transcribe 2.0 is a no-code-change upgrade for batch and live transcription, with pricing held at $0.10 per batch audio hour and $0.20 per streaming hour. The company reports a drop in word error rate from 20.6% to 6.8% on an internal multilingual test, while also claiming the top position among 32 streaming models on Artificial Analysis. Version 2.0 remains opt-in for now, but is expected to...

  1. 01

    Version 1.0 remains the default; SpaceXAI plans to switch defaults soon and deprecate it in the following weeks.

  2. 02

    The API supports diarization, timestamps, confidence scores, eight-channel transcription, key-term biasing, and formatting controls.

  3. 03

    Claims cover difficult audio such as accents, poor phone connections, overlapping speakers, and language switching across dozens of languages.

SpaceXAI has released Grok Voice Transcribe 2.0, promising twice the accuracy of its previous speech-to-text model at unchanged rates. The immediate offer is straightforward: developers can use it for recorded audio or real-time streams. The harder question is how broadly its performance claims will translate, since its headline comparison and several detailed results are supplied by the company itself.

The model is listed as available in SpaceXAI’s Speech-to-Text API. Its release notes say customers can select either version 1.0 or 2.0, with version 1.0 still the default. SpaceXAI says 2.0 will soon take that role and that it plans to deprecate 1.0 in the following weeks; customers wanting the older version during the transition can pin it.

For existing integrations, SpaceXAI says the new model’s accuracy improvement requires no code changes. Both batch processing and streaming remain supported, which puts the same release in two common workflows: transcribing completed recordings and producing text from live audio.

SpaceXAI positions 2.0 for conditions that frustrate transcription systems, including poor phone connections, competing speakers, accents, and spoken account details. It says the model automatically detects languages, follows a switch in language within one recording, and covers dozens of languages.

The company also says 2.0 leads 32 streaming models for accuracy on the public Artificial Analysis leaderboard and performs better than version 1.0 across four internal sets. Those include customer-support telephony, conversations with Grok, spoken credentials, and short multilingual commands. The public leaderboard claim offers one external reference point, but the release does not provide independent validation of the company’s broader internal comparisons.

Transcript controls included in the API

  • Word-level timestamps and confidence scores, plus speaker labels through diarization.
  • Independent transcription for up to eight audio channels.
  • Key-term biasing, formatting for items such as numbers and email addresses, filler-word removal, and speaker-turn detection.

SpaceXAI lists batch transcription at $0.10 per audio hour and streaming at $0.20 per hour, with diarization, timestamps, and key terms included. Keeping that price schedule makes this a replacement decision rather than a new budget line for current users, provided its claimed gains apply to their audio.

SpaceXAI says Atlassian’s Loom now uses the model to transcribe every video after finding it more accurate than its previous solution. That is a named deployment, not a measure of how the model will perform across other recordings, languages, or production setups. For developers, the near-term decision is practical: test the new default candidate on the messy calls, commands, and recordings their products actually receive.

Sources

  1. docs.x.aiRelease Notes | SpaceXAI Docs
  2. x.aiIntroducing Grok Voice Transcribe 2.0

Loading discussion...

YOUR READING SPACE

Notifications

SpaceXAI Releases Grok Voice Transcribe 2.0 at Its Existing API Prices | Superpower Daily