Read all nine tool guides as text
Every entry is readable without JavaScript. Features are from the linked documentation; strengths, limitations and first workflows are our editorial interpretation. Reviewed October 8, 2026.
Creator studio + APIElevenLabs
A broad creator workflow for narration, voices, dialogue and dubbing. Choose the model family as carefully as the voice.
- Models
- v4 / v4 Turbo; v3; Multilingual v2; Flash v2.5
- Direction
- Inline audio tags on supported v3/v4 models
- Languages
- v4: 90+; coverage differs for other families
- Voice identity
- Library, design and cloning; access depends on feature and plan
- Dialogue / dubbing
- Multi-speaker generation; separate dubbing product
- Access
- Web editor and API; commercial terms depend on plan
Strengths
- A practical starting point for creators who prefer a web interface.
- Several model families let you separate performance from fast delivery.
- Dubbing is available alongside speech generation.
Limitations
- Audio tags are model-specific; a voice can behave differently after a model switch.
- Dubbing v2 website output is automatic; transcript editing/regeneration via API requires Enterprise.
- Credits, allowances and temporary promotions complicate price comparisons.
A first workflow to try
- Pick v4 for an expressive short narration and audition two library voices.
- Use one performance instruction, generate a short take, then listen for skipped words.
- Lock the voice/model, regenerate only the changed section and assemble in your editor.
Example input: [reassuring] You are right on time. [pause] We can take this one step at a time.
Cost & access: Listed v4 API rate: $0.08/1K characters before the temporary $0.022 promotion ending October 12, 2026. Our planner uses the regular rate, not the promotion. Check plan allowances and rights.
Streaming speech APICartesia Sonic
An API-oriented option for conversational speech and repeatable deployment. Sonic 3.6 offers a dated snapshot as well as an automatically updated alias.
- Model
- Sonic 3.6; dated snapshot 2026-08-27
- Languages
- 44 documented languages
- Direction
- Model-specific emotion, speed and volume controls
- Voice identity
- Voice library and consent-based cloning
- Delivery
- Streaming API; playground for auditioning
- Deployment
- Hosted; enterprise deployment options require a separate agreement
Strengths
- Streaming suits responses that begin before the whole passage is ready.
- Dated versions make your own regression checks reproducible.
- Locale and normalization controls help structured speech.
Limitations
- It is an API workflow rather than a complete audiobook or video editor.
- Provider latency claims are not end-to-end conversation measurements.
- Commercial use, credits and concurrency depend on your plan.
A first workflow to try
- Audition a voice with your actual product vocabulary.
- Pin a dated model snapshot and test your chosen locale.
- Stream short turns and test interruptions in the application, not just the playground.
Example input: A short support response: “I found your booking. Your reference is A, one, zero, four.” Test the same line with the intended locale.
Cost & access: Credit-based self-serve plans; inspect the current character conversion, concurrency and commercial terms. Do not compare a credit count directly with another vendor’s credits.
Speech infrastructure APIDeepgram Aura
Useful to evaluate when spoken reminders, identifiers and support conversations matter more than theatrical performance. Speech recognition and agents are separate services.
- Model
- Aura-2 voice catalog
- Direction
- Voice selection and text preparation; inspect supported controls
- Voice identity
- Named voice IDs from the documented catalog
- Delivery
- REST and streaming speech endpoints
- Related products
- Separate transcription and voice-agent APIs
- Language check
- Match the exact voice ID to its supported language
Strengths
- The published examples make numbers and tricky vocabulary tangible.
- Speech and transcription can live within one provider stack.
- Streaming can reduce waiting for the entire output.
Limitations
- A launch-page preference study is conducted by the provider, not by us.
- Voice synthesis alone does not listen, reason or handle interruption.
- Catalog language support is not interchangeable with transcription language support.
A first workflow to try
- Choose an Aura-2 voice appropriate to the listener and language.
- Test dates, amounts, names and your most frequent identifiers.
- Integrate short streamed responses, then measure the entire request-to-playback path.
Example input: An appointment reminder with “March fourth”, “three oh five in the afternoon” and a letter-by-letter reference code.
Cost & access: Character-based synthesis; transcription and full voice-agent usage have separate rates. Check the current pay-as-you-go tier and language/model.
Instruction-driven speech APIOpenAI speech
A focused speech endpoint for apps and scripted content. Our original studio examples use one pinned model, with the script and instructions published.
- Our example model
- gpt-4o-mini-tts-2025-12-15
- Direction
- Natural-language instructions; speed parameter
- Request limit
- 4,096 input characters per speech request
- Formats
- MP3, Opus, AAC, FLAC, WAV and PCM
- Voice identity
- Built-in voices; custom voices are eligibility-gated
- Live conversation
- The speech endpoint differs from a Realtime conversation API
Strengths
- Separate text and delivery instructions make controlled lessons easy to reproduce.
- Output formats suit both editing and application playback.
- A dated model makes the provenance of these examples explicit.
Limitations
- This endpoint is not a complete creator timeline or dubbing editor.
- Instructions guide a performance; they are not a guarantee of exact timing or stress.
- Actual billing is token-based; the per-minute price is an estimate.
A first workflow to try
- Write a short, plain-language script and choose a built-in voice.
- Add a concise tone/pacing instruction and generate two takes.
- Listen with the transcript, correct pronunciation and export the chosen take.
Example input: Text: “You are right on time.” Instructions: Warm and reassuring; gentle pace; speak to one worried person.
Cost & access: The pricing page estimates gpt-4o-mini-tts at $0.015 per generated minute. Actual charges use text and audio tokens; retries create additional output.
Directed speech + dialogue APIGoogle Gemini TTS
A route for style-directed narration and dialogue. Google’s current guide documents both speaker configuration and per-turn direction.
- Direction
- Style instructions and speech metadata in the current interface
- Dialogue
- Up to two prebuilt speakers in a single multi-speaker request
- Custom dialogue
- Custom voices require separate turns in the documented workflow
- Output
- Audio output; inspect WAV/PCM handling for assembly
- Languages
- Check the supported language list for the chosen model
- Model / interface
- Pin the exact model and SDK/API version from the current guide
Strengths
- Useful to explore a scripted exchange between two speakers.
- Turn-level direction fits dialogue rather than a single narration style.
- The documentation includes single- and multi-speaker workflows.
Limitations
- Two-speaker generation is not an unlimited cast in one request.
- Custom-voice turns need separate synthesis and assembly.
- Different Gemini/Cloud interfaces expose different controls and billing.
A first workflow to try
- Create a two-person exchange with short, clearly labeled turns.
- Assign prebuilt voices and a restrained style for each turn.
- Check voice separation and silence, then assemble and caption the output.
Example input: HOST: “What changes when we slow down?” GUEST: “The listener gets a moment to think.” Keep the two styles calm and distinct.
Cost & access: Model-specific API pricing; inspect text-input and audio-output units for your selected interface. A Gemini API allowance is not a Google Cloud TTS allowance.
Cloud speech synthesisGoogle Chirp 3 HD
A cloud voice catalog with documented text and pronunciation controls. The examples below show what a real pause or pronunciation override changes.
- Model
- Chirp 3: HD, Google Cloud Text-to-Speech
- Direction
- Documented SSML; voice controls include preview features
- Pacing
- Documented speaking-rate control, 0.25–2.0
- Pronunciation
- IPA/X-SAMPA overrides in supported locales
- Pauses
- Pause markup uses the markup field, not plain text
- Access
- Cloud project, billing and synthesis API
Strengths
- Explicit control can be useful for names, codes and instructional pacing.
- The voice catalog lets you inspect locale-specific options.
- Published examples include the same phrase with and without a pause.
Limitations
- Preview controls have language and launch-stage limitations.
- Pause duration can vary with context and some tags may be ignored.
- Cloud setup takes more effort than an all-in-one creator editor.
A first workflow to try
- Choose the exact locale and voice ID.
- Test a phrase with a difficult name and a deliberate pause.
- Use supported pronunciation controls and validate the actual audio before scaling.
Example input: Compare “Let me take a look, yes, I see it.” with the same text using [pause long] in the documented markup field.
Cost & access: Published Chirp 3 HD rate: $30 per million characters after the listed free allowance. Planner excludes allowances, taxes and other Cloud services.
Cloud speech + SSMLMicrosoft Azure Speech
A route for teams that need structured control over pronunciation and pacing within an existing Microsoft speech workflow.
- Direction
- SSML for pronunciation, rate, pitch, pauses and more
- Styles
- Speaking styles and roles depend on the selected voice
- Access
- Speech SDK, REST and voice-specific configuration
- Output
- Select an audio format appropriate to the playback channel
- Voice / language
- Verify the exact voice’s locale and supported features
- Workflow
- Cloud resource and credentials; editor or app integration
Strengths
- SSML makes the intended structure explicit.
- SDK and REST routes support integration into existing systems.
- Voice-specific styles can support instructional speech.
Limitations
- Not every voice supports every SSML element or style.
- Markup support does not guarantee the pronunciation you intended.
- Resource region, voice availability and the selected tier need checking.
A first workflow to try
- Choose a supported neural voice and inspect its style list.
- Add one pronunciation override and one pause in SSML.
- Synthesize a short test, review by ear and export at the required audio format.
Example input: An SSML learning example: <speak version="1.0" xml:lang="en-US"><voice name="en-US-JennyNeural">Take a breath.<break time="500ms"/>Then continue.</voice></speak>
Cost & access: Cloud billing depends on the voice tier and region. Check the Speech pricing page for your resource; synthesis, custom voice work and other services differ.
Open-weight local modelKokoro-82M
A small open-weight speech model for people who want to operate their own runtime. The model card is the authoritative project entry point.
- Size
- 82 million parameters
- Weights license
- Apache 2.0
- Model card release
- v1.0: January 27, 2025
- Voice workflow
- Select supplied voices and language pipeline
- Deployment
- Local runtime or a third-party host
- Control
- Text/phoneme preparation and pipeline settings; no promised arbitrary acting prompts
Strengths
- Your deployment can run without sending each script to a hosted speech API.
- Open weights allow experimentation with the synthesis pipeline.
- A useful alternative when hosting control matters more than a creator suite.
Limitations
- Local operation still costs compute, setup and maintenance.
- A third-party hosted endpoint has its own pricing and data handling.
- Open weights do not provide an unlimited cloning or dialogue editor.
A first workflow to try
- Start from the author’s model card and installation instructions.
- Choose a language pipeline and supplied voice; synthesize a short script.
- Inspect pronunciation and runtime speed on your own hardware.
Example input: Try one 30-word tutorial introduction, then change the supplied voice while keeping text and runtime settings fixed.
Cost & access: No hosted vendor generation fee when run locally; hardware and operation still cost money. Third-party API prices are separate. License applies to the published weights.
Retiring service / historical referenceHume Octave
An instructive example of voice design and acting controls, but not a recommended dependency for a new project: Hume has announced a voice API sunset.
- Status
- TTS and EVI access ends November 13, 2026, 12:01 a.m. EST
- Data
- The notice says account data is deleted after November 13
- Acting direction
- Description field: Octave 1; Octave 2 support listed as coming soon
- Other controls
- Speed and trailing silence supported across the documented models
- Learning value
- Distinguishes voice identity from delivery instructions
- New-project fit
- Excluded from our recommended starting points
Strengths
- The archived concepts help explain identity, intention and timing.
- Version-specific docs expose why a control should not be generalized across models.
Limitations
- Announced access deadline makes a new production integration unsuitable.
- Export needed material and confirm the notice with Hume before the closure.
- Old articles and demos can outlive the product they describe.
A first workflow to try
- Read the retirement notice before taking any action.
- If you are an existing customer, export needed audio and review your migration requirements.
- Audition alternatives with your own scripts, not a vendor leaderboard alone.
Example input: A performance brief can specify a calm voice, a measured pace and silence after an utterance; the supported fields vary by Octave version.
Cost & access: Historical reference only. Do not budget a new production project around an API with an announced closure.