Audio that starts before the sentence ends
Grok TTS streams chunks as they are generated, so playback begins while the rest of the paragraph is still being synthesised.
Grok TTS turns text into natural, controllable speech - streaming audio, expressive delivery, and production formats from one request.
Give it a sentence and it gives back a voice with intent: emphasis where it matters, pauses where they belong, and a read that stays consistent from the first word to the ten-thousandth.
Studio
Write a line, pick a voice, and set the delivery. The studio mirrors the API one to one, so anything you dial in here maps directly to the request you ship.
Grok TTS reads the room. [pause] Ask for warmth and you get warmth. Ask for a warning and the warning lands. [serious] Every sentence arrives with the intent you wrote into it.
Capabilities
Everything below is a request parameter or a documented behaviour. Nothing depends on a human in the loop.
Grok TTS streams chunks as they are generated, so playback begins while the rest of the paragraph is still being synthesised.
Inline tags control laughter, pauses, emphasis, and register. The voice follows your script instead of flattening it.
Upload a few seconds of clean speech and keep it as a reusable voice. Prosody carries over; the licence stays yours.
128 languages share the same voice identity, so a line recorded in English keeps its character in Japanese or Portuguese.
Tempo, sample rate, and output format are request parameters, not post-processing. What you ask for is what you receive.
Stable latency under sustained load, deterministic output for the same input, and usage you can watch per key.
Voices
Six of them are below. Every voice speaks all 128 languages with the same identity, and every voice can be re-created from a short reference recording of your own.
Warm, unhurried, and steady across long passages.
Low register with deliberate pacing and weight on key nouns.
Bright and quick, comfortable with jokes and sharp turns.
Neutral and efficient, tuned for guidance rather than drama.
Clear articulation and even rhythm for step-by-step content.
Confident and forward-leaning with a tight, punchy cadence.
Languages and formats
Latency
A request carries the script, the voice, the format, and the tempo. The response starts streaming as soon as the first chunk is ready, so your player is never waiting on the whole paragraph.
POST /v1/audio/speech
{
"voice": "aria",
"input": "Grok TTS reads the room. [pause] Ask for warmth and you get warmth.",
"format": "mp3",
"speed": 1.0
}How it works
Paste a sentence, a chapter, or a system prompt. Add delivery tags where the read should change.
Pick one of 32 built-in voices or a clone you have trained. Voices stay consistent across languages.
Request MP3, OPUS, WAV, or PCM and start playing on the first chunk instead of waiting for the file.
Track usage per key, watch remaining credits, and top up from the pricing page when volume grows.
Use cases
Narration for product videos and social cuts without a booking.
Long-form chapters with a voice that stays identical for hours.
Low-latency replies inside a live conversation.
Read any article, ticket, or document aloud on request.
Re-record a script in another language with the same character.
Turn alerts and digests into audio people actually finish.
Pricing
Signing in with Google includes free credits. Monthly plans add generation capacity; annual plans cost less per month.
For a first paid run and validating one workflow.
For recurring narration and iterative script testing.
For continuous generation and higher throughput.
Questions
Grok TTS is a text-to-speech engine for products that speak: narration, assistants, notifications, and dubbing. It returns natural, controllable speech and streams audio so playback can start almost immediately.
Sign in with Google, generate your first line, and hear the difference delivery control makes.