Grok TTS

Speech that sounds like it means it.

Grok TTS turns text into natural, controllable speech - streaming audio, expressive delivery, and production formats from one request.

Give it a sentence and it gives back a voice with intent: emphasis where it matters, pauses where they belong, and a read that stays consistent from the first word to the ten-thousandth.

~190 mstime to first audio128languages32built-in voices44.1 kHzMP3 · OPUS · WAV · PCM

Studio

Hear it before you wire it up

Write a line, pick a voice, and set the delivery. The studio mirrors the API one to one, so anything you dial in here maps directly to the request you ship.

Script
Delivery tags
Voice
Output format
Tempo
Sign in with Google
Free credits included
Preview

Aria

Grok TTS reads the room. [pause] Ask for warmth and you get warmth. Ask for a warning and the warning lands. [serious] Every sentence arrives with the intent you wrote into it.
  • VoiceAria
  • FormatMP3
  • Tempo1.0x
  • Sample rate44.1 kHz

Capabilities

A voice layer that behaves like infrastructure

Everything below is a request parameter or a documented behaviour. Nothing depends on a human in the loop.

Streaming

Audio that starts before the sentence ends

Grok TTS streams chunks as they are generated, so playback begins while the rest of the paragraph is still being synthesised.

Delivery

Direct the performance, not just the words

Inline tags control laughter, pauses, emphasis, and register. The voice follows your script instead of flattening it.

Cloning

Your voice, from a short reference

Upload a few seconds of clean speech and keep it as a reusable voice. Prosody carries over; the licence stays yours.

Languages

One product, many markets

128 languages share the same voice identity, so a line recorded in English keeps its character in Japanese or Portuguese.

Control

Knobs that map to how it sounds

Tempo, sample rate, and output format are request parameters, not post-processing. What you ask for is what you receive.

Reliability

Built to sit inside a request path

Stable latency under sustained load, deterministic output for the same input, and usage you can watch per key.

Voices

Thirty-two voices, each with a reason to exist

Six of them are below. Every voice speaks all 128 languages with the same identity, and every voice can be re-created from a short reference recording of your own.

Aria

Narration

Warm, unhurried, and steady across long passages.

Atlas

Documentary

Low register with deliberate pacing and weight on key nouns.

Ember

Character

Bright and quick, comfortable with jokes and sharp turns.

Onyx

Support

Neutral and efficient, tuned for guidance rather than drama.

Sage

Instruction

Clear articulation and even rhythm for step-by-step content.

Vale

Advertising

Confident and forward-leaning with a tight, punchy cadence.

Languages and formats

Pick the language and the container

Output formats

  • MP3Compressed, plays everywhere, the default for web playback.
  • OPUSSmallest footprint for real-time streams and mobile clients.
  • WAVUncompressed for a downstream editor or a mastering pass.
  • PCMRaw 44.1 kHz samples for pipelines that handle their own audio.

Languages

EnglishSimplified ChineseTraditional ChineseJapaneseKoreanSpanishPortugueseFrenchGermanItalianDutchPolishRussianUkrainianTurkishArabicHebrewHindiBengaliIndonesianVietnameseThaiFilipinoSwedish+ 104 more

Latency

Fast enough to be part of the conversation

~190 mstime to first audio on a warm connection
< 400 mssteady-state latency for a full sentence
128languages from one shared voice identity
32built-in voices plus your own clones
Pipeline

Text in, audio out, no round trip to a human

A request carries the script, the voice, the format, and the tempo. The response starts streaming as soon as the first chunk is ready, so your player is never waiting on the whole paragraph.

POST /v1/audio/speech
{
  "voice": "aria",
  "input": "Grok TTS reads the room. [pause] Ask for warmth and you get warmth.",
  "format": "mp3",
  "speed": 1.0
}
What you control
  • First audio~190 ms
  • Steady state< 400 ms
  • Sample rate44.1 kHz
  • Delivery tags6 presets
  • Languages128
  • Voices32 + clones

How it works

Four steps from a script to a shipped voice

Write the line

Paste a sentence, a chapter, or a system prompt. Add delivery tags where the read should change.

Choose a voice

Pick one of 32 built-in voices or a clone you have trained. Voices stay consistent across languages.

Stream the audio

Request MP3, OPUS, WAV, or PCM and start playing on the first chunk instead of waiting for the file.

Ship it

Track usage per key, watch remaining credits, and top up from the pricing page when volume grows.

Use cases

Where a good voice pays for itself

Voiceovers

Narration for product videos and social cuts without a booking.

Audiobooks

Long-form chapters with a voice that stays identical for hours.

Voice agents

Low-latency replies inside a live conversation.

Accessibility

Read any article, ticket, or document aloud on request.

Dubbing

Re-record a script in another language with the same character.

Notifications

Turn alerts and digests into audio people actually finish.

Pricing

Start free, then pay for what you generate

Signing in with Google includes free credits. Monthly plans add generation capacity; annual plans cost less per month.

Starter

$9/month

For a first paid run and validating one workflow.

  • Credits150
  • BillingMonthly

Pro

Most popular
$19/month

For recurring narration and iterative script testing.

  • Credits500
  • BillingMonthly

Premium

$49/month

For continuous generation and higher throughput.

  • Credits1,500
  • BillingMonthly

Questions

Answers, before you ask

Grok TTS is a text-to-speech engine for products that speak: narration, assistants, notifications, and dubbing. It returns natural, controllable speech and streams audio so playback can start almost immediately.

Give your product a voice

Sign in with Google, generate your first line, and hear the difference delivery control makes.

View pricing
Grok TTS - Text to speech for products that talk back