Spoken replies

Last updated 2026-08-24

Every answer in the chat has a speaker button. This page says which languages have a voice, how good it is, and what it still cannot do — so that none of it has to be discovered after paying.

What it does

Press the speaker under an answer and it is read aloud. Never automatically — sound starts only when you ask for it, which is also the only behaviour a phone browser permits.

While it plays you get pause, stop, and speed from 0.75× to 2×, with the position and the length shown. Pressing the speaker under a different answer moves the voice to that one.

Playback starts while the rest is still being made, so the wait does not grow with the length of the answer. A reply of thirteen hundred characters — nearly two minutes of speech — begins speaking in about two seconds, the same as a one-line reply.

Which language it speaks

The voice follows the language of the answer, not the language of the menus. An English interface holding a Russian conversation is read in Russian. Most products get this backwards and read the answer through the wrong language's rules.

Numbers are spoken as words rather than spelled out digit by digit, and a block of code is announced as a block of code instead of being read out as a minute of punctuation.

Languages with a voice

Thirty-eight of our forty-six languages have a measured voice. Ten of them are also built for expressive delivery — that part is written and measured, and it is not switched on yet, so every language currently reads clearly rather than with feeling.

🔑 Which of them are switched on at any moment can be fewer, while the machines behind them are brought up. The speaker button appears only where it will actually work — you will not be offered a voice that then fails.

Read aloud, and ready for expression
German · English · Spanish · French · Italian · Korean · Portuguese (Brazil) · Portuguese (Portugal) · Russian · Chinese (Simplified). These ten are the ones expressive delivery will reach first. Today they read the same as the rest.
Reads aloud
Arabic · Azerbaijani · Bulgarian · Catalan · Croatian · Czech · Danish · Dutch · Filipino · Finnish · Greek · Hebrew · Hindi · Hungarian · Indonesian · Japanese · Malay · Norwegian · Polish · Romanian · Serbian · Slovak · Swahili · Swedish · Thai · Turkish · Ukrainian · Vietnamese
No voice
Bengali · Burmese · Kazakh · Pashto · Persian · Urdu · Uzbek · Chinese (Traditional)

Persian, and why it is empty

Persian is the one language on that last line with real usage behind it, and its absence is not an oversight. Every engine we could license was measured against it this month and none reached the quality we require — one reads only fifteen languages cleanly, one does not fit the hardware, one ships broken.

We would rather say so than ship a voice that mangles the language. When something passes, it will appear here.

How the quality was measured

Every language was checked the same way: speak a text, transcribe the audio back with a separate system, and compare it word for word with what was sent. The gate is a five per cent character error rate. All thirty-eight languages are under it, and six of them come back at zero.

🔑 What that measures is whether the words come out right. It does not measure whether the voice sounds natural, which no automatic test can settle. Treat it as a floor rather than a verdict, and judge the rest with your own ears.

What it cannot do yet

No word-by-word highlighting
The engine returns no timings, and estimating them from character counts drifts within a single sentence. We built it, measured it, and switched it off rather than ship something that wanders across the text.
Expression is not switched on
The machinery exists and was measured — the voice can sigh, laugh or drop to a whisper when the reply calls for it. It is not connected to the chat yet, so nothing you hear today carries it. Ten languages are ready for it; the rest need an engine that does emotion and wide language coverage at once, which does not exist under a licence we can use.
Not every language at every moment
See the note above: coverage is what has been measured, and what is switched on can be smaller while machines are provisioned.

How much you get

The free chat reads a small number of replies aloud each day, and paid plans read far more. The exact numbers are on the plans page — kept in one place, so the two can never disagree about what you get.

The paid ceiling is high rather than absent on purpose. An endpoint with no ceiling at all has no floor either if an account is ever shared or sold, and the machines that make the audio are rented by the hour. The number is set far past what one person listens to in a day.

A reply already being spoken is never cut off in the middle when the limit is reached, and the count moves once per playback rather than once per sentence.

What happens to the audio

Nothing is stored. The audio is made for the request and never written to disk on our side. Pressing stop cancels the work on the server too, rather than leaving it running for sound nobody will hear.

This is the same rule the rest of the service follows, and it is the reason there is no library of your spoken replies to browse: there is nothing to browse.