How it works
Neural vocoders are what make modern TTS sound human instead of robotic.
What is TTS?
Text-to-Speech, or TTS, is the technology that reads translated text aloud. It takes the output of the translation engine and generates audio that sounds like a human speaker.
TTS is what lets you play a translated phrase to a taxi driver, a doctor, or a shopkeeper without handing them your phone.
Neural vs. classic TTS
Older TTS sounds robotic because it stitches together pre-recorded snippets. Neural TTS uses deep learning to produce smoother intonation, pauses, and emphasis. The best neural voices are hard to distinguish from real speakers.
When TTS matters most
TTS is critical when you cannot read the script of the target language, when the other person cannot read your screen, or when you want to practice pronunciation while learning a language.
Frequently asked questions
Can I slow down TTS?
Most translation apps let you adjust playback speed, which is helpful for learning or for listeners who need extra clarity.
Does TTS work offline?
Many apps include offline TTS voices, but quality is usually lower than cloud voices. Download the voice pack in advance if you will be off the grid.