Text to Speech Tools

Text-to-speech tooling: the TTS pipeline and the four factors driving naturalness — text normalisation, pronunciation lexicons, prosody control and post-processing — plus a comparison of concatenative versus neural TTS on stability, cloning and controllability.

TTS turns text into spoken audio, and neural models have improved its quality substantially.

Approach Comparison

AspectTraditional concatenative TTSNeural TTS
Synthesis methodPre-recorded phonemes concatenatedEnd-to-end acoustic model
NaturalnessAudible concatenation artifactsClose to human speech
ProsodyRule-driven and stiffLearned jointly from text and audio
Voice cloningImpossiblePossible, given samples
Long-text stabilityGood, since it concatenatesCan suffer prosody breaks
Synthesis speedVery fastSlower, depending on model size
ControllabilityDirectly adjustable via rulesRequires prompt engineering
Typical useClock reading, short numbersAudiobooks, voice assistants

Edge Cases

  • If text normalisation is not handled, even an excellent model mispronounces numbers, abbreviations and symbols.
  • Temperature zero gives reproducible output; above zero it varies per call, so batch generation needs a fixed seed.
  • Neural TTS can break prosody across very long passages, requiring sentence-level splitting and concatenation.
  • A pronunciation lexicon is what makes proper nouns accurate; without one the model merely guesses.

Common Pitfalls

  • Feeding text containing HTML tags or Markdown markers straight to TTS, which then reads the markup aloud.
  • Not fixing the random seed during batch generation, so per-batch differences break manual verification.
  • Synthesising very long text in one pass, producing prosody breaks and volume drops at the end.
  • Ignoring the pronunciation lexicon and mispronouncing proper nouns such as people and brand names.

Related Tools