
Five Tags Fix Most Voice Problems
The short answer is markup, not a different voice. If a generated voice pauses in the wrong place, misreads a number, or mangles a name, SSML tags fix it by wrapping your script in instructions the engine follows literally.
Five tags carry most of the work. Break sets a pause of an exact length such as 250ms, and say-as tells the engine whether a string of digits is a date, a price or a phone number.
The other three handle voice and pronunciation. Sub replaces a spelling, phoneme specifies exact sounds, and prosody adjusts rate, pitch and volume.
The catch is support. Amazon Polly documents emphasis as unavailable on its neural, long-form and generative voices, and prosody as only partially available, so newer voice models honour fewer tags than older ones.
That trade sits at the centre of this guide. The most natural-sounding voices give you the least manual control, and knowing which tags survive on your voice saves an afternoon of confused re-rendering.
Why A Plain Script Gets Read Wrong
A text-to-speech engine guesses at everything you did not specify. It guesses whether 1997 is a year or a quantity, whether Dr is doctor or drive, and how long to rest at a line break.
Those guesses are usually right in ordinary prose. They go wrong precisely where your script is specialised, which means product names, figures, abbreviations and deliberate dramatic pauses.
Punctuation is a weak instrument for this. A comma nudges the pacing, but it cannot request a pause of a specific length, and stacked ellipses produce inconsistent results across voices.
SSML replaces the guess with an instruction. The engine still generates the audio, but the timing and the interpretation of odd strings are yours to set.
This is why editing a script for a voice tool feels different from editing for a reader. You are not writing punctuation for a human, you are writing directions for a synthesiser.
Pauses Are The Tag You Will Use Most
- ● Break takes time or strength
- ● Time values are predictable
- ● Two lengths cover most scripts
The break tag controls pausing between words. Google documents a time attribute taking values such as 250ms or 3s, and a strength attribute with the values x-weak, weak, medium, strong, x-strong and none.
Time values are the practical choice. Strength values vary in how each voice interprets them, while a 400ms request produces a comparable gap across renders.
Two lengths cover most scripts. A short pause after a clause and a longer one between sections give a script structure without sounding staged.
The none value is worth knowing about. It suppresses a pause the engine would otherwise insert, which is how you stop a voice from breaking mid-phrase after an abbreviation.
Placement matters more than duration. A pause before an important number lands as emphasis, while the same pause after it sounds like hesitation, and our guide to AI voice pacing and pauses covers that rhythm in more depth.
Telling The Voice What A Number Actually Is
The say-as tag exists because digits are ambiguous. Google documents interpret-as values including currency, telephone, date, time, cardinal, ordinal, fraction, unit, characters, spell-out and verbatim.
The differences are audible. As a cardinal, 1997 is one thousand nine hundred ninety-seven, and as a date it becomes nineteen ninety-seven, which is what a script about a year needs.
Phone numbers and reference codes need spell-out or characters. Without it, an order number gets read as an enormous single figure that no listener can transcribe.
Currency handles the symbol properly. Writing the amount inside a currency say-as produces dollars and cents rather than a symbol read out of order or skipped entirely.
There is even an expletive value that bleeps the enclosed word. It is niche, but it beats editing a beep into the audio afterwards.
What Each Tag Controls And Where It Breaks
- ● Say-as fixes numbers and dates
- ● Sub fixes names permanently
- ● Prosody wants full sentences
| Tag | What it controls | Typical use | Where it breaks |
|---|---|---|---|
| break | Pause length between words | Section beats, dramatic timing | Overuse makes speech sound stilted |
| say-as | How a string is interpreted | Dates, currency, phone numbers | Wrong interpret-as sounds worse than none |
| sub | Substitutes an alias for pronunciation | Brand names, acronyms | Alias must be respelled per language |
| phoneme | Exact sounds in IPA or X-SAMPA | Personal and place names | Requires phonetic notation to write |
| prosody | Rate, pitch and volume | Slowing a dense passage | Google advises using it around a full sentence |
| emphasis | Stress on a word or phrase | Highlighting one term | Polly lists it as unavailable on neural voices |
| audio | Inserts a recorded clip | Stings, jingles, sound beds | Google limits clips to 240 seconds and 5MB |
Read the last column before writing any of these into a long script. Most SSML frustration comes from a tag that is technically valid and silently unsupported on the chosen voice.
The audio row is the one people miss. Being able to drop a recorded clip into synthesised speech removes an entire editing step for intros and stings.
The Tags Your Voice Model Quietly Ignores
Support is per voice family, not per platform. Polly publishes a table showing which tags work on neural, long-form and generative voices, and the gaps are significant.
Emphasis is the clearest example. Polly lists it as not available across neural, long-form and generative voices, so a script leaning on emphasis will simply not sound emphatic.
Prosody is listed as only partially available on those newer voices. Rate changes may apply while pitch changes do not, which produces results that look like a bug in your script.
Failures are not always graceful. Polly states that using unsupported tags in standard, neural or long-form formats produces an error, so a long script can fail outright rather than degrade.
The practical rule is to test the tag set on your exact voice first. A thirty-second sample with every tag you intend to use costs almost nothing and prevents a full re-render.
Pronunciation Fixes That Outlast One Script
The sub tag is the easiest win. It replaces the written text with an alias for pronunciation, so a brand spelled one way can be read another without changing the visible script.
Phoneme is the heavier tool. It takes IPA or X-SAMPA notation and specifies the exact sounds, which is what stubborn surnames and place names usually need.
Neither is worth applying inline forever. Most platforms support a pronunciation lexicon or dictionary, which applies the same fix across every script automatically.
Recurring names belong in that dictionary. If a name appears in every episode of a series, fixing it once at the account level beats tagging it in each script.
Multilingual work multiplies the problem. Our look at using one AI voice across multiple languages explains where a single voice stops sounding native, which is usually the point aliases stop helping.
Which Tags Fit Your Project
- ● Match tags to your voice model
- ● Newer voices honour fewer tags
- ● Test before a long render
You narrate long-form videos: rely on break and say-as. Consistent pacing and correctly read figures matter more over twenty minutes than any single stressed word.
You produce short ads or promos: you want emphasis and prosody, so check support before choosing a voice. A neural voice that ignores emphasis is the wrong tool for that job.
You publish anything with phone numbers or codes: say-as with spell-out is non-negotiable. A listener who cannot transcribe the number gets nothing from the segment.
You run a series with recurring names: move pronunciation fixes into a lexicon. Inline sub and phoneme tags in every script become a maintenance problem within a month.
You localise into several languages: expect to redo aliases per language. A sub alias is spelled for one pronunciation system and rarely survives translation intact.
You are choosing a voice right now: let tag support influence the choice. Our guide to picking an AI voice for a channel covers the other factors that matter over a long run.
Testing SSML Without Burning Your Character Budget
Build one short test script containing every tag you plan to use. Thirty seconds of audio reveals which tags your voice honours and which it ignores or rejects.
Change one thing at a time after that. Adjusting a break length and a prosody rate together makes it impossible to tell which change produced the improvement.
Keep a working snippet library. Once you know your standard pause lengths and your alias spellings, reusing them is faster than rewriting the markup for each script.
Watch the billing model while testing. Most engines charge by characters synthesised, and markup can count toward that, so short test renders matter more than they first appear.
If costs are the deciding factor rather than control, our breakdown of how AI voice pricing works explains what those character counts translate into monthly.
FAQ
What is SSML in a text-to-speech tool?
SSML is a markup language that wraps your script in tags a text-to-speech engine reads as instructions. Instead of hoping a voice pauses correctly, you write a break tag and the engine inserts the pause you asked for.
How do you add a pause to an AI voice?
Use the break tag, which accepts a time value such as 250ms or 3s in Google Cloud Text-to-Speech, or a strength value from x-weak through x-strong. Time values are predictable, so most people settle on a few standard lengths.
Do all AI voices support every SSML tag?
Not all of them. Amazon Polly lists the emphasis tag as unavailable on neural, long-form and generative voices, and marks prosody support as partial, so the newer the voice model, the fewer tags it honours.
How do you make an AI voice read numbers correctly?
Use say-as with an interpret-as value. Google documents values including currency, telephone, date, cardinal, ordinal, fraction, characters and spell-out, which covers most numbers that get read the wrong way.
What is the difference between the sub and phoneme tags?
The sub tag substitutes a plain spelling for pronunciation, and the phoneme tag specifies exact sounds using IPA or X-SAMPA. Sub is easier to maintain, so reach for it first and keep phoneme for names it cannot fix.
What happens if you use a tag your voice does not support?
Polly states that unsupported tags produce an error in standard, neural and long-form formats, so a script can fail rather than degrade. Test a short sample with your exact voice before running a long script through it.
Sources
- SSML reference (Google Cloud Text-to-Speech) — checked 2026-09-08
- Supported SSML tags (Amazon Polly) — checked 2026-09-08
Some links may be affiliate links. We may earn a commission at no extra cost to you.
This article was written with AI assistance. It is researched and fact-checked, not based on personal hands-on testing unless explicitly stated.
Comments
Post a Comment