Why does my writing sound robotic when it is read aloud?
This is the fix that matters most, and it is not a setting. Written prose carries things that speech cannot: subordinate clauses, parenthetical asides, written abbreviations, and lists formatted with bullets.
Those all read badly aloud. An aside that is invisible on a page becomes a stumble in audio, because the narrator has to hold the main clause, deliver a subordinate thought, and return without the listener knowing where they are.
Short sentences solve most of it. One idea each, with a deliberate full stop where a speaker would breathe. Read the script aloud before generating it: anything you stumble over, the voice will stumble over too.
How do I stop numbers and symbols being read out literally?
A model reads what is written, not what is meant. That produces the classic giveaway: a percentage sign read as "percent symbol", or an ampersand read as "and" where you wanted "and".
Write numbers as you want them heard. "20 percent" rather than "20%". Spell out an abbreviation on first use and abbreviate afterwards, because an unexplained acronym read in full sounds like a stutter.
The same applies to dates, currency, and times. "March 3" is read differently from "3 March" depending on locale, and writing it the way you want it said removes the ambiguity entirely.
Which AI voice sounds least robotic?
The test is whether you can listen for ten minutes without noticing it is synthetic. Voices with strong regional character, exaggerated inflection, or a narrow pitch range all fatigue faster, because listeners perceive them as effortful even when they cannot say why.
Register matters more than accent. A mid-range voice sits comfortably for a whole paragraph; a very high or very low voice draws attention to itself. The catalogue covers American and British English in both registers, so the choice is about intent rather than availability.
Match the voice to the audience you already have. A voice that changes between videos reads as inconsistent, and viewers notice that long before they can name it.
What speed and pause settings sound most natural?
The default of 1.0x is a presenting pace. Natural conversation runs slower, and 0.9x to 0.95x usually removes the machine quality more than any other single change.
Paragraph pauses matter just as much. A short pause between paragraphs reads as considered; a long one reads as a section break. Raising the pause is the cheapest way to make a flat passage sound structured.
The opposite problem is over-pausing, which reads as hesitation. If the voice sounds like it is searching for words, the pause is too long rather than too short.
What will not fix a robotic-sounding voice?
No setting compensates for a script written to be read on screen. Normalising text helps with the symbols, but it cannot restructure a sentence built for the eye.
Post-processing in an audio editor is a last resort. Compression, de-essing, and EQ all make the audio cleaner, but they do not make the prosody less synthetic, because the problem is in how the sentence was planned.
Generating the same text repeatedly with different voices rarely helps either. If a passage sounds wrong, the passage is wrong. Change the script and the voice becomes a secondary question.