All guides

Local AI line · stop 13 of 14 · 24 min · members

Local speech synthesis, including the languages the default models cannot handle

Running speech synthesis on your own machine, including the languages the default models cannot handle.

Free with an account

Sign in to read.

Membership is free: an account opens all 86 script pages. The Lab, Studio Canvas and the paid guides need the $99 pass, paid once. Already signed in on this browser? The page opens by itself.

01

Why local

Voiceover is iterative, and per-request pricing fights that.

Running it yourself changes how many takes you are willing to try.

A narration script goes through many revisions — a rewritten line, a different pace, a changed emphasis. When each generation costs money and a network round trip, you stop iterating sooner than you should.

Running locally removes both frictions. Generate fifty variations of a line to find the reading that works, at no marginal cost, without anything leaving the machine.

That last point matters for client work. A script under embargo, or one containing details a client would rather not send to a third party, stays on your disk.

02

The language trap

The default voice usually does not speak your language.

And it will attempt it anyway, producing something confidently wrong.

Most local speech setups ship with a default engine trained primarily on English. Given text in another language it does not refuse — it applies English pronunciation rules, and the result is fluent-sounding nonsense.

The symptom is characteristic: correct rhythm, wrong sounds. Accented characters get mangled, and language-specific vowels come out as their nearest English approximation.

The fix is to select an engine or voice trained for the target language rather than adjusting settings on the default one. This is a model choice, not a configuration, and no amount of parameter tuning substitutes.

03

Getting a good read

Punctuation is your direction.

These systems take timing cues from the text, so write for them.

Commas produce short pauses, full stops longer ones, and paragraph breaks longer still. Writing a script with correct punctuation is most of what you can do to control pace.

Where a pause is needed that punctuation does not naturally give, break the line and generate the halves separately, then place them with the gap you want. This is more reliable than any pause markup and it gives you exact control.

Generate line by line rather than as a block. Individual lines can be re-rolled without regenerating the whole narration, and assembling them in an editor lets you set the timing against picture.

04

Voice cloning

Technically straightforward, and not always yours to do.

The consent question is the important part, not the method.

Cloning a voice from a sample is now easy and works from a surprisingly small amount of clean audio. The technical requirements are modest: clear recording, no background noise, a few minutes of varied speech.

The question that matters is whether you have permission. Cloning your own voice is unproblematic. Cloning a client's presenter requires their explicit agreement, in writing, for the specific use.

Cloning anyone else's voice without consent is not something to do, regardless of how easy it has become. Treat a voice as belonging to its owner and get agreement first.

05

Quality

Where synthetic voices still give themselves away.

Three patterns, and two of them are fixable in the edit.

Uniform pacing. A human speeds up and slows down; synthesis tends towards even delivery. Fixable by generating in short segments and adjusting the gaps yourself.

Wrong emphasis. The system does not know which word carries the meaning. Fixable by rephrasing so the important word falls where the reading naturally stresses it.

Flat emotion on long passages. Harder to fix. For anything requiring genuine performance, a person is still the answer.

Synthetic voice works well for informational narration, explainers and temporary tracks. For a piece that depends on the delivery, it does not.

06

Practical use

Scratch track first, decide about the real one later.

The highest-value use is the one before the final voice exists.

Cutting picture against a synthesised scratch track lets you build the whole edit before booking anyone. Timing, pacing and length are all resolved, and the script is finalised against picture rather than guessed at.

Then either the scratch is good enough to keep, or you record a person against a locked edit, which is a much shorter session than recording against an unfinished one.

Either way the edit was not waiting on the voice, which is usually the thing holding up the whole piece.