Generate your voiceover

With a voice chosen, this is where your script comes alive. Generate the whole audio in one go, listen back, then fine-tune only the lines that need it. It's also how you fill the gaps in a human voiceover using a cloned voice. Not only gaps either: a line you'd rather say differently can be replaced the same way, without sending anyone back to a microphone.

If you haven't picked a voice yet, or you'd like to try a different one, start with Find the right voice for your script.

Generate the whole script

Open the AI Voice section in the side panel and click Generate All. That's the best first move, because it gives you audio for every line at once and lets you hear the whole thing before deciding what needs work.

A loading indicator appears next to each line while its audio is being generated.

When it finishes, press play in the timeline along the bottom to hear the full script, or use the play button next to a line to hear just that one.

Set how the voice performs

Just under the voice you've chosen, three presets control how your script gets delivered:

  • Neutral reads it straight, holding an even tone throughout.
  • Dynamic puts more energy and movement into the delivery.
  • Varied gives the widest range, shifting tone and pace more across a line.

Each preset is really a saved combination of the advanced settings further down. Picking one gives you a sensible starting point, and from there you can open those settings and change whatever you like.

Tempo controls how quickly the voice speaks as the audio is generated. That's different from the playback speed control, which only changes the speed of audio that already exists. Lower values give slower, more deliberate delivery. Higher values sound faster and more energetic.

Fine-tune with advanced settings

Getting a voice exactly right takes a little practice, and a first pass that isn't quite there is completely normal. If the presets aren't getting you close, open Advanced settings and work from there.

Model is the ElevenLabs model used to synthesize the speech, and different models trade off quality, speed and language support. Solid picks the model your chosen voice was built in, so it's worth leaving as it is unless you want to experiment.

Stability controls how consistent the voice stays between generations. Lower values open up a broader emotional range. Higher values are steadier, but can flatten into something monotone.

Similarity controls how closely the AI sticks to the original voice it's replicating.

Style exaggeration amplifies the original speaker's style. Anything above zero uses more processing, so generation takes a little longer.

Speaker Boost pushes similarity to the original speaker further. It also adds slightly to generation time.

Regenerate a single line

Select the line you want to redo and click Generate in the side panel. That regenerates only that line and leaves the rest of your script alone. The Update button on the line itself does exactly the same thing, so use whichever is closer to hand.

If you're still not quite happy with how it sounds, it may be the voice rather than the settings. Some voices simply come out more natural than others, so going back and trying a different one often gets you there faster than endlessly tweaking the one you started with.

Worth knowing: changes to the presets, tempo or advanced settings only apply to audio you generate afterwards. Anything already generated stays exactly as it is until you regenerate it.

Go back to an earlier take

Every line keeps its own history. Click the history icon next to a line to hear previous generations and bring one back, which is handy when you decide you preferred an earlier take after all.

You can also open History at the bottom of the side panel to work through older versions in one place.

Fix one sentence without redoing the rest

Remember from Write and edit hooks, body and CTA variants that pressing Enter inside a script line splits it in two. Once audio has been generated, that split applies to the audio as well, which gives you a quick way to keep the part of a take you liked and regenerate only the part you didn't.

Say a line reads "Hello. My name is Frederic. And today the weather is great. I am twenty one years old." The opening sounds perfect, but the last sentence comes out wrong. Rather than regenerating the lot, click just before I am twenty one years old, press Enter to move it onto its own line, then generate that line on its own. The part you liked stays untouched.

Keep your splits at sentence boundaries though. If you split out a single word and regenerate it, that word is generated as its own sentence, so its pronunciation and flow won't match the words either side of it and the join tends to sound unnatural.

With your voiceover done, next up is navigating the timeline to build the video around it.