Add a voiceover
Goal: add spoken narration to a video. ReelBolt turns written lines into speech with a speech service (Fish Audio) and mixes it into the video. There are two ways, depending on the kind of video:
- Edited footage: you write the lines and when each starts; the
VideoCompilestep mixes them in. - Promo from code: the
ScriptwriterAgentalready writes narration for each scene; aVoiceoverstep speaks it and theAuthorAgentputs each line into its scene.
ReelBolt speaks exactly the text it is given: yours, the scriptwriter's, or (with the footage
templates below) a narration writer's. No promo template includes a voiceover; you add a
Voiceover step to a promo workflow. Three footage templates do include one; see
Letting an AI write the narration for footage.
Prerequisites for a voiceover
- A speech provider: an admin adds one under Admin, Inference Providers with capability
SpeechSynthesis— kindFishAudiowith a Fish Audio API key, or kindOpenAICompatiblefor the bundled free local speech server — and marks it default (or the step names it withproviderId). - A voice id from Fish Audio (its "reference id" for the voice you want), used as
defaultVoiceId. - Licence note: Fish Audio's hosted service has no documented non-commercial restriction; running Fish Audio's open models yourself is covered by a separate licence that needs a paid agreement for commercial use.
Voiceover for edited footage
Start from any footage workflow, for example the Highlight edit from footage template
(video-derush-edit). Add a Voiceover step before the compile step and switch voiceover on in the
compile step. Example layout:
| # | Label | Step type | Agent |
|---|---|---|---|
| 1 | Analyze source video | VideoAnalyze | VideoTransform |
| 2 | Decide which spans to keep | Agent | VideoStoryEditor |
| 3 | Synthesize narration | Voiceover | VideoTransform (not an AI) |
| 4 | Compile edited video | VideoCompile | VideoTransform |
| 5 | Review edit quality | ReviewLoop back to step 2 | VideoReviewAgent |
Step 3 (voiceoverConfigJson):
{
"source": {
"kind": "Inline",
"lines": [
{"text": "This is how we built our first prototype.", "startSec": 2.0},
{"text": "Three months later, it shipped.", "startSec": 41.5, "voiceId": "OTHER-VOICE-ID"}
]
},
"defaultVoiceId": "YOUR-FISH-AUDIO-VOICE-ID"
}
startSecis measured on the original, uncut video's clock. ReelBolt moves each line to the right moment in the edited video. A line whose start was cut out of the edit is dropped.voiceIdper line is optional; otherwisedefaultVoiceIdis used.- Text in brackets like
[laughs]or(pause), links and control characters are removed before speaking; a line that becomes empty is skipped and never billed.
Step 4 — because steps were inserted, point at the editor by number:
{"version": 1, "decision": {"from": "Step", "stepOrder": 2}, "analysisStepOrder": 1,
"enableVoiceover": true, "voiceoverStepOrder": 3}
Note that voiceoverStepOrder is a plain step number, not a {"from": ...} reference. Voiceover
needs "mode": "Reencode" (the default). Narration is not turned down; if background music is also
on, the music does not rise while narration plays.
Voiceover for a promo made from code
In the promo templates the ScriptwriterAgent writes narration for every scene. To have it spoken:
- Add a
Voiceoverstep after theScriptwriterAgentstep and before theAuthorAgentstep. - Use the scriptwriter as its source:
{"source": {"kind": "ScriptScenes", "scriptwriterStepOrder": 8},
"defaultVoiceId": "YOUR-FISH-AUDIO-VOICE-ID"}
- Run the workflow. The
AuthorAgentpicks up the spoken lines by itself, puts each into its scene, and renders the promo with narration. Scenes without narration text get no audio.
In Promo video from your code (quick-win-promo) the scriptwriter is step 8, so the new step
goes in as step 9 and the later steps move down by one: Director becomes 10, Author 11 and the
review step 12. Update the review step's loopTargetStepOrder to 11 so it still sends work
back to the Author. In Promo video from your code (lighter) the scriptwriter is step 9. Always
set scriptwriterStepOrder to the scriptwriter's actual number.
The two narrated promo templates, Quick vertical reel (cinematic-feature-spotlight) and
Narrated launch trailer (story-arc-launch-trailer), already carry this step, right after the
scriptwriter and before the director, so you only need to set the voice.
What to check after adding a voiceover
- The
Voiceoverstep result lists each line (vo0,vo1, ...) with its status (ok,failedorskipped) and its spoken length.meta.allLinesFailed: truemeans nothing was spoken (usually a provider or voice id problem). - Edited footage: the
VideoCompileresult has a voiceover section reporting how many lines were applied, which were dropped and why (for example a line whose clip was cut out). Each line plays over the clip it was written for, even in an edit made from many clips. If some lines were dropped, the run shows a warning; if more than half were dropped, the review step sends the edit back for another try instead of passing it. - Promo: the
AuthorAgentreports which scenes received narration in its output'smetadata.audionote. A promo with noVoiceoverstep says "No audio: this workflow has no narration step, so the video is silent" there, rather than shipping silence without a word.
Pitfalls with voiceover
- No speech provider: the step cannot speak anything. Ask an admin to add a Fish Audio or
OpenAI-compatible provider with capability
SpeechSynthesis; other provider kinds cannot be used for speech. - Wrong step numbers after inserting the step: check
voiceoverStepOrder,scriptwriterStepOrder, the compile step'sdecision, and every review loop'sloopTargetStepOrder. In Promo video from your code (lighter), later agent steps also read a hand-picked list of earlier steps by number (contextSelectedPriorSteps); renumber those lists too. ScriptScenesplaced after the Author: the Author only uses narration that already exists when it runs; put theVoiceoverstep before it.- Lines overlapping: allowed; they play over each other. Space your
startSecvalues by each line's length. - A cloned voice that will not work: a
projectVoiceIdwhose consent was revoked or whose retention ended fails the step on purpose (VOICE_CONSENT_MISSING). Clone the voice again with fresh consent, or use a library voice. - No lip sync: pickups and dubs change the sound only. The mouth on screen is the original
take's, and the result says so (
lipsync.applied: false). - Not in this version: re-cutting footage to fit the narration (the review asks for shorter words instead), and lip sync.
Letting an AI write the narration for footage
The Footage edit with narration and captions template (video-derush-edit-voiceover) adds a
NarrationWriter agent that writes the lines for the edit, each tied to a shot, pause or spoken
phrase rather than to a time, and a Voiceover step with "kind": "NarrationPlan" that speaks
them, with karaoke captions on. The editor cuts first and the writer follows it. When a line is
longer than the moment it is tied to, the result reports the overflow and the review step asks the
writer for shorter words; the picture is never re-cut to the voice. See
Templates.
To have the story written first and the clips chosen and ordered to fit it, use the
Narrated product story template (narrated-story): the same NarrationWriter runs before the
editor, from your brief and what the analysis saw in each clip, and the editor keeps the clips each
line is tied to, in the script's order. See
Templates.
Captions from the narration
On the Compile step, open Narration & captions and switch on Burn captions (in JSON,
"enableCaptions": true). The words burnt into the picture are the words that were actually
spoken (after ReelBolt removed anything it does not read aloud), timed by the measured speech.
Pick a style:
- Subtitle (default) shows a line or two at a time on a see-through box. It works everywhere.
- Karaoke lights up each word in your spoken-word colour the moment it is said.
- Bold shows big, bold words two or three at a time, white with a heavy outline and no box — the look most Instagram Reels and TikToks use. On a vertical video it sits about two-thirds of the way down, above the buttons and caption the apps draw over the bottom of the screen. It works without a transcription provider too. The Fast social reel style picks it automatically.
- Punch shows one big, bold, outlined word at a time in the middle of the frame, the look most short-form videos use.
Karaoke and Punch need a transcription provider so ReelBolt can hear where each word falls;
without one they are skipped with the reason no_word_alignment, and Subtitle still works.
Pick a font. The editor shows a sample of every font so you can see it before you run. Leave it on Automatic to get the font made for your style (Inter for Subtitle, Montserrat for Karaoke, Anton for Punch), or choose Inter, Montserrat, Poppins, Bebas Neue, Anton or DejaVu Sans. Every font is built in and licensed for commercial use.
Size and position. Leave Automatic text size on: captions are then sized for the style and for phones, larger on vertical (9:16) videos. Turn it off to set a size yourself, as a percentage of the video's height. Position is the bottom of the frame (or the middle for Punch) unless you choose Bottom, Middle or Top. Captions always keep clear of the edges, and on vertical videos they stay out of the bottom fifth and the top of the frame, where TikTok, Reels and Shorts put their own buttons and text.
Colours. Choose the text colour, the box behind it (dark or light see-through, solid black, or no box with outlined text instead), and for Karaoke the spoken-word colour. Punch never uses a box.
Captions are drawn last, so they never fade with the picture. The compile step's result has a
captions section that says which font, position and size were used; if it shows
degradeReason: "libass_unavailable", the video engine was missing its caption renderer and plainer
captions were drawn instead. Captions saved before automatic sizing existed (size 4) now get the
automatic size too.
Narration at the end of the video
The last line of narration always finishes. If it would still be playing when the picture ends,
ReelBolt holds the last frame of the video a little longer (until a quarter of a second after the
line ends) instead of cutting the voice off. A fade-out at the end of the video
(programAudioFadeOutMs) fades the music, the original sound and the sound effects, never the
narration, so a sentence spoken during the fade is heard in full.
The compile step's result tells you when this happened: voiceover.extendedForNarrationSec is how
many seconds the last frame was held, and the video's length (outputDurationSec) includes them.
Cloning a voice (with consent)
The project's Voices tab lists the voices a Voiceover step can use by projectVoiceId:
- Library voices are voices the speech provider already offers; register one by its id.
- Cloned voices are built from recordings of a real person that are already in the project. Cloning needs a Fish Audio provider. Before ReelBolt sends anything, you must tick the consent statement word for word and name the person; the statement, who ticked it and when are kept with the voice. Choose how long the voice is kept (up to two years). Revoke consent stops every step from using the voice immediately; Delete removes the voice from the provider, the recordings from the project and the entry itself, and reports what it removed. Voices past their retention date are deleted automatically.
A step that names a cloned voice without valid consent fails; it never falls back to another voice.
If your Fish provider is a Fish speech server you run yourself (its address is not
api.fish.audio), cloning works a little differently, and ReelBolt notices this on its own:
- Use exactly one recording, and it must be a .wav file. A clear 10–30 seconds of the person speaking works best.
- The server needs to know what is said in the recording, so ReelBolt writes it down first using your default Transcription provider. Without one, cloning stops with a message asking you to set one up. If that provider is a hosted service rather than a local one, the recording is sent there too.
- Consent, the retention period, Revoke consent and Delete work exactly as above. Delete also removes the voice from your server.
Pickups in the presenter's own voice
The Fix flubbed lines in a talking-head video template (video-derush-edit-pickups) needs
talking-head footage with clear speech. It lets a
PickupPlanner agent name the sentences the presenter stumbled over and write the corrected
sentence for each. Set projectVoiceId in its Voiceover step to a consented cloned voice of the
presenter, and the compile step mutes the original phrase and plays the pickup in its place.
The picture does not change.
Dubbing into another language
The Dubbed video in another language template (video-derush-edit-dub) needs footage with
speech. It translates the narration plan
with a NarrationTranslator agent (name the language in the request box when you run it), speaks
the translation, replaces the original speech under every line ("voiceoverMode": "ReplaceDialogue") and burns subtitles of what was spoken. To merely turn the original speech
down under narration instead, use "voiceoverMode": "DuckDialogue".
A free, local speech service
An admin can register the bundled local speech server as an OpenAI-compatible provider with
capability SpeechSynthesis (endpoint http://whisper:8000/v1, model speaches-ai/Kokoro-82M-v1.0-ONNX, voice ids such
as af_heart). It costs nothing and sends no audio to a vendor, but cannot clone voices.