Voiceover
Text-to-speech narration for ReelBolt workflows — a deterministic, non-LLM StepType.Voiceover
step that synthesizes one WAV per configured script line through a Fish Audio provider, and a
VideoCompile mix path that lays those WAVs onto the compiled video's audio track next to the
dialogue, background music and sound effects.
Phase 2 adds narration planning on top of that spine: a NarrationWriter agent that writes the
words and anchors each line to an offered cut-anchor id, deterministic word alignment over each
synthesized line, and a deterministic fit solver whose overflow facts reach the review loop. The
spine is unchanged — an agent chooses opaque ids and prose, and code measures and mixes; nothing on
this path judges audio with a model.
This document is the reference for that feature as built in phase 1, plus phase 2's narration
planning. For the surrounding workflow engine (step types, executors, inference-provider resolution
in general) see CLAUDE.md; for the video-editing feature the compile path extends, see
docs/video-editing.md. What phase 1 left out, and which of it phase 2 now delivers, is enumerated
at the end, under
Phase 1 scope, and what phase 2 delivers.
Table of Contents
- The shape
- The Fish Audio provider
- Narration planning (
NarrationWriter) - Config reference
- Word alignment
- The fit solver and the review loop
- Step output and error codes
- Artifacts, caching and where the audio lives
- The mix path (
VideoCompile) - Narration and the end of the program
- The promo pipeline path (
ScriptScenes) - Text sanitization
- Per-line artifact resolution
- Admin surfaces
- Phase 1 scope, and what phase 2 delivers
- Phases 3 and 4: what landed after narration planning
- Story-first narration and the narrated templates
- Local fish-speech servers and endpoint rules
The shape
Two things were added, and nothing else: a step that produces narration audio, and a flag on the existing compile step that consumes it.
┌────────────────────┐ ┌──────────────────┐ ┌────────────────────┐
│ StepType. │ │ object storage │ │ StepType. │
│ Voiceover │─────▶│ voiceover/ │─────▶│ VideoCompile │
│ │ │ lines/ │ │ │
│ deterministic, │ │ {sha256}.wav │ │ EnableVoiceover │
│ Fish Audio TTS │ └──────────────────┘ │ VoiceoverStepOrder│
└────────────────────┘ └────────────────────┘
emits {view, meta} resolves each line
+ a JSON artifact to an OUTPUT-timeline
moment, then amixes
There is no LLM anywhere on this path. VoiceoverStepExecutor sends text to a provider and probes
what comes back; VideoCompileStepExecutor reads the step's own output JSON and builds ffmpeg
arguments from it. No audio content is ever judged by a model, the same discipline
docs/video-editing.md applies to VideoAnalyze.
The step type is StepType.Voiceover (ReelBolt.Shared/Data/Models/Enums.cs). Like
VideoAnalyze/VideoCompile, it is a deterministic step whose WorkflowStep.AgentDefinitionId FK
is satisfied by the existing built-in AgentType.VideoTransform placeholder row — no new
AgentType was introduced for it.
Phase 2 adds exactly one thing upstream of that diagram and changes nothing inside it: a
NarrationWriter agent step that decides what the narration should say, plus the
VoiceoverSourceKind.NarrationPlan source kind that feeds that plan into the same step. The step
itself stays deterministic — the model writes words and names offered ids, and everything that
measures, mixes or positions audio is still code. A NarrationWriter step is an ordinary Agent
step; AgentType is appended, never reordered.
That split is the point of the design, so state it plainly: a NarrationPlan line carries no
start time and cannot carry one, and a plan cannot know where the picture will land. The step
therefore declares the line positions itself — narration-first, back-to-back in plan order — and
the compile step measures the picture against the result afterwards. Those are two different
measurements made by two different components, neither reading the other; see
The fit solver and the review loop.
The Fish Audio provider
InferenceProviderKind.FishAudio is a provider kind that serves exactly one capability,
InferenceProviderCapability.SpeechSynthesis. Both halves of that pairing are enforced at the API
boundary (InferenceProvidersController.IsUnsupportedCombination): a FishAudio row
with any other capability is rejected with a 400, and any other kind with SpeechSynthesis is
rejected with a 400. SpeechSynthesisClientFactory throws the same NotSupportedException as a
second line of defence, the way TranscriptionClientFactory does for Anthropic.
Wire format
FishAudioSpeechSynthesisClient (ReelBolt.Shared/Inference/FishAudioSpeechSynthesisClient.cs)
calls POST https://api.fish.audio/v1/tts with Authorization: Bearer <key> and a JSON body.
Every fact below was checked against Fish Audio's primary OpenAPI schema rather than against the
feature request that proposed them, and several of the request's original assumptions turned out to
be wrong — the corrections matter, because code written against the wrong shape would fail at the
boundary and look like a credential problem.
- Model selection is a request HEADER (name:
model), not a body field. The client sets it withhttpRequest.Headers.Add("model", model). - The model enum is
s1,s2-pro,s2.1-pro(default),s2.1-pro-free,drama-3-preview. There is noopenaudio-s1value; the request's guess was wrong. reference_idselects a stored voice. Phase 1 usesreference_idonly.format,sample_rate,latencyare top-level body fields. ReelBolt sendsformat: "wav",sample_rate: 48000,latency: "normal".speed,volumeandnormalize_loudnessare nested under aprosody: { … }object, NOT top-level request fields — another correction. ReelBolt sendsprosody: { speed: 1, volume: 0, normalize_loudness: true }. (A top-levelnormalizefield does exist, but means something different: text normalization for numbers, English/Chinese only — not audio loudness.)- Expressiveness tag syntax depends on the model and is not uniformly open-ended.
- The S2 family (
s2-pro,s2.1-pro,s2.1-pro-free) anddrama-3-previewuse[bracket]syntax over open-ended natural language (e.g.[whispers sweetly]). s1instead uses(parenthesis)syntax over a fixed, documented vocabulary of 64 expressions (24 basic + 25 advanced + 5 tone + 10 audio-effect) — not open-ended.- Phase 1 ships no free-text emotion/delivery control, so this matters only to the text
sanitizer, which must therefore strip both
[...]and(...)forms. It does — see Text sanitization.
- The S2 family (
- Pricing is exactly $15.00 per M UTF-8 bytes for
s2.1-pro/s2-pro/s1, and $0.00 fors2.1-pro-free. This is why the step's byte budget is measured over sanitized text (VoiceoverStepConfig.MaxBytes), not over an estimated duration. - There is no official .NET/C# SDK — Fish Audio ships Python and JavaScript only — so the plain
HttpClientinFishAudioSpeechSynthesisClientis the intended integration, not a shortcut. Because phase 1 usesreference_idalone (no zero-shotreferencesarray), a JSON body is sufficient and no MessagePack NuGet package is needed at all, keeping the repository's "add no new NuGet dependency if avoidable" constraint satisfied. MessagePack remains an option Fish Audio advertises; it is simply not required for what phase 1 sends.
Licence: two separate things
The licensing position is more favourable than "the free tier is non-commercial", and the two things that phrasing conflates are actually separate:
Fish Audio's hosted API, including the free
s2.1-pro-freetier, has no documented non-commercial restriction — it's rate/SLA-limited only. Self-hosting Fish's open-weight models is licensed under the Fish Audio Research License Agreement: free for research/personal use, commercial self-hosting requires a separate paid license from Fish Audio directly.
Concretely:
- The hosted API — including the $0
s2.1-pro-freetier — carries no documented non-commercial restriction. Fish's own documentation recommendss2.1-pro-freefor "testing, prototyping, development, and smaller businesses"; the only thing it lacks is an SLA (latency/uptime) guarantee. There is no "free plan = non-commercial" clause for API access. The framing "the free plan is non-commercial" is verified incorrect and must not be used. - What is non-commercial-only is self-hosting Fish's open-weight models (
fishaudio/fish-speech). Its current licence — checked directly against the repository'sLICENSE, dated March 2026 — is the Fish Audio Research License Agreement, notCC-BY-NC-SA 4.0(that was an older, 2024-era version per the repository's own changelog). The current agreement is free for Research/Non-Commercial use; any Commercial Purpose, including internal business use, requires a separate paid licence from Fish Audio ([email protected]).
ReelBolt talks to the hosted API, so the second bullet only becomes load-bearing if someone points
an Endpoint at a self-hosted fish-speech instance. The admin form
(web/components/admin/InferenceProviderForm.tsx) renders the notice above whenever Kind is set
to FishAudio, so the distinction travels with the provider row rather than living only here.
Resolution
Speech synthesis resolves through IInferenceProviderResolver.ResolveSpeechSynthesisAsync, and —
unlike transcription and vision — the choice is health-aware (ProviderHealthSelector, shared
library):
VoiceoverStepConfig.ProviderId (always honoured, healthy or not) → the single enabled row with
IsDefault = true AND Capability = SpeechSynthesis, unless its last connection test failed
(LastTestOk == false; a never-tested default is trusted) → the enabled SpeechSynthesis row whose last
test passed, most recently tested first → the default anyway (it may have recovered) → none.
This exists because a deployment's default (local-fish-tts) had failed its test while a working row
(qa-local-fish) sat unused, and every narrated video was synthesized against the broken one with no
warning. The resolved provider carries SelectionReason, LastTestOk and PassedOverDefaultName, and
the Voiceover step turns them into top-level output warnings (the C1 shape, present only when
non-empty): voice_provider_fallback when a healthier row was used instead of the default, and
voice_provider_unhealthy when the row that narrated had failed its last test. The resolver caches rows
for Inference:ProviderCacheSeconds (60 s), so a re-test takes effect within a minute. Health-aware
choice is deliberately not applied to chat, video generation (another row is another vendor's bill)
or embeddings (another model is another vector space).
GET /api/v1/inference-providers/status?capability=SpeechSynthesis reports the same choice to any
signed-in user — { capability, providerName, lastTestOk, lastTestAt, healthy, usingFallback, message },
never an endpoint, model, key or raw error — so the step settings can warn before an hour-long run
(web helper getProviderStatus). Other capabilities report their default row.
There is deliberately no legacy-config fallback, for the same reason transcription and vision have
none: silently sending text to a chat deployment would produce a confusing 404 rather than a clear
"no speech provider configured". The explicit-id branch also re-checks
Capability == SpeechSynthesis before honouring the id — a chat provider id authored directly into
a VoiceoverConfigJson blob (bypassing the UI, which only offers speech rows) falls through to the
default instead of being handed to the speech client factory.
The resolved provider is a ResolvedSpeechSynthesisProvider — its own type, not the shared
ResolvedInferenceProvider, since a chat, a transcription and a speech-synthesis resolution are
never interchangeable. It carries a CacheKey hashed over Kind|Endpoint|ModelName|SHA256(ApiKey)| TimeoutSeconds so the same configuration reuses one client, and it overrides ToString() so a
stray LogError("… {Provider}", provider) cannot print the API key.
Narration planning (NarrationWriter)
Phase 1 narrated text a human wrote. Phase 2 adds the agent that writes it: AgentType.NarrationWriter,
an LLM agent whose output is a plan, not audio — NarrationPlanOutput, the seventh schema in the
house's id-anchored, number-free family alongside VideoEditDecisionOutput,
MotionGraphicsPlanOutput, MusicPlanOutput, SfxPlanOutput, ColorGradePlanOutput and
GeneratedShotPlanOutput. Those six plus NarrationPlanOutput are the invariant-test files that
guard such a schema; RoomConveneGateInvariantTests also carries the name but is not one of them —
it guards the decision gate, not an agent output.
It is appended last in AgentType and adds no step type of its own: a narration plan is an
ordinary Agent step. Enums in this repository are appended, never reordered.
Numbers in the spoken text. The schema carries no numeric PROPERTY, but the spoken text may
contain digits, and the prompt (both copies) asks for years and large numbers as digits ("1989",
"14 mila"): Fish reads digits with the right stress, while a long number spelled as one word
("millenovecentottantanove", "quattordicimila") was stressed wrong by every Italian voice tried. A
brief that says how to write numbers wins. The prompt used to ask for every quantity as words.
The schema cannot carry a number
public sealed record NarrationPlanOutput(List<NarrationLine> Lines, string PlanRationale);
public sealed record NarrationLine(string AnchorId, string Text, string Delivery, string Reason);
Every property on both types is a plain string or a List<NarrationLine>. AnchorId names an
offered cut-anchor id — an s{n} shot, g{n} gap or t{n} transcript span that the
VideoAnalyze step's bounded view actually offered, the same vocabulary SfxCue.AnchorId and the
graphics planners draw from. Delivery is prose describing how the line is read; it is deliberately
not an enum-mapped setting and carries no rate, pitch or dB value. An empty Lines list is a
valid outcome — a silent stretch is allowed, and the Voiceover step completes on one rather than
failing; see An empty plan is a successful step.
Stated as the guarantee it is: a narration plan is structurally incapable of naming a timestamp or a coordinate. There is no field it could put one in, so a line that wanted to say "0:07" or "1280,720" has to name an offered anchor id instead and let the deterministic half do the arithmetic.
NarrationPlanOutputInvariantTests is the guarantee, and it is a reflection test rather than a
review habit: it enumerates the schema's properties and fails on any
int/long/float/double/decimal/TimeSpan/DateTime/DateTimeOffset property anywhere in
it, and on any string property whose name is time-shaped (Sec/Seconds/Ms/Millis/
Time/Timestamp/Duration/Frame/Frames/Offset/Start/End). A NarrationPlanOutput that
grows a double, or a string that merely sounds like a time, fails the wave rather than a review.
It also pins both property sets and round-trips camelCase JSON, so a renamed member is a test failure
rather than a silently ignored field.
Tool scope: read-only, deliberately
ToolGroupCatalog.GroupsFor(AgentType.NarrationWriter) is exactly [ToolGroup.ProjectRead, ToolGroup.WorkflowControl] — the minimal read-only set VideoStoryEditor has, with no
SandboxAuthoring, no SandboxRender and no ProjectWrite. An agent that writes prose about a
picture has no reason to touch a file or start a render, and its prompt does contain media-derived
content (on-screen text, transcript spans), which is the one place in this repository where
read-only tools are the point rather than a default. The grant is written as its own explicit arm
in the catalog rather than left to a default, so a future default change cannot silently widen it.
NarrationWriterAgent.DefaultPrompt and its DatabaseSeeder.BuiltInAgents entry are byte-identical,
asserted by a [Fact] in VideoStoryEditorPromptConsistencyTests — the test run is the proof, not a
visual diff. That duplication is what every built-in agent carries, and it is why a prompt edit
changes both sides in one commit.
Config reference
VoiceoverStepConfig (ReelBolt.Shared/Workflows/VoiceoverStepConfig.cs), serialized to
WorkflowStep.VoiceoverConfigJson (voiceover_config_json, jsonb, nullable):
public sealed record VoiceoverStepConfig(
VoiceoverSource Source,
Guid? ProviderId = null,
string? Model = null,
string? DefaultVoiceId = null,
long MaxBytes = 2_000_000);
public sealed record VoiceoverSource(
VoiceoverSourceKind Kind,
IReadOnlyList<VoiceoverLine>? Lines = null,
int? ScriptwriterStepOrder = null,
int? NarrationWriterStepOrder = null);
public sealed record VoiceoverLine(
string Text,
double StartSec,
string? VoiceId = null,
string? AnchorId = null);
public enum VoiceoverSourceKind { Inline, ScriptScenes, NarrationPlan }
Three members, and the record is append-only in three separate places. VoiceoverSourceKind's
source file says so in its own doc comment: every member is serialized into
voiceover_config_json (jsonb) and read back by the execute path and by the frontend's
TypeScript mirror, so a member inserted before ScriptScenes would silently reinterpret every
stored Inline or ScriptScenes config. NarrationPlan is therefore appended last, and the
two members that preceded it keep their ordinal values. VoiceoverSource.NarrationWriterStepOrder
is appended last for the same reason, and is a distinct member rather than a reuse of
ScriptwriterStepOrder — the two name different agents with different output schemas, and
overloading one member would make a stored config's meaning depend on which kind it happens to
carry. VoiceoverLine.AnchorId is appended last on the line record so an Inline line stored
before it existed still deserializes with the same first three values.
Two JSON facts follow from VoiceoverStepConfig.JsonOptions (Web defaults plus a
JsonStringEnumConverter): properties are camelCase, and the enum member is written as the
PascalCase literal "NarrationPlan" — not "narrationPlan". That is what the committed
template, the tests and the frontend mirror all write:
{"source":{"kind":"NarrationPlan","narrationWriterStepOrder":3}}
The same JsonOptions instance is used by the executor at run time and by the Inference API's
save-time validation, so the two cannot drift into reading one config two ways.
| Member | Meaning |
|---|---|
Source.Kind | Inline (lines embedded in the config), ScriptScenes (lines read from an earlier Scriptwriter step's output) or NarrationPlan (lines read from an earlier NarrationWriter step's NarrationPlanOutput). |
Source.Lines | The Inline lines. Each has Text, a StartSec on the source/picture timeline, and an optional per-line VoiceId that overrides DefaultVoiceId. |
Source.ScriptwriterStepOrder | Required when Kind = ScriptScenes — the StepOrder of the step whose output is a ScriptwriterOutput. |
Source.NarrationWriterStepOrder | Required when Kind = NarrationPlan — the 1-based StepOrder of the NarrationWriter step whose output is a NarrationPlanOutput. |
VoiceoverLine.AnchorId | The opaque cut-anchor id the line was written for (s{n}/g{n}/t{n}), carried only by a NarrationPlan line — no other source kind can declare one. It is never spoken and never affects synthesis or the content-addressed artifact key; it is threaded through to the fit facts so an overflow stays attributable to the moment the line was written for. |
ProviderId | Optional per-step speech-synthesis provider override, resolved as above. |
Model | Optional model name; wins over the provider row's ModelName, whose value in turn wins over the client's hardcoded s2.1-pro default. |
DefaultVoiceId | Fish Audio reference_id used for any line that does not carry its own VoiceId. |
MaxBytes | Budget in UTF-8 bytes of sanitized text, checked before any provider call. Default 2,000,000. |
The step is registered as IStepExecutor/StepType.Voiceover in
WorkflowEngine/WorkflowEngineServiceCollectionExtensions.cs — the composition root was extracted
there and Program.cs delegates to it — and StepCachePolicy marks StepType.Voiceover cacheable
(StepCachePolicy.cs): it is deterministic, non-LLM, and its only side effects are a
WorkflowEngine-owned scratch file and content-addressed objects of its own in storage. The
content-addressed half is what makes a cached output replayable — see
Step-result caching and the content-addressing invariant.
Inline and ScriptScenes
Inline is the general case: the workflow author writes the narration and its start times directly
into the step config. Nothing is resolved against a cut decision on this path — those times are the
author's.
ScriptScenes exists for the existing promo pipeline, where a ScriptwriterAgent step has already
produced a ScriptwriterOutput with per-scene Voiceover text and a StartTime. The executor
finds the step named by ScriptwriterStepOrder in the step-output history, deserializes it as
ScriptwriterOutput, and converts every scene with non-empty Voiceover text into a
VoiceoverLine whose StartSec is that scene's StartTime and whose VoiceId is null (so
DefaultVoiceId applies). Scenes with empty/whitespace voiceover are skipped — which is exactly the
convention the staging side relies on, since empty scenes consume no line index.
None of the three kinds re-cuts picture to match narration. Phase 1 adds narration audio to the existing picture; phase 2 measures whether it fits and makes that measurement readable downstream (The fit solver and the review loop); moving the picture itself is phase 3's "audio-first timing".
NarrationPlan
NarrationPlan is the model-authored path, and it is the one kind whose lines arrive without a
position. The executor finds the step named by NarrationWriterStepOrder in the step-output
history, deserializes it as NarrationPlanOutput, and turns each plan line into a VoiceoverLine
carrying that line's Text plus its opaque AnchorId — and nothing else, because the plan has
nothing else to give:
{"source":{"kind":"NarrationPlan","narrationWriterStepOrder":3}}
Lines whose text is empty or whitespace are dropped here, exactly as the ScriptScenes arm drops an
empty scene voiceover (the sanitizer would skip them anyway) — and such a line takes its anchor id
with it, because a line that was never emitted carries no fit to attribute. Every line that is
kept carries its anchor, and the fit pass reads that anchor rather than inventing one: a fabricated
id would claim an anchor the writer never named.
A plan line carries no start, so the step declares one. NarrationPlanOutput is structurally
incapable of carrying a start time — every property is a plain string, and
NarrationPlanOutputInvariantTests bans a property even named for a time — and the voiceover step
holds no picture timeline to place a line against. So after the synthesis loop,
AssignNarrationPlanStarts walks the emitted lines and assigns each one a start back-to-back in
plan order: line 0 at 0.0, each next line at the previous line's declared start plus that line's
measured durationSec (0 for a skipped or failed line, which synthesized nothing, and 0 for a
non-finite or negative duration, which cannot advance a position). The result is deterministic and
never overlapping.
Two consequences are worth stating, because both are load-bearing:
- The declared start written to
view.lines[].startSecand the declared window the fit pass reads come from one assignment, so they cannot disagree. Because the next line's start is the previous line's start plus its measured duration, aNarrationPlanline's declared window is exactly its own duration and its declared-window overflow is structurally zero. That is not a bug to fix here: the declared-window number answers "did the line overrun the slot it declared", and a narration-first line declares exactly the slot it occupies. Whether the narration overruns the shot it lands in is the compile step's separate measurement. - This path deliberately does not read the analysis step, require one, or invent an analysis-step reference in its config. Placing narration against the picture would be a new cross-step dependency for a number the compile step already produces.
Inline and ScriptScenes are untouched by the placement pass: their starts come from the
config/script, and their fit numbers stay byte-identical.
Word alignment
A synthesized line is a WAV with no notion of where its words sit inside it. Phase 2 recovers that:
an IVoiceoverAlignmentService (Services/Voiceover/ in the WorkflowEngine) aligns each synthesized
line by re-transcribing its own WAV.
Why ITranscriptionClient, and not Fish's native timestamps
Fish Audio advertises per-word timestamps through POST /v1/tts/stream/with-timestamp (SSE) and its
live WebSocket twin, and phase 1 explicitly left the choice between those and a second ASR round-trip
to whoever scoped phase 2. It is settled by measurement, not preference: the self-hosted fish-tts
returns
POST http://localhost:8180/v1/tts/stream/with-timestamp
-> 404 {"statusCode":404,"message":null,"error":"Not Found"}
That is the server the whole local pipeline runs against, so a Fish-native alignment path would be
permanently dead code on every self-hosted deployment — and phase 4 adds a local TTS provider kind,
which makes self-hosted the direction of travel rather than an edge case. Meanwhile
ITranscriptionClient (Local Whisper) already returns a real
words[{text, startSec, endSec}] array on this stack, through a provider-agnostic seam that already
exists and is already exercised by VideoAnalyze.
So phase 2 aligns by re-transcribing the line WAV through ITranscriptionClient, and Fish's native
endpoints stay unused. If a later wave wants them for hosted-Fish deployments they are a fast path
behind the same interface, never a replacement — the self-hosted case has to keep working.
ASR word boundaries are the recognizer's, not the synthesizer's, and that is the right source here for the same reason it will be right for captions: it measures what a listener would actually hear, which is exactly what a caption must match.
How it runs, and how it fails
AlignAsync(projectId, storageKey, cancellationToken) resolves the provider with
IInferenceProviderResolver.ResolveTranscriptionAsync — the documented precedence, with no legacy
fallback — downloads the line WAV the step already stored, and calls ITranscriptionClient with word
timestamps requested. It runs on both paths that produce an ok line: fresh synthesis and the
object-store cache-hit path, so a repeat run does not silently lose its words.
What leaves the host. Alignment is a host egress, and it is unconditional: for every ok
voiceover line the step downloads that line's synthesized WAV and POSTs the audio to the
deployment's default Transcription provider row (Capability = Transcription, IsDefault = true),
which in the normal deployment is a hosted third party. The payload is the narration audio the step
itself just produced — text that was, on the ScriptScenes arm, derived from media-derived material
— and the returned transcript is what becomes the line's words.
Two properties follow from how the row is resolved, and both are worth stating plainly:
- There is no per-step override.
ResolveTranscriptionAsyncis called with a null id, so unlikeSpeechSynthesis— where the step's ownProviderIdis tried first — a workflow author cannot point alignment somewhere else, and cannot leave it unset for a given step. The destination is whatever row the deployment has markedIsDefault = true AND Capability = Transcription; an admin changes it by editing that row, not by editing the workflow. - A deployment with no
Transcriptionrow configured sends nothing. Resolution returns none, the path recordsNO_TRANSCRIPTION_PROVIDERand degrades — so alignment is never silently redirected to some other provider, and a stack with no transcription default simply has no word timings.
This is the same provider row, reached by the same ITranscriptionClientFactory, that the shipped
VideoAnalyze step already sends the project's own source-video audio to, so it widens an
accepted external-dependency flow rather than opening a new kind of one. It is nevertheless a second
reachable production call site through that factory. docs/decision-models.md's "What leaves the
host" is the register for the Decision call sites and does not inventory this one — this section
is that inventory for voiceover.
Alignment is the third soft-failure seam on this path and behaves like the others: it degrades with
a recorded reason and never fails the step. A missing provider, a failed download, a throwing client
and an empty word list each return "not applied" with a distinct reason from a closed set —
NO_TRANSCRIPTION_PROVIDER, DOWNLOAD_FAILED, TRANSCRIPTION_FAILED, NO_WORDS. Nothing escapes as
an exception except a genuine caller cancellation, which is re-thrown, the same way
GuardrailScreener treats one: a cancelled run must not be recorded as a degraded line.
Two deliberate omissions. Alignment is not part of the step-cache key — the whole output JSON is
cached, so a hit already carries the words of the run it cached. And a skipped or failed line is
never aligned: there is no audio to align.
What lands in the output
Additively, per line (Step output and error codes):
words—[{ "text", "startSec", "endSec" }], present only when alignment applied.alignmentReason— the named reason, present only when it did not.
Both are omitted rather than null when they do not apply, the same convention storageKey already
uses, so "field absent" keeps its existing meaning.
meta.alignment rolls the whole step up: { applied, alignedLines, unalignedLines, reasons: { <reason>: <count> } }. It is always present once lines are produced, never null, and computing it
never throws.
The fit solver and the review loop
A narration line is anchored to a shot, and it has to be speakable in that shot. Phase 1 probed a line's duration and reported it; phase 2 measures it twice — against the slot the line declared, on the voiceover step, and against the shot it landed in, on the compile — and turns both answers into facts the review loop can act on.
Deterministic code, never an agent
No model computes a duration anywhere on this path, and neither half of the fit is an agent. The overflow is arithmetic over data the pipeline already has — the line's declared or mapped start, the duration probed from its WAV, and the words that were actually spoken — so there is nothing for a model to judge. An agent here would also be unfalsifiable: a narration that overruns its slot is a measurement, not an opinion, and the review loop's whole lever is that it can threshold a number.
There are two fits, deliberately, and they answer different questions:
The NarrationFitSolver | The compile's picture-window facts | |
|---|---|---|
| Where | StepType.Voiceover — ReelBolt.Shared/Workflows/NarrationFitSolver.cs | VideoCompileStepExecutor — WorkflowEngine/Services/Video/NarrationFitFacts.cs |
| Question | "does the narration overrun the slot it was written for?" | "does it overrun the shot it landed in?" |
| Window | The declared narration window [StartSec_i, StartSec_{i+1}) | The containing output placement, after seam overlaps |
| Unit | Words (and seconds, as evidence) | Seconds only |
| Node | view.lines[].fit + meta.fit on the voiceover step | voiceover.fit + per-line keys on the compile step |
VideoCompileStepExecutor did not reimplement the solver's arithmetic and NarrationFitSolver
does not read, require or reimplement the picture window. NarrationFitSolver lives in
ReelBolt.Shared/Workflows rather than in the WorkflowEngine's own tree precisely because both
services may read a fit report and Services/Video/** is a different department's surface. A
consumer must not conflate the two, and neither type may grow a "read the other" path.
The declared window on the producing step (NarrationFitSolver)
NarrationFitSolver.Solve(IReadOnlyList<NarrationFitLine>?) returns a NarrationFitReport, a closed
serializable record:
public sealed record NarrationFitReport(
IReadOnlyList<NarrationFitLineResult> Lines,
int MeasuredLineCount,
int OverflowLineCount,
int OverflowWords,
double OverflowTotalSec,
double MaxOverflowSec,
int UnknownWindowLineCount);
public sealed record NarrationFitLineResult(
string Id, string? AnchorId, double? WindowSec, double PlayDurationSec,
int WordCount, int? AlignedWordCount, double WordsPerSecond,
double OverflowSec, int OverflowWords, bool Fits, string WindowSource);
Every member of NarrationFitLineResult is always present in the serialized form — this is a
closed record whose key set the compile's passthrough reads unconditionally (overflowWords in
particular), and a null already means exactly one thing per member: no window, no anchor declared, or
no alignment applied.
The rules, from the committed code:
- The window is
[this line's start, the next declared line's start)—NarrationFitLine'sWindowStartSec/WindowEndSec, both of which the step leaves null on the last declared line, which has no successor to end at. Such a line carriesWindowSec: nullon the result, is reported asWindowSource: "none"and is counted inUnknownWindowLineCountrather than being given an invented duration. A windowless line never contributes overflow and always counts as fitting — an unknown window is an absence of evidence, never a defect to report as one. The other literal is"declared_starts"; both are a closed set, and the usable-window test is that both ends are present, finite, andend > start. WordCountcounts whitespace-separated tokens in the sanitized text — what was actually spoken, which is what makes the number comparable to a duration.NarrationFitSolver.CountWordsis public for exactly that reason: the step that has the sanitized text must count it here rather than with a second, subtly different split.OverflowSecismax(0, playDurationSec - windowSec), rounded to 3 decimals;OverflowWordsisceil(overflowSec * wordsPerSecond). Ceiling, not rounding: a line that overruns by a fraction of a second still overruns by at least one spoken word, and "0 words over" would read as a fit.WordsPerSecondis the line's own rate and is0.0— never NaN, never Infinity — whenever it is undefined.- No alignment is required to solve.
AlignedWordCountis carried through only as evidence;WordsPerSecondandOverflowWordsare identical on a deployment with no ASR provider and on one with a perfect recognizer. - Never throws. A null list, a null element, a non-finite duration or a negative word count all
produce a well-formed report: non-finite durations normalize to
0, negative counts clamp, andEmpty— every count0, both second totals0.0, never null and never NaN — is returned for an empty input.
The step attaches it in one pass over the finished line list, after synthesis (a duration is only
known once a line has been synthesized or probed), and from the declared line list: a failed line
still occupies its declared slot, so the window of the line before it still ends where that failure
was declared to begin — but the failed line itself gets no fit key at all, because it synthesized
nothing to measure. Fit rides in the output JSON rather than in the step-cache key, so a step-cache
hit re-serves the fit computed by the run it cached, exactly like the alignment words.
The window is the containing shot, not the gap to the next line
The compile step measures every resolved line against the output placement its start landed in — the shot the line starts in — extended across every following kept shot that starts no line of its own (a line may carry over a cut), and reports per line:
| Key | Meaning |
|---|---|
windowStartSec / windowEndSec | The containing placement's start, and the end of the last shot the line may carry into. |
windowSec | windowEndSec - outputStartSec — the picture time left from the line's own start to the end of that window. |
carriedShots | Present only when the window was extended: how many following shots it spans. |
overflowSec | max(0, playDurationSec - windowSec), rounded to 3 decimals like its sibling keys. Zero means it fits. |
fits | overflowSec <= 0. |
Measuring instead against [line start, next line start) would report "fits" for a line that runs
straight past the cut into the next line's shot — precisely the defect these facts exist to surface.
The carry stops at the first shot that another line was WRITTEN FOR — the shot its anchor starts
in, not the second its audio happens to start: narration plays back to back, so a line pushed back
behind a long one must not let that long one claim the picture it was anchored to. A line followed
by another line's anchor in its own starting shot never carries. An overrun of at most 100 ms (NarrationFitFacts.OverflowToleranceSec) is the
WAV's trailing breath, not a word over the next shot, and reports overflowSec: 0: a 3.019 s opener
over a 3.0 s shot had capped a reel's review at 4.
Lines that carry over a cut
A beat-cut reel holds a shot for under two seconds, and one spoken sentence takes three or four: with
the window limited to the starting shot, every such line overflowed and the narration cap held every
review at 4 however good the edit was. Speaking over a cut is ordinary editing; speaking over the
NEXT line's picture is the defect, so only that still counts as an overflow. That is why the window is
the containing placement, and why the lookup belongs on OutputTimeline: it is the only thing that
knows where a span landed once seam overlaps are accounted for.
The two shapes are complementary rather than redundant, and a NarrationPlan step is where that
becomes visible: the step's own declared-window overflow is structurally zero (a narration-first line
declares exactly the slot it occupies — see
NarrationPlan), so the compile's picture fit is the one that reports an actual
overrun. On an Inline or ScriptScenes step, where the author or the script declared the starts
independently of the measured durations, both can be non-zero.
How the overflow reaches the ReviewLoop
A review loop's lever is VideoReviewAgent's score against the compile step's own output JSON, which
is already in the pipeline history the reviewer is handed — so the fact has to be in the node the
reviewer reads. It is: voiceover.fit, a node-level roll-up carried, like every other key on that
node, into both outputSummary["voiceover"] and edl["voiceover"]. The voiceover step's own
meta.fit and its two promoted integers are the same fact one step earlier — gate on whichever step
your Conditional/ReviewLoop sits behind.
The reviewer's prompt does carry the narration clause, so a ReviewLoop over VideoReviewAgent
acts on overflow twice over, and the two mechanisms are complementary rather than redundant.
In the prompt. VideoReviewAgent's fact list gained a voiceover clause describing the per-line
picture windows (with fitSource/solverFit), the voiceover.fit roll-up, and the rule the whole
check turns on: the containing output placement is the window a listener actually hears, so a line
that runs past the end of its own shot is an overflow, never a fit, however much room there was
before the next line began. It scores no higher than 4 when overflowLineCount > 0, names
maxOverflowSec and the lineId of each overflowing entry, and quotes the narration text so the
retry knows which line to shorten. It also states the only two remediations the upstream agents can
follow — rewrite the line shorter for the same anchor, or keep a longer span — and forbids telling
any agent to "re-time", "shift", "extend" or "move" a line, because the plan carries no timestamp.
That clause's "not measurable" list names both degradation reasons the node can carry, and its
unknownWindowLineIds sentence describes a partial measurement. On the shipped executor the list is
narrower than the prose: the no-fit-node case is live (a compile with EnableVoiceover: false), of
the two reasons only no_resolved_lines is producible, and the partial measurement it describes does
not occur — see the degraded shape. The instruction is
still correct and still load-bearing: it is what stops the reviewer capping a fit that was never
computed. It simply has fewer live cases than it enumerates, and a reviewer must not report
window_not_found as something to expect.
In code, because a prompt is a request. There are three deterministic caps, applied together
by ReviewLoopStepExecutor.ApplyNarrationCaps. They are separate Math.Mins, so the order is
immaterial; each can only lower a score, never raise one, none throws, and none rewrites anything —
the reviewer's own output JSON is persisted verbatim, so the capped value drives the loop decision
only. Every constant is 4, below every video template's MinScore (8), which is what makes a cap
iterate the loop rather than merely score lower.
ApplyNarrationFitCap— a line overran its window.min(score, NarrationOverflowScoreCap)when the prior step's output carries a top-levelvoiceover.fitwithapplicableliterallytrueand a positiveoverflowLineCount; the score unchanged otherwise. It is inert to a fit that was never measured:applicable: false, a missingfitnode andreason: voiceover_not_enabledare all "not measurable, therefore never a defect". A partial measurement is not exempted — a non-emptyunknownWindowLineIdsmeans fewer lines were measured, and the lines that were measured still count. That list is empty by construction on the shipped executor, so this is a statement about the cap's definition rather than a case a reviewer meets.ApplyNarrationDeclaredLostCap— narration was declared and never reached the picture.min(score, NarrationDeclaredLostScoreCap)on the other half of the silence story: the compile rendered a green, narration-less video, and because that node has no measurable fit itsfit.applicableisfalse, so the overflow cap above is inert to it by design. Without this cap a green reviewer score advances the loop over a video with no narration at all.ApplyNarrationMostlyDroppedCap— more than half the narration is missing.min(score, NarrationMostlyDroppedScoreCap)when the top-levelvoiceovernode haslineCount > 0and(lineCount - appliedLineCount) * 2 > lineCount. This is the gap between the other two: the declared-and-lost cap needs zero resolved lines and the fit cap only measures the lines that did resolve, so a compile that placed 1 of 8 lines (the multi-clip bug fixed in Multi-clip narration) passed review green over ~60 s of silence. It reads the same declared-vs-applied pair the compile's ownnarration_mostly_droppedwarning is computed from, so the warning the customer sees and the loop decision cannot disagree. Exactly half missing is not "mostly" and is left to the reviewer (the compile still warnsnarration_dropped).
IsDeclaredAndLostNarration accepts either of two machine-readable forms on the top-level
voiceover node, and only after two gates: applied must be literally false, and
appliedLineCount must be literally 0. (That pair is what makes the cap inert for every partial
resolve — one where appliedLineCount is positive — and for the empty plan below, which declares
nothing at all.)
- The declared-vs-applied key pair —
lineCount > 0withappliedLineCount == 0, the form the contract names and the one the cap is keyed on.lineCountis the declared count: every parsedview.linesentry carrying a line id, counted before the usability filter, so askipped/failedline carryingdurationSec: 0still counts as asked-for.appliedLineCountis the resolved count, so the two are equal only when nothing was lost. This form fires on every arm where lines were declared and none reached the picture. - The compile's terminal outcome token —
reason: "all_lines_unavailable", whichVideoCompileStepExecutorsets if and only if at least one line was declared, not one of them resolved, and theamixprobe succeeded. It is a closed terminal value the rest of this codebase already decides on, not an inference from prose, and on that arm it is carried alongside form 1 rather than instead of it.
Neither form is a rename, a reshape or a change to media's node, and both are live: the key pair
stopped being a duplicate of itself when Finish() was corrected to report the declared count, which
is the declared-vs-resolved distinction its own comment always said it was for.
The amix-unavailable arm is capped too. A container whose ffmpeg build exposes no amix
normalize option reports lineCount > 0, appliedLineCount == 0, unavailable: true with an
explanatory reason. The cap fires on it exactly as it does on a planner-side total loss. A loop
cannot fix an environmental cause, and that is not the point: a bounded (MaxIterations: 3) loop that
ends with the run not reporting green is the desired outcome, because the alternative is a green
render with no narration — the single failure this whole path exists to prevent. A wrong-but-loud
report beats a wrong-and-silent one. The node's own reason and unavailable fields are what let a
human tell an environment problem from a planner problem, which is why the reviewer's feedback must
name the cause rather than only the missing narration.
What the cap is inert for, now. Exactly two shapes: a compile with no voiceover node at all
(narration switched off, EnableVoiceover: false) and a Voiceover step that declared nothing —
the legitimate empty NarrationPlan, whose lineCount is 0 and which therefore carries neither
form. A node reporting lineCount > 0 with appliedLineCount == 0 is capped whatever the cause, and
a fit.applicable: false node is not an exemption: of the two reasons that node can carry, only
no_resolved_lines is producible, and zero resolved lines with lines declared is precisely the shape
this cap exists for.
Both caps are also honoured by the calibrated decision gate's early stop, which is the one path
that returns before the reviewer runs and therefore before either cap is applied:
NarrationCapsWouldFire probes the caps themselves (with UncappedProbeScore, int.MaxValue, chosen
so the predicate cannot drift from the caps it describes) and, when one would fire, escalates to the
reviewer instead of accepting. That costs the gate nothing but the escalation it already treats as
its safe failure mode, and it short-circuits before the outbound provider call, so a step the caps
will reject never costs a provider round trip. It also means the caps cannot be declined by turning
that gate on. The gate's own fact whitelist is deliberately not widened with a voiceover fact —
the set of things that leave the host to a third-party Decision provider is not being increased —
so the fix lives on the local, inert side.
A Conditional/ReviewLoop that would rather not depend on the prompt can still gate on the node
explicitly (overflowLineCount / maxOverflowSec).
{ "applicable": true, "measuredLineCount": 3, "overflowLineCount": 1,
"overflowTotalSec": 1.8, "maxOverflowSec": 1.8, "unknownWindowLineIds": [] }
maxOverflowSec is the worst single overflow and overflowTotalSec the sum; overflowLineCount is
how many lines missed their window. unknownWindowLineIds is where a resolved line whose containing
placement could not be found would be named, so a partial measurement is reported rather than silently
averaged into the totals — on the shipped executor it is empty by construction (below), so every
number here is a total over all resolved lines. It is present — possibly empty — only on the
applicable: true shape; the degraded shape below has no such key. applicable is true only when
at least one line was actually measured.
A compile that cannot measure fit degrades to a different shape: { "applicable": false, "reason": <reason> }, carrying no measuredLineCount, no totals and no unknownWindowLineIds.
On the shipped executor reason is always no_resolved_lines — nothing resolved, so there was
nothing to measure. The other value, window_not_found ("lines resolved but none had a containing
placement"), survives in BuildVoiceoverFitNode as a defensive arm only, unreachable from the
mapped path: containment is looked up with the line's raw MapToOutputSec second while the
reported outputStartSec stays the 3-decimal rounded one, and MapToOutputSec returns
placement.OutputStartSec + (sourceSec - span.SnappedStart) for a span whose bounds are exactly what
OutputTimeline.Build sets that placement's OutputEndSec from — so a second it produced is always
inside its own placement and the lookup cannot fail. A reviewer will not meet this reason; do not
write it up as one to expect.
Rounding for reporting while looking up raw is itself load-bearing, and is why the arm closed:
testing containment on the rounded value dropped a line that rounded onto its shot's exclusive end out
of measuredLineCount while leaving overflowLineCount at 0 — a real overflow hidden behind a
clean-looking node. It still returns StepStatus.Completed with valid JSON, exactly like every other
soft-failure path on that node. A missing measurement must never become a failed render.
Optional upstream passthrough
The passthrough now has a producer: the Voiceover step's own view.lines[].fit, the
NarrationFitSolver verdict described above. A line node may carry that fit object contributed
upstream, copied verbatim under the key solverFit, with fitSource set to "solver". The
object the compile reads is that line's own upstream fit key; it re-emits it here under
solverFit. Without one, fitSource is "compile_measured". The passthrough is tolerant by
construction: it is never required, a malformed one is ignored rather than failing the compile, and
it never changes the compile-measured overflowSec/fits — the two numbers sit side by side on
the same line node, each labelled by which measurement it is.
fitSource is written on every resolved line; solverFit only when an upstream one existed. The
node-level voiceover["fit"] keeps its name: fit at the top level of that node is the roll-up, and
the per-line passthrough is deliberately not called fit for exactly that reason.
Step output and error codes
VoiceoverStepExecutor never throws: every failure mode is representable in the JSON it returns,
because output_json is jsonb and a thrown exception would otherwise be a persisted-format
problem rather than a step result.
Successful output:
{
"view": {
"lines": [
{ "id": "vo0", "startSec": 0.0, "status": "ok", "durationSec": 3.42,
"storageKey": "projects/{projectId}/agentFiles/voiceover/lines/{sha256hex}.wav",
"words": [ { "text": "ReelBolt", "startSec": 0.12, "endSec": 0.71 } ],
"fit": { "id": "vo0", "anchorId": null, "windowSec": 6.0, "playDurationSec": 3.42,
"wordCount": 7, "alignedWordCount": 7, "wordsPerSecond": 2.05,
"overflowSec": 0.0, "overflowWords": 0, "fits": true,
"windowSource": "declared_starts" } },
{ "id": "vo1", "startSec": 6.0, "status": "failed", "durationSec": 0 }
]
},
"meta": { "totalLines": 2, "failedLines": 1, "skippedLines": 0, "allLinesFailed": false,
"alignment": { "applied": true, "alignedLines": 1, "unalignedLines": 0,
"reasons": {} },
"fit": { "measuredLineCount": 1, "overflowLineCount": 0, "overflowWords": 0,
"overflowTotalSec": 0.0, "maxOverflowSec": 0.0,
"unknownWindowLineCount": 0 },
"overflowWords": 0, "overflowLineCount": 0 }
}
Line ids are positional over the resolved line list — vo{index}, zero-based — and durationSec is
the probed duration of the synthesized WAV, not an estimate.
An ok line carries its full storageKey: the content-addressed key described under
Per-line artifact resolution, not a path a consumer reconstructs.
The property is omitted, never null, on a line that has no artifact (skipped/failed) — and
that absence is precisely the signal a consumer reads as "fall back to the legacy execution-scoped
key".
An ok line may also carry the two word-alignment fields described in
Word alignment: words ([{ "text", "startSec", "endSec" }]) when alignment
applied, or alignmentReason (the named reason) when it did not. Both follow the same
omitted-rather-than-null convention as storageKey, so an absent field keeps meaning "not
available" rather than "present and empty".
Every ok and skipped line then gains fit, the NarrationFitSolver verdict described in
The declared window on the producing step.
A failed line carries no fit key at all — it synthesized nothing to measure — and that absence is
how a consumer tells "not measured" from "measured and fitting".
meta carries four counters — totalLines, failedLines, skippedLines and allLinesFailed —
and all four are present on every envelope, success or failure. Once the step has produced lines it
also carries:
alignment— the word-alignment roll-up described in Word alignment.fit— the completeNarrationFitReporttotals, always present (never null) once lines are produced, so "nothing overflowed" reads as zeros rather than as an absent key:measuredLineCount,overflowLineCount,overflowWords,overflowTotalSec,maxOverflowSecandunknownWindowLineCount. Note the last member: it is a count here, not the compile node'sunknownWindowLineIdsarray.overflowWordsandoverflowLineCount— the same two integers promoted tometaitself, deliberately duplicated so aConditionalstep can gate on a bare scalar comparison instead of walkingview.lines[]. The nestedmeta.fitobject stays the complete report; this pair is the one that has to be cheap to read.
A step that produced no lines at all never reaches this shape — it returns through Failure(),
whose envelope is unchanged. One exception, described next.
An empty plan is a successful step
A well-formed NarrationPlan declaring zero lines completes the step. The planner is the only
thing that knows whether an edit wants narration, and both NarrationWriterAgent's prompt ("if the
edit needs no narration at all, output an EMPTY lines list rather than filling the silence") and
NarrationPlanOutput's own remark say so. Failing the step here made a model that obeyed its prompt
redden a correct run — and after MaxStepRetries, fail the whole execution, on a video that is not
wrong.
This is not a general weakening of NO_LINES, and the two kinds differ for a reason:
NarrationPlan— a decision. Nothing in the plan's emptiness is a mistake; it is an editorial choice made by the agent that was asked to decide. The step did its job and narrates nothing.InlineandScriptScenes— a misconfiguration. There, the lines come from the author (aVoiceoverstep configured with none) or from a referencedScriptwriterOutputwhose scenes carry no voiceover. Nothing decided silence; something was left out.NO_LINESremains exactly as it was for those two, and phase 1's semantics for them are deliberately untouched.
The empty-plan step returns a Completed envelope of the ordinary shape with every count at zero —
view.lines: [], totalLines/failedLines/skippedLines at 0, allLinesFailed: false, and
alignment and fit present rather than omitted so a downstream Conditional reading
meta.fit.overflowLineCount does not have to special-case the empty plan. No speech provider is
resolved, no line is synthesized, no artifact is uploaded: the step completes by doing nothing.
What distinguishes it, for a machine, is one extra key:
meta.emptyReason: "narration_plan_empty"— present only on this path, and absent rather than null or false on every other Completed envelope. "The writer chose silence" is therefore a key existence check, never an inference from a reason string.
It is added because the envelope otherwise cannot say whether the empty list was a decision or an
absence, and a consumer must never have to guess that. The two silences are also already separable
without it — this step is Completed with allLinesFailed: false, whereas narration that was
declared and every line of which was lost is Completed with allLinesFailed: true (or Failed)
and additionally trips the review cap in
How the overflow reaches the ReviewLoop — but neither
of those keys says why the list is empty, which is what this one does.
The third status: skipped
A line whose sanitized text is empty is never sent to the provider and therefore never billed.
Sanitization runs before the byte budget and before any provider call, so a line consisting only of a
[...]/(...) tag, a bare URL, control/format characters, or whitespace reduces to the empty string
— and such a line is emitted with "status": "skipped" and durationSec: 0, counted in
meta.skippedLines.
The load-bearing part is that the entry is retained in view.lines rather than dropped. Both
consumers — VideoCompileStepExecutor.ResolveVoiceoverAsync and
ReactRemotionSandboxTools.StageVoiceoverAudio — pair scenes to voiceover lines positionally, so
removing an entry would shift every later line and silently pair a later scene with an earlier line's
audio. A skipped line holds its slot precisely so that cannot happen; it simply has no storageKey
for a consumer to resolve.
meta.allLinesFailed
meta.allLinesFailed is true only when at least one line was emitted AND every emitted line
failed — the step was asked to narrate and produced nothing usable. The executor also logs at
Error in that case, naming the failed-line count and the execution id, so that a green step which
synthesized nothing is not the only thing an operator sees.
Stated plainly, and without softening it: such a step still returns StepStatus.Completed by
design. "Never throw, degrade one line, never fail the whole step" is the contract, and it is
deliberately unchanged. A workflow author who needs a hard failure must therefore gate on
meta.allLinesFailed with a Condition/ReviewLoop step — the executor will not fail the step for
them.
Failure output keeps the same shape with an empty lines array and a meta that is zeroed but still
carries all four keys ({ "totalLines": 0, "failedLines": 0, "skippedLines": 0, "allLinesFailed": false }), and moves the reason into ErrorDetails as CODE: message. The codes are:
| Code | When |
|---|---|
CONFIG_INVALID | VoiceoverConfigJson is absent, is not valid JSON, or deserializes to null. |
NO_LINES | The resolved line list is empty on a source kind where that is an author misconfiguration: no Inline lines, or a referenced ScriptwriterOutput whose scenes carry no non-empty voiceover. The NarrationPlan arm deliberately does not return this code for an empty plan — see An empty plan is a successful step. |
SCRIPT_SOURCE_INVALID | Kind = ScriptScenes without a ScriptwriterStepOrder. |
SCRIPTWRITER_NOT_FOUND | No step output in history at that StepOrder. |
SCRIPTWRITER_OUTPUT_INVALID | The referenced output is not deserializable as a ScriptwriterOutput. |
NARRATION_SOURCE_INVALID | Kind = NarrationPlan without a NarrationWriterStepOrder. |
NARRATION_STEP_NOT_FOUND | No step output in history at that NarrationWriterStepOrder. A distinct code from SOURCE_INVALID, deliberately: "the step you named is not in this run's history" is an operator-visible wiring mistake (wrong step order, or the named step sits after this one), and reporting it as "unknown source kind" would send whoever triages it looking at the enum instead of at the workflow. The ScriptScenes arm's SCRIPTWRITER_NOT_FOUND is the same code for the same reason. |
NARRATION_OUTPUT_INVALID | The referenced output is not deserializable as a NarrationPlanOutput, deserializes to null, or is a well-formed document that is not a plan — an explicit JSON null for lines, or a null entry inside it. Those two shapes parse perfectly, so without their own check they would dereference and escape as UNEXPECTED_ERROR; they get the arm's own code instead. A plan that is merely empty is none of these — it succeeds. |
SOURCE_INVALID | An unknown VoiceoverSourceKind. |
BYTE_BUDGET_EXCEEDED | Sanitized text totals more than MaxBytes — the provider client is never invoked. |
PROVIDER_NOT_RESOLVED | ResolveSpeechSynthesisAsync returned null (or threw; the exception is logged and treated as null). |
UNEXPECTED_ERROR | Any other exception that escaped to the executor's top level. |
A per-line synthesis failure is not a step failure: that line is recorded with
"status": "failed" and durationSec: 0, and its count goes into meta.failedLines — one of the
two per-line counters, the other being meta.skippedLines — after which the step still completes.
Only the guardrails above (bad config, byte budget, unresolvable provider) fail the step outright.
This is the same soft-failure discipline the compile step applies to music, graphics and SFX.
Two of those per-line refusals are named, and both mean the vendor's bytes were rejected before
anything was stored, so a bad payload cannot poison the content-addressed key:
VOICEOVER_PAYLOAD_NOT_WAV is a payload that is not a RIFF/WAVE container at all, and
VOICEOVER_PAYLOAD_NOT_DECODABLE is one whose container looks right but which ffprobe could not
decode to a positive duration with an audio codec. Those two are the only ones; the literal appears in
the engine log for that line, while the line's own output entry stays the generic
"status": "failed".
Ordering is deliberate: all lines are sanitized and the UTF-8 budget is summed before the provider is even resolved, so a config that would blow the budget costs nothing.
Artifacts, caching and where the audio lives
Two kinds of object come out of a Voiceover step, both written through
IProjectFileWorkspace.UploadArtifactAsync:
| Object | Name passed to UploadArtifactAsync | Content type |
|---|---|---|
| One WAV per successfully synthesized line | voiceover/lines/{sha256hex}.wav | audio/wav |
| The step's output JSON | voiceover/{executionId}/artifact.json | application/json |
Per-line WAVs are content-addressed: the file name is a hash of what was actually synthesized,
so no execution id appears in the path. The hash input is the literal tag voiceover-line-v2
followed by sanitizedText, effectiveVoiceId, effectiveModel, providerId and endpoint — six
fields joined with \n, hashed with SHA-256 over their UTF-8 bytes and rendered as lowercase
hex. The five value fields are length-prefixed — each written as value.Length + ":" + value rather
than joined by a delimiter — so the material is injective by construction: two distinct field tuples
can never hash to one name; the voiceover-line-v2 tag is a fixed literal and is not
length-prefixed itself. Lowercase because that hex string is a URL path segment. Four of the six
fields are resolved rather than taken verbatim from the step config.
effectiveVoiceId is line.VoiceId ?? config.DefaultVoiceId — the voice actually sent, not merely
the config's default. effectiveModel is
FishAudioSpeechSynthesisClient.ResolveEffectiveModel(config.Model, provider.ModelName):
config.Model when non-empty, else the provider row's ModelName when non-empty, else the literal
s2.1-pro — and since config.Model is null by default, the resolved model is normally the
provider row's. providerId and endpoint are the resolved provider row's ProviderId and
Endpoint verbatim. Provider name, kind, timeout and API key are deliberately not part of the
material: none of them changes the audio, and the key stays free of secrets. Only the step's own
JSON artifact stays execution-scoped, under voiceover/{executionId}/artifact.json.
UploadArtifactAsync writes a bare artifact — it creates no ProjectWorkspaceFile row — and
prefixes the caller's layout with the category, so the object that actually lands in the bucket is
projects/{projectId}/agentFiles/voiceover/…. The full key handed to consumers for a line is
therefore projects/{projectId}/agentFiles/voiceover/lines/{sha256hex}.wav, and that is what both
consumers download the line by (see
Per-line artifact resolution). The key returned for the step's own
JSON is what the step records as WorkflowStepResult.ArtifactStorageKey. This is the correct column
for a non-playable artifact (see docs/video-editing.md — the column is kept separate from
OutputStorageKey precisely so the execution UI never mistakes a JSON artifact for a render).
There is no in-memory per-line cache
VoiceoverStepExecutor is registered as a singleton, but it now holds no mutable state at all.
A previous design kept a process-wide ConcurrentDictionary<string, double> keyed by
SHA-256("{text}|{voiceId}|{model}|") mapping to the probed duration, and a hit skipped both
the synthesis and the artifact upload. That dictionary is deleted, not scoped.
The replacement is the object store itself. Before synthesizing a line, the executor asks
IProjectFileWorkspace.ArtifactExistsAsync, which is a HEAD (metadata) request and never a
download. A hit means "a previous run uploaded exactly these bytes", so the line is not re-sent to
Fish Audio; the object is still downloaded to scratch and ffprobed, because the duration has to
be re-measured as this execution's output (the entry the step emits is ok with a freshly probed
durationSec and the same full storageKey). A repeat run therefore pays a HEAD, a download and
one ffprobe per line — never a second Fish Audio request. Re-synthesis is bought only when the
content, the effective voice, the effective model or the provider identity genuinely changed, since
those are the only inputs to the key.
Why deleting it, rather than scoping it, was the fix. The old key's shape was the bug, not
the cache's lifetime. The artifact key was built from Execution.Id, so a cache hit emitted output
naming objects under another execution's prefix — and, for a line the current run's own upstream
had never produced, no object at all. Compile then dropped every line as line_file_not_found
while the workflow still reported success: the second run rendered with no narration and said
nothing about it. The execution id is also deliberately absent from the step-cache key, so a
step-cache hit reopened the identical hole. Content-addressing removes the whole class of bug — the
key no longer depends on which execution uploaded the object, so a hit verified within the current
execution is a verified claim that the exact object the output names exists, and no decision can
leak across executions.
Because an object written under a content-addressed key is what every future execution will find
by HEAD and trust as the real audio for those words, the executor validates an upload with two
independent gates before it stores anything: the payload must carry a canonical RIFF/WAVE container
header, and ffprobe must decode it to a positive duration with an audio codec. A failure —
reported as the per-line reason VOICEOVER_PAYLOAD_NOT_WAV or VOICEOVER_PAYLOAD_NOT_DECODABLE —
fails that one line and the bytes are never stored, so a bad vendor payload cannot permanently
poison the cache.
Step-result caching and the content-addressing invariant
StepCacheKeyInputs.VoiceoverConfigJson (StepCacheKeyInputs.cs) is part of the step-cache key,
appended by StepCacheKeyBuilder.cs and populated by WorkflowExecutorService.cs, exactly like
every sibling config blob — without it, two distinct Voiceover steps in one project hashed
identically and the second replayed the first's lines as if they were its own.
StepCacheKeyInputs.SchemaVersion is v2; the bump is what makes an entry written under the
older key shape an orphan rather than a false hit.
Content-addressing is precisely what makes caching this step safe. An execution id is not in the
step-cache key, so a Voiceover step's cached output can legitimately be replayed by a later
execution — that is the point of the cache. Because the per-line storageKey values inside that
replayed output contain no execution id either, they still name objects that exist regardless of
which execution uploaded them, and both consumers resolve lines from those keys. Had the per-line
keys stayed execution-scoped, every cache hit would have replayed keys the new execution never
wrote, and the failure would again have been silent.
The mix path (VideoCompile)
VideoCompileStepConfig gains two members (ReelBolt.Shared/Workflows/VideoCompileStepConfig.cs):
bool EnableVoiceover = false, // false is byte-identical to the pre-voiceover compile path
int? VoiceoverStepOrder = null // which step's output to take lines from
Both defaults preserve the previous behaviour exactly: EnableVoiceover = false means no voiceover
is even looked for and the generated ffmpeg command is unchanged. As with the rest of the compile
config, applying audio requires Mode = Reencode.
VideoCompileStepExecutor.ResolveVoiceoverAsync performs the whole resolution, and like the
music/graphics/SFX/inserts paths it is purely soft-failure: a missing or malformed voiceover
step, an unresolvable line, or an ffmpeg build whose amix lacks normalize all degrade to "no
voiceover applied", never to a failed compile. The voiceover node it produces therefore exists on
every path and reports why it did nothing:
{
"enabled": true,
"applied": true,
"unavailable": false,
"appliedLineCount": 2,
"lineCount": 2,
"lines": [ { "lineId": "vo1", "outputStartSec": 2.0, "playDurationSec": 3.0,
"windowStartSec": 0.0, "windowEndSec": 6.0, "windowSec": 4.0,
"overflowSec": 0.0, "fits": true, "fitSource": "compile_measured" } ],
"dropped": [],
"overlap": true,
"headroom": { "applicable": true },
"fit": { "applicable": true, "measuredLineCount": 1, "overflowLineCount": 0,
"overflowTotalSec": 0.0, "maxOverflowSec": 0.0, "unknownWindowLineIds": [] }
}
The two counts on that node are not the same fact. lineCount is the declared count — every
parsed view.lines entry carrying a line id, counted before the usability filter — and
appliedLineCount is the resolved count, so the two are equal only when nothing was lost. A node
reading lineCount > 0 with appliedLineCount == 0 is stating that the producer was asked for
narration and the render about to be produced carries none of it, which is the shape
the declared-and-lost cap reads.
The non-fatal reason values are voiceover_not_enabled, voiceover_step_not_found,
voiceover_output_empty, voiceover_output_invalid_json, voiceover_lines_not_found,
voiceover_lines_extraction_failed, no_valid_voiceover_lines, and — for a build whose amix
does not expose normalize — unavailable: true with an explanatory reason. Lines that cannot be
used individually are listed in dropped with line_file_not_found, line_cut_away,
anchor_cut_away, line_past_program_end or line_download_or_probe_failed (an anchored drop also
names its anchorId).
Two additions to that node report the outcome an operator needs, as opposed to explaining a lookup that went wrong:
reason: "all_lines_unavailable"— set when the step declared at least one line and not one of them resolved. The declared count is taken from everyview.linesentry that carries a line id, before the usability filter, so a step whose lines all carrieddurationSec: 0scores as "was asked and delivered nothing usable" rather than as "was asked for nothing". Like every other reason on this node it is a report, not a failure: the compile still returns a valid result.voiceover["missingArtifactLineIds"]— present only when at least one line's artifact object was genuinely absent from the bucket, and naming exactly those line ids. It is kept separate fromdropped(which also carries cut-away and probe failures) because this one array is the operator's "the narration is missing from storage" signal. TheNoSuchKey/NotFoundpath behind it logs at Warning — it used to beInformation, a level nobody reads, which is how a completely failed TTS produced a green, narration-less render with no signal anywhere — and it names the full storage key alongside the line id, because the key is the only thing an operator can hand to the object store to find out what happened.
What resolution does per line:
- Place the line on the output timeline (
PlanVoiceoverLinePlacements, before anything is downloaded). A line carrying ananchorIdthe artifact knows is placed at the first surviving second of that anchor on the anchor's own source clip — see Multi-clip narration. A line without one (Inline/ScriptScenes) maps its declared sourcestartSecwithOutputTimeline.MapToOutputSec(startSec, 0), exactly as before. A line whose start was cut away is dropped asline_cut_away(unanchored) oranchor_cut_away(anchored) rather than being silently clamped to zero. - Take the object key from the line's own
storageKey— anokline always carries its full content-addressed key — download the WAV to scratch andffprobeit; a file with no audio stream or zero duration is rejected.storageKeyis omitted, not null, onskipped/failedlines, so the read is aTryGetProperty; when it is absent the executor rebuilds the legacy execution-scopedvoiceover/{executionId}/{lineId}.wav, which is the only object a step result persisted before content-addressing can mean. The download enforces theprojects/{projectId}/scope, so a key carried in a step output can never be coerced into reading another project's object. - Build a
ResolvedVoiceoverLinecarrying the output start, the duration recorded by the voiceover step, and a fullGainLinear = 1.0. Voiceover is deliberately not attenuated: unlike an SFX cue (bounded, and a deliberate accent), narration is the message. - Append the line's output window to a list that is handed to music planning as
additionalNoLiftWindows, soMusicMixPlanner.PlanLiftWindowswill not lift the music bed while narration is playing.
Overlapping lines are detected and set overlap: true plus a log line, but are otherwise allowed —
they mix, and the compile never fails over them.
Multi-clip narration
A NarrationPlan line's declared startSec is a back-to-back layout the Voiceover step made
without a picture (see NarrationPlan). The compile used to read it as a time
on source 0 (VoiceoverSourceIndex), which is harmless for a one-clip edit and fatal for a multi-clip
one: in an 8-clip narrated story every line written for clips 2..8 either landed over the wrong clip
or fell past clip 1's end and was dropped as line_cut_away — 7 of 8 lines, one sentence over 67 s,
while the run reported Completed.
The compile now places an anchored line the way SFX cues and overlays are placed: its anchorId
(s{n}/g{n}/t{n}) is resolved through the analysis artifact to (start, end, sourceIndex), and
OutputTimeline.MapWindowToOutput(start, end, sourceIndex) gives the first second of that anchor
that survived the cut on its own clip. Anchored lines are then laid out in output order, and a
line that would start before the previous anchored line has finished is pushed back to that line's
end — two narration lines are never spoken over each other — reported on the line node as
shiftedSec. A line pushed to the end of the program is dropped as line_past_program_end.
Anchored line nodes additionally carry anchorId and sourceIndex; unanchored nodes keep their
exact previous shape.
Everything downstream follows the placed start: the mix's adelay, the caption cues (which are cut
from the placed lines), the duck/replace windows, the music no-lift windows, voiceover.headroom
(the shot under the line is looked up on the line's own clip), voiceover.wer, and the
picture-window fit below, which for an anchored line is measured against whatever placement is on
screen at its start (OutputTimeline.TryGetContainingPlacement(outputSec), any clip). Pickups follow
the same rule: a pickup's t{n} segment is mapped through its own source clip.
OutputTimeline.MapToOutputSec also no longer gives up at the first span of a clip that starts after
the requested second, so a clip that returns later in the cut with an earlier part of itself still
maps.
Picture-window fit facts
Every resolved line is also measured against the output placement its start landed in, additively on
the same node: windowStartSec, windowEndSec, windowSec, overflowSec and fits — written only
on a line whose containing placement was found — plus fitSource on every resolved line, and
solverFit on the lines whose upstream output carried one. Nothing pre-existing on the node changes —
the new keys are appended after every key that was already there, so a pre-wave consumer reading
lineId/outputStartSec/playDurationSec sees exactly what it saw before.
Nothing in this block can fail the compile: the window lookup is a list walk and the arithmetic is
total. unknownWindowLineIds stays the reported home for a line whose containing placement could not
be found, but it is empty by construction — the lookup runs on the raw MapToOutputSec second,
which is always inside its own placement, so no line that the mapper placed can miss it. It remains on
the applicable: true shape for a caller that supplies a second from somewhere else. A compile with
nothing measurable records { "applicable": false, "reason": "no_resolved_lines" } — a shape with no
unknownWindowLineIds at all — and still returns a valid result.
The fit solver and the review loop has the meanings, the
degradation reasons and how the roll-up reaches the compile's own output.
EnableVoiceover = false remains byte-identical to the pre-voiceover compile path — no extra
ffmpeg inputs, no new filtergraph, and no voiceover node at all — so the fit facts are reached only
by a compile that was already asking for narration.
The ffmpeg fragments
VoiceoverMixFilterBuilder (WorkflowEngine/Services/Video/SfxMixFilterBuilder.cs) mirrors
SfxMixFilterBuilder's role:
BuildVoiceoverLineBranch(inputIndex, lineIndex, line)emits one line's branch:[{inputIndex}:a]atrim=end={duration},asetpts=N/SR/TB,aformat=sample_rates=48000:channel_layouts=stereo,volume={gain},adelay={ms}|{ms}[vo{k}].asetptsnormalizes the branch to PTS 0 regardless of container start-time weirdness,aformatis mandatory becauseamixrequires matching sample rate and channel layout across inputs, andadelay(integer milliseconds, one value per channel) is the one and only place a line's timing enters the filtergraph — and it was computed server-side, from the step's own recorded start time, never from anything a model wrote.BuildVoiceoverMixStage(baseLabel, lineCount, finalLabel)emits the mix:{base}[vo0][vo1]…amix=inputs={n+1}:duration=first:dropout_transition=0:normalize=0{final}.
Those three amix options are load-bearing, for the same reasons music and SFX spell them out:
normalize=0 (without it amix divides every input's level by the input count, quietly turning the
dialogue down), duration=first (pins the mixed length to the base input, so a line near the end
can never extend the file), and dropout_transition=0 (no gain re-ramp when a line's branch ends
before the base does — which most do).
Labels chain in a fixed order: music first, then voiceover, then SFX. With voiceover present the
pre-voiceover audio becomes [abase] and the voiceover mix stage becomes the new final label (or
[abase] again when SFX is also present, which then mixes into it). Each resolved line is added as
one extra -i input, in exactly the order the branch calls assumed.
When the source has no audio stream and there is no music bed, the base the lines mix into is a
synthesized program-length silence (anullsrc) — narration over silent b-roll used to ship with no
audio track at all. See video-editing.md "Program audio on silent footage".
Where it surfaces
outputSummary["voiceover"]— the node above is copied into the compile step's output summary, so a reviewer can see which lines landed and why others did not.edl["voiceover"]— the same node is written into the compile EDL artifact when voiceover was attempted, alongside the music/graphics/SFX/color-grade nodes.
Narration and the end of the program
Two things used to cut the last narration line of a VideoCompile short, and both are fixed in the
compile step itself (VideoCompileStepExecutor, both encode paths):
- The program audio fade-out faded the narration.
ProgramAudioFadeOutMswas the last audio stage, after the voiceover mix, so a line still being spoken inside the fade window was faded mid-word. Now, whenever a tail stage exists (a program audio fade, a seam audio ramp, or the hold below), the narration is mixed after it: the base program (source dialogue, music, sound effects) goes through the tail into an internal[anar]label, and the voiceover mix stage[anar][vo0]…amix…[aout]runs last. The fade still shapes music, dialogue and effects exactly as before; it never touches a line. With no tail stage nothing moves, so a compile without a program fade is byte-identical. - A line running past the last picture frame was cut off. The voiceover
amixisduration=first(pinned to the program), so a line ending after the picture simply stopped. NowComputeNarrationHoldSecmeasures how far the last line's end, plusNarrationTailSec(0.25 s), runs past the program — a generated cold open in front of the body moves every body-timed line by its own length — and the compile holds the last picture frame for exactly that long:tpad=stop_mode=clone:stop_duration={hold}as the first stage of the picture tail (before the program fade, after censoring, so the held frame is the finished picture; captions are drawn after it), andapad=whole_dur={program+hold}as the first stage of the audio tail, so the base the narration is mixed over lasts as long as the picture. The video and audio program fades are re-resolved against the lengthened program, so the fade-out still finishes on the last frame — now a quarter second after the line ends.
A hold of 0 (every line fits) changes nothing. When there is a hold:
outputSummary.outputDurationSecis the lengthened program;outputSummary.voiceover.extendedForNarrationSecand, when a program fade is configured,outputSummary.programFade.extendedForNarrationSecreport the hold in seconds (both also in the EDL);- the step reports "Holding the last frame … so the narration can finish" while it runs.
The held frame is the program's own last frame, inserts and overlays included — the hold is applied after every picture stage except captions. Music is not extended: under a hold it ends where the picture would have, its own fade-out included, and the line finishes over silence (or over the held source audio's padding). There is no configuration member for any of this: cutting a line off was never a choice anyone wanted.
The promo pipeline path (ScriptScenes)
For the existing promo pipeline (Scriptwriter → Director → Author), narration reaches the
video without any new template or agent. The chain is:
- A
StepType.Voiceoverstep withSource.Kind = ScriptScenesandScriptwriterStepOrderpointing at the pipeline'sScriptwriterAgentstep. The executor reads each scene'sVoiceovertext andStartTime, synthesizes one WAV per narrated scene into the content-addressed keyvoiceover/lines/{sha256hex}.wav, and records each line's probed duration and its fullstorageKeyin the step output — that key, not a reconstructed path, is what the staging side reads. AuthorAgentcallsStageVoiceoverAudio(its prompt instructs it to call the tool once, step 14).ReactRemotionSandboxTools.StageVoiceoverAudiofinds the latest CompletedStepType.Voiceoverresult and the latest CompletedScriptwriterAgentresult for the same execution, walks the scenes in order, skips whitespace-only voiceovers, skips lines whose status is not"ok", and copies each remaining line's WAV into the sandbox atvoiceover/scene-{n}.wav— 0-indexed by scene position, so the mapping between scene index and voiceover-line index survives empty scenes. Each line's object is resolved from the emittedstorageKey, with the legacy execution-scopedvoiceover/{executionId}/{lineId}.wavused only as a fallback when a step result carries none. It returns{ "staged": ["voiceover/scene-0.wav", …], "failed": [] }.- The
failedarray is additive, andstagedis unchanged in both shape and contents, soAuthorAgent's existing parsing of the result is unaffected. An entry is{ lineId, scenePosition, reason: "object_not_found" }, added when the line's object is genuinely missing from the bucket.NoSuchKey/NotFoundis no longer swallowed by the generic catch — that catch previously logged at Error next to every other failure while the tool still answered{"staged":[]}, so a promo rendered without narration and nothing in the tool result said why. The missing-object path is now a Warning that names the full key, the line id and the scene position, and reports itself to the caller. AuthorAgentthen passesvoiceoverSrc="voiceover/scene-{n}.wav"for exactly the scene positions that appear instaged, and omits the prop for every other scene. An emptystagedarray is the normal case for a promo with no narration, and the prompt says to omit the prop everywhere then.sandbox/template/src/root.tsximplements the receiving half: each ofSceneWordmark,ScenePillsandSceneCtatakes an optionalvoiceoverSrc?: stringand renders{voiceoverSrc && <Audio src={staticFile(voiceoverSrc)} />}; theLaunchPromocomposition exposesvoiceoverSrc0/voiceoverSrc1/voiceoverSrc2and passes one per<Sequence>. The file's own comment documents thevoiceover/scene-{n}.wavconvention and points here.
The tool lives in ToolGroup.SandboxAuthoring (AgentToolProvider.cs), so it is reachable by
exactly the agents that already have sandbox authoring — AuthorAgent among them — and by nobody
else.
Text sanitization
Every line's text is passed through VoiceoverTextSanitizer.Sanitize(text, int.MaxValue) before the
byte budget is measured and before anything is sent to Fish Audio. The executor still passes
int.MaxValue as the length cap: MaxBytes is the real budget, and silently truncating narration
by byte count would produce audio that no longer matches the script.
The pass order is: cap the raw input at 4096 → strip Unicode Cc control and Cf format
characters → strip [...] (non-greedy) → strip (...) (non-greedy) → strip https?://\S+ URLs →
collapse runs of whitespace → trim → truncate by Unicode text element (grapheme cluster, via
StringInfo) so a surrogate pair or a combining mark is never split mid-character.
Three guarantees on that order are load-bearing, and each one is enforced by
VoiceoverTextSanitizerTests:
- The raw input is capped at
VoiceoverTextSanitizer.HardMaxRawChars= 4096 characters BEFORE any regex runs. The bracket patterns are non-greedy and backtracking, so on a line with many unmatched open brackets and no closer they cost O(n²) — a 1.28M-character line of[measured 28.5 s of CPU on the engine shared by every tenant. The cap keeps those patterns off an unbounded string in the first place; a match timeout alone would only have turned a slow line into a thrownRegexMatchTimeoutException, i.e. a failed step. The consequence, stated honestly: a line longer than 4096 characters is silently truncated to 4096, and the remainder is never spoken. 4096 UTF-16 code units is roughly 700 words — several minutes of continuous speech, for a line whose whole point is to be anchored to a start time — so no legitimate line approaches it, but the truncation is silent by design, andMaxBytesis measured against what survives it. - Control (
\p{Cc}) and format (\p{Cf}) characters are stripped BEFORE the bracket tags. This ordering was the bug. When the bracket patterns ran first,[whis\npering]was untouched by them (.does not match\n), and the control-character strip then simply deleted the newline — reconstituting the literal tag[whispering]in the text handed to Fish Audio. Stripping first yields[whispering], which the bracket patterns then remove normally.Cfis a separate category and invisible by definition: U+200B, U+202E, U+FEFF, U+00AD, U+2066 and U+2069 all used to reach the TTS request verbatim. - Both bracket patterns are
RegexOptions.Singleline, and every staticRegexin the class carries a finite 100 ms match timeout.Singlelinemakes.match\n, so a tag split by a literal newline is still matched as one tag — defense in depth over the ordering fix, not a substitute for it. The timeout is finite on every pattern, because aRegexconstructed without one silently gets the infinite timeout back, which is how the super-linear behaviour above went unnoticed.
Both bracket forms are stripped because both are meaningful to Fish Audio and both would otherwise
be spoken when they came from model- or media-derived text: [...] is the S2-family /
drama-3-preview emotion syntax, (...) is s1's fixed-vocabulary equivalent. Stripping only the
square form — which is what the feature request originally specified — would have left (excited)
audible on any s1 provider row. Since phase 1 offers no emotion control at all, no legitimate
config can be harmed by removing them.
Per-line artifact resolution
The producer and both consumers agree on one key per line, and that agreement is what makes a
cached or replayed Voiceover step work at all:
- The producer (
VoiceoverStepExecutor) uploads each line under the content-addressed relative pathvoiceover/lines/{sha256hex}.wav, where the hash is lowercase SHA-256 over the six-field material described above, and records the full key —projects/{projectId}/agentFiles/voiceover/lines/{sha256hex}.wav— on the line in its output asstorageKey. That property is omitted, not null, onskipped/failedlines, which is exactly the signal a consumer uses to tell "this line has an artifact" from "this line does not". VideoCompileStepExecutor.ResolveVoiceoverAsyncreads each line's ownstorageKeyand downloads it. It does not list project files: noProjectWorkspaceFilerow is ever created for a bare artifact, so aListFilesAsyncmatch on a file whose key contains the line id could never find one, dropped every line asline_file_not_found, and still reported the compile Completed. Only when a step result carries nostorageKey— i.e. output persisted before content-addressing — does it rebuild the legacy execution-scopedvoiceover/{executionId}/{lineId}.wav, which is the only object such a result can mean.ReactRemotionSandboxTools.StageVoiceoverAudioresolves each scene's line the same way: the emittedstorageKeyfirst, the legacy execution-scoped key only as a fallback.
Because the content-key formula changed, per-line WAVs written under the old formula are now
ordinary cache misses: nothing migrates them, so the next run re-synthesizes those lines once and
uploads them under their new keys. That is a one-time vendor cost, paid per line, after which the
objects are content-addressed as described above. The legacy execution-scoped fallback in the two
bullets above is unchanged — it is still the only key a storageKey-less step result can mean.
Four properties of this contract are worth remembering before changing anything about it. First, the
key contains no execution id, which is why the step is safe to cache and why a cache hit replayed
in a later execution still names objects that exist. Second, the key is derived from what was
actually synthesized — effectiveVoiceId, not the config's raw DefaultVoiceId — so two configs
that differ only in default voice cannot collide on one object. Third, its value fields are
length-prefixed rather than joined by a delimiter, so the key is injective by construction: two
distinct field tuples can never share one object. Fourth, it covers the project's provider identity
and the model actually sent, so re-pointing the step at another provider row — or editing that
row's model — produces different objects instead of silently reusing another configuration's audio.
Admin surfaces
InferenceProvidersControllercarries the kind/capability pairing rules described above, theSpeechSynthesismember of the accepted-capability list, and aSpeechSynthesisbranch in both test endpoints.RunSpeechSynthesisTestAsyncsynthesizes a one-word"ok"request through the saved (or unsaved) provider and validates the result by parsing the WAV RIFF header directly — no ffmpeg, no new dependency — treating "audio with zero duration" as a failed test. Like every other test branch it never lets a bad endpoint, model or key escape as a 500.web/lib/types/inference-provider.tsmirrors the enums, including theFishAudiokind with itsSpeechSynthesis-only note and a pointer to this document.InferenceProviderFormexposesKindandCapabilityas fields and renders the Fish Audio licence notice for aFishAudiorow.- The workflow builder ships a
VoiceoverStepConfigeditor (Inline line list with per-line text/start/voice; a Scriptwriter-step picker forScriptScenes; a NarrationWriter-step picker forNarrationPlan; provider id, model, default voice, byte budget), aVoiceoverNodefor the flowchart, and a JSON schema inweb/lib/schemas/workflow-step-configs.tswhoseVoiceoverSourcedefinition pinskind: ['Inline', 'ScriptScenes', 'NarrationPlan']and addsnarrationWriterStepOrder(integerornull, minimum 1). ANarrationPlansource with no step chosen is reported as unconfigured rather than merely empty —getVoiceoverSourceError/NARRATION_PLAN_STEP_REQUIRED_ERRORinweb/lib/utils/voiceover-validation.ts— so the builder says so instead of the step failing at execution time. Step orders are 1-based, so0counts as unconfigured too.
The two halves of that kind now agree, and it is worth being precise about where, because for one
commit window they did not. The frontend mirror (frontend 029b8998) accepts and validates
NarrationPlan; the engine half is committed too (backend ba1160f1) — the enum member, the config
member, the executor arm and the save-time arm below — so a config the builder accepts is a config
the engine can execute. The validation is deliberately duplicated on both sides rather than left to
one: the builder blocks early for the author's sake, and the save-time validator plus
NARRATION_SOURCE_INVALID remain the authority, since the API is reachable without the builder.
Provider-network hardening
A provider row's Endpoint is operator-supplied, which makes every path that dials it an SSRF
surface. These are the checks that make such an endpoint safe to call.
IsDisallowedEndpoint gates Create, Update, and the unsaved POST /test path. The third is the
one that was missing: "test an unsaved config" finalized its endpoint — from the caller, or inherited
from the saved row when the stored key is reused — and went straight to a client, so the one test
path that persists nothing was also the one where the server could be aimed at 169.254.169.254,
127.0.0.1 or a Docker bridge address and make the request on the caller's behalf. The check now
sits after that id-inheritance has settled the endpoint and before any client is constructed, so
a rejected endpoint never reaches a factory at all. (POST /{id}/test does not re-run it — that
endpoint comes from a row that already cleared the same gate when it was written.)
What the predicate refuses is an address, not a string. It rejects anything that is not an absolute
http/https URL, anything whose host cannot be resolved, and any resolution that yields a
non-public address. The blocklist covers loopback (127/8, ::1), the unspecified addresses
(0.0.0.0, ::), link-local (169.254/16, fe80::/10), the RFC 1918 ranges (10/8, 172.16/12,
192.168/16), CGNAT shared space (100.64/10), IPv6 unique-local (fc00::/7) and the deprecated
site-local fec0::/10. ::ffff:127.0.0.1 is classified as its IPv4 self rather than being waved
through by the IPv6 branch, and an address in a family the code does not understand fails closed. An
IPv6 literal arrives from Uri.Host still wearing its brackets, which IPAddress.TryParse will not
accept, so the brackets are stripped before parsing — otherwise a literal would fall through to DNS,
fail to resolve, and be rejected for entirely the wrong reason (a public IPv6 literal wrongly with
it).
A hostname is checked against every address it resolves to, not just the first
(IsDisallowedAddressSet). A name carrying one public and one private record — split-horizon DNS, or
a deliberately poisoned entry — has to be refused, because which record a given connect actually uses
is not something the caller controls.
Two client-side limits bound what a hostile endpoint can do once a request is under way.
SpeechSynthesisClientFactory.BuildFishAudio builds its SocketsHttpHandler with
AllowAutoRedirect = false: the endpoint was validated once, at write time, and a 3xx from it would
hand the request to a second host that was never validated at all — the "point the vendor at
169.254.169.254 and let the vendor's own client fetch it" move. Refusing the redirect turns such a
response into an ordinary non-success status the caller can report. On the response side,
FishAudioSpeechSynthesisClient buffers a success body behind MaxAudioResponseBytes (32 MiB,
roughly nine minutes of 48 kHz 16-bit mono WAV), checking both the declared Content-Length and the
bytes actually arriving — a hostile endpoint can simply omit or understate the former, so the
streaming check is what makes the cap real, and the size message names the limit without carrying any
part of the payload. An error body is read only up to MaxErrorBodyBytes (4 KiB), with the remainder
left unread on the wire, and the text that reaches the exception message is truncated independently,
since a JSON message field is exactly as attacker-controlled as the bytes around it and is
interpolated into a logged HttpRequestException.
This is the speech-synthesis path only. Whether the sibling chat, transcription, vision and video-generation factories set redirect or read limits of their own was not examined, and is not asserted either way.
How a Voiceover step is created and persisted
voiceoverConfigJson is a first-class field on CreateWorkflowStepRequest and
WorkflowStepResponse (Controllers/Dto/ProjectDtos.cs), not a blob smuggled through some other
field. The WorkflowsController read paths project it, and both WorkflowEditDiffService — which
diffs it, so editing a Voiceover step's config is a real, reviewable diff — and
WorkflowTemplateProvisioningService carry it exactly like every sibling config.
StepConfigSaveValidator now has a StepType.Voiceover arm, so a config-bearing Voiceover step
with no config is rejected at save time instead of failing mid-run with CONFIG_INVALID.
That arm is also domain-aware for exactly one kind: a config that deserializes to
VoiceoverSourceKind.NarrationPlan is checked further, and a missing, null, zero or negative
narrationWriterStepOrder is rejected at save time with
Step 4 (Voiceover): VoiceoverConfigJson source narrationWriterStepOrder is required and must be at
least 1 when the source kind is NarrationPlan.
The same reasoning as the missing-config check: the executor hard-fails exactly that config with
NARRATION_SOURCE_INVALID, so saving it only defers the failure to a real execution run. The arm is
deliberately additive — Inline and a ScriptScenes config (including one whose
ScriptwriterStepOrder is absent, which the executor rejects as SCRIPT_SOURCE_INVALID) stay
savable exactly as they were, and an unrecognised kind string is left to the executor's
SOURCE_INVALID rather than becoming a new save-time rejection. The message goes through the shared
step-error prefix, because that prefix is what every save path renders into one 400 body.
On the frontend, a newly added Voiceover step is assigned the VideoTransform placeholder agent in
both the add path (FlowchartBuilder.handleAddStep) and the change path (StepCard) — without that
it could not be saved at all, since WorkflowStep.AgentDefinitionId is non-nullable. Its editor is
reachable from the flowchart node: clicking the node selects it and opens the StepConfigPanel
drawer, whose Voiceover branch renders the type-specific VoiceoverStepConfig editor. The node
itself is keyboard reachable — role="button", tabIndex={0}, an aria-label, and an
Enter/Space key handler forwarding to the click.
The editor warns before a run rather than after it. VoiceProviderStatusNotice
(web/components/workflows/) calls getProviderStatus('SpeechSynthesis')
(GET /api/v1/inference-providers/status?capability=SpeechSynthesis, readable by every signed-in
user and carrying no endpoint or key) and shows a red Narration will probably fail box when
healthy is false, or a yellow Using a backup voice service box when usingFallback is true,
preferring the server's own message. It renders nothing while loading, on an error (an older API
without the endpoint) or for a healthy provider. The Voice section links to the project's
?tab=voices page; model, default voice id, byte budget, pronunciation and casting sit behind the
shared AdvancedSection disclosure (config-kit.tsx), which opens itself when the assistant's
proposal or a rejected save names a field inside it.
Phase 1 scope, and what phase 2 delivers
Phase 1 shipped the spine: a deterministic Voiceover step, content-addressed per-line audio, a
soft-failing compile mix path, and two source kinds. It was deliberately narrow. Phase 2 was scoped
to add four things on top of that spine, and all four are now in the wave's committed branches:
the planning agent and its id-anchored schema, word alignment over the synthesized lines, the fit
measurement — which landed as two halves, the deterministic NarrationFitSolver on the producing
step and the picture-window facts on the compile — and the NarrationPlan source kind plus the
template that wires the whole narration path together. This section names each with the branch and
commit it was verified against, and the next section covers what phases 3 and 4 added on top.
Delivered by phase 2
-
NarrationWriter— phase 1 narrated text a human wrote (Inline) or text an existingScriptwriterAgentstep had already produced (ScriptScenes); nothing turned a picture edit into narration. Phase 2 adds the agent and its id-anchored schema — see Narration planning. Committed onaiforge/voiceover-complete/backend(3e1f1c3e). -
Word alignment — nothing mapped synthesized words back onto audio. Phase 2 records per-line word timings by re-transcribing each synthesized WAV through
ITranscriptionClient, and degrades to absent words rather than to an error — see Word alignment. Committed onaiforge/voiceover-complete/backend(54df4d0d). Phase 1 noted that Fish Audio exposes word-level timestamps throughPOST /v1/tts/stream/with-timestamp(SSE) andwss://api.fish.audio/v1/tts/live/with-timestamp, and guessed they would be "simpler and cheaper" than a re-transcription round trip. That guess does not hold for this deployment: the self-hostedfish-ttsreturns404onPOST /v1/tts/stream/with-timestamp, which is why phase 2 aligns through the transcription client and why neither Fish endpoint is consumed. -
The picture-window fit facts — nothing measured narration against the picture; durations were probed and reported, and no number reached the compile's output. Phase 2 measures every resolved line against the shot it starts in, reports
overflowSec/fitsper line plus the node-levelvoiceover.fitroll-up, and carries that roll-up into the compile step's own output JSON — so aConditionalorReviewLoopstep downstream of the compile can gate onoverflowLineCount/maxOverflowSec. It is deterministic code, never an agent — no model computes a duration anywhere on this path — see The fit solver and the review loop. Committed onaiforge/voiceover-complete/media(3a74fa4f, keys renamed in349a773c, tests ind403c7a0). -
The
NarrationFitSolver— the same question asked in words, on the producing step:NarrationFitSolver(ReelBolt.Shared/Workflows/NarrationFitSolver.cs) measures every line against its declared narration window and emits per-linefitplus themeta.fitreport and the promotedmeta.overflowWords/meta.overflowLineCount— see The declared window on the producing step. Deterministic code, never an agent: no clock, no RNG, no I/O, no ffmpeg. Committed onaiforge/voiceover-complete/backend(8cfcce4f, T-203). -
The
NarrationPlansource kind — the piece that lets aVoiceoverstep actually consume aNarrationWriterstep's plan:VoiceoverSourceKind.NarrationPlanappended last,VoiceoverSource.NarrationWriterStepOrder,VoiceoverLine.AnchorId, the executor arm with its three narration failure codes, theNarrationPlansave-time validation arm, and the placement pass that declares the line starts — seeNarrationPlan. Committed onaiforge/voiceover-complete/backend(ba1160f1, T-204). On the frontend,aiforge/voiceover-complete/frontend(029b8998) has the matching builder, schema and client-side validation. -
The reviewer's narration clause and its deterministic cap — the compile node carried
voiceover.fitand the voiceover step carriedmeta.fit, but nothing downstream read them:VideoReviewAgent's prompt listed only the pre-wave facts, so a loop that had to react to overflow had to gate on the step output itself. Phase 2 adds avoiceoverclause to the prompt (and its byte-identical seeder copy) andReviewLoopStepExecutor.ApplyNarrationFitCap—min(score, 4)on a measured positive overflow, inert otherwise — so the fact is acted on whether or not the model heeds the instruction. See How the overflow reaches theReviewLoop. Committed onaiforge/voiceover-complete/backend(ac96479c). -
The empty-narration contract — a well-formed
NarrationPlancarrying zero lines used to be aNO_LINESfailure, so aNarrationWriterdoing exactly what its own prompt instructs reddened a correct run and, after retries, failed the execution. It is now a successful, emptyVoiceoverstep, marked bymeta.emptyReason: "narration_plan_empty".NO_LINESdeliberately remains a failure forInlineandScriptScenes, where an empty resolved list is an author misconfiguration rather than a decision — see An empty plan is a successful step. The same commit gives{"lines": null}and anullentry insidelinestheir ownNARRATION_OUTPUT_INVALIDinstead of escaping asUNEXPECTED_ERROR. Committed onaiforge/voiceover-complete/backend(caf937c0). -
The declared-and-lost narration cap — the overflow cap reacts only to a measured overrun, and a compile that applied zero narration lines has no measurable fit at all, so a green review score advanced the loop over a narration-less render.
ApplyNarrationDeclaredLostCap(min(score, 4)) now caps exactly that shape — theamix-unavailable arm included, since a narration-less render is a defect whichever side caused it — and is inert only for a compile with novoiceovernode and for the empty plan.NarrationCapsWouldFiremakes the calibrated decision gate's early stop escalate rather than accept when either cap applies — see How the overflow reaches theReviewLoop. Committed onaiforge/voiceover-complete/backend(caf937c0). -
The declared count on the compile's node, and the raw window lookup —
lineCountwas a copy ofappliedLineCount, which made the key pair the cap is keyed on unsatisfiable on the very path it was written for (all_lines_unavailablereported0/0). It now reports the declared count, so both cap forms are live; the picture-window containment lookup also moved onto the rawMapToOutputSecsecond, which stops a line that rounds onto its shot's exclusive end from being dropped out ofmeasuredLineCountand leavesunknownWindowLineIdsempty by construction. No key renamed or reordered; on every path where declared equals resolved the emitted JSON is unchanged. Committed onaiforge/voiceover-complete/media(0924b7cc,6a62f2fc). -
The
video-derush-edit-voiceovertemplate — the twelfth entry inWorkflowTemplateCatalog.csand the first catalog entry that wires narration end to end, in six steps:VideoAnalyze(Source: ProjectFile) →Agent(VideoStoryEditor)→Agent(NarrationWriter)(AgentInputContextMode: FullWorkflow) →Voiceover({"source":{"kind":"NarrationPlan","narrationWriterStepOrder":3}}) →VideoCompile(enableVoiceover: true,voiceoverStepOrder: 4,decision: Step 2,analysisStepOrder: 1) →ReviewLoop(VideoReviewAgent)looping back to step 2.AutoCreateOnProject: false, like every other derush template. Committed onaiforge/voiceover-complete/backend(ba1160f1, T-204).Why six steps and not five. The catalogue makes
VideoCompileStepConfig.Decisiona required positional member, so a template with noVideoStoryEditorstep before its compile step cannot build at all — the narration steps are inserted between the story editor and the compile, not in place of either. TheDecision/Voiceoverrefs are explicitStepOrders rather thanPreviousfor the same reason every sibling template spells them out:Previousrelative to the compile step would resolve to the voiceover step's output (step 4), not the story editor's decision (step 2).Modeis left at itsReencodedefault, whichEnableVoiceoverrequires.
Phases 3 and 4: what landed after narration planning
Phase 2 (Narration planning through The fit solver and the review loop) made narration a decision an agent takes and a measurement code makes. Phases 3 and 4 build on that spine without changing its contract: every addition below is an opt-in whose default is byte-identical to the phase-2 behaviour, every model-authored output stays id-anchored and number-free, and every soft stage degrades with a recorded reason rather than failing a compile. Each subsection names the code that owns it.
Captions
VideoCompileStepConfig.EnableCaptions (default false, byte-identical off, requires
Mode = Reencode — CAPTIONS_REQUIRE_REENCODE otherwise) burns the applied narration into the
picture during the same encode. CaptionStyle is Subtitle (a chunk of text per cue, split at word
boundaries at CaptionMaxCharsPerCue), Karaoke (the same chunks, with the words already spoken
redrawn left-anchored in CaptionHighlightColor as each word begins) or Punch (one word at a time,
larger and centred).
Two facts about where the words come from are load-bearing:
- The text is the Voiceover step's own sanitized, actually-synthesized text — the
textkey everyokline now carries — never the narration plan's raw text. If the sanitizer dropped a[tag]or a URL, the caption does not say it either; a caption that says words the narrator never spoke is a correctness bug, not a cosmetic one. - The timing is the measured word alignment (
words[]on the line, from Word alignment), mapped onto the output timeline from the line's resolved start.Subtitleworks without words too — each chunk then owns a share of the line's play duration proportional to its characters — butKaraokeandPunchneed words and degrade per line withno_word_alignmentwhen there are none. A deployment with no Transcription provider therefore gets Subtitle captions and no Karaoke, and the node says so.
The cues are cut by CaptionCueBuilder (WorkflowEngine/Services/Video/CaptionFilterBuilder.cs)
and burned through libass: AssSubtitleBuilder (WorkflowEngine/Services/Video/AssSubtitleBuilder.cs)
writes one ASS script, captions.ass, into the step's scratch space at encode time — laid out for the
canvas actually being encoded, PlayResX/PlayResY equal to its size — and the picture stage is a
single ass=filename='…':fontsdir='…' filter. Captions are the last picture stage of all — after
censoring and the program fade — so a caption is never dipped or faded with the picture it annotates.
Both encode paths (single-source and segmented) carry the stage.
The text never enters the filter string, the same discipline as drawtext's textfile= +
expansion=none: the filter names two paths and nothing else, each escaped with
EscapeFilterPath inside single quotes. Inside the script, spoken text goes through
AssSubtitleBuilder.EscapeText, because libass reads override syntax anywhere in event text: {/}
become parentheses (no override block can open), a word joiner (U+2060, invisible) follows every
backslash (no \N/\n/\h can form) and every control character becomes a space (a newline would
end the Dialogue: line and let the rest parse as a new script line). Before that, the text passes
CaptionTextSanitizer — the overlay allowlist plus the punctuation narration needs (commas,
apostrophes, quotes, colons, semicolons, an ellipsis), which the overlay sanitizer strips. Golden tests
pin all of it: AssSubtitleBuilderTests.
Presets. Every number in the script is computed in C# from the canvas, the style and the config:
| Style | Look | Automatic font | Automatic size (% of the canvas's short side, landscape / portrait) | Automatic position |
|---|---|---|---|---|
Subtitle | One or two lines on a box (BorderStyle 3, the box in CaptionBoxColor); with box none, outlined text plus a soft shadow | Inter Bold | 5.5% / 6.5% | Bottom |
Karaoke | The same chunks as ONE event each, every word carrying a \k duration equal to the gap between its measured start and the next word's, so libass flips it to CaptionHighlightColor exactly when it is spoken | Montserrat Bold | 6.5% / 8% | Bottom |
Punch | One word at a time, centred, heavy outline and shadow, never a box | Anton | 11% / 13% | Centre |
Bold | Up to two words (≤ 12 characters) a cue, heavy outline and shadow, never a box, no highlight — the Reels/TikTok look; word timings used when present, else shared by length like Subtitle | Poppins | 10.5% / 14% | Bottom, at 28% of the height on a portrait canvas (10% landscape) — above the caption, handle and buttons the platforms draw |
On a 1920x1080 canvas that is 59 / 70 / 119 px; on 1080x1920 it is 70 / 86 / 140 px. A cue that would
need more than two lines (one for Punch) inside the safe width is shrunk for that cue alone with an
\fs override — libass wraps only at spaces, so a long single Punch word would otherwise run off a
9:16 frame. Portrait means taller than wide.
Size. CaptionFontSizePct 0 means automatic, and so does 4: that was the default every config
saved before the automatic size existed carries, and it rendered ~4%-of-height captions that were too
small to read on a phone. Any other value keeps its old meaning — a percentage of the frame height,
clamped 2..12, Punch doubled within the clamp. The step output says which applied (fontSize: "auto"
or "explicit").
Position and safe margins. CaptionPosition (Auto | Bottom | Center | Top, appended after
TargetDurationSec) maps to ASS alignment 2/5/8. Side margins are 8% of the width. A landscape canvas
keeps 8% clear at the top and bottom; a portrait canvas keeps the bottom 20% and the top 12% clear,
where Reels/TikTok/Shorts draw their own caption, buttons and header.
Fonts. CaptionFont (Auto | Inter | Montserrat | Poppins | BebasNeue | Anton |
DejaVuSans) is a closed enum — never a path or a family string. The files are vendored in
inference/fonts/captions (SIL Open Font License 1.1, provenance and hashes in its UPSTREAM.md), and
the WorkflowEngine Dockerfile copies them to /usr/share/fonts/reelbolt-captions, the default of
VideoEditing:CaptionFontsDir. CaptionFontCatalog resolves a font against that directory: a missing
file (or a missing directory) falls back to DejaVu Sans and records fontFallback: "font_not_installed"
rather than letting fontconfig silently substitute something else. Single-weight display faces (Anton,
Bebas Neue) are never asked for bold, so libass never smears a synthetic bold over them.
Colours are validated by CaptionColor: white, black, yellow or #RRGGBB, each optionally
@opacity (0..1) — e.g. [email protected]; CaptionBoxColor also takes none. Anything else falls back to
that field's default and is listed in colorFallbacks. Before this, caption colours reached the
drawtext filter string unvalidated; they now pass the same check on both renderers.
The drawtext degrade. When the ffmpeg build has no ass filter (probed once per process from
ffmpeg -filters, exactly like the drawtext probe), captions fall back to the older drawbox/drawtext
chain (CaptionFilterBuilder.BuildFilterChain: per-cue textfile= scratch files, half-open
gte(t,a)*lt(t,b) windows, Karaoke as a highlighted prefix redrawn over the base text), still at the
automatic size when the size is automatic, and the node records renderer: "drawtext" plus
degradeReason: "libass_unavailable". Alpine's ffmpeg is built with --enable-libass, so the shipped
image takes the libass path.
The compile node is captions: {enabled, applied, style, renderer, degradeReason?, font, fontFallback?, position, fontSize, colorFallbacks?, cueCount, lineCount, cues[], dropped[], reason?}, present only
when EnableCaptions is on. Degrade reasons: voiceover_not_enabled, no_applied_voiceover_lines,
drawtext_unavailable (neither renderer exists), no_word_alignment, no_cues;
max_caption_cues_exceeded truncates rather than drops (MaxCaptionCues, default 400). The tests are
CaptionCompileTests (the EnableCaptions=false filtergraph is the phase-2 voiceover graph, byte for
byte; the libass graph and script; the drawtext degrade; the colour allowlist), AssSubtitleBuilderTests
(golden scripts per preset, escaping, layout, fonts, and a real-ffmpeg render that runs when ffmpeg has
libass) and CaptionCueBuilderTests. The dashboard's compile editor shows each font as a live sample,
served from web/public/fonts/captions as small WOFF2 subsets — no font CDN.
Keeping captions clear of the picture's own text
With CaptionPosition = Auto, Subtitle and Karaoke captions sit in the lower third — which, on a
screen recording or a tracked app, is where the app's own titles are: a narrated launch reel put
"No timeline wrangling" straight over the UI's "Timeline editor" label. Automatic placement now
checks each kept shot that carries a caption (CaptionPlacementProbe): one frame from the middle
of the shot, fitted to the canvas exactly as the encode fits it (crop focus included), is
edge-mapped at a fixed 640 px width (edgedetect), and the mean of the edge map is read in the band
the caption would cover — two lines at the layout's size, above the bottom margin — and in the same
band under the top margin. When the bottom band reads at least 8 and at least twice the top, the
captions over that shot are drawn top-centre at the top title-safe margin instead ({\an8} and the
event's own MarginV; every other cue keeps the style untouched).
Measured on the launch reel's sources (bottom vs top): a UI timeline 12.9 vs 3.1 and a monitor over
a keyboard 16.8 vs 6.1 move up; a desk shot 0.2 vs 8.3 and a phone held in front of a face 8.0 vs
6.6 stay down. Two small ffmpeg calls per captioned shot; a frame that cannot be measured leaves its
captions at the bottom. Punch (centred) and an explicit Bottom, Center or Top are never moved.
The captions node reports movedToTopCount (cues) when any moved.
Over a DESIGNED composition (a source the analysis marks IsComposition, i.e. a Remotion render
with its own typography) the probe reads three frames (a quarter, half and three quarters through
the shot) and keeps each band's busiest value, since such a render animates its type in and out.
When both bands are busy (CaptionPlacementProbe.BothBusy: each at least the busy threshold), the
cues over that shot are NOT drawn — two layers of text on one frame is unreadable, and a render that
burned its own subtitles over the narration showed exactly that — while the narration still plays;
captions.hiddenOverDesignedTextCount reports how many. Footage keeps the single middle frame and
is never hidden.
Pronunciation dictionary
VoiceoverStepConfig.PronunciationDictionary (Off default, Auto) sends Fish Audio a pronunciation
dictionary with each line. Three rules keep it deterministic and first-party
(Shared/Workflows/PronunciationDictionaryBuilder.cs):
- Candidate terms come from the code analysis, never from prose: dependency names, framework,
project type and component names read leniently from any
DependencyAnalysisOutput/CodeStructureOutput/ComponentInventoryOutputin the run's history. "go" in ordinary narration is never read as Golang unless the analysis named it. - Phonemes come from a fixed table (
PronunciationDictionaryBuilder.Table, IPA). A term the table does not know contributes nothing — failing closed, because the phoneme alphabet the engine expects is verified only for what the table carries. The table is code, not a config knob, like the colour-grade and SFX tables. - An entry goes only with a line whose text contains the term as a whole word. Fish matches keys
by plain substring, so three-letter acronyms are sent case-sensitively (
APInever rewrites "rapid") and nothing shorter than three characters is sent at all.MaxPronunciationEntries(default 16) drops the entries that appear later in the line, never the whole dictionary.
The wire shape was verified against api.fish.audio/openapi.json on 2026-10-04 and corrects a
phase-1 note: Fish caps the request at 3 dictionaries, each with up to 5,000 items (not "3
entries"); ReelBolt sends one inline dictionary (pronunciation_dictionary: [{items: [{key, value, case_sensitive}]}]), omitted entirely when empty so the phase-1 request body is byte-identical. The
dictionary joins the per-line content key only when non-empty, so lines without one keep their cached
objects. meta.pronunciation records mode, analysisTermCount, entryCount and the keys per line.
The OpenAI-compatible speech contract has no pronunciation field; the entries are ignored there, never
smuggled into the text.
Review facts: voiceover.headroom and voiceover.wer
Both are always present on the compile's voiceover node whenever voiceover was attempted, and carry
applicable: false with a reason (never an absent key, never a throw) when they cannot be computed
(WorkflowEngine/Services/Video/NarrationIntelligibilityFacts.cs):
headroom— per applied line, the line WAV's own mean RMS (power-averaged over 250 ms windows throughWavRmsSampler) minus the analyze step's measuredAudio.RmsDbfsof the shot the line starts over:{applicable, measuredLineCount, meanHeadroomDb, minHeadroomDb, lines[]}. Not measurable when the line bytes are not a 16-bit PCM WAV or the shot carries no audio descriptor. A line whose clip is muted (MuteSourceAudio), has no audio stream, or whose dialogue is replaced (ReplaceDialogue) plays over no dialogue and is left out; when every line is, the node is{applicable: false, reason: "no_dialogue_audio_in_output"}. Measuring the muted take had held muted reels at a review score of 5 for "buried" narration nobody could hear.wer— per line with word alignment, the Levenshtein word error rate of what the recogniser heard against the sanitized text the step asked for, over normalized tokens:{applicable, measuredLineCount, meanWer, maxWer, lines[]}.no_word_alignmentotherwise. A recogniser writes an invented name however it likes, so a reference word of six letters or more matches the recogniser's spelling of it — one token, or two adjacent tokens joined — when they differ by at most a fifth of its letters (AlignSpelling): "ReelBolt" heard as "Real Bolt" is not a mispronunciation (it was scored WER 0.4–0.5 on every launch reel and capped the review), while "Real boat" still is. Short words compare exactly. A word of six letters or more also matches a sound-alike spelling, one whose consonant skeleton is the same once voiced and unvoiced pairs (b/p, d/t, g/k, v/f, z/s) are merged and vowels dropped, with at least four consonants: "Real bold" is how a recogniser writes the spoken "ReelBolt" (r-l-p-l-tboth ways).
VideoReviewAgent's prompt (both byte-identical copies) names both: below roughly 6 dB of headroom
or above roughly 0.25 WER on any line, score no higher than 5 and name the line; the headroom remedy
is a dialogue treatment on the compile step (below), never a rewrite.
Audio-first timing (promo pipeline)
Nothing re-cuts footage to narration — that stays out of scope by design — but the promo pipeline
now cuts its picture to the voice. StageVoiceoverAudio (the Author's staging tool) returns, next
to staged/failed, a timing array: one entry per staged scene with the narration's measured
durationSec and durationInFrames at the fps the entry names. The AuthorAgent and
DirectorAgent prompts (seeder copies byte-identical; the Director's is now pinned by the consistency
test too) set a narrated scene's length from that measurement instead of the script's guess, and
leave scenes with no narration — and promos with no Voiceover step — exactly as before.
Voices and consent
A project has a voice library (project_voices, owned by the Inference API, migration
AddProjectVoices; mapped read-only by the WorkflowEngine). Two sources:
Library— a voice the provider already offers, registered by its vendor id. No person in the project is its subject, so no consent is captured and it never expires.Cloned— a voice built from the project's own recordings through Fish Audio'sPOST /model(IVoiceCloningClient, implemented byFishAudioSpeechSynthesisClient; multipart,visibility: private,train_mode: fast;DELETE /model/{id}to remove it) — or, when theFishAudiorow points at a self-hosted fish-speech server, through that server's reference library (see "Cloning on a self-hosted fish-speech server" below). This is biometric data, and the row carries the facts as columns, not free text:ConsentAttestedByUserId,ConsentAttestedAt, the verbatimConsentStatement,ConsentSubjectName,ConsentRevokedAt/ConsentRevokedByUserId,RetentionUntil, and the referenceproject_filesids.
One read-path rule, ProjectVoice.IsSelectable(now): a cloned voice is usable only while consent
is attested, not revoked, and within retention. The API's selectable flag, the web picker (which
offers only selectable voices) and the engine's IProjectVoiceResolver all apply it, so a
VoiceoverStepConfig.ProjectVoiceId naming an unusable voice fails the step —
VOICE_CONSENT_MISSING / VOICE_NOT_FOUND — and never falls back to DefaultVoiceId: a step that
asked for a cloned voice must not quietly speak in another one. A voice is resolvable only within the
project that owns it.
Endpoints (ProjectVoicesController, owner-scoped like every project resource; a voice in another
project is 404):
| Method | Path | Description |
|---|---|---|
GET | /api/v1/projects/{projectId}/voices | List, with selectable and unselectableReason. |
POST | /api/v1/projects/{projectId}/voices | Register a library voice. |
POST | /api/v1/projects/{projectId}/voices/clone | Clone. consentStatement must equal RequiredConsentStatement verbatim and consentSubjectName must be given, or 400 and nothing is read or sent; recordings must be this project's audio/video files (≤ 20, ≤ 50 MiB); retentionDays 1..730 (default 365). 400 cloning_not_supported for a provider without a cloning API; a self-hosted fish-speech row adds 400 too_many_reference_files (more than one), reference_file_not_wav, transcription_provider_required and reference_has_no_speech, all before anything is sent to it. |
POST | /api/v1/projects/{projectId}/voices/{id}/revoke-consent | Unselectable everywhere from now on; the attestation stays on record. |
DELETE | /api/v1/projects/{projectId}/voices/{id} | Deletes the vendor model (best effort), the reference recordings (object + project_files row) and the row; the response reports each outcome rather than assuming it. |
Retention is enforced, not described: ProjectVoiceRetentionService sweeps every six hours and
deletes cloned voices past RetentionUntil, or revoked more than 30 days ago, the same way DELETE
does. The web Voices tab shows the consent state on every row, requires the attestation checkbox
before a clone, and reports what a delete actually removed.
A local, zero-cost TTS server
OpenAICompatible rows may now serve SpeechSynthesis: OpenAICompatibleSpeechSynthesisClient
speaks OpenAI's POST {base}/audio/speech (verified against the OpenAI OpenAPI specification —
model, input ≤ 4096 chars, voice required, response_format: wav), which the bundled
whisper (speaches) container serves at http://whisper:8000/v1 with the Apache-2.0 kokoro voices.
Register it as a SpeechSynthesis row of kind OpenAICompatible; a blank API key sends no
Authorization header and reads nothing from the environment — a hosted endpoint then answers
401 loudly rather than spending the host's OPENAI_API_KEY. The kind is exempt from the
private-address classification (it already was, for the Transcription case), which is what makes
http://whisper:8000/v1 reachable through the product API — DM-033's narrowest answer. No cloning
on this kind: only Fish exposes one.
Pickups and dubs
Two more things the compile step can do under a narration line, both requiring Mode = Reencode:
VoiceoverMode(OverMusicdefault — the phase-1 mix, byte-identical;DuckDialogue;ReplaceDialogue) decides what happens to the source dialogue under each applied line: a deterministic keyframedvolumeexpression (DialogueGate, ramped byVoiceoverDuckRampMs, depthVoiceoverDuckDb) applied to the dialogue cut before any mix, so neither narration nor music is gated with it. The node records it undervoiceover.dialogue.ReplaceDialogueis the basis for dubs.- Pickups (
EnablePickups+PickupVoiceoverStepOrder):AgentType.PickupPlannernames the transcript segments the presenter flubbed and the corrected sentence for each (PickupPlanOutput— its own schema with its own invariant test, soVideoEditDecisionOutput's rushcut invariant stays untouched); aVoiceoverstep withVoiceoverSourceKind.PickupPlanre-records them (point it at a consented cloned voice); the compile step resolves each line'sanchorIdto thatt{n}segment's source window from the analysis artifact — never from a number the model wrote — maps it through the cut, mutes the original take for exactly that window and plays the pickup from its start. Nodepickups:{lineCount, appliedLineCount, lines[] (with overflowSec), dropped[], lipsync}. Templatevideo-derush-edit-pickups. - Dubs:
AgentType.NarrationTranslatorre-emits aNarrationPlanOutputin the target language named in the run's request — same lines, same anchors, translated words — so the existingNarrationPlanarm speaks it;VoiceoverMode = ReplaceDialogueremoves the original dialogue under every line andEnableCaptionsburns Subtitle captions from the words actually spoken.VoiceoverStepConfig.Languageis recorded on the output as provenance; the speech providers detect the language from the text. Templatevideo-derush-edit-dub.
Lipsync is not available. The #93 lipsync purpose this plan depended on does not exist, so a
pickup or a dub keeps the original take's picture — mouth and all. The compile node states it
(pickups.lipsync: {applied: false, reason: "lipsync_not_available"}) rather than implying a match,
and the templates' descriptions say so in plain words.
Voice casting
VoiceoverStepConfig.VoiceCasting (Off default — asks nothing; Decision) asks the configured
Decision provider one Choice question through the existing decision gate (site VoiceCasting,
recorded in decision_observations like every other gate call): the options are the project's
selectable voice ids, the state is the lines' text plus each candidate's own description — never a
vendor voice id. The answer is used only when the gate accepted it and it names an offered voice;
a declined gate, no provider, no selectable voices or an answer outside the set keeps the configured
voice and says why under meta.voice.casting. A single candidate is cast without a call.
What is still not here
- Re-cutting picture to narration in the derush pipeline. Audio-first timing is the promo pipeline's; for footage the picture is cut by the story editor and narration is measured against it.
- Lipsync for pickups and dubs — see above; stated on the node, not papered over.
- Narration on the starter promo —
quick-win-promoandlean-context-promostill gain narration only by inserting aVoiceoverstep (see the user guide's voiceover recipe); the two narrated promo templates are described under Story-first narration and the narrated templates. - Fish's native word timestamps (
/v1/tts/stream/with-timestamp) stay unused; alignment is the transcription client's, for the reasons under Word alignment. - Pronunciation on non-Fish kinds — the OpenAI speech contract has no dictionary field.
Story-first narration and the narrated templates
A customer's verdict on the derush-and-narrate pipeline was "no story line to follow, random clips".
The structural cause: in video-derush-edit-voiceover the VideoStoryEditor picks spans first and
the NarrationWriter then writes a line per kept clip, so the script describes clips instead of
telling a story. The narrated-story template ("Narrated product story",
WorkflowTemplateCatalog.NarratedStoryTemplateKey) inverts the order without a new agent type:
VideoAnalyze (Source: ProjectFile, Vision: Optional)
→ Agent(NarrationWriter) FullWorkflow — sees only the analysis view and the brief
→ Agent(VideoStoryEditor) FullWorkflow — sees the view AND the narration plan
→ Voiceover {"kind":"NarrationPlan","narrationWriterStepOrder":2}
→ VideoCompile decision: Step 3, voiceoverStepOrder: 4, enableCaptions: true, captionStyle: Karaoke
→ ReviewLoop(VideoReviewAgent) back to step 2 (the story), MinScore 8, MaxIterations 3
Why no new agent. NarrationPlanOutput is already a script that names the picture under each
line (an offered s{n}/g{n}/t{n} anchor plus prose) and carries no time, so it satisfies the
rushcut invariant unchanged. A new AgentType would have meant an enum append, a seeded row, tool
scoping, drift guards and a second schema for the same shape. Instead both prompts gained a section
for this order: the NarrationWriter's "Story first" section (no edit decision in the input
means "write the script from the brief, one story, lines in the order they are heard, each anchored
to the shot that best carries it, not empty"), and the VideoStoryEditor's "When the narration
script came first" section (keep every anchored shot, list Keep spans in the script's order). Both
copies of each prompt stay byte-identical (VideoStoryEditorPromptConsistencyTests).
Ordering is real across clips, not within one. VideoCompileStepExecutor sorts and coalesces
spans only within a run of consecutive same-source spans; spans from different src clips keep the
list order the editor wrote. So the story's beat order maps onto the decision's span order for a
multi-clip project. Two moments of the same clip always play in source order — reordering within
one clip would need the normalisation pass to stop sorting a same-source run, i.e. a compile change
(it currently treats that as spanNormalization.reordered). The narration is placed narration-first
(back-to-back in plan order), which matches a picture cut in the same order; resolving each line
against its anchor shot's own source clip is the compile step's job.
Captions are on in both footage narration templates. video-derush-edit-voiceover and
narrated-story set enableCaptions: true with captionStyle: "Karaoke". Karaoke needs word
timings from a Transcription provider; without one each line degrades with no_word_alignment
and the captions node says so. Both templates, like every template, now set
RequiresUserInput: true so the brief is asked for at run time.
The narrated promo templates. cinematic-feature-spotlight ("Quick vertical reel") and
story-arc-launch-trailer ("Narrated launch trailer") gained a Voiceover step with
{"source":{"kind":"ScriptScenes","scriptwriterStepOrder":N}} directly after the Scriptwriter and
before the Director, so audio-first timing applies and the Author's StageVoiceoverAudio finds the
lines; their review loops now target the Author at its new position (step 9 in both). They fail
with PROVIDER_NOT_RESOLVED on a deployment with no SpeechSynthesis provider, which their
descriptions state. WorkflowTemplateCustomerFacingTests pins all of this, plus catalog-wide facts:
every template is savable through StepConfigSaveValidator, every Voiceover step reads an
earlier step of the agent its source kind expects, and every review loop's target sits inside the
window it re-runs.
Local fish-speech servers and endpoint rules
A FishAudio row may point at a self-hosted fish-speech server (for example
http://host.docker.internal:8180). Like OpenJev, the kind is exempt from the private-address
SSRF classification (admin-only writes; the scheme must still be http(s)). inference and
workflow-engine carry extra_hosts: host.docker.internal:host-gateway so that name resolves.
The client requests sample_rate: 44100 — the cloud API rejects 48000 for WAV ("Supported sample
rates: 8000, 16000, 24000, 32000, 44100"); the local server ignores the field. The admin test button
now surfaces the vendor's own error text (e.g. a cloud 402 Insufficient API credit, an account
balance problem, not a client bug).
Cloning on a self-hosted fish-speech server
A self-hosted fish-speech server has no /model endpoint; it keeps its own reference library
instead (tools/server/views.py): POST /v1/references/add (multipart id, audio, text),
GET /v1/references/list and DELETE /v1/references/delete (JSON {"reference_id": …} — the
handler is registered for DELETE despite the path). Each reference is one directory holding
sample.wav beside a sample.lab transcript, and POST /v1/tts with reference_id conditions
synthesis on that pair — the same request field the hosted API uses, so a voice cloned either way is
spoken by the unchanged synthesis path. Every response is msgpack unless the request sends
Accept: application/json, which the client always does.
Detection. FishAudioSpeechSynthesisClient.DetectCloningApiAsync treats a row whose endpoint host
is api.fish.audio as hosted without a request. Any other endpoint is probed once with
GET /v1/references/list: a 2xx JSON object with a reference_ids array means self-hosted, any other
HTTP answer means the hosted /model API (e.g. a proxy in front of api.fish.audio). The answer is
remembered for the life of the cached client; a probe that could not connect is not remembered.
What changes for a clone. IVoiceCloningClient.GetRequirementsAsync reports the backend's
constraints before any recording is read: the hosted API takes 1–20 recordings of any container and
transcribes them itself; a self-hosted server takes exactly one WAV (it stores the upload as
sample.wav whatever it is) plus its transcript, which it cannot produce. ProjectVoicesController
checks the count first, then the RIFF/WAVE header, then transcribes the recording through the
default Transcription provider (ResolveTranscriptionAsync(null), no per-request override) and
refuses an empty transcript. The reference id is generated (reelbolt- + a GUID, which always
matches the server's ^[a-zA-Z0-9\-_ ]+$), never user-supplied, and becomes the voice's
RemoteVoiceId.
Consent is unchanged. The verbatim statement, subject name, retention, revocation and the single
read-path rule apply exactly as above, and every check runs before the recording is downloaded. The
one new egress is the transcription call: when the default Transcription row is a hosted service,
the reference recording goes there as well as to the fish-speech server. DELETE and the retention
sweep remove the server-side reference through DELETE /v1/references/delete (404 counts as
already gone).
Server-side caveats. References live in the server's references/ directory, so mount it as a
volume or they vanish with the container. A POST /v1/tts naming an unknown reference_id does not
fail: the server creates an empty directory and speaks in its default voice, which is why ReelBolt
never sends an id for a deleted or unselectable voice (the engine's IProjectVoiceResolver fails the
step first). And the server must be able to decode reference audio: a fish-speech image whose
torchaudio is 2.9 or newer without torchcodec installed accepts the upload but answers every
reference-conditioned synthesis with HTTP 500 ("TorchCodec is required for load_with_torchcodec").