Voice and subtitles

Subtitles

Captions from your recording, from a script, from a file — or placed one at a time by hand.

Captions sit in a fixed band on the screen. The camera can zoom anywhere and the band stays exactly where you put it.

Four ways to get cues

  • From a recording. Transcribed from what you actually said, with word-level timing.
  • From a script. A one-way copy of your narration lines.
  • From a file. Drop in an .srt or .txt. SRT keeps its real timestamps; TXT is estimated from text length.
  • By hand. Press u, click into the chart, type.

Importing a file replaces the whole cue list, so it asks first.

Reading a row

Each row has a start and a duration in seconds.

Seconds, not columns, and that is the important part: a cue measured off a spoken word is pinned to a moment in an audio file. Raising the candle tempo must not slide a caption off its syllable.

Leaving the duration empty estimates it from the text length. A new cue starts at a few seconds; clear the field at any time to go back to the estimate.

Long lines are never truncated

A cue longer than the band's line limit becomes successive cards rather than being cut. Two cards are never on screen at once, and a cue always ends by the time the next one starts.

Editing on the timeline

Open the timeline dock and cues become draggable chips:

GestureEffect
DragMove the cue
Drag an edgeTrim it
Double-clickSplit at the nearest word boundary
MMerge with the next cue
DeleteRemove it

Snap candidates are the take's own word boundaries first, then neighbouring cue edges. Snapping to what was actually said is what lands a caption on a syllable rather than near it.

Script edits do not reflow captions

Generating from the narration script is a one-way copy. Editing the script afterwards leaves your subtitles alone — fix the wording in the cue list.

This is deliberate: a caption is often shorter than the sentence spoken, and re-syncing on every script edit would silently undo that.

They are not chart elements

A cue has no chart position, so it cannot be clicked or dragged on the canvas and does not appear in the element list. The band as a whole is draggable on the canvas — that is a project-level setting, covered under how the band looks.

A transcript is your own speech

Content checks run over generated scripts as a hard gate, but over a transcript only as an advisory. Blocking a caption would mean stopping you from captioning something you already said out loud, which is not ours to do.