Voice and subtitles

Writing a script and generating a voice

Paste a script, one line per beat, and have each line spoken — then let the chart make room for it automatically.

Text becomes a voice. This is one of two directions — the other is recording your own, and a project has one or the other, never both.

Press u or open Settings → Voice & Subtitles.

Writing the script

Paste your script into the panel. One line per beat — roughly one sentence each. Each line becomes one spoken clip and one unit of timing.

Lines are spread across the candles by text length, then stay hand-movable. A long line gets more chart than a short one, which is usually right and always adjustable.

Generating the voice

Pick a voice from the allowlist and generate. Each line is synthesised separately, which is what makes a single typo cheap to fix — only that line is re-synthesised.

Clips are cached against the text that produced them. Edit a line and it shows as stale rather than silently keeping the old audio; recolouring a word with [[word|#hex]] markup does not make it stale, because the markup is stripped before the text is spoken.

Speaking room

The chart makes room for the voice automatically. Holds are derived on every read from the current tempo and the script — never stored — so changing the candle speed recomputes them rather than leaving the video padded for a speed it no longer runs at.

Gap seconds in the panel controls how much breathing room sits between lines. That is independent of the candle tempo: one is how fast the chart moves, the other is how fast you talk.

Holds you placed by hand are merged in, taking whichever is longer where the two overlap — so a pause you set for emphasis is never shortened by the automatic calculation.

Anchoring things to a line

Any element can anchor its timing to a narration line plus an offset. Re-record or regenerate that line and everything anchored after it reflows automatically.

This is the best timing option whenever the project has narration — see how timing works.

Subtitles from the script

Generating subtitles from the script is a one-way copy. Editing the script afterwards leaves your subtitles exactly as they are; fix the wording in the cue list.

That is deliberate. A caption is not always the sentence you spoke — it is often shorter, and re-syncing on every script edit would undo that work silently.

Music

Add a track and set how far it ducks under the voice. What you hear while editing is an approximation; the export mixes it sample-accurate against the same clock as the picture.

Bring your own music

No music ships with the app. Every candidate source failed the "may we redistribute this file inside a product we sell" test, and the reasons are recorded per source rather than re-litigated. Uploads stay in your browser and never reach our servers, and the rights responsibility for what you upload is yours.

When synthesis is unavailable

If the deployment has no speech key configured, the panel still works for writing and timing a script — only the synthesis is disabled. Recording, trimming, the timeline and export all continue to work; only automatic subtitles bow out.