Voice and subtitles

Recording your own voice

Record one continuous take against the playing video, then fix a sentence without re-recording the rest.

Your voice becomes the soundtrack, and the captions come out of what you actually said.

A project has either a recording or a written script, never both. Switching asks once and then clears the other.

Recording a take

Hit Record. You get a microphone prompt, then a 3-2-1 count, then the video plays from the start while your microphone runs.

It always starts at second zero

That is not a convenience, it is the design. Because the take begins exactly where the video begins, every word's timestamp is its video time — there is no drift offset to store, and a punch-in can splice by absolute seconds.

The countdown is drawn in the page, never on the canvas. A "3" must not be able to reach an exported frame.

What a take does not do

A take produces no automatic speaking room. You already spoke to a video playing at this tempo; deriving holds afterwards would stretch the picture out from under your own voice.

That is the whole reason a take is one continuous clip and not a list of lines. If you want the chart to wait for you, place pauses first and then record.

Trimming

Drag the ends of the take. The trim is stored as an offset into the source rather than re-encoding the file, which is what keeps it a two-way door — you can always drag it back out.

Punch-in: fixing one sentence

Shift-drag on the timeline to mark the piece you want to replace. You get a two-second pre-roll, then capture starts exactly at the mark.

The replacement is crossfaded in over about 30 milliseconds at each seam, overlapping rather than appending — which is what keeps the resulting shift exact to the frame, so everything after the splice stays where it belongs.

Only the replaced piece is re-transcribed. A caption straddling a seam is trimmed back to that seam rather than stretched over audio nobody spoke.

The audio is committed before transcription, so a failed transcription costs you captions, never the performance.

Uploading instead

Upload is a second door into the same take. A file you drop in runs through the same transcription and behaves identically from there.

Music underneath

Add a track and set how far it ducks under your voice. What you hear while editing is an approximation; the export mixes it against the same clock as the picture.

Element sounds deliberately do not duck the music — see element sounds.

Language

The transcription language comes from the project's narration language and is sent explicitly, never auto-detected. A German take opening with an English product name would otherwise come back transcribed as English.

Where the audio lives

In your browser, beside the project. It is never uploaded except momentarily for transcription, and it does not travel inside an exported .json — see projects and files.