Voice and subtitles

How the caption band looks

Styling the fixed band captions sit in, and where it sits.

The caption band is one object. Style it once and every cue uses it.

The band is an element in the overlay layer — it is anchored to the frame, like the title and the progress bar. Its controls are also reachable under Voice & Subtitles → Appearance, beside the cues themselves, because "how do my subtitles look" and "what do my subtitles say" are usually the same sitting.

The band is an overlay element; the cues are not. Cues belong to the soundtrack, which is why a saved look never carries somebody else's sentences.

The band does not move

Captions sit in a fixed position on the screen, not on the chart. Zoom, pan and camera moves all leave the band exactly where it is.

That is the difference between a subtitle and a text element. If the words describe something in the picture, they want to be text. If they are what is being said, they want to be a caption.

Controls

FieldNotes
Typeface, size, weightSize is a fraction of the frame, so it survives an aspect change
ColourThe text itself
BackgroundFill and opacity behind the band
PaddingSpace inside the box
Max linesBeyond this, a cue becomes successive cards
PositionWhere the band sits on the frame, as a fraction of it

You can also drag the band directly on the canvas. Whatever you drag, it cannot leave the frame. To drop captions from one video without deleting them, use the band's "not in this video" switch — the eye only clears your working canvas.

Sizing for vertical

On 9:16 the band competes with the platform's own interface — captions, the action rail, the nav bar. Turn on the matching safe area outline (Settings → Appearance → Theme) to see the zones while you position it.

Those outlines are estimates, and they are drawn in the export preview but never in the exported file. See watermark and safe areas.

Two-card cues

A long cue splits into cards rather than truncating. If you are getting more splits than you want, the fix is usually shorter cues rather than more lines: three lines of caption on a vertical video is already most of the readable width.