Why word timing matters
Caption tracks that show a whole sentence at once are fine for subtitles on a film. On a phone in a feed they read as a wall of text that appears and vanishes. Word-level timing changes the experience: the current word lights up as it is spoken, the eye tracks along with the voice, and the viewer stays locked to the clip. That only works if the timing is accurate to a fraction of a second, which is why the clipper transcribes with per-word timestamps rather than per-sentence ones.
How the captions are built
The transcript is split into short cards of at most four words or about twenty characters, breaking on sentence punctuation and on pauses longer than roughly two thirds of a second. Each card is rendered as an image with a heavy bold typeface, white fill and a dark outline, so it stays readable over bright or busy footage. In the highlight style, one image is rendered per word with that word tinted yellow, and the images are swapped in and out on the word timings. In the clean style, one image per card is shown in sentence case without the highlight.
Cards are placed in the lower third of the frame, above the area most platforms cover with their own interface, and centred horizontally. Long words shrink the type a little rather than overflowing the frame.
Burned in, on purpose
The captions are composited into the video frames, not attached as a subtitle track. That is deliberate. Subtitle tracks are stripped by some platforms, ignored by others and rendered in a different font everywhere. A burned-in caption looks the same in every player and survives every re-upload and re-share. The trade-off is that you cannot edit the text after rendering; if the transcript got a word wrong, re-run the job.
Choosing a caption style
Highlight is the default and suits energetic content: podcasts, commentary, fitness, anything where the delivery is punchy. Clean is sentence-case white with an outline, no highlight, and suits lessons, interviews and calmer material. None turns captions off for the rare case where a platform insists on generating its own. All three styles keep the progress bar.
- Highlight: uppercase, current word in yellow, maximum readability.
- Clean: sentence case, white with outline, no highlight.
- None: no caption cards; progress bar only.
Languages and accuracy
Transcription handles most widely spoken languages and renders captions in the language that was spoken, with correct accents and punctuation. Accuracy depends on audio quality: a clear microphone in a quiet room is close to perfect, a laptop mic in a café will miss words. Since the same transcript drives the moment picker, cleaner audio produces better clips as well as better captions.