AI Shorts Creator

Auto captions for shorts

Auto captions for shorts, timed to the word

Most shorts are watched muted. The captions are the audio. They need to be large, correctly timed and impossible to miss, and they need to be in the file.

The 8-hour sleep rule is wrong

podcast · 34s

Compound interest in 40 seconds

lesson · 40s

What a real market morning looks like

vlog · 31s

The warm-up everyone skips

fitness · 36s

Why word timing matters

Caption tracks that show a whole sentence at once are fine for subtitles on a film. On a phone in a feed they read as a wall of text that appears and vanishes. Word-level timing changes the experience: the current word lights up as it is spoken, the eye tracks along with the voice, and the viewer stays locked to the clip. That only works if the timing is accurate to a fraction of a second, which is why the clipper transcribes with per-word timestamps rather than per-sentence ones.

How the captions are built

The transcript is split into short cards of at most four words or about twenty characters, breaking on sentence punctuation and on pauses longer than roughly two thirds of a second. Each card is rendered as an image with a heavy bold typeface, white fill and a dark outline, so it stays readable over bright or busy footage. In the highlight style, one image is rendered per word with that word tinted yellow, and the images are swapped in and out on the word timings. In the clean style, one image per card is shown in sentence case without the highlight.

Cards are placed in the lower third of the frame, above the area most platforms cover with their own interface, and centred horizontally. Long words shrink the type a little rather than overflowing the frame.

Burned in, on purpose

The captions are composited into the video frames, not attached as a subtitle track. That is deliberate. Subtitle tracks are stripped by some platforms, ignored by others and rendered in a different font everywhere. A burned-in caption looks the same in every player and survives every re-upload and re-share. The trade-off is that you cannot edit the text after rendering; if the transcript got a word wrong, re-run the job.

Choosing a caption style

Highlight is the default and suits energetic content: podcasts, commentary, fitness, anything where the delivery is punchy. Clean is sentence-case white with an outline, no highlight, and suits lessons, interviews and calmer material. None turns captions off for the rare case where a platform insists on generating its own. All three styles keep the progress bar.

  • Highlight: uppercase, current word in yellow, maximum readability.
  • Clean: sentence case, white with outline, no highlight.
  • None: no caption cards; progress bar only.

Languages and accuracy

Transcription handles most widely spoken languages and renders captions in the language that was spoken, with correct accents and punctuation. Accuracy depends on audio quality: a clear microphone in a quiet room is close to perfect, a laptop mic in a café will miss words. Since the same transcript drives the moment picker, cleaner audio produces better clips as well as better captions.

Process

Three steps between a recording and a week of posts

  1. Step 1

    Upload the long version

    Drop in an MP4, MOV or WebM up to 15 minutes and 500 MB. It goes straight to private storage and is wiped 24 hours later.

  2. Step 2

    We listen to the whole thing

    The audio is transcribed with word-level timing, then a language model reads the transcript and marks the three to five stretches that stand on their own.

  3. Step 3

    Each moment becomes a short

    Every pick is cut at a sentence boundary, reframed to 9:16, captioned word by word and finished with a progress bar. You get MP4s, not a project file.

Options

Opinionated defaults, a few switches that matter

Two ways to go vertical

Centre crop keeps a single speaker large in frame. Blur-pad keeps the full widescreen picture, floating over a softened copy of itself, for slides and gameplay.

Captions that follow the voice

Pick bold uppercase with the spoken word lit up, a quieter sentence-case style, or none. Timing comes from the transcript, not a guess.

Lengths that fit the feed

Every short lands between 20 and 60 seconds, snapped to where sentences actually start and stop. Set a shorter ceiling if your audience scrolls fast.

Answers

Auto captions for shorts: common questions

Can I edit the caption text before rendering?
Not in the current version. The transcript is rendered as recognised. Re-running the job on the same upload within 24 hours is the quickest fix for a one-off error.
Which font is used?
A heavy geometric sans-serif chosen for legibility at phone size. Font choice is fixed so every short from the service reads consistently.
Who owns the captioned shorts?
You do. The captions are a transcription of your own words over your own footage; the terms of service spells out the grant. Copyright law varies by country, so a lawyer is the right source for anything beyond everyday use.
Do captions cover the speaker’s face?
They sit in the lower third, which is clear in most centre-cropped talking shots. With blur-pad framing they sit below the letterboxed picture entirely.

Also clip from

Captions should take zero minutes

Upload a recording and get back shorts with the words already on screen, timed to the voice.

Upload a video