AI caption generator

Animated captions, burned in on the exact word being spoken

Upload a video and get subtitles rendered into the pixels — because Reels, Shorts, and TikTok autoplay muted and platform captions cannot be styled. Timing comes from Whisper's word-level timestamps and the caption file is built locally from them, with no model guessing at line breaks, so the highlight lands on the word rather than near it. Eight named presets, three safe-area positions, no watermark, priced in rupees.

Caption presets
8
Safe-area anchors
3
Timing source
Per word
Watermark
None

The eight presets, previewed honestly

These are not mock-ups — every preview and every cell below is drawn from the same preset definitions the renderer hands to libass, so the colours, stroke weights, font stacks, and line lengths are exactly what gets burned into your export. The second word is highlighted in each preview to show what the spoken word looks like. Pick a preset per clip and re-export as many times as you like.

  • Thisisthe

    Hormozi

    Chunky uppercase white with a heavy black stroke and a yellow spoken word. Safe on any footage.

    3 words per line · bottom anchor

  • Thisisthehook

    Beast

    Extra-bold white with a very thick stroke and a green pop on the spoken word.

    4 words per line · bottom anchor

  • Thisisthehookthey

    Karaoke

    Sentence case that fills word by word from white to cyan as it is spoken.

    5 words per line · bottom anchor

  • Thisisthehookthey

    Minimal

    Clean sentence case, thin stroke, the spoken word simply brightens. Fades in.

    5 words per line · bottom anchor

  • Thisisthehook

    Boxed

    White uppercase on a solid black block — the podcast-clip subtitle bar.

    4 words per line · bottom anchor

  • Thisisthe

    Neon

    Electric cyan with a magenta spoken word and a coloured glow behind it.

    3 words per line · bottom anchor

  • Thisisthehooktheyremember

    Podcast

    Large centred sentence case over two lines, with a soft blue spoken word.

    6 words per line · center anchor

  • Thisisthe

    Bounce

    Each word scales up and settles as it lands, with a hot orange highlight.

    3 words per line · bottom anchor

Previews are a CSS approximation of libass output: `-webkit-text-stroke` stands in for the ASS outline and a blurred text-shadow for Neon’s glow. Motion is described in the table rather than animated here.

The eight burned-in caption presets, with resting and spoken-word colours, font stack, words per line, motion, and keyword emphasis
 Text → spoken wordFont stackWords / lineMotionKeyword emphasis
Hormozi#FFFFFF → #FFD400Anton, Impact3 · UPPERCASEpopYes — #3BE88A
Beast#FFFFFF → #00E676Montserrat, Poppins4 · UPPERCASEbounceYes — #FFE44D
Karaoke#FFFFFF → #00D8FFPoppins, Montserrat5 · Sentence casekaraoke fillNo
Minimal#D9DEE6 → #FFFFFFHelvetica Neue, Inter5 · Sentence casefadeNo
Boxed#FFFFFF → #FFE066Montserrat, Poppins4 · UPPERCASEfadeNo
Neon#00F0FF → #FF2FD0Poppins, Montserrat3 · UPPERCASEpopNo
Podcast#FFFFFF → #8AB4FFMontserrat, Poppins6 · Sentence casenoneNo
Bounce#FFFFFF → #FF6B35Poppins, Montserrat3 · UPPERCASEbounceNo
Font stacks are ordered preferences: the renderer probes them in order and falls back to a system sans (Helvetica, Arial, Liberation Sans, DejaVu Sans, Noto Sans) when the display face is not installed. Keyword emphasis recolours long, rare, or shouted words in a third colour. Words per line drives the caption grouping, not just where the line wraps.

Positioning

Where the captions sit

Bottom

Lower third, clear of the TikTok/Reels action bar

Center

Vertically centred — the Reels/Shorts talking-head look

Top

Upper third, clear of the status bar and caption overlay

No black box

How the captions are actually generated

Three steps, only the first of which involves a model.

  1. 1. Word-level transcription

    Whisper transcribes your audio with a timestamp on every individual word. Long audio is split and stitched back onto a single timeline, so a two-hour recording is fine. The detected language is stored with the job.

  2. 2. Local grouping

    Words are grouped into caption lines locally, using the preset's own words-per-line setting — 3 for Hormozi and Neon, 6 for Podcast. No model call, so nothing drifts out of sync, and switching preset regenerates the grouping rather than reflowing the old one.

  3. 3. Burned in with FFmpeg

    An ASS subtitle file is emitted with one event per word state, which is what gives frame-exact control over the colour and scale of the word being spoken, then libass burns it into an H.264 MP4 — 1080×1920, 1080×1080, or your source size, with AAC audio at 128 kbps and faststart.

Straight answer

What this tool does not do

Read this before you sign up

Eight presets, not a style editor. You choose a preset and an anchor. You cannot change a font, add your own colours, or save a brand kit — those settings live in code, not in the UI.

Devanagari falls back to a system font. Whisper transcribes Hindi and many other languages accurately, and captions are built from those word timestamps, so the timing is right. But every preset lists Latin display faces such as Anton, Montserrat, and Poppins first, so non-Latin scripts render through whatever the host substitutes unless a matching family is installed. The clip-scoring and hook-writing prompts are also written in English and behave best on English speech.

1080 on the long edge. Vertical and square exports are 1080×1920 and 1080×1080; only the Original preset keeps your source resolution. There is no 4K or upscaling path.

Captions on top of a straight cut. No face tracking, no punch-in zooms, no transitions, no b-roll, no speaker labels. The 9:16 default fills the frame with a blurred copy of your video precisely so that no crop can push your speaker out of frame.

Burn-in needs libass. Captions are rendered by FFmpeg's subtitle filter. On a build without libass, clips still export — just without captions — and the studio tells you so rather than failing silently.