AI caption generator
Animated captions, burned in on the exact word being spoken
Upload a video and get subtitles rendered into the pixels — because Reels, Shorts, and TikTok autoplay muted and platform captions cannot be styled. Timing comes from Whisper's word-level timestamps and the caption file is built locally from them, with no model guessing at line breaks, so the highlight lands on the word rather than near it. Eight named presets, three safe-area positions, no watermark, priced in rupees.
- Caption presets
- 8
- Safe-area anchors
- 3
- Timing source
- Per word
- Watermark
- None
The eight presets, previewed honestly
These are not mock-ups — every preview and every cell below is drawn from the same preset definitions the renderer hands to libass, so the colours, stroke weights, font stacks, and line lengths are exactly what gets burned into your export. The second word is highlighted in each preview to show what the spoken word looks like. Pick a preset per clip and re-export as many times as you like.
- Thisisthe
Hormozi
Chunky uppercase white with a heavy black stroke and a yellow spoken word. Safe on any footage.
3 words per line · bottom anchor
- Thisisthehook
Beast
Extra-bold white with a very thick stroke and a green pop on the spoken word.
4 words per line · bottom anchor
- Thisisthehookthey
Karaoke
Sentence case that fills word by word from white to cyan as it is spoken.
5 words per line · bottom anchor
- Thisisthehookthey
Minimal
Clean sentence case, thin stroke, the spoken word simply brightens. Fades in.
5 words per line · bottom anchor
- Thisisthehook
Boxed
White uppercase on a solid black block — the podcast-clip subtitle bar.
4 words per line · bottom anchor
- Thisisthe
Neon
Electric cyan with a magenta spoken word and a coloured glow behind it.
3 words per line · bottom anchor
- Thisisthehooktheyremember
Podcast
Large centred sentence case over two lines, with a soft blue spoken word.
6 words per line · center anchor
- Thisisthe
Bounce
Each word scales up and settles as it lands, with a hot orange highlight.
3 words per line · bottom anchor
Previews are a CSS approximation of libass output: `-webkit-text-stroke` stands in for the ASS outline and a blurred text-shadow for Neon’s glow. Motion is described in the table rather than animated here.
| Text → spoken word | Font stack | Words / line | Motion | Keyword emphasis | |
|---|---|---|---|---|---|
| Hormozi | #FFFFFF → #FFD400 | Anton, Impact | 3 · UPPERCASE | pop | Yes — #3BE88A |
| Beast | #FFFFFF → #00E676 | Montserrat, Poppins | 4 · UPPERCASE | bounce | Yes — #FFE44D |
| Karaoke | #FFFFFF → #00D8FF | Poppins, Montserrat | 5 · Sentence case | karaoke fill | No |
| Minimal | #D9DEE6 → #FFFFFF | Helvetica Neue, Inter | 5 · Sentence case | fade | No |
| Boxed | #FFFFFF → #FFE066 | Montserrat, Poppins | 4 · UPPERCASE | fade | No |
| Neon | #00F0FF → #FF2FD0 | Poppins, Montserrat | 3 · UPPERCASE | pop | No |
| Podcast | #FFFFFF → #8AB4FF | Montserrat, Poppins | 6 · Sentence case | none | No |
| Bounce | #FFFFFF → #FF6B35 | Poppins, Montserrat | 3 · UPPERCASE | bounce | No |
Positioning
Where the captions sit
Bottom
Lower third, clear of the TikTok/Reels action bar
Center
Vertically centred — the Reels/Shorts talking-head look
Top
Upper third, clear of the status bar and caption overlay
No black box
How the captions are actually generated
Three steps, only the first of which involves a model.
1. Word-level transcription
Whisper transcribes your audio with a timestamp on every individual word. Long audio is split and stitched back onto a single timeline, so a two-hour recording is fine. The detected language is stored with the job.
2. Local grouping
Words are grouped into caption lines locally, using the preset's own words-per-line setting — 3 for Hormozi and Neon, 6 for Podcast. No model call, so nothing drifts out of sync, and switching preset regenerates the grouping rather than reflowing the old one.
3. Burned in with FFmpeg
An ASS subtitle file is emitted with one event per word state, which is what gives frame-exact control over the colour and scale of the word being spoken, then libass burns it into an H.264 MP4 — 1080×1920, 1080×1080, or your source size, with AAC audio at 128 kbps and faststart.
Straight answer
What this tool does not do
Read this before you sign up
Eight presets, not a style editor. You choose a preset and an anchor. You cannot change a font, add your own colours, or save a brand kit — those settings live in code, not in the UI.
Devanagari falls back to a system font. Whisper transcribes Hindi and many other languages accurately, and captions are built from those word timestamps, so the timing is right. But every preset lists Latin display faces such as Anton, Montserrat, and Poppins first, so non-Latin scripts render through whatever the host substitutes unless a matching family is installed. The clip-scoring and hook-writing prompts are also written in English and behave best on English speech.
1080 on the long edge. Vertical and square exports are 1080×1920 and 1080×1080; only the Original preset keeps your source resolution. There is no 4K or upscaling path.
Captions on top of a straight cut. No face tracking, no punch-in zooms, no transitions, no b-roll, no speaker labels. The 9:16 default fills the frame with a blurred copy of your video precisely so that no crop can push your speaker out of frame.
Burn-in needs libass. Captions are rendered by FFmpeg's subtitle filter. On a build without libass, clips still export — just without captions — and the studio tells you so rather than failing silently.
Caption a real video before you pay anything
New accounts get 10 credits. That is enough to take a short video through transcription, clip selection, and an export with captions burned in — unwatermarked, with no card on file and nothing to cancel.