The short answer: run the video through a speech-to-text model that returns word-level timestamps, then burn the resulting text into the frame as styled captions. No typing, no timeline, no SRT wrangling. A 60-second clip is captioned in under a minute. Below: how it works, what separates good captions from unwatchable ones, and how to do it in Indian languages where most tools quietly fail.
Why burned-in captions are non-negotiable
Short-form video autoplays muted. A viewer scrolling a feed decides whether to stay within about a second and a half, and in that window the only thing communicating with them is what is visibly on screen. A clip without captions is a silent film with no title cards.
Beyond the sound-off problem, burned-in captions:
- Survive re-uploads. Platform auto-captions are stripped the moment you download and cross-post.
- Are styleable. A caption is a design element, and a good one carries your look across every platform.
- Hold attention mechanically. Word-by-word highlighting gives the eye something to track, which is why the style is everywhere.
- Make the video usable by deaf and hard-of-hearing viewers, which is simply the right default.
How automatic captioning works
- Audio extraction — the audio track is pulled out of the video file.
- Transcription — a speech model converts audio to text. The thing that matters is whether it returns a timestamp per word rather than per sentence.
- Grouping — words are chunked into lines of two to four so they fit a vertical frame without covering the speaker.
- Styling — font, colour, stroke, shadow, the highlight colour for the currently spoken word, and any entry animation.
- Burn-in — the styled text is rendered into the video frames themselves during export.
Step 2 is where cheap tools cut corners. Without per-word timing you cannot do karaoke highlighting at all, and captions drift out of sync within a few seconds.
What makes captions readable
Two to four words per line
Full sentences at the bottom of a 9:16 frame force the eye to travel and cover too much picture. Short chunks, swapped rapidly, read as rhythm.
Heavy stroke, not a background box
White text with a thick black outline stays legible over any footage — bright, dark, busy, or moving. A solid background box is safer still but costs you a lot of frame.
Position clear of the platform UI
The bottom sixth of the frame is where TikTok and Reels stack usernames, captions and buttons. Put your text in the lower third but not the very bottom — or use centre placement for talking-head footage, which is the look most Reels use now.
A highlight colour that means something
The spoken word in yellow or green is not decoration; it paces the viewer. Pick one accent and keep it consistent so the style becomes recognisably yours.
Captions in Hindi, Gujarati, Tamil and other Indian languages
This is where most Western tools break, in two distinct ways. First, script rendering: a caption engine without the right font falls back to boxes, and Devanagari, Gujarati, Tamil, Telugu, Kannada, Malayalam, Gurmukhi, Bengali and Odia each need their own typeface. Second, transcription: a model tuned on English will hear Gujarati speech and write it in Devanagari, or simply translate it to English — neither of which is a caption.
AI Clip Cutter handles both. SarvamAI transcribes in the speaker's own language and script for ten Indian languages, and because Sarvam returns coarse timing, Whisper is run over the same audio purely as a clock so each word lands where it was actually spoken. Script -specific font stacks are selected automatically from the detected language.
Captioning a clip you already have
You do not need to cut anything to use the captioning. If you already have a finished reel, the caption editor transcribes it and captions it whole, with no scoring and no cutting, for clips up to two minutes. Try it at the caption generator.
- Upload the finished vertical clip.
- Wait for the transcript — seconds for a short clip.
- Fix any names the model misheard. Editing the text does not break the timing.
- Choose a preset, set the position, and export.
Choosing a caption style
AI Clip Cutter ships 18 presets. In practice they fall into four jobs:
- Bold (Hormozi, Beast, Punchline) — chunky uppercase with a heavy stroke. Highest stopping power, right for advice, hot takes and business content.
- Clean (Minimal, Podcast, Headline) — quieter, lets the footage lead. Right for interviews and brand accounts.
- Kinetic (Karaoke, Bounce, Motion, Typewriter) — motion per word. Right for fast, high-energy edits.
- Cinematic (Cinematic, Cinetop, Ember, Softfocus) — restrained and atmospheric, for narrative or documentary footage.
Pick one and stay with it for a month. Consistency is worth more than any individual style.
Frequently asked questions
Can I edit the text the AI transcribed?
Yes. Proper nouns, brand names and technical terms are where every speech model slips. Edited captions replace the originals at export and keep their word timings.
Do burned-in captions hurt reach?
No. Platforms cannot read burned-in text for indexing, so pair the clip with a written description — but the retention lift from on-screen text dwarfs the indexing question.
Is there a watermark on the free tier?
No. There is no watermarking anywhere in AI Clip Cutter, free or paid. Exports are 1080p H.264 MP4.
How much does captioning cost?
One credit per rendered clip, plus one credit per minute of source analysed. Signup gives you 10 free credits, and purchased credits never expire.
Caption one clip tonight
Take a reel you already posted without captions, run it through the caption generator, and repost it. Same footage, same idea, measurably different retention. Your free credits cover it.