Your podcast, cut into clips worth posting
AI Clip Cutter transcribes your long-form video with word-level timestamps, scores every candidate moment from 0 to 10, and exports the winners as 9:16 clips with captions burned in — each one with the reason it was picked, so you stay the editor.
No card required. Analysis costs 1 credit per minute of source video, each export 1 credit — nothing runs until you press the button.
“She reframes the whole argument in one sentence here — a self-contained hook that needs no setup.”
Illustration of the clip studio — scores and reasons come from your own video.
Three steps, and you approve every one
- Step 1
Upload your long video
Drop in a podcast, interview, or webinar — mp4, mov, webm, or mkv, up to 1024 MB. It streams straight to disk, ffprobe reads the duration and dimensions, and anything without an audio track is rejected before it costs you a credit.
Nothing runs automatically. Analysis waits for you to press the button.
- Step 2
The AI finds the best moments
Audio is downmixed to 16 kHz mono and transcribed by Whisper with word-level timestamps. Candidate windows of 18–75 seconds are cut only at sentence boundaries or real pauses, then scored 0–10 by GPT-4o on four engagement dimensions.
Every pick arrives with a plain-English reason, so you can overrule it.
- Step 3
Export with captions burned in
A straight cut with a 0.75 s pad either side, reframed to 9:16 with a blurred background or a centre crop, then encoded H.264 + AAC at 1080×1920. Captions are burned in word by word through libass.
Trim points stay editable — captions are rebuilt against the new origin.
8 burned-in presets, previewed honestly
Each card below is drawn from the same preset table the ASS writer hands to libass, so the stroke weights, colours and words-per-line are the ones your export will use. The spoken word is highlighted in real time because captions are built from Whisper's word timestamps, not guessed.
- Thischangedeverything
Hormozi
Chunky uppercase white with a heavy black stroke and a yellow spoken word. Safe on any footage.
3 words per line · bottom anchor
- Thischangedeverythingfor
Beast
Extra-bold white with a very thick stroke and a green pop on the spoken word.
4 words per line · bottom anchor
- Thischangedeverythingforus
Karaoke
Sentence case that fills word by word from white to cyan as it is spoken.
5 words per line · bottom anchor
- Thischangedeverythingforus
Minimal
Clean sentence case, thin stroke, the spoken word simply brightens. Fades in.
5 words per line · bottom anchor
- Thischangedeverythingfor
Boxed
White uppercase on a solid black block — the podcast-clip subtitle bar.
4 words per line · bottom anchor
- Thischangedeverything
Neon
Electric cyan with a magenta spoken word and a coloured glow behind it.
3 words per line · bottom anchor
- Thischangedeverythingforushonestly
Podcast
Large centred sentence case over two lines, with a soft blue spoken word.
6 words per line · center anchor
- Thischangedeverything
Bounce
Each word scales up and settles as it lands, with a hot orange highlight.
3 words per line · bottom anchor
Previews are a CSS approximation of libass output. Captions can be anchored at the bottom, centre or top; the preset fonts are Latin display faces, so non-Latin scripts render in whatever family the render host substitutes.
Built to be checked, not trusted blindly
Word-level Whisper timestamps
Captions land on the word, not the sentence. Groups break on punctuation or a pause of 0.35 s or more, and each line holds until the next one starts, so captions never flicker off during a natural gap.
Ranked, non-overlapping picks
Up to 48 candidate windows are kept, spread evenly across the whole timeline rather than only the opening minutes, then scored and combined as 30% hook, 25% self-contained, 25% density, 20% emotion. The best non-overlapping set wins.
A stated reason per clip
Every selected moment carries the model's plain-English reason plus its four scores, shown beside the clip. You are reading an argument you can reject, not a black-box list.
Trim points are editable numbers
In and out points are fields, not a fixed decision. Saving a new trim regenerates the caption groups against the new origin without re-encoding — a render credit is only spent when you export.
Credits, charged in rupees
1 credit per minute of source video to analyse, 1 credit per exported clip, both shown on the button before you press it. Credit packs are billed in INR through Razorpay, and a failed pipeline refunds automatically.
What it deliberately does not do
- No face tracking — the default 9:16 blurs a background copy instead of cropping
- No b-roll, no zooms, no transitions: clips are straight cuts
- No speaker diarisation
- 1080×1920 is the reframing ceiling — there is no 4K preset
- Scoring and hook prompts are English-tuned
Questions worth asking first
What kind of video works best?
- Speech-led long-form: podcasts, interviews, webinars, talking-head uploads. Candidate windows are constrained to 18–75 seconds, so the source needs at least one continuous stretch of speech that long. Music-only or near-silent audio produces nothing, and the app tells you so instead of charging you.
Do I have to trust the AI's choices?
- No. Every clip shows the reason the model gave for picking it plus its four scores, and the in and out points are editable numbers. Saving a new trim regenerates the caption groups against the new origin without re-encoding — you only spend a render credit when you export.
Can I change how the captions look?
- There are eight presets and three vertical positions (bottom, centre, top). Changing the preset also regenerates the caption grouping, because words-per-line is part of the style — a three-word Hormozi line and a six-word podcast line need different chunking.
What does it deliberately not do?
- There is no face tracking, so the centre-crop preset will not follow a speaker who sits off to one side — the default blurred 9:16 avoids the problem by never cropping. There is no b-roll, no zooms, no transitions, and no speaker diarisation. Clips are straight cuts, and 1080×1920 is the ceiling.
Which languages does it handle?
- Whisper transcribes many languages and the detected language is stored on the job, but the scoring and hook-writing prompts are written in English, so results are strongest on English-language source video.
How am I charged?
- Credits. Analysis costs 1 credit per minute of source video and each export costs 1 credit, both shown on the button before you press it. Credits are deducted before work starts and refunded automatically if the pipeline throws. Credit packs are billed in INR through Razorpay.
Run one real episode through it
Signing up grants 10 credits — enough to analyse a short video and export a few clips. No card, and nothing renews on its own.