Long video → captioned shorts

Your podcast, cut into clips worth posting

AI Clip Cutter transcribes your long-form video with word-level timestamps, scores every candidate moment from 0 to 10, and exports the winners as 9:16 clips with captions burned in — each one with the reason it was picked, so you stay the editor.

No card required. Analysis costs 1 credit per minute of source video, each export 1 credit — nothing runs until you press the button.

Thischangedeverything

“She reframes the whole argument in one sentence here — a self-contained hook that needs no setup.”

Hook9
Self-contained8
Density7
Emotion8

Illustration of the clip studio — scores and reasons come from your own video.

How it works

Three steps, and you approve every one

  1. Step 1

    Upload your long video

    Drop in a podcast, interview, or webinar — mp4, mov, webm, or mkv, up to 1024 MB. It streams straight to disk, ffprobe reads the duration and dimensions, and anything without an audio track is rejected before it costs you a credit.

    Nothing runs automatically. Analysis waits for you to press the button.

  2. Step 2

    The AI finds the best moments

    Audio is downmixed to 16 kHz mono and transcribed by Whisper with word-level timestamps. Candidate windows of 18–75 seconds are cut only at sentence boundaries or real pauses, then scored 0–10 by GPT-4o on four engagement dimensions.

    Every pick arrives with a plain-English reason, so you can overrule it.

  3. Step 3

    Export with captions burned in

    A straight cut with a 0.75 s pad either side, reframed to 9:16 with a blurred background or a centre crop, then encoded H.264 + AAC at 1080×1920. Captions are burned in word by word through libass.

    Trim points stay editable — captions are rebuilt against the new origin.

Caption styles

8 burned-in presets, previewed honestly

Each card below is drawn from the same preset table the ASS writer hands to libass, so the stroke weights, colours and words-per-line are the ones your export will use. The spoken word is highlighted in real time because captions are built from Whisper's word timestamps, not guessed.

  • Thischangedeverything

    Hormozi

    Chunky uppercase white with a heavy black stroke and a yellow spoken word. Safe on any footage.

    3 words per line · bottom anchor

  • Thischangedeverythingfor

    Beast

    Extra-bold white with a very thick stroke and a green pop on the spoken word.

    4 words per line · bottom anchor

  • Thischangedeverythingforus

    Karaoke

    Sentence case that fills word by word from white to cyan as it is spoken.

    5 words per line · bottom anchor

  • Thischangedeverythingforus

    Minimal

    Clean sentence case, thin stroke, the spoken word simply brightens. Fades in.

    5 words per line · bottom anchor

  • Thischangedeverythingfor

    Boxed

    White uppercase on a solid black block — the podcast-clip subtitle bar.

    4 words per line · bottom anchor

  • Thischangedeverything

    Neon

    Electric cyan with a magenta spoken word and a coloured glow behind it.

    3 words per line · bottom anchor

  • Thischangedeverythingforushonestly

    Podcast

    Large centred sentence case over two lines, with a soft blue spoken word.

    6 words per line · center anchor

  • Thischangedeverything

    Bounce

    Each word scales up and settles as it lands, with a hot orange highlight.

    3 words per line · bottom anchor

Previews are a CSS approximation of libass output. Captions can be anchored at the bottom, centre or top; the preset fonts are Latin display faces, so non-Latin scripts render in whatever family the render host substitutes.

Features

Built to be checked, not trusted blindly

Word-level Whisper timestamps

Captions land on the word, not the sentence. Groups break on punctuation or a pause of 0.35 s or more, and each line holds until the next one starts, so captions never flicker off during a natural gap.

Ranked, non-overlapping picks

Up to 48 candidate windows are kept, spread evenly across the whole timeline rather than only the opening minutes, then scored and combined as 30% hook, 25% self-contained, 25% density, 20% emotion. The best non-overlapping set wins.

A stated reason per clip

Every selected moment carries the model's plain-English reason plus its four scores, shown beside the clip. You are reading an argument you can reject, not a black-box list.

Trim points are editable numbers

In and out points are fields, not a fixed decision. Saving a new trim regenerates the caption groups against the new origin without re-encoding — a render credit is only spent when you export.

Credits, charged in rupees

1 credit per minute of source video to analyse, 1 credit per exported clip, both shown on the button before you press it. Credit packs are billed in INR through Razorpay, and a failed pipeline refunds automatically.

What it deliberately does not do

  • No face tracking — the default 9:16 blurs a background copy instead of cropping
  • No b-roll, no zooms, no transitions: clips are straight cuts
  • No speaker diarisation
  • 1080×1920 is the reframing ceiling — there is no 4K preset
  • Scoring and hook prompts are English-tuned
FAQ

Questions worth asking first

What kind of video works best?

Speech-led long-form: podcasts, interviews, webinars, talking-head uploads. Candidate windows are constrained to 18–75 seconds, so the source needs at least one continuous stretch of speech that long. Music-only or near-silent audio produces nothing, and the app tells you so instead of charging you.

Do I have to trust the AI's choices?

No. Every clip shows the reason the model gave for picking it plus its four scores, and the in and out points are editable numbers. Saving a new trim regenerates the caption groups against the new origin without re-encoding — you only spend a render credit when you export.

Can I change how the captions look?

There are eight presets and three vertical positions (bottom, centre, top). Changing the preset also regenerates the caption grouping, because words-per-line is part of the style — a three-word Hormozi line and a six-word podcast line need different chunking.

What does it deliberately not do?

There is no face tracking, so the centre-crop preset will not follow a speaker who sits off to one side — the default blurred 9:16 avoids the problem by never cropping. There is no b-roll, no zooms, no transitions, and no speaker diarisation. Clips are straight cuts, and 1080×1920 is the ceiling.

Which languages does it handle?

Whisper transcribes many languages and the detected language is stored on the job, but the scoring and hook-writing prompts are written in English, so results are strongest on English-language source video.

How am I charged?

Credits. Analysis costs 1 credit per minute of source video and each export costs 1 credit, both shown on the button before you press it. Credits are deducted before work starts and refunded automatically if the pipeline throws. Credit packs are billed in INR through Razorpay.