The short answer: upload the long video to an AI clip cutter, let it transcribe and score every moment, pick the 3–5 segments it ranks highest, reframe them to 9:16, burn in word-level captions, and export. A 30-minute episode takes about 5 minutes of your attention instead of half a day. The rest of this guide is the workflow in detail — including the part most creators get wrong, which is not the cutting but the choosing.
Why long-to-short is the highest-leverage thing you can do
You already paid for the hard part. The research, the guest booking, the recording, the hour of talking — that cost is sunk the moment you publish the long video. Shorts made from that footage cost you almost nothing extra and reach an audience the long video will never touch, because Reels, Shorts and TikTok push content to people who have never heard of you, while a 45-minute upload is shown mostly to people who already subscribed.
The catch has always been labour. Scrubbing a timeline for good moments, trimming each one, reframing 16:9 footage so the speaker is not cropped out, and captioning word by word is roughly 30–45 minutes per clip by hand. Five clips a week is a part-time job. That maths is why most creators start a Shorts habit and abandon it by week three.
The workflow, end to end
Step 1 — Start from footage that has moments in it
AI cannot invent a highlight. Before you upload anything, ask whether the video contains at least a few self-contained 30-second ideas: a strong opinion, a specific number, a story with a beginning and an end, a disagreement, a reversal of something the audience believes.
Footage that clips well:
- Interviews and podcasts, especially unscripted ones
- Q&A sessions and AMAs — the question is a built-in hook
- Webinars where the speaker gives concrete numbers or steps
- Tutorials with discrete, nameable steps
- Panels and debates, where the disagreement is the clip
Footage that clips badly:
- Slow-build narrative where nothing makes sense without the preceding ten minutes
- Screen-share-only walkthroughs with no speaker on camera
- Heavily music-bedded videos — the transcript is thin, so there is little for a model to score
Step 2 — Transcribe with word-level timestamps
Everything downstream depends on this. A transcript with only sentence-level timing is enough to cut a clip roughly, but not enough to make captions that highlight each word as it is spoken — and that highlight is what keeps a viewer watching a muted video. Word-level timestamps come out of speech-to-text models like Whisper, which is what AI Clip Cutter uses, with SarvamAI available for Indian-language sources.
Step 3 — Score moments rather than skim them
This is the step humans do badly and models do consistently. You get bored, you skip ahead, and you tend to pick the moment you remember rather than the moment that plays well cold. A scoring pass splits the transcript into candidate windows — 18 to 75 seconds is the useful band — and rates each one on four axes:
- Hook potential — does the first line make a stranger stop scrolling?
- Self-containedness — does it make sense with zero context?
- Information density — is something actually said, or is it filler?
- Emotional tone — is there surprise, conviction, humour, tension?
A clip that scores well on all four is a clip worth exporting. A clip that scores high on density but low on self-containedness usually needs its start moved 15 seconds earlier so the question is included.
Step 4 — Reframe without decapitating anyone
16:9 source into a 9:16 frame means losing about 60% of the width. There are three ways to handle it:
- Blurred background — the full frame sits in the middle of a blurred, scaled copy of itself. Nothing is cropped, nobody loses a head, and it reads as deliberate. The safest default for two-person interviews.
- Centre crop — tighter and more immersive, but it fails the moment the speaker is off-centre.
- 1:1 square — a compromise that performs well in LinkedIn and Facebook feeds, less so on TikTok.
Step 5 — Burn in captions
Short-form plays muted by default. Captions are not an accessibility nicety here, they are the soundtrack. Burned-in captions also survive re-uploads and cross-posting in a way that platform auto-captions do not. If you want the detail on styles, positions and timing, we wrote a whole guide on adding captions to video automatically.
Step 6 — Write the hook, keep the clip
The clip the AI chose is usually right. The title it suggests is a starting point. Spend your saved time here: rewrite the first on-screen line, because that line decides whether the other 40 seconds are ever watched. Fifteen formulas that work are in our hook formula guide.
How long each step actually takes
- Upload — 1–3 minutes for a typical 30-minute episode
- Transcribe + score + select — roughly 1–2 minutes of processing per 10 minutes of source
- Your review — 3–5 minutes to read the picks and their stated reasons
- Export — under a minute per clip
Call it 15 minutes of wall-clock time and 5 minutes of your own attention for four clips. Manually, the same four clips are a two-to-three hour afternoon.
Doing it with AI Clip Cutter, concretely
- Sign in with Google — new accounts get 10 credits, no card.
- Open the clip cutter workspace and drop in an MP4, MOV or WebM.
- Wait for analysis. Analysis costs 1 credit per minute of source video, so a 10-minute upload is 10 credits.
- Read the selected clips. Each arrives with a title, three alternative hook lines, and a plain-English reason it was picked.
- Pick a caption preset from the 18 available, choose bottom, centre or top placement, and set the format: 9:16 blurred, 9:16 cropped, 1:1 square, or original size.
- Export. 1 credit per rendered clip, no watermark, 1080p.
Mistakes that waste the effort
- Posting the clip exactly as exported. Rewrite the hook line. Always.
- Clips that are too long. Under 60 seconds keeps you eligible for every surface. If the idea needs 90 seconds, it is two clips.
- Starting with pleasantries. Cut the "so, yeah, I think" preamble. The clip starts at the claim.
- Publishing four clips on one day. Four clips is four days of posting. Spread them.
- Ignoring the reason text. If the tool tells you why it picked a segment and the reason is weak, that is a signal to skip it.
Frequently asked questions
How many shorts can one long video realistically produce?
Three to five good ones from a 30–60 minute episode. Tools that promise fifteen are counting the filler. AI Clip Cutter defaults to four per video for exactly this reason.
Do I need editing skills?
No. The pipeline handles cutting, reframing and captioning. What you bring is judgement about which clip represents you well and what the first line should say.
Will Shorts made this way get flagged as reused content?
Repurposing your own footage into a new format with new framing and captions is standard practice and is not what reused-content policies target. Uploading someone else is.
What does it cost to try?
Nothing. Ten free credits on signup covers a 10-minute source video and one exported clip. Paid credit packs start at ₹399 and never expire.
Start with one video
Take the last long video you published — one you already know has a good moment in it — and run it through AI Clip Cutter with your free credits. Compare what it picked with what you would have picked. That comparison tells you more in five minutes than any more reading will.