What You'll Need Before You Start
Almost nothing. You need a recorded episode — the full-length version, not an edited highlight cut. Supported formats include mp4, mov, webm, and mkv. The file needs an audio track, obviously, since transcription drives the entire scoring process. There's no minimum length, but you'll get more clip candidates from a 45-minute conversation than from a 10-minute solo segment. The maximum upload size is 1024 MB, which covers the vast majority of podcast recordings at standard quality.
That's it. No account needed to start. No software to install. No timeline editor to learn.
Step 1 — Upload Your Full Episode
Upload the complete, unedited episode. Not a 10-minute highlight cut, not a trailer. The full file.
This matters for one reason: the AI finds candidate windows spread across the entire timeline. If you upload only part of the episode, you're artificially limiting the candidate pool. Some of the most self-contained, shareable moments happen 40 minutes into a conversation, after the guest has settled in and stopped performing. Those moments only appear if you give the tool the whole episode.
Upload takes roughly 1 to 3 minutes depending on your connection speed. Nothing runs automatically once the upload completes. The tool waits for you to trigger the analysis — which means you can upload several episodes and batch-process them.
This step typically takes under 5 minutes including the upload wait.
Step 2 — Let the AI Find and Score the Best Moments
Once uploaded, the tool transcribes the episode using a Whisper-style model that assigns a timestamp to every individual word. It then sweeps the full timeline looking for candidate windows — stretches of speech between 18 and 75 seconds long that begin and end at sentence boundaries or natural pauses. Roughly 48 candidates are identified and scored.
Each candidate is scored on four signals. Hook strength rates whether the opening words would stop a scroll. Information density measures how much non-redundant value is packed into the window. Self-containedness scores whether the clip makes sense to someone who knows nothing about your show or your guest. Emotional tone detects whether the moment makes you feel something — curiosity, surprise, validation, or mild tension.
The final set you're shown is the highest-scoring non-overlapping clips — meaning the tool doesn't hand you five variations of the same 90-second exchange. You get ten clips spread across the episode.
For a 60-minute podcast, this analysis typically completes in 3 to 6 minutes.
For a deeper explanation of exactly how each score is calculated, the article on how AI scores podcast clips covers the full breakdown.
Step 3 — Review the Picks and Read the Reasons
This is the step most people skip, and it's the most important one.
Each clip comes with a score breakdown and a plain-English reason for why it was selected. Something like: "Opens with a counterintuitive claim about hiring that doesn't require setup. High hook and self-containedness scores. Emotion score is moderate — the tension builds but doesn't fully resolve within the clip window."
Read that. Then watch or listen to the clip. Then decide whether you agree.
The AI doesn't know your audience. It doesn't know that clip number three would confuse anyone who hasn't heard your previous three episodes. It doesn't know that the host sounds tired in clip seven even though the words are technically strong. It doesn't know that the story in clip nine has been posted to your feed twice already this month.
You do. The AI's job is to surface candidates efficiently. Your job is to evaluate them with context the tool doesn't have. The reason per clip is what makes that evaluation possible — it gives you something to agree with or reject, rather than just a ranked list you either accept or don't.
A good pass through ten clips with their reasons takes about 5 to 10 minutes.
Step 4 — Trim to Taste
Every clip has editable in and out points. Drag to trim.
The most common trim is cutting a half-second of silence from the front. Sometimes the AI's window starts a beat before the hook really lands, and trimming those 0.4 seconds tightens it noticeably. The second most common trim is cutting the tail — the AI's window might end a sentence after where you'd prefer it to end, with a cleaner pause or a more punchy final word earlier.
One thing to know: captions are rebuilt against the new origin after every trim. You don't need to redo anything manually. The word-level timestamps recalibrate to whatever in and out points you set, so the caption timing stays exact.
No re-encoding happens until you export. Trimming is non-destructive and instant.
This step takes about 5 minutes for a full set of ten clips if you're being selective. Some clips won't need trimming at all.
Step 5 — Pick a Caption Style and Export
Caption style selection is a real decision, not just an aesthetic one. The style you pick changes how a clip performs on a given platform.
For high-energy hooks and punchy podcast moments going to TikTok or Reels, a bold uppercase style with a highlighted spoken word reads as native to those feeds. For LinkedIn or YouTube Shorts where the viewer profile skews more professional, a cleaner minimal style performs better. If your host has a distinct rhythmic delivery, a karaoke-style fill that tracks syllable timing makes the captions feel like part of the performance rather than an overlay.
For a full guide to all 18 available styles and when to use each, the caption styles guide covers every preset in detail.
Once you've picked a style, export. Output format is 9:16 at 1080x1920. Captions are burned in word-by-word. Background options include blurred background and center crop. The export is a final render — nothing to re-encode or convert afterward.
Export time per clip is typically 30 to 90 seconds at standard quality.
How to Get 10 Clips Instead of 3
The most common complaint from new users is getting only 3 or 4 usable clips from a 60-minute episode. Here's why that happens and how to fix it.
The first reason is source length. A 15-minute episode produces fewer candidates than a 60-minute one — not because the AI is penalizing short content, but because there are genuinely fewer non-overlapping windows. Longer episodes give the algorithm more material to find the best non-overlapping set. If you're consistently getting few clips, try uploading full episodes rather than pre-edited highlights.
The second reason is evaluating only the top of the list. The ranked output is a starting point, not a prescription. The seventh-ranked clip might be better for LinkedIn than the top-ranked clip that scores high on hook but low on self-containedness. Scroll through all ten rather than stopping at the first three.
The third reason is not varying hook types. If your best clips all open with a bold claim, you'll saturate your audience with one format. Look for variety across the set — a question-opener, a story-opener, a statistic-opener, a contrast-opener. The scoring doesn't weight diversity, but you should.
Common Mistakes When Clipping a Podcast
The most expensive mistake is posting clips that need setup. A clip that opens with "...and that's exactly what I was saying earlier" is only useful to the people who already heard the preceding 20 minutes. It will perform poorly with cold audiences on TikTok or Reels regardless of how high it scores on hook strength, because the hook is a reference, not a claim. Check the self-containedness score and, more importantly, check the opening words yourself.
Ignoring the self-containedness score entirely is a related mistake. It's the signal most directly tied to how a clip performs with new viewers. A clip that makes no sense without context might still score well on hook and emotion — the first sentence can be punchy and the delivery can be electric — but it will bleed potential followers who don't have the context to follow what's being said.
Over-trimming past the hook is the third one. If you trim aggressively to remove what feels like a slow start, you can accidentally cut the setup line that makes the hook land. The hook and its one-sentence setup are usually inseparable. Trim silence, not content.
Frequently Asked Questions
How do I turn a long video into short clips? Upload the full video, let AI score candidate windows on hook, density, self-containedness, and emotion, review the scored clips with their plain-English reasons, trim to taste, pick a caption style, and export 9:16 clips with burned-in captions.
How long does it take to clip a podcast with AI? For a 60-minute episode: 1 to 3 minutes upload, 3 to 6 minutes analysis, 5 to 10 minutes review, 5 minutes trimming, 5 to 15 minutes export for ten clips. Total: under 30 minutes for a full set of ten clips, compared to 4 to 6 hours manually.
How many clips can I get from one podcast episode? The tool surfaces up to ten non-overlapping clips per episode. Longer episodes produce more candidates to choose from. A 60-minute podcast typically yields between 6 and 10 usable clips depending on content density.
What video formats can I upload? mp4, mov, webm, and mkv. Maximum file size is 1024 MB. Audio track is required — video-only files cannot be transcribed.
Do I need to edit the clips after the AI makes them? You should review each clip and trim where needed, but many clips require no changes. The in/out point editing is optional. Caption timing recalibrates automatically after any trim, so there's no manual caption work.
Can I turn a YouTube podcast into shorts? Currently the tool accepts file uploads. If you want to clip a YouTube podcast, download the episode as an mp4 first, then upload it. Set honest expectations: the current input path is file upload, not a URL or YouTube link.
Upload your first episode — 10 free credits, no card required.