You recorded an hour-long webinar last month. It sits in a folder, already paid for, watched by the people who attended and by nobody since. An AI clip maker is the tool category that promises to fix exactly that. Most of what it promises is real. One part of it quietly is not.
What does an AI clip maker actually do?
An AI clip maker takes one long recording and returns several short, vertically framed, captioned clips without anyone scrubbing a timeline. It transcribes the audio, splits the transcript into candidate segments, scores those segments, sets in and out points, crops the frame to follow whoever is talking, burns in captions, and exports each clip at the dimensions its destination feed expects.
Seven stages, and six of them are genuinely mechanical. Machines do mechanical work cheaply, in parallel, without getting bored on clip nineteen.
- Transcribe — timed text, speaker-labelled.
- Segment — split at topic boundaries, not silences.
- Score — rank candidates. This is the stage that disappoints.
- Set the cut — snap to sentence edges.
- Reframe — crop to a vertical frame that tracks the subject.
- Caption — render the text into the picture.
- Export — a master plus per-destination derivatives.
Automatic highlight detection finds loud moments, not interesting ones
Automatic highlight detection ranks moments by the signals that are cheap to compute from a media file: audio energy, laughter, speech density, shot changes and keyword hits. Those signals reliably find the loudest ninety seconds in a recording. They do not find the most useful ninety seconds. A quiet, precise answer to an expensive question scores near zero and is usually the best clip in the file.
My take: this is not a model quality problem that gets solved next quarter. Loudness is measurable and interest is not, so any scorer reading the raw signal drifts toward volume and applause. Ask a panel host which thirty seconds mattered and they will name the part where the room went quiet.
The fix is not a better scorer. So what do you change instead? The input: constrain the search before it ever runs.
- Give it the questions, not the file. Write down the six questions your buyers ask, then have the pipeline retrieve the segment that answers each one.
- Over-generate, then cut hard. Selection is cheap at the shortlist stage and ruinous after captions are burned in.
- Score the transcript, not the waveform. A segment that is self-contained on the page is usually self-contained on screen.
- Mark the moments live. Someone typing timestamps into the chat remains the highest-precision detector available.
The clip you would never pick is the one the tool loves
When we build these pipelines for clients, the first review round almost always rejects the top-scored candidate and promotes something from the bottom of the list. That is a signal to change the input, not the model. We now feed the scorer a written list of target questions and treat unconstrained ranking as a fallback.
How do you reframe to vertical without decapitating the speaker?
Reframing a 16:9 recording into a 9:16 vertical clip throws away roughly two thirds of the picture width, so the crop must follow the subject rather than sit mid-frame. Automatic speaker tracking handles one seated presenter well. It struggles with two-shots, whiteboards and shared screens.
The arithmetic explains most of the ugly output you have seen. Hold the height constant and the width you keep is nine sixteenths divided by sixteen ninths, a little under a third of the frame. If your speaker sat left of centre because the slides occupied the right, a centred crop returns a shoulder and an elbow. If two people share a sofa, a tracking crop swings between them on every exchange and induces motion sickness in seconds. And if the substance lives in the slide rather than the face, cropping to the face deletes the content and keeps the packaging. Reframing is a shooting decision disguised as a post-production one.
Record with the crop in mind and the automation works: subject roughly centred, nothing load-bearing in the outer thirds, one speaker on camera at a time, slides captured separately. When the source will not crop, letterbox it inside a vertical canvas and use the space above and below for a headline and captions. That looks deliberate. A decapitated presenter never does.
Captions, hooks and the first two seconds
Captions and the opening two seconds decide whether a repurposed clip gets watched at all. A clip that opens with pleasantries loses the scroll before the point arrives, and a clip playing silently without captions is unreadable to a muted viewer and inaccessible to a deaf one. Cut the greeting. Start on the claim. Put words on screen from frame one.
Captions are a baseline requirement rather than a nice touch: the W3C's Understanding SC 1.2.2: Captions (Prerecorded) places them at Level A and notes they should identify who is speaking and carry meaningful non-speech sound, not just the words. Automatic captions get names, jargon and Arabic-English code-switching wrong often enough that a human pass is not optional.
- Burned-in and sidecar, both. Pixels are not machine-readable, so burn captions and still upload a caption file where the platform takes one.
- Two to four words per line. Long lines on a vertical frame collide with the interface.
- Treat the bottom third as missing. Every vertical feed overlays controls there.
- Fix names first. A misspelled client or product name burned into the picture is the caption error that costs you something.
Where clips go next is a separate discipline from making them, and the formats, sequencing and posting cadence are covered in our guide to AI video for social media marketing. Clips with no distribution plan are smaller files in the same folder.
How many usable clips should one recording produce?
One hour of recorded conversation typically contains a handful of genuinely self-contained moments, not dozens. An AI clip maker will happily generate thirty candidates from that hour; the number surviving review against your brand and your buyer's real questions is far smaller. Plan for a shortlist you would defend in front of the client.
Budget review time rather than generation time. Generation is a per-minute cost that keeps falling; review is a person with a shortlist, and that cost is flat. If the pipeline produces more candidates than anyone can review, you have not automated the work. You have moved it.
Record for clipping and the yield changes before any processing
The change that has helped the teams we work with most is not a tooling change. It is asking the presenter to restate the question before answering, and to pause for a beat between topics. Both give the segmenter clean boundaries and every clip a built-in opening line, at no cost, before a frame is processed.
Which recordings repurpose well, and which fight you
Interviews and panels repurpose best, because someone is asking questions and that imposes structure a segmenter can find. Slide-heavy webinars and screen-recorded demos repurpose worst, because the meaning sits in a wide frame that a vertical crop destroys. Audio-led podcasts land in between: the words are excellent and the pictures are static, so you are really producing captioned audiograms.
| Source recording | What you can realistically pull | Where it fights you |
|---|---|---|
| Interview or two-person panel | Question-and-answer pairs that stand alone | Cross-talk confuses speaker tracking |
| Conference talk, one presenter | Opinions, definitions, memorable lines | Room audio and a distant camera hurt both captions and crops |
| Podcast, static cameras | Whole arguments, lightly trimmed | No visual variety, so caption design carries the clip |
| Slide-led webinar | Spoken asides between slides | Substance lives in a 16:9 slide that will not crop |
| Product demo or screen recording | Short single-action moments | Interface text becomes unreadable at vertical scale |
| Customer conversation | Specific, quotable problem statements | Consent and confidentiality, which no tool checks |
One caveat I will not soften. Never let a clip pipeline publish regulated, legal, medical or pricing content without a named human approver, because an automatic cut can end a sentence one clause early and reverse its meaning.
Cost per derived clip is the only number that matters
Cost per derived clip decides whether repurposing is worth doing, and it is the easiest number in video to calculate honestly. The footage is sunk cost, already shot and already paid for. That leaves three inputs: processing cost per minute of source, human review minutes per surviving clip, and yield, meaning how many of the batch you would genuinely publish.
Run it on your own last recording before you buy anything. Most teams find processing is trivial and review is the budget, which changes the criteria: you want the pipeline that wastes least of a reviewer's afternoon, not the one generating the most clips.
If you produce one long recording a quarter, do not build a pipeline. Clip it by hand, or use a free self-serve generator — ours sits at free-video-generator.techvisionera.com. Tech Vision Era has been building software since 2010, and we build and operate AI video generation pipelines for clients as code they own, but that engagement only earns its keep with a steady stream of source footage behind it.
So the decision in front of you is not which AI clip maker to sign up for. It is whether your next recording gets produced as a clip source — restated questions, one speaker on camera, a written list of questions each clip must answer — or shot as you always have and handed to software to rescue. Make that choice first, because the second one barely matters afterwards.