Turn a long video into
short clips worth posting.
Drop a podcast, stream or interview. You get vertical clips for TikTok, Reels and YouTube Shorts — cut on the line they open and close on, reframed to 9:16, captioned, with the edit already written down.
Checking your account…
- One recording a day, free
- Up to 2 hours · 500 MB
- Cut and captioned in your browser
- No watermark
One recording in,
a field of clips out.
Each clip is cut on the line it opens and closes on, cropped to 9:16 from a measured face track, and captioned from the same timings — then handed over with an edit sheet timed to the second.
Models cannot reliably timestamp
a long recording.
Ask one for the seconds and it will answer confidently and be wrong, and the further into the file you go the worse it gets. Most clipping tools ask anyway. This one is built the other way round.
What a model is asked
Which moment is worth posting, what it actually says, what to caption it, and what the edit should do. Judgement and language — the things it is genuinely good at.
What is measured instead
Every number. Where the words fall in the audio, where the face is in each frame, where the pauses are, how long the clip runs. None of it is a model's answer.
Six steps, and each one
depends on the last
The order matters. Captions are timed against the boundary the cut really used, and the framing follows a track measured on the clip that was really made.
Read the recording once
The whole file is read on the server in a single pass — audio, transcript and timings together. Nothing is transcribed six times, and nothing depends on which chat you happen to have signed in.
Find the moments worth posting
The loudness envelope narrows a twenty-minute recording down to a shortlist of candidates. A model then reads what is actually said in each one and ranks them — so a moment is chosen for the line in it, not for being the loudest.
Cut on the words, not on a guess
Every clip is named by the line it opens on and the line it closes on. The seconds come from measuring the audio for those words. Ask a model to timestamp a long recording directly and it invents arithmetic — this never asks it to.
Keep the speaker in frame
A face track is measured across the clip frame by frame, weighted towards whoever is facing the camera. A 9:16 crop then follows the real position. Asking a model where somebody was standing picked the reflection on a mirror shot; measuring does not.
Burn in captions that land
Captions come from the same measured timings the cut used, word by word, in five looks including a karaoke-style highlight. A short caption holds only into free time, so it never sits on top of the next one.
Hand you the edit
Each clip arrives with a sheet: the b-roll, the punch-in, the text on screen, the sound effect, the intro and outro card — each one timed to the second, so you can drop them straight into CapCut instead of rewatching to find the beat.
Four things that are true of
every clip it makes
Measured, not guessed
Every number in a finished clip — the in point, the out point, where the speaker is, when a caption appears — comes from measuring the file. Models are asked what a moment is worth and what it says, which is what they are good at.
Run it once, ever
The cuts, the captions and the edit sheet are written down with the workflow. Reopening it asks for nothing and re-cuts nothing, so a clipping run you paid for is a clipping run you keep.
Pieces Omni will take
Flow edits video up to ten seconds at a time. A longer clip is divided on its own pauses — never mid-word — so every piece is something Flow will actually accept, and the joins fall where a listener expects them.
No API keys
Clipping runs inside the same extension as everything else in Studio, against the accounts you already have. There is no key to paste and none stored in the extension.
A clip is the raw material.
The edit is the video.
Every cut arrives with a director's sheet — what to add, and the second to add it at.
| Move | What it does | When it lands |
|---|---|---|
| B-roll | Covers a line that is being told rather than shown | On the noun, held about 1.5–2s |
| Punch-in | Tightens the frame so a line reads as the point | On the emphasised word |
| Text on screen | Puts the number or the claim where it cannot be missed | As it is said, not after |
| Sound effect | Marks the beat a viewer would otherwise scroll past | 0.2–0.4s, on the hit |
| Intro / outro card | Frames the clip for somebody who arrived mid-thought | Before the first word, after the last |
| Whoosh on a seam | Covers a join where two pieces meet | Across the cut itself |
Before you upload
two hours of anything.
Is the AI clip maker free?
Yes. A free account clips one recording a day; Pro raises that to ten. Reading a video runs a model over the whole file, which is the part that costs money — the cutting, reframing and captioning all happen in your own browser and cost nothing.
What happens to my video?
It is uploaded once so the model can read it — audio, transcript and timings in a single pass — and it is deleted when the reading finishes. Every clip is then cut, cropped and captioned locally in your browser, so the footage is not sent anywhere a second time.
How long can the recording be?
Up to two hours and 500 MB, in MP4, MOV, WebM or MKV. A twenty-minute recording is read in well under a minute; the encoding afterwards depends on your machine and how many clips you asked for.
Does it add captions automatically?
Yes, burned into the picture, word by word, in five looks including a karaoke-style highlight. The cue times come from the same measured reading the cut used, so they follow the voice rather than drifting from it.
Will the clips be vertical for TikTok, Reels and Shorts?
9:16 by default, with 1:1, 4:5 and 16:9 available. The crop follows a face track measured across the clip frame by frame, so the speaker stays in shot instead of walking out of a fixed centre crop.
Which browsers does it work in?
Any browser with WebCodecs: Chrome and Edge 94 and up, Safari 16.4 and up, Firefox 130 and up. The page checks before you upload anything and says so if it cannot encode, rather than failing after the reading is paid for.
How does it decide which moments are worth posting?
The loudness envelope shortlists candidates from the audio, then a model reads what is actually said in each one and scores it out of 100. You set the bar. Nothing is chosen for being loud — that only decides where to look.
Do I need the Chrome extension?
Not for this page. AutoFlow Studio runs the same pipeline on a canvas where clips become nodes you can re-cut, re-frame and re-caption without paying for the reading again, and every run here can be downloaded as a Studio workflow.
Stop scrubbing a two-hour recording
looking for the good bit.
Clip one here for free, or run the same pipeline on a canvas in AutoFlow Studio, where every clip becomes a node you can re-cut without paying for the reading again.