How the Clip Finder chooses clips
There is no mystery model behind the scores. The Clip Finder builds many candidate windows, each about as long as you asked, and scores every one with a short list of signals you can see. The windows start and end where a sentence does, or at a pause when there is no transcript, so a clip never opens or stops in the middle of a word.
A transcript scan combines five signals. The numbers are how many of the 100 points each one is worth.
- Opening line, 25 points. A question, a number, or a hook word such as secret, mistake, never, best, nobody or how to in the first line. A start on a connecting word like so or but counts against it.
- Speech density, 20 points. How many words a second are spoken compared with the rest of the file.
- Energy and peaks, 20 points. How loud the voice is compared with the rest of the file, and how many emphatic bursts there are.
- Few long pauses, 20 points. Pauses of nearly a second or more inside the clip cost points.
- Clean ending, 15 points. A clip that ends on a complete sentence scores in full, one cut off mid-sentence scores nothing.
A quick scan has no words to read, so it scores speech density (30 points), energy and peaks (35), few long pauses (20) and how clear the pauses are around the start and end (15).
Scores compare a clip with the other windows in the same file. That makes them useful for picking the best moments of a recording, and meaningless for comparing two different recordings. They say what stands out, not how a clip will perform online.
The top clips never overlap, and each pick is nudged away from the ones already chosen so the results spread across the whole file instead of bunching up in one great five minutes.
Transcript scan or Quick scan?
Use a transcript scan whenever there is speech. It runs an open speech model (Whisper) on your own device, finds every word with its timing and can tell a strong opening from a weak one. The first time, it downloads the model, from about 41 MB for the Fast one (more when it runs on a graphics card, and for the Balanced and Accurate ones), and your browser keeps it afterwards. Choose the spoken language yourself for better results.
Use a quick scan for music, very long files or when you just want a fast look. It reads only the audio track, so it needs no download and finishes much sooner, but it can only judge energy, pauses and how continuous the speech is.
Hook words are matched in English only. Question marks and digits count in every language the speech model supports.
Fast cuts and precise cuts
Compressed video stores a complete picture only every so often, at a keyframe. A copy that skips re-encoding has to begin on one, so a Fast clip may start a little before the moment you chose. The clip editor shows where it would begin, and the finished list shows where each one did.
When the nearest keyframe is too far from the start, usually because the original has long gaps between keyframes, that clip is cut precisely instead so it doesn't open with unrelated footage. Precise always re-encodes at a bitrate matched to the original, so the cut is exact.
Audio has no keyframes, so a fast audio clip is exact to within about 25 milliseconds. An MP3 saved as MP3 or an AAC file saved as M4A is copied untouched. Any other combination is converted.
After you have your clips
For social video, Make it vertical sends a clip straight to the Repurposer, which reframes it for phones. Auto Subtitles can caption a clip, Trim Video fine-tunes one frame by frame and the Video Compressor shrinks it for chat apps.