Which size goes where
Every platform has a shape it was built around, and a video in the wrong shape gets thin bars, a heavy crop or a blurry stretch. These are the four shapes the repurposer makes:
- 9:16 vertical, 1080×1920 for TikTok, Instagram Reels and YouTube Shorts, which fill a phone held upright.
- 1:1 square, 1080×1080 for feed posts on Instagram and Facebook. It looks fine almost everywhere.
- 4:5 portrait, 1080×1350 for Instagram and Facebook feeds, where it takes up more of the screen than a square.
- 16:9 landscape, 1920×1080 for YouTube's main player, websites and presentations.
Choose 720p for faster exports and smaller files: the sizes become 720×1280, 720×720, 720×900 and 1280×720. If your source is smaller than the output, the tool warns you that the picture will be enlarged.
How face following works
Before anything is encoded, the tool looks at about four small frames per second and finds faces with MediaPipe BlazeFace, Google's open-source face detector. It runs in your browser, and your frames never leave the tab. The detector is a small download of about 3.5 MB that your browser keeps for next time.
- One main face. It follows the largest and most persistent face. If it leaves the frame for a while, the next most prominent face takes over.
- A calm camera. Tiny movements are ignored, pans have a speed limit, and the crop eases into place, so the picture doesn't jitter. When the subject jumps to a new spot, as it does after a cut, the crop slides across quickly instead of snapping.
- No face for a moment? The crop stays where it was. If no face is found anywhere, the crop is centered.
Cropping some shapes throws away most of the picture. A 16:9 video cut to 9:16 loses about two thirds of it, and so does a vertical video cut to 16:9. With Follow faces, the tool crops those shapes only when it found a face that fits comfortably inside the crop. Otherwise it keeps the whole picture on a blurred background and tells you which shapes it handled that way. Choose Center or Fit with blurred background yourself if you'd rather decide.
Face following works best with clear, fairly large faces: talking heads, vlogs, interviews and presenters. It follows one face at a time and doesn't know who is speaking, and small faces in wide shots can be missed.
Captions that fit every shape
With captions on, the speech is transcribed once on your device with an open-source Whisper model, then burned into every version in a bold style that highlights each word as it's spoken. They are sized and placed for each shape: on 9:16 they sit about 70% of the way down, above the bottom of the screen where TikTok, Reels and Shorts place their own buttons and text, and on wide frames they sit lower.
You also get an SRT file with standard subtitle lines, handy for uploading to YouTube or opening in an editor. Want to fix a misheard word before burning captions in? Use Auto Subtitles, which has a full editor.
Before you post
Apps draw their own buttons and captions over vertical video. Open a finished 9:16 file in the safe zone checker to see what gets covered. If a file is too big for where you're posting, run it through the video compressor.