Skip to content
cliplimeClipLime

Repurpose one video for every platform

One video in, every social shape out. The crop follows the face, captions fit each shape, and nothing is uploaded.

Runs on your device, nothing is uploaded

or drop it here, or paste from the clipboard

MP4, MOV, M4V, WebM or MKV · one video in, every social shape out

At a glance

What it does
Turns one video into vertical, square, portrait and landscape versions in a single run, with the crop following the main face and optional captions.
Shapes
9:16 (1080×1920), 1:1 (1080×1080), 4:5 (1080×1350) and 16:9 (1920×1080), each in 1080p or the faster 720p.
Framing
Follow faces (the default), a plain center crop, or the whole picture over a blurred background.
Captions
Optional bold word-by-word captions, sized and placed for each shape, plus an SRT file.
Saves
One MP4 per shape (H.264 video, AAC audio, original frame rate), or all of them in one ZIP.
Privacy
Your video is processed on your device and never uploaded. Only the face detector (about 3.5 MB) and the speech model are downloaded, once.
Price
Free, no watermark.

How to do it

  1. 1

    Open a video

    Drop in a landscape or vertical clip. The shape your video already has is greyed out, since there's nothing to reframe.

  2. 2

    Pick your shapes

    Tick 9:16, 1:1, 4:5 and 16:9 as you need them, and choose 1080p or 720p. Keep Follow faces for talking-head videos, or switch to Center or Fit with blurred background.

  3. 3

    Decide on captions

    Leave them on to add bold captions to every version. Pick the spoken language and accuracy. The speech model downloads once and runs on your device.

  4. 4

    Create and download

    Each shape is rendered one after another. Download them one by one, or all together as a ZIP with the SRT file.

Which size goes where

Every platform has a shape it was built around, and a video in the wrong shape gets thin bars, a heavy crop or a blurry stretch. These are the four shapes the repurposer makes:

  • 9:16 vertical, 1080×1920 for TikTok, Instagram Reels and YouTube Shorts, which fill a phone held upright.
  • 1:1 square, 1080×1080 for feed posts on Instagram and Facebook. It looks fine almost everywhere.
  • 4:5 portrait, 1080×1350 for Instagram and Facebook feeds, where it takes up more of the screen than a square.
  • 16:9 landscape, 1920×1080 for YouTube's main player, websites and presentations.

Choose 720p for faster exports and smaller files: the sizes become 720×1280, 720×720, 720×900 and 1280×720. If your source is smaller than the output, the tool warns you that the picture will be enlarged.

How face following works

Before anything is encoded, the tool looks at about four small frames per second and finds faces with MediaPipe BlazeFace, Google's open-source face detector. It runs in your browser, and your frames never leave the tab. The detector is a small download of about 3.5 MB that your browser keeps for next time.

  • One main face. It follows the largest and most persistent face. If it leaves the frame for a while, the next most prominent face takes over.
  • A calm camera. Tiny movements are ignored, pans have a speed limit, and the crop eases into place, so the picture doesn't jitter. When the subject jumps to a new spot, as it does after a cut, the crop slides across quickly instead of snapping.
  • No face for a moment? The crop stays where it was. If no face is found anywhere, the crop is centered.

Cropping some shapes throws away most of the picture. A 16:9 video cut to 9:16 loses about two thirds of it, and so does a vertical video cut to 16:9. With Follow faces, the tool crops those shapes only when it found a face that fits comfortably inside the crop. Otherwise it keeps the whole picture on a blurred background and tells you which shapes it handled that way. Choose Center or Fit with blurred background yourself if you'd rather decide.

Face following works best with clear, fairly large faces: talking heads, vlogs, interviews and presenters. It follows one face at a time and doesn't know who is speaking, and small faces in wide shots can be missed.

Captions that fit every shape

With captions on, the speech is transcribed once on your device with an open-source Whisper model, then burned into every version in a bold style that highlights each word as it's spoken. They are sized and placed for each shape: on 9:16 they sit about 70% of the way down, above the bottom of the screen where TikTok, Reels and Shorts place their own buttons and text, and on wide frames they sit lower.

You also get an SRT file with standard subtitle lines, handy for uploading to YouTube or opening in an editor. Want to fix a misheard word before burning captions in? Use Auto Subtitles, which has a full editor.

Before you post

Apps draw their own buttons and captions over vertical video. Open a finished 9:16 file in the safe zone checker to see what gets covered. If a file is too big for where you're posting, run it through the video compressor.

Questions people ask

What shapes and sizes does it make?

9:16 vertical (1080×1920), 1:1 square (1080×1080), 4:5 portrait (1080×1350) and 16:9 landscape (1920×1080). Choose 720p instead and they come out at 720×1280, 720×720, 720×900 and 1280×720.

How does it keep a face in frame?

It looks at about four small frames per second, finds faces with an open-source detector that runs in your browser, follows the largest and most persistent one, and smooths the path so the crop doesn't jitter. When the face is out of sight the crop stays put, and if there are no faces it stays centered.

What if my video has no faces?

Shapes that need only a light crop are centered. Shapes that would lose most of the picture, such as 16:9 to 9:16, use a blurred background instead. You can also pick Center or Fit with blurred background yourself and it will be used for every shape.

Is my video uploaded anywhere?

No. Face finding, transcribing and rendering all happen on your device. What gets downloaded is the face detector (about 3.5 MB) and, if captions are on, the speech model (for the default accuracy, about 77 MB on a processor or 206 MB on a graphics card, and more for the Accurate setting). Each is downloaded once and cached by your browser.

Will the sound and frame rate stay the same?

Yes. The audio is kept as AAC, and each version keeps the frame rate of your original video.

How long does it take?

It depends on your computer and the length of the video. Each shape is its own pass over the video, so three shapes take roughly three times as long as one. 720p is quicker, and a graphics card speeds up the captions.