Skip to content
cliplimeClipLime

Publish your podcast episode in one go

Drop an episode and get a mastered MP3, transcript, chapters, audiograms and thumbnails, all made on your device.

Everything runs on your device, your episode is never uploaded

or drop it here, or paste from the clipboard

MP3, WAV, M4A, FLAC, MP4, MOV, MKV or WebM · audio or video · long episodes welcome

At a glance

What it makes
A mastered MP3, a transcript (TXT, SRT and VTT), suggested chapters, audiogram videos in 16:9 and 9:16, and thumbnail picks when the episode is a video.
Opens
MP3, WAV, M4A, FLAC and OGG audio, or MP4, MOV, MKV and WebM video. You can add a square JPG or PNG as cover art.
Loudness
−16 LUFS, the level Apple Podcasts asks for, or −14 LUFS for Spotify and YouTube. A limiter aims to keep true peaks under −1 dBTP, and the saved MP3 is measured again so the results show both numbers.
Transcript
Whisper speech recognition running on your device, in 30+ spoken languages and three model sizes.
Audiograms
Up to 3 minutes from any start time, as 1280×720 (16:9) and 720×1280 (9:16) MP4, or 1080p, with optional burned-in captions.
Privacy
Your episode is decoded and processed in your browser. Only the open-source speech model is downloaded, once.
Price
Free, no sign-up, no watermark.

How to do it

  1. 1

    Drop in your episode

    Choose an audio or video file, give it a title, add a square cover if you have one, and tick the outputs you want.

  2. 2

    Pick the details

    Choose the spoken language and speech model, the loudness target, and where the audiogram clip starts and how long it runs.

  3. 3

    Make everything in one run

    Each output has its own step with live progress. You can stop at any time and keep whatever already finished.

  4. 4

    Check and edit

    Play the master, skim the transcript, fix the suggested chapter titles, choose a thumbnail frame and type its title.

  5. 5

    Download

    Take each file on its own, or everything at once as one ZIP.

Everything an episode needs, from one file

Publishing a podcast usually means five small jobs in five different tools. Podcast Studio does them in one run, from a single audio or video file, and you only tick what you need.

  • Mastered audio: an MP3 at a consistent podcast loudness, so listeners don't reach for the volume between shows.
  • Transcript: plain text for show notes and the web, plus SRT and VTT caption files for YouTube and video players.
  • Suggested chapters: timestamps for your YouTube description, found from the pauses and topic changes in the episode.
  • Audiogram videos: a clip with your cover, a moving level meter and optional captions, in 16:9 for YouTube and 9:16 for Shorts and Reels.
  • Thumbnail picks (video episodes): six frames to choose from, then a 1280×720 PNG or JPG with a big title on it.

Everything runs in your browser. The episode is never uploaded, and nothing is sent anywhere except the one-time download of the open-source speech model.

How the mastering works

Podcast apps play shows at very different volumes unless the audio is mastered. ClipLime measures the integrated loudness of the whole episode (ITU-R BS.1770, the method behind LUFS), then applies one fixed gain to reach the target. A look-ahead limiter catches only the peaks that would otherwise pass the −1 dBTP ceiling, so the voice keeps its natural dynamics.

The audio is read in small pieces instead of being loaded whole, so the length of an episode doesn't limit mastering. Mono recordings are saved as stereo, because the −16 LUFS figure is meant for stereo files: a mono file measured as one channel would need about 3 dB less to sound equally loud.

MP3 encoding removes a little level, usually a few tenths of a LU. So the saved file is measured again, and if it landed noticeably off target the gain is corrected and the MP3 is made once more. The loudness shown on the results page is measured on the saved MP3 itself. To check any file yourself, use the loudness meter.

How the chapters are chosen

Two signals decide where a chapter could begin: a long pause in the audio, and a change of vocabulary, where the words just after a point share few topic words with the words just before it. Boundaries are kept at least two minutes apart in a normal episode, and closer together in a short clip so that YouTube's minimum of three chapters still fits.

The first chapter always starts at 0:00, and each title is simply the first few words of the section's opening sentence. That makes them a starting point, not a summary, so the page labels them as suggested and lets you edit every time and title before you copy them. The checker uses the same rules as the chapter maker: start at 0:00, at least three chapters, ten seconds or longer each.

Long episodes and slower computers

Speech recognition is the slow part, and its speed depends on your computer and the model you pick. A graphics card with WebGPU is much faster than a processor, and the Fast model is quicker than Balanced, which is quicker than Accurate. The page shows an estimate for each model, and after your first run it uses the speed your own computer reached.

The transcript step keeps the whole episode in memory as speech-ready audio, about 0.25 GB per hour, so a three-hour show needs around 0.75 GB. A phone or an older laptop may not have that to spare, and a very long show is better transcribed in parts. Mastering does not have this limit.

Audiograms are limited to three minutes so they encode quickly. Pick the moment you want to share, or use the clip finder to look for it first.

Questions people ask

Is my episode uploaded anywhere?

No. Your episode is decoded, measured, transcribed and encoded inside your browser. The only download is the open-source Whisper speech model, which is fetched once and then cached by your browser.

How long does the transcript take?

It depends on your computer. A graphics card running WebGPU is much faster than a processor, and the Fast model is quicker than the Accurate one. Before you start, the page shows an estimate for each model, and once you have run it, the estimate uses the speed your computer actually reached.

Which loudness target should I pick for a podcast?

Use −16 LUFS, which is what Apple Podcasts asks for and a sensible level for spoken word anywhere. Choose −14 LUFS if the episode is going mainly to Spotify or YouTube. Either way, a peak limiter aims to keep true peaks under −1 dBTP, and the results page shows the peak it measured.

Can I edit the chapters and the transcript?

The chapter times and titles are editable on the page, and the YouTube description block updates as you type. The transcript files are plain text, SRT and VTT, so any editor can fix them. To fix captions line by line against the video, use Auto Subtitles.

How long an episode can it handle?

Mastering reads the file in small pieces, so length is not a limit. The transcript holds the episode in memory as speech audio, about 0.25 GB per hour, so a three-hour show needs around 0.75 GB, which a phone or an older laptop may not have to spare. For those, transcribe the episode in parts.

Will the audiograms work on YouTube, Shorts and Reels?

Yes. They are MP4 files with H.264 video and AAC audio. The 16:9 version is 1280×720 (or 1920×1080) and the 9:16 version is 720×1280 (or 1080×1920), at 30 frames per second, up to three minutes long.