Everything an episode needs, from one file
Publishing a podcast usually means five small jobs in five different tools. Podcast Studio does them in one run, from a single audio or video file, and you only tick what you need.
- Mastered audio: an MP3 at a consistent podcast loudness, so listeners don't reach for the volume between shows.
- Transcript: plain text for show notes and the web, plus SRT and VTT caption files for YouTube and video players.
- Suggested chapters: timestamps for your YouTube description, found from the pauses and topic changes in the episode.
- Audiogram videos: a clip with your cover, a moving level meter and optional captions, in 16:9 for YouTube and 9:16 for Shorts and Reels.
- Thumbnail picks (video episodes): six frames to choose from, then a 1280×720 PNG or JPG with a big title on it.
Everything runs in your browser. The episode is never uploaded, and nothing is sent anywhere except the one-time download of the open-source speech model.
How the mastering works
Podcast apps play shows at very different volumes unless the audio is mastered. ClipLime measures the integrated loudness of the whole episode (ITU-R BS.1770, the method behind LUFS), then applies one fixed gain to reach the target. A look-ahead limiter catches only the peaks that would otherwise pass the −1 dBTP ceiling, so the voice keeps its natural dynamics.
The audio is read in small pieces instead of being loaded whole, so the length of an episode doesn't limit mastering. Mono recordings are saved as stereo, because the −16 LUFS figure is meant for stereo files: a mono file measured as one channel would need about 3 dB less to sound equally loud.
MP3 encoding removes a little level, usually a few tenths of a LU. So the saved file is measured again, and if it landed noticeably off target the gain is corrected and the MP3 is made once more. The loudness shown on the results page is measured on the saved MP3 itself. To check any file yourself, use the loudness meter.
How the chapters are chosen
Two signals decide where a chapter could begin: a long pause in the audio, and a change of vocabulary, where the words just after a point share few topic words with the words just before it. Boundaries are kept at least two minutes apart in a normal episode, and closer together in a short clip so that YouTube's minimum of three chapters still fits.
The first chapter always starts at 0:00, and each title is simply the first few words of the section's opening sentence. That makes them a starting point, not a summary, so the page labels them as suggested and lets you edit every time and title before you copy them. The checker uses the same rules as the chapter maker: start at 0:00, at least three chapters, ten seconds or longer each.
Long episodes and slower computers
Speech recognition is the slow part, and its speed depends on your computer and the model you pick. A graphics card with WebGPU is much faster than a processor, and the Fast model is quicker than Balanced, which is quicker than Accurate. The page shows an estimate for each model, and after your first run it uses the speed your own computer reached.
The transcript step keeps the whole episode in memory as speech-ready audio, about 0.25 GB per hour, so a three-hour show needs around 0.75 GB. A phone or an older laptop may not have that to spare, and a very long show is better transcribed in parts. Mastering does not have this limit.
Audiograms are limited to three minutes so they encode quickly. Pick the moment you want to share, or use the clip finder to look for it first.