The first time I edited a podcast properly, I was staring at 43 minutes of raw tape: a guest who said “right?” like a nervous tic, an eight-minute tangent about espresso grinders that went nowhere, and a door slamming shut at minute 18. Doing that in a traditional NLE would have eaten a Saturday. In Descript it took under two hours, and I never once scrubbed a waveform.
What follows is the order of operations I’ve settled on after a couple of dozen episodes. It assumes spoken-word content — podcast, interview, tutorial, talking head — and that you’d rather think like an editor than a button-pusher. If you’re new to the tool, it’s worth understanding why Descript treats a video like a text document, because this entire workflow rests on that one idea.
Step 1: Set up the project so the transcript becomes your editing surface
Start with one composition per episode, not one per scene. Import your audio and video together and check the transcription language before anything else. A six-minute transcription that comes back in the wrong language is six minutes you never get back.
Then do the boring bit that saves hours later: fix names, jargon and product names in the transcript immediately. Every correction you make flows through to captions, search and any clips you export down the line. I keep a running list for my show — my guest’s surname spelled the way their own website spells it, plus a handful of technical terms — and I paste it in before I watch a single frame.
Speaker labels matter too. Descript guesses who’s talking and usually gets it right, but a mislabelled guest turns your speaker-based edits into a mess. Ten seconds of clicking now beats thirty minutes of confusion halfway through.
Step 2: The first pass is reading, not watching
Read the transcript at normal reading speed instead of listening at 1.5x. It’s faster, and it forces you to judge the content rather than the performance. Anything that doesn’t earn its place gets deleted outright, and the audio goes with it.
That espresso tangent? Gone. Not trimmed, deleted — all eight minutes of it, in about four seconds of work. The episode was instantly tighter and nobody noticed the missing section, because it never belonged in the rough cut.
Drop markers instead of stopping
The temptation is to fix problems the moment you spot them. Don’t. When something needs a second look, drop a marker and keep reading. I’ve finished read-throughs with 14 markers left; clearing them afterwards took 20 minutes, whereas interrupting myself each time would have cost me the thread of the whole story. Mark:
- Sentences you’ll want to patch or re-record later
- Moments that need a visual — a screen recording, a chart, a demo
- Anything factually shaky that has to be checked before publish
Step 3: Strip filler words and dead air in one pass
Filler word removal is the biggest single time-saver in the tool, but the default settings are greedy. I run it on “um”, “uh” and “you know”, and leave “like” and “right” alone unless the speaker is genuinely leaning on them. Those words are often part of somebody’s rhythm, and stripping them makes a guest sound like a robot reading a statement.
The word gap setting is where people go too far. Crank it and you’ll tighten every pause, including the ones that let a joke land or a heavy answer breathe. I start around 0.35 seconds, listen to the result, then back it off if the conversation starts feeling frantic. On one recent episode this single pass removed 212 fillers plus roughly nine minutes of dead air, taking the runtime from 43:12 down to 38:50 — exactly the range I wanted for the feed.
Step 4: Fix the audio before you judge the edit
Everything sounds worse before it sounds better. If your room has a hum, a laptop fan or passing traffic, clean it now so you aren’t misjudging your own cuts later. Studio Sound handles most of it in one click, and it’s genuinely good at rescuing mediocre recordings — a USB mic in a carpeted room comes back sounding like it was cut in a booth.
What it won’t fix: a plosive that clips, or a level spike when someone laughs into the mic. Those you handle by hand, pulling the volume down for that moment only. Resist the urge to flatten the whole track; a little dynamic range is what stops a two-person show sounding like an audiobook. If you use a music bed, keep the voice on top — roughly -6 dB peaks on speech, bed sitting 15 to 20 dB below it.
Step 5: Patch mistakes by typing, not re-recording
Fluffed a line? Descript’s voice cloning lets you type the corrected sentence and generate it in your own voice, without setting up a mic again. I’ve used it to fix a date I said wrong, a name I mangled, and one sentence where a plane flew over mid-clause. Twenty seconds each, no retake.
Where it falls apart is scale. Whole new paragraphs, or any line delivered with a laugh, come back flat — and listeners notice within a sentence. My rule is one word to one sentence, always mid-thought, never opening a segment. Anything longer, I re-record the paragraph for real.
Step 6: Build the visuals without leaving the timeline
Scenes and layouts cover most of it. For an interview I run a stacked layout: two speakers side by side with the active one slightly larger, switching to full screen when the conversation heats up. For a screen-share tutorial, a small presenter box in the corner is plenty.
B-roll is where it gets fun. Drag a clip in and the timeline handles placement; edit a sentence in the transcript and the visual follows the words. When stock footage doesn’t fit, I generate an image instead — the same prompt discipline that makes AI image generation reliable applies here. Set the aspect ratio to 16:9 before you generate, or you’ll spend the next ten minutes re-framing everything.
Step 7: Export once per destination, then stop
Every episode gets exactly two exports: the master video and an audio-only file for the podcast host. Save both as presets so you aren’t rebuilding settings each week — YouTube 1080p for video, WAV or high-bitrate MP3 for audio, loudness around -16 LUFS so your show isn’t noticeably quieter than everyone else’s in a queue.
Captions are free because the transcript already exists. Export a caption file alongside the video rather than burning them in; upload it to YouTube and let the platform index it properly. Before publishing, run a short export of the first 90 seconds and watch it on your phone. Treat that like a staging deploy before a production release — the same discipline that keeps a first API call from becoming a production incident. Half my export mistakes only show up through a phone speaker or in a vertical crop.
Make episode two take 40 minutes instead of four hours
Here’s the part most walkthroughs skip. The real win isn’t editing this episode faster — it’s that the next one begins from a finished setup. Once the first project looks and sounds right, save it as a template: layouts, export presets, intro sequence, lower-third style, music levels. Next week you import two files and jump straight to the read-through.
Worth locking down in that template:
- Naming conventions for files and compositions, so you can find episode 41 in eight months
- A short pre-publish checklist — captions exported, loudness checked, transcript spellings corrected
- Only the three or four layouts you actually use, nothing aspirational
It’s the same principle behind any practical, step-by-step workflow in a tool you use daily: the first run teaches you the steps, and every run after that executes the system you built. I now spend about 35 minutes on an episode that used to take an afternoon, and nearly all of it goes to the one thing that genuinely matters — deciding what stays in the story.

