How to record a podcast remotely
Recording the call gives you whatever survived the internet. Recording each end locally gives you the original. Almost everything else about remote podcasting is downstream of that one choice.
8 minute read
The one structural decision: where the audio is captured
Remote podcasting contains exactly one decision that changes the outcome, and the rest is detail. Either you record what comes out of the call, or every participant records themselves on their own machine and you assemble the files afterwards. The second approach is called a double-ender, from the days when both ends of a phone interview ran their own tape machine, and it is what nearly every good-sounding remote show does. The difference is not cosmetic. A recording of the call is a recording of whatever survived the trip across the internet: already compressed, already processed, already missing whatever the network lost. A local track is the original microphone signal, captured a few inches from someone's mouth and never transmitted in real time. You cannot turn the first into the second later.
What the call itself does to your audio
Video-conferencing software is engineered for intelligibility on a bad connection, not for archival quality, and it makes that trade continuously and invisibly. It compresses, it suppresses background sound, it cancels echo, and it quietly lowers quality when bandwidth tightens. You can see the trade admitted in Zoom's own description of high-fidelity music mode, which it says disables echo cancellation and post-processing and eliminates audio compression — a premium feature defined entirely by what it stops doing to you. None of those processes are effects you can reverse in an edit. Noise suppression that shaved the front off a word has discarded that audio, not hidden it. Recording the call bakes every one of those real-time decisions in permanently, made on your behalf by an algorithm optimising for a conference call.

Network damage is not an editing problem
The gap between the two approaches is widest at precisely the moment you care about. When a guest's connection stumbles, the call recording captures what the call got: a robotic syllable, a quarter-second of silence where a word should be, half a sentence simply absent. No plugin recovers audio that never arrived, because there is nothing there to repair. Meanwhile that same guest's microphone was feeding their own computer the whole time, at full quality, completely unaware that the network was struggling. This is why remote-recording platforms capture on each participant's device rather than recording the stream. It also means the live call can be as ropey as it likes without affecting the episode, which takes some getting used to.
Latency, and why remote conversations feel slightly wrong
There is a measurable reason remote interviews feel off even on a good line. Tanya Stivers and colleagues, publishing in PNAS in 2009, timed turn-taking across ten typologically diverse languages and found the same pattern everywhere: speakers avoid both overlapping talk and silence, responses peak within roughly 200 milliseconds of a question ending, and average gaps in all ten languages sat within 250 milliseconds of the cross-language mean. Now add a connection. ITU-T Recommendation G.114 treats one-way delay under 150 milliseconds as essentially transparent for interactivity, with 150 to 400 milliseconds merely acceptable. Even at the transparent end, a round trip adds around 300 milliseconds — longer than the gap people are unconsciously listening for. That is why you keep talking over each other, and no amount of politeness fixes it.
Drift, the failure that only shows up at minute forty
Two computers recording the same conversation are running two independent clocks, and those clocks are never exactly equal. One samples fractionally fast, the other fractionally slow, and the error accumulates: the tracks line up perfectly at the top of the episode and slide apart until answers land before the questions that prompted them. Nobody notices in the first ten minutes, which is exactly why it gets discovered late, usually during the edit. Two habits deal with it. Start every session with a sync point visible on a waveform — a countdown and one clap is enough — so you have something to line the files up against. And use a tool built for it: Jason Snell's Double Ender aligns local files to a reference recording and compensates for drift by patching or removing silence, rather than leaving you nudging clips by hand.
Backups, in the order you will actually need them
Assume one leg of this fails, because eventually one does. The strongest arrangement is a platform that records locally and uploads while you are still recording. Zencastr describes exactly this: each participant's audio and video is captured to their own browser storage, then progressively uploaded in pieces during the session, so the call is not disrupted and nobody waits for a large transfer at the end. Progressive upload matters because the classic disaster is a guest closing the tab before their file has left their machine. Underneath that, keep the call recording as well — poor audio is still a perfectly good reference for syncing and for reconstructing a lost passage. And a phone recording a voice memo on the desk has rescued more episodes than any premium plan.
The guest is the weak link, and that is your job
Your guest has not done this before and will not read a long brief. Send five lines the day before: wired headphones so their speakers do not leak back into the microphone, the most wired connection they can manage, everything else on the machine closed, phone on silent, and the browser their recording platform actually supports. Then say the sixth thing out loud at the end of the session rather than in writing: stay in the room until the upload finishes. Because local recordings live on the guest's device before they live anywhere else, the minute after you stop recording is the most fragile minute of the whole process. Say it before you start and again before anyone leaves.
What clean tracks buy you afterwards
Isolated per-speaker tracks are worth the setup for their own sake, but the real payoff lands further downstream. Separate audio means you can cut one person's cough without touching the other person's answer, and it means every clip you pull later starts from full-quality speech rather than something the internet chewed. If the video is recorded locally too, you have a usable full-resolution frame per speaker instead of a conferencing grid. That is what makes the clipping pass quick rather than painful: FrameOS takes the finished recording, surfaces the moments worth posting, reframes to vertical while following whoever is speaking, and captions them for silent feeds. 300 credits for 3 days · no card. The decisions you made before the interview started are what decide whether any of it is possible.
FAQ
What is a double-ender podcast recording?
Each participant records their own microphone and camera locally on their own machine while the conversation happens over a call. Afterwards the separate files are assembled into one timeline. The call carries the conversation but is not the recording, so network problems affect what you hear live and not what you publish.
Is it OK to record a podcast on Zoom?
It works, but you are recording a processed, compressed stream. Zoom's own high-fidelity music mode is described as disabling echo cancellation and post-processing and eliminating compression, which tells you what the normal path does. If you use Zoom, record each participant to a separate track and treat the file as a backup and sync reference rather than your master.
Why do my remote podcast tracks drift out of sync?
Because each machine has its own sample clock and no two are exactly equal. A tiny per-second error accumulates until, an hour in, the tracks are visibly misaligned. Record a shared sync point at the start, and use a syncing tool such as Double Ender that corrects drift by patching or trimming silence instead of relying on a single alignment at the top.
How do I stop a remote guest's audio from dropping out?
You cannot stop the call dropping out, but you can stop it reaching your episode. Record locally on the guest's device so their microphone signal never depends on the connection. Then reduce the risk to the upload: use a platform that uploads progressively during the session, and keep the guest in the room until their file has finished transferring.
Sources
Related pages
Keep reading
More Podcasts →The best time to post on Instagram, according to four studies
Buffer says Thursday 9am. Later says 5am. Sprout says early afternoon. They analysed billions of posts and still disagree — here is why, and what to do about it.
8 minute readThe best time to post on TikTok, according to five studies
Buffer analysed 7.1 million posts and rates Saturday the best day. Sprout Social analysed 2 billion engagements and says avoid the weekend entirely. Both are right, for different accounts.
8 minute readThe best time to post on LinkedIn, according to five studies
Buffer's 4.8 million posts point to Wednesday at 4pm and rank Tuesday among the worst days. Sprout Social's 2 billion engagements call Tuesday the best day of the week. Here is why.
8 minute readTurn one long video into a week of posts
300 credits for 3 days · no card.
Start free