Skip to content

An in-person interview can get away with a single microphone picking up both voices. A remote one can't — mixed together, a host talking over a guest becomes a muddy overlap that no editor can cleanly separate afterward. Getting a remote interview recorded properly means keeping the two sides apart from the first second, which is a recording decision made before the call starts, not something fixed in post.

38:00
Microphoneyou, in the room
System audioeveryone on the call
The host's microphone and the guest's voice, arriving through the call's own audio, kept as two separate tracks for the whole interview.

Why two tracks matters more for a podcast than a meeting

For a meeting, a single combined recording is often fine — nobody's re-cutting it. For a podcast episode, the recording is raw material for an edit: trimming a long pause, removing a false start, balancing levels between a host who speaks quietly and a guest who doesn't. All of that is far easier, and often only possible at all, when each voice sits on its own track rather than blended into one. A recorded interview with the two tracks kept separate is a recording an editor can actually work with; one that isn't often isn't worth having recorded to begin with.

Step 1 — get the guest on a call the ordinary way

Nothing about this setup asks the guest to change how they join. Zoom, Google Meet, a phone call routed through the Mac's own audio — whatever's already comfortable for them works, because recording reads the audio the Mac is already playing rather than requiring a specific platform or a plugin installed on the guest's end. There's no bot joining as a visible participant and nothing appearing on the call for the guest to notice, which matters for an interview subject who might otherwise feel like they're being recorded by a third party rather than talking to the host.

Step 2 — start recording before the "so, tell me about yourself" moment

The habit that saves the most retakes is starting the recording as part of joining the call, not once the conversation has warmed up. There's no way to recover audio from before recording started, and the first genuine, unrehearsed answer to an opening question is exactly the kind of moment that's expensive to lose. Pressing record a minute before dialing in costs nothing; losing the best line of the interview to a late start costs the whole point of doing an interview at all.

Step 3 — the two tracks build themselves once recording starts

The host's microphone captures the room; the call's own audio — the audio the Mac is playing, which is where the remote guest's voice actually comes from — is captured separately. Neither track needs to be set up individually; both start the moment recording does, and both keep recording for the length of the call without either one needing attention partway through.

Step 4 — the interview proceeds exactly like any other call

From the guest's side, nothing distinguishes this from any other video or phone call — there's no separate app open, no visible recording indicator specific to a third-party tool, and no participant on the call besides the host and the guest. That matters for getting a natural, unselfconscious interview rather than one where the subject is performing for an obvious recording setup.

Step 5 — a transcript is ready by the time the call ends

Because transcription runs segment by segment while the call happens, the interview is already text by the time it's over — useful for pulling a quote for show notes, checking a name or a statistic was said correctly, or finding the exact moment a particular story started without scrubbing back through forty minutes of audio. Turning that transcript into something publishable is a separate step from producing it, but having the text ready immediately shortens everything downstream of the recording itself.

Step 6 — hand the raw tracks to an editor, because that's the next job

This is where it's worth being precise about what this workflow does and doesn't cover. Recording, transcribing and summarizing an interview is the whole of it — there's no timeline, no cutting, no leveling, and nothing that produces a finished, publishable episode. Coii AudioNotes against Descript covers that distinction directly: if the recording is going to be edited and published, a dedicated editing tool is the next stop, and this one hands off two clean, separated tracks and a transcript to make that edit easier, rather than attempting the edit itself.

What "no bot" actually buys an interview specifically

For a podcast interview, a visible meeting participant recording the call changes the dynamic in a way it might not for an internal team meeting — a guest who notices a bot has joined is a guest who's now aware they're being recorded by something beyond the host, which can visibly change how candid an answer is. Recording without anything joining as a participant means the only thing different about a recorded interview, from the guest's side, is that it's being recorded at all — which the host tells them directly, the way any interview should start.

Getting consistent levels between two very different setups

A common problem with remote interviews is a host on a good microphone next to a guest on a laptop's built-in one, with a volume gap that's awkward to listen to. Keeping the two on separate tracks doesn't fix that gap by itself, but it makes fixing it possible afterward — an editor can raise the guest's track and lower the host's independently, which is exactly the adjustment that's impossible once two voices have already been mixed down into one.

Briefing the guest before the call starts

A short note before recording begins — "this call is recorded, the audio stays on my Mac, I'll send you a transcript afterward if you'd like one" — does more for a natural interview than any technical setup. Guests who know what's happening with the recording tend to relax into the conversation faster than ones left wondering; the two-track, no-bot setup described above means there's nothing visibly unusual on the call to explain, which makes that one honest sentence at the start sufficient rather than the beginning of a longer disclosure.

Handling a guest with a genuinely bad connection

Not every remote guest calls in on a strong, stable line — a spotty hotel wifi or a phone call from somewhere with weak signal is a real and common scenario for a podcast booking someone's time from wherever they happen to be. Recording still captures whatever the Mac's system audio actually plays, which means a guest's dropouts or compression artifacts on their end show up in the recording exactly as they were heard live — this setup doesn't fix a bad connection, it only faithfully captures the call as it happened. Where possible, asking a guest to call in from the most stable connection available to them is worth doing before the interview, since nothing downstream can recover audio quality the call itself never had.

Recording a backup line for a guest who might drop

For interviews where losing the connection mid-conversation would be a real loss — a guest who's hard to re-book, or a conversation that's clearly going somewhere good — some hosts ask a guest to also record their own side locally on their own device, as a phone voice memo or similar, purely as insurance against a dropped call. That's a decision made entirely outside this workflow, on the guest's own device, and worth mentioning upfront rather than discovering only after a call has actually dropped that no backup existed.

What happens with a three-person conversation

Not every interview is one host and one guest. A panel discussion or a co-hosted interview with two hosts and one guest still separates by voice rather than by which side of the call someone is on — voice print matching identifies each speaker from how they sound, so a three-way conversation still comes back with each person's lines correctly attributed, whether they're in the room or joining remotely.

Recording a series, not just one episode

For a podcast produced on a regular schedule, each interview is its own separate recording, transcript and set of minutes — there's no cross-episode linking or a shared show timeline the app maintains on its own. That's a smaller feature set than a production tool built specifically for running a show, and it's worth knowing upfront: this handles the recording and the transcript for one conversation at a time, cleanly, and leaves the show-level organization to whatever workflow already runs the rest of the production.

Using the transcript to write show notes, not just pull quotes

Beyond a single pulled quote, the full transcript is often the fastest starting point for writing show notes — a rough outline of the topics covered, in the order they came up, without needing to re-listen to the whole episode to remember what was discussed and when. Skimming the transcript for the handful of moments worth timestamping in the published notes takes a fraction of the time re-listening would, and because the text is searchable, finding the specific point where a particular topic started is a search rather than a scrub through the audio.

What a series of interviews looks like after a few months

Recording a regular interview show this way builds an archive that's more than just audio files sitting in a folder — every past guest has a transcript attached, searchable the same way a single meeting's would be. Coming back months later to check whether a particular topic was already covered with a different guest, or pulling a specific phrase said in episode twelve for a promotional clip, works the same way finding an old meeting does: a text search across files that were never behind an account requiring the show's original recording setup to still be intact.

What this setup doesn't need from the guest, ever

Across the whole process, the guest never installs anything, never creates an account, and never has their side of the call touch a server beyond whatever the call platform itself already uses. That's a meaningfully lower bar for a guest to clear than a dedicated remote-recording tool that asks them to join through a browser plugin or a separate recording client before the interview can even start — the tradeoff being that this records what the Mac can hear rather than a higher-quality feed captured at the guest's own end and uploaded separately afterward.

Scheduling a run of interviews back to back

Recording several guest interviews in one day — common when a show batches recording sessions ahead of a run of episode releases — needs nothing extra beyond starting and stopping a new recording for each guest. Every interview stays a separate entry in the archive, with its own tracks, transcript and minutes, rather than getting appended to a single running file; nothing about finishing one interview requires closing the app or resetting anything before the next guest calls in.

Recording on more than one Mac for a co-hosted show

A licence covers up to three Macs, which matters for a show with more than one host recording from different locations — each host's Mac keeps its own separate archive, so a co-host recording their own side isn't the same as a single shared recording of the whole conversation. For a show that wants one definitive recording rather than two partial ones, it's worth deciding upfront which single Mac is doing the actual recording for the episode, with everyone else simply on the call.

What it costs, what it runs on

None of the above requires more than the app itself: $19 once, a 30-day trial with no card and no account, running on macOS 13 Ventura or later, Apple Silicon or Intel. What a podcaster's workflow looks like day to day beyond one interview covers the rest of it — prepping, recording a series, and keeping a searchable archive of every guest who's been on the show.

Questions

Does the guest need to install anything to be recorded this way?
No. The call happens in whatever the guest already joins with — Zoom, Meet, a phone line routed through the Mac — and nothing about their side of it changes.
Are the host and guest recorded on separate tracks?
Yes. The microphone in the room and the call's own audio — which is where the remote guest's voice comes through — are captured as two separate tracks from the start.
Can this replace a dedicated podcast editing tool?
No. This records and transcribes the conversation; it has no timeline, no editing and nothing to publish. The recording it produces is the raw material an editor works from afterward.
Does recording this way require the guest's call quality to be re-uploaded anywhere?
No — nothing about either track is sent anywhere during or after the call. What's captured locally is what exists; there's no separate upload step for either side.