An interview recording that only captures one side clearly is worse than useless — it's a transcript with half the conversation missing exactly where the good quote was. The setup that avoids this is different depending on whether you're sitting across a table from someone or talking to them through a screen, and getting it right the first time matters more here than almost anywhere else, because you rarely get to redo an interview.
That's the constraint that makes interview recording a slightly different problem from recording an ordinary meeting. In a meeting, missing a few seconds of someone talking over someone else is a minor annoyance you can usually reconstruct from context. In an interview, the moment you lose might be the one honest, unguarded answer the whole conversation was for — and a second attempt at that same question a week later rarely produces the same response. Setting the recording up correctly before you start is worth the extra two minutes it takes.
In person: one microphone, two voices
If you and the person you're interviewing are in the same room, your Mac's built-in microphone — or an external one plugged in, if you're using one — is the only input that matters. Start the recording before the conversation begins rather than partway through; there's no way to retroactively capture the opening exchange once it's passed. Coii AudioNotes transcribes as the conversation happens, segment by segment, and tells the two voices apart by comparing them against each other rather than needing either of you to be enrolled ahead of time.
Positioning matters more than most people expect. A microphone equidistant from both of you, or slightly favoring whichever voice tends to speak more quietly, gives the clearest separation later. If one person trails off at the end of sentences — a common habit in someone who's more comfortable listening than being recorded — sitting them a little closer to the Mac than you'd otherwise think necessary is worth it.
Over a call: two tracks instead of one
A remote interview is a different problem, because the two voices don't arrive the same way. Your own voice comes in through your microphone; the person you're interviewing comes in through your speakers, as system audio from whatever call software you're both using. Recording both of those as one combined track is how you end up with a muddy interview where an overlapping word from either side becomes unintelligible.
Coii AudioNotes keeps these as two separate tracks from the start — your microphone and the system audio, captured apart rather than blended — which is what makes it possible to record both sides of a call clearly without a bot joining as a visible participant. Nothing has to be invited into the call; the recording happens the same way it would if you simply asked your Mac to listen to what it can already hear.
Getting a transcript you can actually quote from
The value of recording an interview at all, rather than just taking notes live, is being able to pull an exact quote later instead of a paraphrase you're not fully sure about. A few things make that transcript more reliably quotable:
Let the interview breathe before you start asking hard questions. The first minute or two of clean, uninterrupted speech from each voice gives the automatic speaker separation the clearest signal to work from for the rest of the recording — a chaotic, interrupting opening gives it less to go on.
State names out loud early, even if you already know them — "Thanks for making time, Dr. Alvarez" gives you a natural point in the transcript to rename the placeholder speaker label to something you'll recognize later.
Don't rely on memory for exact phrasing. If a quote matters enough that you'll want to use it verbatim, that's precisely the situation a transcript is for — checking the transcript against your memory of a good line usually turns up a slightly different, better sentence than the one you remembered.
Note timestamps for the strongest moments as you go. A quick typed mark at the point where the subject says something unusually candid or precise makes that section easy to jump back to directly, rather than searching for it once the interview is transcribed in full.
Two professions, two different pressures
A journalist working an interview is usually under some version of time pressure — a deadline that doesn't move, a source who's only available for twenty minutes, an editor who wants the quote today. What matters most in that setting is speed from recording to usable transcript: because transcription happens on the Mac while the interview is still running, there's no separate upload-and-wait step between hanging up and having text in front of you to pull a quote from.
An academic conducting fieldwork interviews is usually optimizing for something closer to the opposite: fewer interviews, each one long, often covering sensitive personal material that a participant agreed to share under specific research ethics terms. There, the property that matters most is that nothing about what a participant said is uploaded to a third party as a side effect of transcribing it — the recording and the transcript stay on the researcher's own machine, which is often a condition of the ethics approval in the first place, not just a nice-to-have.
What happens to the recording afterward
Nothing about an interview recorded this way leaves your Mac unless you choose to send it somewhere yourself. There's no upload step, no cloud copy created as a side effect of transcribing, and no server involved in generating the summary either — the language model that turns the transcript into minutes is bundled with the app and runs locally, the same machine that recorded the conversation in the first place. For a journalist or a researcher whose interview subjects were promised confidentiality, that's not a minor detail; it's the difference between a promise you can actually keep and one that depends on a third party's data practices you don't control.
Choosing between the two setups when a call could go either way
Some interviews are ambiguous until the last minute — a subject who was supposed to come into the office but ends up dialing in instead, or a call that starts on video and finishes with everyone in the same room a week later for a follow-up. The good news is that switching between the two recording setups described above doesn't require reconfiguring anything meaningful: an in-person conversation uses the microphone, a call uses the microphone plus system audio, and the app handles either without you needing to pick a mode ahead of time. What matters is checking, before you press record, which situation you're actually in — because a call recorded as if it were in-person will capture your own voice clearly and miss the other side of the conversation entirely.
How this compares to other ways people record interviews
The default option most people reach for first is a phone propped up on the table, recording to a voice memo, or a laptop's built-in recorder app with no transcription attached at all. Both work as a basic safety net, but neither gets you further than an audio file — you still have to listen back to the whole thing to find anything, and neither one separates the two voices for you. A dedicated interview transcription setup closes that gap by producing text you can search and scan while the audio file sits underneath as the reference copy, which is the difference between an interview you can act on the same day and one you have to schedule an hour to relisten to.
The other common alternative is a cloud transcription service you upload a recorded file to after the fact. That can work well for a single one-off interview, but it adds a wait — upload time, processing time, and in some cases a queue behind other users' files — between finishing the conversation and having a transcript in hand. Transcribing as the interview happens, on the same machine that's recording it, removes that gap entirely: by the time you say goodbye, the transcript is already there.
A short pre-interview checklist
Running through this before a subject sits down, or before you dial in, takes under a minute and avoids the two most common interview-recording mistakes:
- Confirm which setup applies — in-person microphone, or call with system audio — based on how the interview is actually happening today, not how it was originally scheduled.
- Grant microphone and screen recording permissions in System Settings ahead of time rather than discovering a permission prompt mid-interview.
- Do a five-second sound check at the start: say a sentence, glance at the transcript appearing, and confirm both your voice and the other person's are showing up as expected before the real conversation begins.
- Ask, and note the answer, about anything the subject wants kept off the record — a recording captures everything indiscriminately, so that boundary has to be tracked separately, by you, rather than assumed.
Why the two-track approach matters more for interviews than for meetings
In a group meeting, if the recording quality dips for a few seconds, there are usually several other voices and enough surrounding context to reconstruct roughly what was missed. An interview doesn't have that redundancy — it's two voices, and if either one is unclear at the exact moment they said something important, there's no third person's recap to fall back on. That's the practical reason two separate tracks matter more here than almost anywhere else in this app's other uses: your microphone captures your own voice cleanly regardless of what's happening on the call, and the system audio track captures the other person cleanly regardless of your own room noise, so a problem on one side never degrades the other.
This also means the two tracks fail independently rather than together, which is worth knowing as a troubleshooting fact. If your own audio sounds muffled but the other person is perfectly clear, the issue is almost certainly your microphone or its positioning, not the recording setup as a whole — and the fix is local to your side of the call, not something to worry about having lost on the other end.
After the interview: turning the transcript into something usable
Once the conversation ends, the transcript is already sitting there rather than being something you wait on, but there's usually still a step between "raw transcript" and "the piece I'm writing" or "the notes I'm filing." Renaming the subject's speaker label from a placeholder to their actual name takes a few seconds and makes the rest of the read-through far easier to follow. From there, reading the transcript once with the recording's structure fresh in mind — what was the opening question, where did the conversation turn, what was the strongest answer — is usually faster than reading it cold a week later, when the context of why you asked a particular follow-up has faded.
If the interview also produced a generated summary, treat it as a fast index into the transcript rather than a replacement for reading the transcript itself: a summary is built to compress a conversation down to its main points, and a strong quote is often found in a sentence that wasn't important enough to make the summary but is exactly the line you want anyway.
Before you hit record
Two things are worth checking before an interview starts rather than after: that macOS has granted the app microphone access (and screen recording access, which is what enables system audio capture for a call), and that you know which of the two setups above applies to today's conversation. Getting the setup right before the interview starts is the only step you don't get a second chance at — everything after that, from speaker labels to minutes, can be corrected once the transcript exists. Testing this once during the 30-day trial, on a low-stakes call before it matters, is worth the ten minutes it takes.