Skip to content

A recruiter running eight screening calls in a day isn't short on conversation — the problem is that conversation number four and conversation number seven start to sound like the same call by the time either one needs to be written up. Typing while listening costs attention the candidate can hear in the silence; not typing means trusting memory to hold eight distinct sets of answers until end of day.

What a bot-based notetaker adds to a candidate call

Otter and Fireflies both solve the write-up problem by joining the call as a named participant, which works fine for an internal meeting where nobody minds a bot on the roster. A candidate call is a slightly different room: someone who is already nervous, meeting the company for the first time, and a named bot logging in a few seconds after they do is one more unfamiliar thing to register before the actual questions start. Granola avoids the bot but is still a cloud subscription, transcribed and summarised on infrastructure the recruiter doesn't control.

Reading the call's own audio instead of joining it

Coii AudioNotes records the microphone and the system audio the Mac is already carrying during the call, as two separate tracks, rather than logging in as anything. Nothing appears on the candidate's screen, and the transcript and summary are both produced on that same Mac afterward.

2 speakers
  • 00:07:18YouTell me about a time a project's scope changed halfway through.
  • 00:07:26CandidateAt my last role, a client doubled the deliverables two weeks before launch.
  • 00:07:35YouHow did the team handle that?
A screening call transcribed on the recruiter's own Mac, with nothing on the candidate's screen to notice.

From a screening call to a note the next reviewer can trust

A language model bundled inside the app turns the transcript into a short recap right after the call ends — what stood out, what to follow up on, what's still unclear — while the specifics are still the ones actually said rather than the ones half-remembered eight calls later. That recap is what a hiring manager reading it two days after the call actually gets, instead of a one-line note typed in the ninety seconds between interviews.

Screening call — Senior Engineer req

Decisions

  • Advance to the technical round

Actions

  • Schedule the technical interview for next week

Open questions

  • Confirm reference availability before the offer stage
A recap written on the Mac immediately after the call, from a transcript that was never uploaded to produce it.

Telling two voices apart without a feed from the platform

A bot inside a call separates speakers using the individual audio feed the meeting platform hands it per participant. Reading the Mac's own microphone and system audio instead means that feed isn't available, so speaker separation works by voice print — matching each voice against the others already in the same recording. A recruiter and a single candidate on a call are exactly two distinct voices, which is the case this works best against.

What it deliberately doesn't do

There is no ATS integration — nothing here files a transcript against a requisition or a candidate record automatically, and no scorecard gets filled in on its own. There is no shared workspace a hiring manager can browse without being sent the recap directly; each recording lives on the Mac that made it, across up to three machines on one licence. And it runs on macOS only, so a recruiting team split across Windows and Mac laptops won't get a consistent tool from this one alone.

What actually changes between calls

Nothing about the interview itself changes — same platform, same questions, same conversation. What changes is what happens in the ninety seconds between one call ending and the next one starting: instead of scrawling a half-sentence note before the next candidate joins, there's already a transcript and a short recap of the call that just ended, ready to be read later rather than reconstructed from memory at the end of the day.

Who this genuinely isn't built for

A larger recruiting team running structured interviews through a shared ATS, where a hiring manager needs to pull up a scorecard and a transcript from the same screen without anything being sent manually, has a real workflow this tool doesn't attempt to replace. This is written for the version of recruiting where the calls are mostly one person's to run and write up, not a team already living inside a shared pipeline tool built around scorecards and stage tracking.

Setting it up is smaller than the decision to switch

The first launch asks for two macOS permissions — microphone access and permission to record system audio — the same prompt any Mac app requesting audio capture triggers, and that's the whole of it. No account to register, no workspace to configure, nothing to connect before the first screening call can be recorded. The bigger part of moving off a bot-based tool is usually remembering to press record instead of trusting a bot that used to join automatically — a habit that takes a few days of back-to-back calls to fully replace.

Recording is a decision, not a default

A bot that auto-joins every interview on a calendar captures everything without anyone remembering to start it, which is a real convenience for a recruiter with a packed day and no consistent habit of hitting record. Moving to a tool that has to be started manually shifts that responsibility onto the recruiter — a genuine cost worth weighing honestly against the awkward-first-impression problem, not something to wave away because the rest of the case favours switching.

The interview where this actually matters most

A first-round screening call is, for a lot of candidates, the first real interaction with the company at all — and a named bot logging in a few seconds after they do is a small, avoidable thing to have to explain before the actual conversation starts. Running that same call with nothing but the recruiter's own name on the roster keeps the candidate's attention on the questions instead of the software, which matters most on exactly the call meant to make a good first impression.

Why a transcript settles more than a memory does

A hiring manager reading a two-line recap two days after a screening call sometimes wants more than the summary — the exact way a candidate answered a specific question, not just the recruiter's paraphrase of it. A transcript sitting alongside the recap answers that without a second call, and it only exists if the interview was actually recorded rather than summarised from memory in the gap before the next candidate joined.

What a full pipeline makes visible over time

A single interview's recap is useful the same day; an archive of them becomes useful differently once a requisition has run through a dozen candidates, when the question is which of them actually said something specific about a skill or a timeline, rather than which one the recruiter simply remembers best. Reconstructing that from memory across a full pipeline isn't realistic past the first handful of calls; a transcript sitting on the recruiter's own Mac, searchable by keyword, turns it into something that can actually be checked.

What actually changes across a full requisition

Nothing about how a screening call runs changes — same questions, same platform. What changes is what's available two weeks into a requisition, when a hiring manager asks how a specific candidate compared to another on a particular question, and the answer depends on whether both calls are still just memory or actual transcripts sitting on the recruiter's Mac, searchable in seconds rather than reconstructed from a one-line note typed between calls.

Who this genuinely isn't built for, said a second way

An independent recruiter or a small agency where the calls are mostly one person's to run and write up is exactly who this is built for. A larger in-house team running structured interviews through a shared ATS, with scorecards a whole hiring panel needs to see, has a real reason to stay with tools built around that infrastructure — not because a single call is handled worse here, but because the value of a shared pipeline tool comes specifically from the sharing this tool deliberately doesn't do.

The first weeks after switching

The first few screening calls after moving off a bot-based tool tend to feel slightly unfamiliar for one specific reason: the habit of glancing at the participant list to confirm the bot joined before the questions start. That habit fades within a week or two of back-to-back calls, and what replaces it is simpler — press record, run the interview, and find a transcript and recap waiting on the Mac once it ends.

A screening day, before and after

A day of eight candidate calls used to mean either typing shorthand notes mid-conversation, which costs attention a nervous candidate can often sense, or trying to write each one up from memory in the gap before the next call starts — by which point candidate four and candidate six have usually started to blur together. With a transcript and recap produced automatically per call, the write-up step becomes reading back a short summary rather than reconstructing eight separate conversations from memory at the end of the day.

What a candidate remembers about the call

Candidates rarely remember the exact questions asked, but they do remember how the call felt — rushed, distracted, or attentive. A recruiter typing notes mid-call inevitably reads as at least a little distracted, however unfair that impression is. Recording instead of transcribing live keeps that attention on the conversation itself, which matters as much for the candidate's impression of the company as it does for the accuracy of the eventual write-up.

Who this is written for, plainly

Not a large in-house team standardising interviews across a hiring panel through a shared ATS — that's a recruiting-ops decision, not a single app's job. This is written for the independent recruiter or small agency whose screening calls currently get either typed live, at the cost of attention a nervous candidate can sense, or reconstructed from memory between one call and the next, and who wants a transcript worth trusting without a bot on the roster to get it.

What it costs, what it runs on

Coii AudioNotes is $19, paid once, for three of a recruiter's own Macs, with a 30-day trial needing no card and no account, and every feature unlocked from the first launch rather than gated behind a tier an independent recruiter would have to justify upgrading into. It runs on macOS 13 Ventura or later, Apple Silicon or Intel. Against Otter's $8.33–$19.99-a-month tiers or Granola's $14-a-month Business tier, the $19 is paid once rather than every month for as long as screening calls keep happening. Full comparisons here and here. A wider round-up of interview transcription tools for Mac covers the rest of the field a recruiter's search usually turns up and is worth reading before settling on a single tool for candidate calls, and private meeting recorders covers adjacent tools built around the same no-upload requirement. The broader case for a Mac-native meeting notes app, what an AI notetaker without the bot changes, and what skipping the bot specifically does for a candidate call all cover ground this page doesn't repeat, and each is worth a look before settling on a single tool for recording interviews.

Questions

Does a bot join the candidate call to record it?
No. It reads the microphone and system audio the Mac already has during the call, so nothing new appears on the participant list a candidate would notice.
Can it tell the recruiter's voice apart from the candidate's?
Yes, by voice print — matching each voice against the others already in the same recording, without needing an individual feed from the call platform.
Does it connect to an applicant tracking system?
No. There is no integration with an ATS, a scheduling tool or a scorecard system. The recording and transcript stay on the Mac that made them.
Does it work if back-to-back interviews leave no time to review notes between calls?
Yes — the transcript is produced live during the call and the summary is generated right after it ends, so each interview's specifics are captured before the next one starts, without the recruiter having to type anything mid-call.