Skip to content

A think-aloud usability session only works if the moderator is actually listening for the moment a participant hesitates, back-clicks, or says something under their breath that matters more than whatever they said when asked directly. Typing a running log competes for the same attention, and a research day that runs five sessions back to back rarely leaves room between them to write each one up properly before the next participant joins.

What a session actually asks of the person running it

The read comes from watching, not from the transcript — where a cursor hesitates before a click, the sentence that trails off before a participant decides not to finish it, the sigh between two attempts at the same task. None of that survives being typed live by the same person who's supposed to be noticing it. Voice Memos will capture the audio and do nothing else with it: no speaker labels, no searchable text, no summary waiting when the session ends.

Where a bot or a cloud tool changes the room

Otter joins a video call as a named participant, which is a reasonable trade for an internal team meeting and a strange thing to explain to a participant who was recruited externally, is unfamiliar with the company, and now sees an unfamiliar name log in alongside the moderator. Granola avoids that specific problem — no bot joins — but the session is still recorded, transcribed and summarised on a cloud service the participant never agreed to and the moderator doesn't control end to end.

Reading the Mac's own audio instead

Coii AudioNotes records the moderator's microphone and the call's system audio as two separate tracks, without joining the call as anything. The transcript is produced segment by segment while the session runs, and the summary is written on the same Mac once it ends.

42:00
Microphoneyou, in the room
System audioeveryone on the call
A 42-minute usability session, the moderator's microphone and the call's system audio kept as two separate tracks.

From a session's audio to something worth reading between interviews

A language model bundled inside the app turns the transcript into a short summary right after the session ends — what came up, what to follow up on, what's still unclear — while the specific way a participant phrased something is still accurate rather than paraphrased from memory an hour later, after two more sessions have happened in between.

Checkout flow — session 4 of 6

Decisions

  • Participant could not locate the saved-cart link without a hint

Actions

  • Re-test the saved-cart link placement with the next two participants

Open questions

  • Is the hesitation specific to this participant or the flow itself?
A summary written on the Mac right after the session, from a transcript that was never uploaded to produce it.

What Grain does that this deliberately doesn't

Worth naming plainly, because it's a real gap. Grain lets a researcher tag a moment or clip a segment live during the call, then build those clips into a shareable reel — the exact workflow a research team uses to hand a stakeholder three minutes of a participant saying the quiet part out loud instead of a forty-minute recording. Coii AudioNotes produces a transcript and a summary, not a library of tagged clips across a study, and a team whose synthesis process runs on highlight reels gets real value from Grain's clipping that this tool doesn't attempt to replace. The full comparison covers the rest of what each one does.

Telling two voices apart without a feed from the platform

A bot inside a call separates speakers using the individual audio feed the platform hands it per participant. Reading the Mac's own microphone and system audio instead means that feed isn't available, so speaker separation works by voice print — matching each voice against the others already in the same recording. A moderator and a single participant are exactly two distinct voices, which is the case this works best against; a session with a second observer on the call adds a third voice the same matching still handles.

What a participant sees on their side of the call

A participant recruited for a study is usually meeting the moderator, and the company, for the first time in that same session. Nothing about how the call looks changes on their end — no additional name on the roster, no recording indicator beyond whatever the call platform itself already shows. That matters more in research than in an internal meeting, because a participant who notices an unfamiliar bot has one more reason to second- guess whether the "think aloud, be honest" instruction at the start of the session is really how the room is set up. The permission to record is still the moderator's to ask for at the start of the call, the same as it would be with any other recording method; this only changes where the audio goes once the participant has agreed to it.

Comparing what session four said against session one

A single session's summary is useful the same afternoon; a full study becomes useful differently once six sessions are done and the actual question is whether the hesitation seen in session four is a pattern or a one-off. Reading through six sets of hand-written notes to check is slow and depends on how consistently each session got written up in the moment. A transcript for each session, sitting on the researcher's own Mac and searchable by keyword, turns that into something checkable in seconds — did three participants independently use the same word for the same confusing step — rather than a synthesis built on how well each session happened to get documented at the time.

What it deliberately doesn't do

There's no tagging, no clip library, and no shared repository a second researcher or a stakeholder can browse without being sent a file directly. There's no integration with a research repository, an insights tool or a recruiting platform — nothing here files a session automatically against a study or a participant record. Each recording lives on the Mac that made it, across up to three machines on one licence, and it runs on macOS only, so a research team split across Windows laptops won't get one consistent tool from this alone.

A day of five sessions

A day of back-to-back usability sessions used to mean choosing between typing shorthand during each one, which costs the attention the session actually needs, or reconstructing each session from memory in the gap before the next participant joins — by which point session two and session four have started to blur. With a transcript and a short summary produced per session, the between-sessions window becomes time to glance at what actually happened rather than racing to write it down before it's gone.

What changes across a full study, not just one session

Nothing about how a session runs changes — same script, same platform, same think-aloud instruction at the start. What changes is what's available once the study is over and a stakeholder asks which participants actually struggled with a specific step, rather than which struggle the researcher happens to remember most vividly from six sessions run over two weeks. A transcript per session, all sitting in the same place on the same Mac, turns that question into something to check rather than something to recall from memory under time pressure the week the readout is due.

What a stakeholder actually reads

The people who requested a study are rarely the people who sat in on any of its sessions, and what they end up reading is whatever synthesis makes it into a readout deck weeks later. A summary written the same afternoon a session happened, from a transcript rather than a memory of it, is a sturdier foundation for that readout than notes taken live and typed up whenever there's time between the next study's sessions — the gap between when something was said and when it was written down is where detail usually goes missing first.

Who this genuinely isn't built for

A research team running a shared repository across a dozen studies and several researchers, where a stakeholder needs to pull a tagged clip without being sent a file, has a real workflow this tool doesn't attempt to replace — that's exactly the job Grain's tagging and clip reels are built for. This is written for the solo researcher or small team whose sessions are mostly one person's to moderate and write up, where a transcript and a summary on their own Mac is the whole of what synthesis needs.

The recruiting slot that runs long, or the one that ends early

A study rarely runs to the scheduled minute. A participant who's articulate and generous with detail can turn a planned thirty minutes into fifty; one who struggles with a task and disengages can end a session in fifteen. Both are fine outcomes for a moderator who isn't also trying to keep a running written log in step with a session moving faster or slower than planned — recording and transcribing happen at whatever pace the conversation actually takes, and the summary written afterward reflects the session that happened rather than the one the schedule assumed would happen. That matters more on a recruiting day with six back-to-back slots than it does for a single interview, because a session that runs fifteen minutes long eats directly into the gap meant for writing up the one before it — a gap this removes the need for in the first place.

Why a transcript sometimes matters more than the summary

A summary is what most of the team reads, but occasionally a specific claim in a readout gets questioned — did a participant actually say the feature was confusing, or is that the researcher's own framing of a longer, more hedged answer. A transcript sitting alongside the summary settles that directly, at the point in the session it happened, rather than asking the researcher to defend a paraphrase from memory in a meeting weeks after the session itself.

Setting it up is smaller than the decision to switch

The first launch asks for two macOS permissions — microphone access and permission to record system audio — the same prompt any Mac app requesting audio capture triggers, and that's the whole of it. No account, no workspace, no recruiting tool to connect before the first session can be recorded. The larger adjustment is usually habit: remembering to press record instead of trusting a bot that used to join the call automatically.

What it costs, what it runs on

Coii AudioNotes is $19, paid once, for three of a researcher's own Macs, with a 30-day trial needing no card and no account. It runs on macOS 13 Ventura or later, Apple Silicon or Intel. Against Otter's $8.33–$19.99-a-month tiers or Granola's $14-a-month Business tier, a solo researcher running a handful of studies a year lands at $19 total rather than a recurring line item that has to be re-justified every renewal. Full comparisons here and here. A wider round-up of interview transcription tools for Mac covers the rest of the field a research-focused search usually turns up, and the adjacent case for journalists recording interviews and recruiters running screening calls covers two nearby jobs built around the same one-on-one conversation. The broader case for a Mac-native AI notetaker covers ground this page doesn't repeat.

Questions

Does a bot join the call to record a usability session?
No. It reads the microphone and the system audio the Mac already carries during the call, so nothing new appears on a participant's screen or on the call's participant list.
Can it tell the moderator's voice apart from the participant's?
Yes, by voice print — matching each voice against the others already in the same recording, without a named audio feed from the call platform. A moderator and one participant are exactly two distinct voices.
Does it build clip reels or tag moments the way Grain does?
No. There is no clipping, tagging or shareable highlight reel. The output is a full transcript and a summary, both produced on the Mac, not a library of tagged moments across a study.
Does it connect to a research repository or insights tool?
No. There is no integration with a repository, a tagging system or a shared workspace. The recording, transcript and summary stay on the Mac that made them, across up to three machines on one licence.