The study calls for six thirty-minute sessions on the same checkout flow, recruited across a single Tuesday with ten-minute gaps between them. The moderator's job in each one is to watch where a participant's cursor hesitates and notice what they mutter under their breath before deciding not to finish the sentence — not to keep a running written log competing for the same attention.
9:00am — session one starts before the participant even joins
Recording begins a minute before the call link goes out, so the moment a participant joins and says hello is already inside the recording rather than lost to a beat of fumbling to hit record. The microphone in the room and the call's system audio are kept as two separate tracks from that first second.
9:31am — the ten-minute gap, spent on nothing except the next session
Ten minutes between session one and session two used to mean a rushed paragraph of notes typed from memory before the details started blending into whatever session two turned out to be. With a transcript already sitting on the Mac from session one, the gap is free for the thing it's actually meant for: reviewing the guide, or just breathing before the next participant joins.
11:20am — session three, where the hesitation actually was
By session three, a pattern the moderator half-noticed in session one starts looking real: participants pause at the same step, right before confirming the saved-cart link. Nothing about running the session changes because of that hunch, but flagging it can wait until the transcript is checked rather than needing to be caught and written down live.
- 00:14:02ModeratorTake your time — what are you looking for right now?
- 00:14:11ParticipantI thought my cart was saved somewhere... I don't see it.
- 00:14:19ParticipantOh — there. I just didn't expect it to be up there.
12:05pm — checking session one against session three without replaying either
Confirming that the hesitation in session three matches what happened in session one doesn't mean re-watching forty minutes of recording. A keyword search across both transcripts for "saved" and "cart" turns up the exact sentence from each session, side by side, in under a minute — the kind of check that's fast enough to run between sessions rather than saved for the end of the week.
1:15pm — session four, and the participant who disengaged early
Not every session runs to plan. Session four's participant struggles with the first task, gives shorter answers, and the whole thing wraps in sixteen minutes instead of the scheduled thirty. The recording and transcript reflect exactly that session — short, and worth reading precisely because it's short — rather than needing to be padded out or explained away in a summary typed up later.
Decisions
- Participant could not complete the saved-cart step without a hint from the moderator
Actions
- Flag session 4 as a possible drop-off case in the readout, not an average one
Open questions
- Is sixteen minutes of disengagement a signal or a one-off?
2:40pm — session five, with a second observer sitting in
A product manager joins session five as a silent observer, then speaks up once, near the end, to ask a follow-up question directly. Three distinct voices now sit in one recording rather than the usual two, and separating them still works the same way — by matching each voice against the others already in that session, not by needing a named feed for each participant from the call platform.
3:50pm — session six, the last of the day
By session six the moderator is tired in the specific way six sessions in one day makes a person tired, and that's exactly when a live written log gets thinnest — the last session of the day is disproportionately likely to be the worst-documented one if documentation depends on energy that's already spent. The last recording gets the same treatment as the first: recorded, transcribed, and summarised on the Mac without anyone needing to have anything left for the write-up.
4:30pm — six transcripts, six summaries, one folder
At the end of the day, six sessions sit in the same place on the Mac, each with its own transcript and its own short summary. None of the six needed to be typed up from memory, and none of the audio, text or summary was uploaded to produce any of it. The question a stakeholder will ask next week — "how many participants actually struggled with the saved-cart step" — is answerable by searching six files rather than reconstructing six sessions from whatever notes survived a long Tuesday.
Why the guide gets checked between sessions, not rewritten after
A study's discussion guide usually needs a small adjustment after the first two or three sessions — a question that lands flat, a task that's phrased ambiguously and needs a tweak before session three repeats the same confusion. That adjustment gets made in the ten-minute gap, informed by listening back to the exact moment a question landed wrong in session one, rather than waiting until all six sessions are done to notice the pattern in a synthesis meeting a week later, by which point three more participants have already hit the same ambiguous task.
The session that almost didn't get moderated at all
Session two very nearly ran without the moderator, who was still finishing the write-up prompt from session one when the participant joined two minutes early. That near-miss is the exact failure this removes: because session one's transcript and summary were already sitting on the Mac rather than half-typed from memory, there was nothing left to finish before session two could start on time. A moderator who's still writing when the next participant joins is a moderator who starts the next session distracted.
What six sessions look like from a stakeholder's side, a week later
The team that commissioned the study wasn't in the room for any of the six sessions, and what they eventually see is a readout deck built from all six. The version of that deck built from six full transcripts, checked against each other by keyword rather than by memory, differs from one built out of whatever survived a long Tuesday of live note-taking — not because the moderator would deliberately misremember anything, but because six sessions in one day is more than anyone reliably holds in exact detail without something to check it against.
What changes if the study runs to twelve sessions instead of six
The gap this closes gets larger, not smaller, as a study grows. Six sessions in one day is already close to what one moderator can hold in memory well enough to write up faithfully; twelve sessions across two days is well past it, and the value of having a searchable transcript for every session — not just the ones that happened to leave a strong impression — grows with every session added rather than staying flat.
What Otter or Grain would have added to today
Otter joining each of today's six calls as a named participant would have produced a broadly similar transcript, and for an internal team meeting that's a reasonable trade — for a participant recruited externally, meeting the company for the first time, an unfamiliar name on the call is one more thing to explain before the session has even started. Grain solves a different problem today's workflow doesn't attempt to replace: it lets a researcher tag and clip a moment live, then build those clips into a shareable reel for a stakeholder who wants three minutes of a participant saying the quiet part out loud rather than six full transcripts. Today's six sessions produced full transcripts and summaries, not a clip library — the full comparison covers what each tool does beyond a single day.
The Friday synthesis meeting, run against six actual transcripts
The team meets Friday to synthesise the week's two study days into a single readout. Synthesis meetings run better against transcripts than against notes, because a disagreement about whether three participants or four actually struggled with the same step gets resolved by searching six files for the word "saved" rather than by two researchers each defending their own recollection of a session the other one didn't sit in on. The meeting still takes judgment — deciding what the pattern means is the researcher's job, not the transcript's — but it stops needing to also settle what was literally said before it can get to that judgment.
The participant who asked for their recording to be deleted
Session two's participant, partway through, asks whether the recording can be deleted once the study wraps — a reasonable question, and one a moderator should be able to answer plainly. The honest answer here is straightforward because the recording never left the Mac in the first place: deleting the file removes the entire thing, transcript and summary included, without needing to also ask a cloud vendor to purge a copy that exists somewhere the moderator doesn't control directly.
The week this stops feeling like extra setup
The first study run this way still feels like a new habit — remembering to start the recording before the participant joins, remembering to check the previous session's transcript during the ten-minute gap instead of skipping straight to the guide. By the second study, checking a transcript before a session becomes as automatic as reviewing the discussion guide always was, and the moment it stops feeling like extra setup is usually the same week a stakeholder asks "how many participants actually said that" and the answer is a keyword search rather than a guess made under time pressure.
What today's workflow deliberately doesn't do
Nothing here tags a moment, builds a clip reel, or files a session automatically against a study or a participant record in a research repository. There's no shared workspace a second researcher can browse without being sent a file directly — each of today's six recordings lives on the Mac that made it, across up to three machines on one licence, and the whole day runs on macOS only. A research team synthesising across a dozen studies with a shared, tagged repository has a real reason to keep that infrastructure; that's not what six sessions in one Mac folder is trying to be.
What a full day of this actually asks of the Mac
Six sessions, each recorded and transcribed while it happens, is the whole of what the day asks of the machine — no batch upload at the end, no queue of files waiting to be processed overnight. The setup that made today possible was decided once, weeks before the study: two macOS permissions granted the first time the app opened, and nothing to reconfigure between one study and the next.
The second study, run the same way
The next study, three weeks later, runs the same recording habit against a different flow entirely — a checkout redesign becomes an onboarding test, the discussion guide changes completely, and the only thing that carries over is the folder structure and the habit of trusting a transcript over a memory of six sessions run in one day. That consistency is what makes comparing studies months apart possible at all: a researcher checking whether this onboarding hesitation resembles last quarter's checkout hesitation is searching two sets of transcripts, not two sets of notes that may or may not have been written the same way.
What it costs, what it runs on
Coii AudioNotes is $19, paid once, covering three of a researcher's own Macs, with a 30-day trial needing no card and no account. It runs on macOS 13 Ventura or later, Apple Silicon or Intel — the setup used for all six of today's sessions. The full case for a user researcher's notetaker covers the argument at more length than a single day can, and the comparison with Otter covers the specific trade-off a bot introduces on a recruited call. For the mechanics behind today's sessions, telling speakers apart in a transcript, what speaker diarization actually is, and turning a transcript into action items each walk through one piece in more depth. A round-up of interview transcription tools for Mac covers the wider field a research-focused search usually turns up, and a consultant's day of client calls shows the same recording habit against a different, less structured calendar.