Skip to content

How-to

Who said what, without you typing a single name

An unlabeled transcript is barely more useful than the recording. Here's how speaker labels get attached automatically, and what to do about a wrong one.

· 11 min read

On this page (9)
  1. How the app tells voices apart
  2. Why voice prints work without any setup
  3. Renaming a speaker once
  4. When a label needs fixing
  5. Why this is different from a meeting bot's transcript
  6. Where correct speaker labels matter most
  7. Getting the most reliable labeling
  8. Larger calls with more than two voices
  9. What this feeds into

A transcript that reads as one long paragraph of unattributed text is not much of an improvement over the raw recording — you still have to listen back to work out who made the commitment on slide four. The part that actually saves time is the label in front of every line: which voice said it, held consistently from the first minute of the recording to the last. That's what this page walks through — how the labeling happens, where it can go wrong, what to do when it does, and why it's worth caring about at all.

How the app tells voices apart

Coii AudioNotes doesn't ask you to record a sample of your own voice before your first meeting, and it doesn't need a directory of who's who. Instead, it listens to the recording itself and builds a voice print for each distinct speaker it hears — a fingerprint of that particular voice, not a name. This process, called speaker diarization, is what separates "one continuous audio file" from "a transcript where every line has an owner." The first time a new voice appears in a recording, it's given a placeholder label. Every later line from that same voice, for the rest of that recording, gets the same label, because the app is matching the sound of the voice, not guessing from context or turn-taking patterns.

This matters for two-track recordings especially. Because your microphone and the system audio are captured as two separate tracks — the room, and the call — the app already knows which sounds came from your own microphone before it even starts telling individual voices in the call apart. That separation is what makes it possible to distinguish "you, at your desk" from "the three other people in the video call" cleanly, instead of trying to untangle everyone from a single mixed-down recording.

The transcription itself runs on your own Mac, segment by segment, while the meeting is still happening — so by the time the call ends, the speaker-labeled transcript is already sitting there rather than something you wait on afterward.

2 speakers
  • 00:01:04Speaker 1Let's start with the renewal numbers before we get to the roadmap ask.
  • 00:01:11Speaker 2Sure — renewal's at ninety-one percent, up from last quarter.
  • 00:01:19Speaker 1That's the number I wanted. Can you send the account list after this?
  • 00:01:24Speaker 2I'll have it over before end of day.
Two voices kept apart automatically from the first minute, before either name is typed in.

Why voice prints work without any setup

It's worth being specific about what a voice print actually is, because it explains both why the labeling works with zero configuration and why it's occasionally wrong. A voice print here is a numerical description of how a particular voice sounds — pitch, cadence, the resonance of someone's speech — built by listening to a stretch of that voice within the current recording. It's a comparison between voices heard in this file, not a match against a library of known people, which is exactly why there's no enrollment step: the app doesn't need to have heard you before, it only needs to hear each distinct voice for a few seconds within today's recording to start telling them apart from each other.

This is also why the app can't recognize a specific person by name on its own — it has no idea who "Priya" is until you tell it, because it was never given a directory of names to match voices against in the first place. That's a deliberate trade: a system that recognized specific people by voice across recordings would need to store a profile of each person somewhere, and there's nowhere for it to store that profile except on a server it doesn't have.

Renaming a speaker once

The placeholder labels — "Speaker 1," "Speaker 2," and so on — are correct in the sense that they're consistently applied, but they're not useful to read back in six months. Click the label once, in the transcript view, and type the real name. That name is retroactively applied to every line from that same voice in the current recording. You don't have to click through and rename each individual line; the labeling and the naming are two different steps, and the second one only has to happen once per person per recording.

If the same colleague shows up in next week's call, the app doesn't remember them from last time — each recording is analyzed on its own, since there's no account, no cloud profile, and nothing about you or your contacts stored anywhere but on your own Mac. That's a small extra step compared to a tool that keeps a directory of speakers across meetings, and it's a fair trade for the alternative not existing: nothing about who you meet with is uploaded anywhere to make that directory possible in the first place.

It also means renaming is a per-recording action rather than a one-time setup. Over a week of five or six calls with overlapping attendees, you'll retype a couple of names more than once. Most people find that a small enough cost once they weigh it against not having a growing contact list of colleagues and interview subjects sitting on a server somewhere they don't control.

When a label needs fixing

Voice separation is good, not infallible, and there are a few situations worth knowing about before they surprise you mid-transcript:

Overlapping speech. When two people talk over each other — the moment somebody jumps in to finish another person's sentence — the transcript has to attribute that stretch of audio to one voice or the other, and it's the one place a fast, quiet interjection can get folded into the wrong speaker's line. It's usually a single short phrase, not a whole paragraph, and it's easy to spot because the sentence reads oddly for that person.

A speaker who changes rooms. If someone steps away from their laptop mid-call and continues on a phone, the acoustic character of their voice changes — different microphone, different room reflections — and the app may treat the second half as a new voice rather than a continuation of the first. Renaming the new label to the same name folds it back together for reading purposes.

A very quiet participant. Someone who mostly listens and contributes one or two short lines gives the voice-print comparison less to work with than someone who talks for ten minutes straight. It still usually resolves correctly, but a one-word interjection from someone who barely speaks is the input most likely to need a manual check.

A speaker who joins late. If a fourth person dials into a call that started as three, the app treats the new voice as a new speaker the moment it first hears them — there's no need to restart the recording or flag the change yourself. The label simply appears at the point they start talking.

None of these require starting over. You're correcting a label on a transcript that already exists in full, on your Mac, rather than waiting on a re-processing step or a support ticket, and none of them require you to re-listen to the whole recording to find the mistake — the mislabeled line is usually obvious from the text alone, because it reads like the wrong person said it.

Why this is different from a meeting bot's transcript

Tools that join your call as a visible participant often lean on the meeting platform's own participant list to assign speaker names — which sounds convenient until the bot is wrong about who's on mute, who joined late under a generic device name, or who is presenting a shared laptop with two people talking near one microphone. Coii AudioNotes isn't in the call as a participant at all: it records what your Mac's microphone and system audio can already hear, the same way you could if you asked, and it works out who's speaking from the audio itself rather than from platform metadata that may or may not match the actual voice.

That also means it keeps working in situations a bot-based tool doesn't reach — an in-person conversation with no video call at all, a phone interview taken on speaker, a lecture with no login for a bot to use. Anywhere your Mac can hear two voices, it can tell them apart, which is the whole reason a tool built this way can be described as an ai notetaker for mac rather than a meeting-platform add-on specifically.

Where correct speaker labels matter most

Some jobs live or die on this. A user researcher running back-to-back interviews needs to know, without doubt, which participant said which thing when they're writing up findings for six sessions at once — a mislabeled quote attributed to the wrong participant is a real error in a research report, not just an inconvenience. A recruiter screening candidates through a day of calls has the same problem in miniature: notes that mix up what one candidate said with what another said are worse than no notes, because they read as confident and are wrong.

In both cases, the fix isn't a different feature — it's the same voice-print separation described above, applied to a day with more interviews in it than a typical meeting day, which is exactly where a manual alternative like typing notes live falls apart first.

Getting the most reliable labeling

A few habits make the automatic labeling more accurate without adding any extra work during the meeting itself:

  • Let each person finish a full sentence early on. The first clean, uninterrupted stretch from a voice is what the app has the most to compare against; a recording that opens with everyone talking at once gives it less to work with for the rest of the call.
  • Keep your Mac's built-in microphone reasonably close if you're the one being recorded in the room — the system audio track for the call side doesn't have this concern, since it's captured directly rather than through a microphone.
  • Rename speakers as soon as you notice the placeholder labels, rather than leaving it until the transcript is long. It takes seconds either way, but doing it early means the rest of your read-through already has real names in it.
  • Avoid running two calls' audio through the Mac at once. If a second call is audible in the background of the one being recorded, the app has two conversations' worth of voices to sort out rather than one, which is the surest way to end up with a label that needs correcting.

Larger calls with more than two voices

Everything above holds for a two-person conversation, but most work calls have more people in them than that, and it's worth knowing how the same mechanism scales. Each additional distinct voice gets its own placeholder label and its own voice print, built the same way as the first two — the app isn't limited to a fixed number of speakers it's expecting, it's building a fingerprint for however many distinct voices actually show up in the recording.

What does change with more people is how much clean audio each voice gets before the app has a confident fingerprint for it. In a two-person interview, each voice speaks for roughly half the recording. In a six-person planning meeting, someone who only says "sounds good, no objection" once in forty minutes gives the system very little to work with, and that's the voice most likely to need a manual rename check afterward. It's rarely wrong in a way that matters — a person who barely spoke usually doesn't have much attributed to them either way — but it's worth knowing which kind of meeting is more likely to need the two-minute cleanup pass and which usually doesn't.

Group calls are also where renaming pays off fastest, because a six-person transcript with "Speaker 3" and "Speaker 5" left unlabeled is genuinely hard to read back, in a way a two-person transcript with unlabeled speakers usually isn't — you can often infer who's who in a one-on-one from context alone. In a group, you can't, and that's exactly the case the rename step exists for.

What this feeds into

A correctly labeled transcript is also what makes the generated minutes useful — a decision or an action item is worth far more with a name attached to who owns it than without one. If you haven't looked at how the transcript turns into that summary yet, that's the natural next step once speaker labeling is working the way you expect.

The whole point of getting this right is that you stop having to remember who said what, and start being able to just read it back. A 30-day trial is long enough to see how the labeling holds up across a real week of your own calls, not just one meeting.

Questions

Does the app need to hear my voice ahead of time?
No enrollment step. It tells voices apart by comparing them against each other inside the recording, not against a profile you set up beforehand.
Can it tell two similar-sounding voices apart?
It compares the acoustic pattern of each voice, not just pitch, so two similar-sounding speakers are usually kept apart. A very short, overlapping remark is the case most likely to be missed.
What if the transcript only shows 'Speaker 1' and 'Speaker 2'?
That's the starting label. Rename it once in the transcript and the same name is used for that voice for the rest of that recording.
Does this work on a call with more than two people?
Yes. Each distinct voice gets its own label; the practical limit is how many people can talk into one Mac's microphone and still be told apart clearly.