Skip to content

Glossary

Diarization is labelling who spoke, not what they said.

Speaker diarization means labelling who spoke in a recording, telling voices apart without needing to know anyone's name. It differs from transcription.

· 3 min read

On this page (6)
  1. Why a transcript needs it
  2. How it's usually done
  3. What it does and doesn't know
  4. Where it matters most
  5. Why it's harder than it sounds
  6. Where the category sits more broadly

Speaker diarization is the process of splitting an audio recording into segments by who is talking and grouping the segments that belong to the same voice — answering "who spoke when," separately from the question a transcript answers, which is "what did they say." A recording run through diarization comes out divided into speaker turns even before any of those turns are transcribed into text.

Why a transcript needs it

A transcript with no diarization applied reads as one continuous block of text, with no marker for where one person stopped talking and another started. For a single voice — a lecture, a voice memo — that's rarely a problem. For a conversation with two or more people, it's the difference between a document that's genuinely useful and one that requires re-listening to the audio anyway just to work out who said which line.

How it's usually done

Two approaches produce it. Some meeting tools get a separate audio feed per participant directly from the video call platform, which makes diarization close to automatic — the platform already knows who's who. Others, including Coii AudioNotes, work from a single recorded track and separate speakers by voice print instead — comparing segments of audio against each other and grouping together the ones that sound like the same person, without needing a dedicated feed or a name attached in advance.

What it does and doesn't know

Diarization on its own answers "how many distinct voices are in this recording, and which segments belong to which one" — it does not, by itself, know anyone's actual name. Labelling a voice "Speaker 1" versus attaching a real name to it are two different steps; some tools ask a user to label speakers once and then remember the voice on future recordings, while others leave the labels generic.

Where it matters most

Diarization earns its keep the moment a recording has more than one consistent voice: a board meeting with several directors, an interview with a host and a guest, a one-on-one between a manager and a report. A board discussion is a case where getting this right matters beyond convenience — knowing precisely which director raised an objection is part of what makes board meeting minutes a reliable record rather than a rough paraphrase. It matters far less for a single continuous voice, like a solo lecture with no discussion segment, where there's only ever one speaker to label.

Why it's harder than it sounds

Two people who sound similar, a lot of crosstalk where voices overlap, or a speaker who trails off and mumbles mid-sentence all make diarization harder to get right, regardless of which approach produces it. Voice-print matching in particular depends on each speaker having said enough, clearly enough, for their voice to be told apart from the others in the same recording — a single one-word answer early in a call is sometimes not enough on its own, and gets attributed correctly only once that speaker has said more.

Where the category sits more broadly

Diarization is one building block inside what an AI notetaker actually does with a recording — transcription turns speech into text, diarization sorts that text by speaker, and a summarising step on top of both is what most notetakers, including Coii AudioNotes compared against Otter, are ultimately built to produce.

Questions

Is diarization the same thing as transcription?
No. Transcription turns speech into text; diarization labels which voice said which part. A transcript with no diarization reads as one unbroken block with no speaker changes marked.
How does an app tell speakers apart without a separate microphone for each one?
By voice print — comparing segments of audio against each other and grouping the ones that sound like the same voice, rather than relying on a dedicated audio feed per speaker.