A voice print is a representation of how a particular voice sounds — pitch, tone, cadence, the qualities that make one speaker's audio distinguishable from another's — built from a segment of recorded speech and used to find other segments in the same recording that sound like the same person. It's the mechanism behind speaker diarization: the step that actually decides "this later segment is the same voice as that earlier one."
What it can and can't tell you
A voice print answers a narrow question well: do these two segments of audio sound like the same person. It doesn't answer who that person is — a voice print on its own carries no name, no identity, nothing beyond "consistent with this other segment in this recording." Some apps let a person label a voice print once, by typing a name against it, and then recognise the same voice in future recordings; without that step, a voice print stays a generic label like "Speaker 2."
How it's built, in outline
The process compares short windows of audio against each other, looking for the acoustic patterns that stay consistent for one speaker and differ from everyone else talking in the same recording — not any single feature but a combination of them, compared across the whole file rather than judged segment by segment in isolation. It needs enough speech to work from: a single short interjection early in a call sometimes isn't enough on its own to build a print that's confidently distinct from the others, and gets attributed correctly only once that speaker has said more.
Why it matters where the recording happens
Building a voice print from an already-recorded file is one thing; doing it without sending that audio anywhere else is another. Coii AudioNotes builds and compares voice prints on the same Mac that made the recording, as one step in transcribing segment by segment while the meeting is still running — there's no upload involved in telling two speakers apart, the same way there isn't for the rest of what it does with the audio. What that means for the recording as a whole, not just the speaker labels, is covered in Coii AudioNotes against Otter.
Where it shows up in the finished notes
A correct voice print is invisible when it works: the transcript just reads correctly, each line attributed to the person who said it, and the generated minutes can say who raised a point rather than leaving it anonymous. Transcript vs minutes covers how that transcript turns into a shorter summary once the speakers are sorted out — voice-print matching is the step that happens before either one is readable.
Where the term gets confused with something else
"Voice print" also shows up in security contexts, describing a biometric check used to confirm someone's identity before letting them into an account or a building. That's a related idea built for a different purpose, with a much higher bar for accuracy and often a deliberate enrolment step. The voice print used to sort speakers in a meeting transcript is a lighter-weight tool solving a narrower problem — telling voices apart within one recording — and was never built or verified for confirming anyone's identity. Treating one as a substitute for the other, in either direction, would be a mistake worth avoiding: neither this page nor do I need an API key for transcription is describing an authentication system.