Techniques & Methods
Speaker Diarization in plain English.
Also known as: diarisation,who spoke when,speaker labelling
The one-sentence version
Working out which speaker said which parts of a recording, so a transcript can be labelled by person.
Speaker diarization answers "who spoke when" in an audio recording. A transcription model turns speech into text; diarization segments that audio by voice and assigns each segment a speaker label, so a meeting transcript reads as a conversation rather than a wall of text. It works by extracting a voice embedding for short windows of audio and clustering windows with similar voices. Deepgram, AssemblyAI, Speechmatics, and the open-source pyannote library all offer it, and it is what meeting tools such as Otter, Fireflies, and Fathom rely on to attribute action items. Typical difficulties are overlapping speech, speakers with similar voices, short interjections, and poor microphones on a shared call, where everyone comes through one channel. Accuracy is usually reported as diarization error rate, and the best systems get most meetings right but still confuse speakers occasionally, which is why meeting apps let you rename and merge speakers afterwards. Naming a speaker (rather than "Speaker 2") requires either a known voice profile or the platform's participant list.