1.3 KiB
ADR 0003: Use pyannote for Speaker Diarization
Status
Accepted, amended by ADR 0011
Context
The project requires speaker diarization to assign transcript segments to speakers.
The first implementation should provide reliable separation of speakers in meetings, while keeping the diarization component replaceable.
Decision
The project will use pyannote.audio as the initial speaker diarization engine.
Speaker diarization will be treated as a separate pipeline step after transcription.
Consequences
- Diarization can be improved or replaced independently from transcription.
- Speaker labels are derived metadata and must not modify the canonical transcript.
- Human correction of speaker names should be supported later.
- The diarization module must hide the concrete engine behind an internal interface.
- Model name, engine version and timestamp should be stored with every diarization result.
Amendment
ADR 0011 makes diarization optional for the MVP and selects pyannote.audio
Community-1 as the current preferred backend. CPU execution remains supported,
with GPU acceleration used when available. Speaker labels remain anonymous
unless a user explicitly confirms a SPEAKER_XX to participant mapping in the
Meeting Context. Automatic speaker-name inference is not allowed.