34 lines
1.3 KiB
Markdown
34 lines
1.3 KiB
Markdown
# ADR 0003: Use pyannote for Speaker Diarization
|
|
|
|
## Status
|
|
|
|
Accepted, amended by [ADR 0011](0011-use-meeting-lab-mvp-backend.md)
|
|
|
|
## Context
|
|
|
|
The project requires speaker diarization to assign transcript segments to speakers.
|
|
|
|
The first implementation should provide reliable separation of speakers in meetings, while keeping the diarization component replaceable.
|
|
|
|
## Decision
|
|
|
|
The project will use `pyannote.audio` as the initial speaker diarization engine.
|
|
|
|
Speaker diarization will be treated as a separate pipeline step after transcription.
|
|
|
|
## Consequences
|
|
|
|
- Diarization can be improved or replaced independently from transcription.
|
|
- Speaker labels are derived metadata and must not modify the canonical transcript.
|
|
- Human correction of speaker names should be supported later.
|
|
- The diarization module must hide the concrete engine behind an internal interface.
|
|
- Model name, engine version and timestamp should be stored with every diarization result.
|
|
|
|
## Amendment
|
|
|
|
ADR 0011 makes diarization optional for the MVP and selects `pyannote.audio`
|
|
Community-1 as the current preferred backend. CPU execution remains supported,
|
|
with GPU acceleration used when available. Speaker labels remain anonymous
|
|
unless a user explicitly confirms a `SPEAKER_XX` to participant mapping in the
|
|
Meeting Context. Automatic speaker-name inference is not allowed.
|