Files

1.3 KiB

ADR 0003: Use pyannote for Speaker Diarization

Status

Accepted, amended by ADR 0011

Context

The project requires speaker diarization to assign transcript segments to speakers.

The first implementation should provide reliable separation of speakers in meetings, while keeping the diarization component replaceable.

Decision

The project will use pyannote.audio as the initial speaker diarization engine.

Speaker diarization will be treated as a separate pipeline step after transcription.

Consequences

  • Diarization can be improved or replaced independently from transcription.
  • Speaker labels are derived metadata and must not modify the canonical transcript.
  • Human correction of speaker names should be supported later.
  • The diarization module must hide the concrete engine behind an internal interface.
  • Model name, engine version and timestamp should be stored with every diarization result.

Amendment

ADR 0011 makes diarization optional for the MVP and selects pyannote.audio Community-1 as the current preferred backend. CPU execution remains supported, with GPU acceleration used when available. Speaker labels remain anonymous unless a user explicitly confirms a SPEAKER_XX to participant mapping in the Meeting Context. Automatic speaker-name inference is not allowed.