Establish initial project architecture

This commit is contained in:
2026-07-08 15:57:35 +02:00
commit bb8a1ccb08
28 changed files with 823 additions and 0 deletions
+25
View File
@@ -0,0 +1,25 @@
# ADR 0003: Use pyannote for Speaker Diarization
## Status
Accepted
## Context
The project requires speaker diarization to assign transcript segments to speakers.
The first implementation should provide reliable separation of speakers in meetings, while keeping the diarization component replaceable.
## Decision
The project will use `pyannote.audio` as the initial speaker diarization engine.
Speaker diarization will be treated as a separate pipeline step after transcription.
## Consequences
- Diarization can be improved or replaced independently from transcription.
- Speaker labels are derived metadata and must not modify the canonical transcript.
- Human correction of speaker names should be supported later.
- The diarization module must hide the concrete engine behind an internal interface.
- Model name, engine version and timestamp should be stored with every diarization result.