69 lines
2.5 KiB
Markdown
69 lines
2.5 KiB
Markdown
# ADR 0002: Use faster-whisper for Transcription
|
|
|
|
## Status
|
|
|
|
Accepted
|
|
|
|
## Context
|
|
|
|
The project requires high-quality speech-to-text transcription for German and English meetings.
|
|
|
|
The transcription component should support local execution where practical and should remain replaceable in the future. The project should not depend on a meeting bot or on a proprietary meeting platform.
|
|
|
|
Practical testing showed that long continuous recordings may suffer from error propagation when the decoder continuously conditions on previously generated text.
|
|
|
|
Independent transcription of smaller audio chunks produced significantly more stable results.
|
|
|
|
## Decision
|
|
|
|
The project will use `faster-whisper` as the initial transcription engine.
|
|
|
|
The preferred first model is `large-v3-turbo`, with `large-v3` as a quality-oriented fallback if required.
|
|
|
|
The transcription pipeline shall process recordings as independent audio chunks rather than as a single continuous recording.
|
|
|
|
The default processing strategy is:
|
|
|
|
```text
|
|
Recording (WAV)
|
|
│
|
|
▼
|
|
Split into 5-minute chunks
|
|
│
|
|
▼
|
|
Independent Whisper transcription
|
|
(condition_on_previous_text = False)
|
|
│
|
|
▼
|
|
JSON transcript per chunk
|
|
│
|
|
▼
|
|
Merge chunk transcripts
|
|
│
|
|
▼
|
|
Canonical transcript
|
|
```
|
|
|
|
Chunk duration shall be configurable, with **5 minutes** as the default value.
|
|
|
|
Each chunk shall be processed independently.
|
|
|
|
The decoder shall not use text generated from previous chunks as context.
|
|
|
|
Each chunk shall produce its own intermediate JSON transcript before merging.
|
|
|
|
The merge step shall preserve timestamps and chunk ordering.
|
|
|
|
## Consequences
|
|
|
|
- Long recordings become more robust.
|
|
- Error propagation across chunk boundaries is prevented.
|
|
- Failed chunks can be retranscribed independently.
|
|
- Parallel processing of multiple chunks becomes possible.
|
|
- Intermediate JSON files simplify debugging and future reprocessing.
|
|
- Chunk size can be tuned later without changing the overall architecture.
|
|
- Transcription can run locally on suitable hardware.
|
|
- NVIDIA CUDA acceleration can be used later on an office AI PC.
|
|
- The transcription module must hide the concrete engine behind an internal interface.
|
|
- Model name, language setting, timestamp mode, engine version, chunk duration and relevant decoding parameters shall be stored with every transcript.
|
|
- The decision can be revisited if another engine provides clearly better quality, speed or deployment characteristics. |