Refine chunked transcription strategy
This commit is contained in:
@@ -10,16 +10,60 @@ The project requires high-quality speech-to-text transcription for German and En
|
||||
|
||||
The transcription component should support local execution where practical and should remain replaceable in the future. The project should not depend on a meeting bot or on a proprietary meeting platform.
|
||||
|
||||
Practical testing showed that long continuous recordings may suffer from error propagation when the decoder continuously conditions on previously generated text.
|
||||
|
||||
Independent transcription of smaller audio chunks produced significantly more stable results.
|
||||
|
||||
## Decision
|
||||
|
||||
The project will use `faster-whisper` as the initial transcription engine.
|
||||
|
||||
The preferred first model is `large-v3-turbo`, with `large-v3` as quality-oriented fallback if required.
|
||||
The preferred first model is `large-v3-turbo`, with `large-v3` as a quality-oriented fallback if required.
|
||||
|
||||
The transcription pipeline shall process recordings as independent audio chunks rather than as a single continuous recording.
|
||||
|
||||
The default processing strategy is:
|
||||
|
||||
```text
|
||||
Recording (WAV)
|
||||
│
|
||||
▼
|
||||
Split into 5-minute chunks
|
||||
│
|
||||
▼
|
||||
Independent Whisper transcription
|
||||
(condition_on_previous_text = False)
|
||||
│
|
||||
▼
|
||||
JSON transcript per chunk
|
||||
│
|
||||
▼
|
||||
Merge chunk transcripts
|
||||
│
|
||||
▼
|
||||
Canonical transcript
|
||||
```
|
||||
|
||||
Chunk duration shall be configurable, with **5 minutes** as the default value.
|
||||
|
||||
Each chunk shall be processed independently.
|
||||
|
||||
The decoder shall not use text generated from previous chunks as context.
|
||||
|
||||
Each chunk shall produce its own intermediate JSON transcript before merging.
|
||||
|
||||
The merge step shall preserve timestamps and chunk ordering.
|
||||
|
||||
## Consequences
|
||||
|
||||
- Long recordings become more robust.
|
||||
- Error propagation across chunk boundaries is prevented.
|
||||
- Failed chunks can be retranscribed independently.
|
||||
- Parallel processing of multiple chunks becomes possible.
|
||||
- Intermediate JSON files simplify debugging and future reprocessing.
|
||||
- Chunk size can be tuned later without changing the overall architecture.
|
||||
- Transcription can run locally on suitable hardware.
|
||||
- NVIDIA CUDA acceleration can be used later on an office AI PC.
|
||||
- The transcription module must hide the concrete engine behind an internal interface.
|
||||
- Model name, language setting, timestamp and engine version should be stored with every transcript.
|
||||
- The decision can be revisited if another engine provides clearly better quality, speed or deployment characteristics.
|
||||
- Model name, language setting, timestamp mode, engine version, chunk duration and relevant decoding parameters shall be stored with every transcript.
|
||||
- The decision can be revisited if another engine provides clearly better quality, speed or deployment characteristics.
|
||||
Reference in New Issue
Block a user