From 31801729ed223daaaa8302e4400f5096207dd1d4 Mon Sep 17 00:00:00 2001 From: Martin Tazl Date: Mon, 20 Jul 2026 12:21:04 +0200 Subject: [PATCH] Refine chunked transcription strategy --- CONTRIBUTING.md | 38 +++++++++++++++++ docs/adr/0002-use-faster-whisper.md | 50 ++++++++++++++++++++-- docs/adr/0008-meeting-storage-layout.md | 7 +++- docs/adr/0010-development-workflow.md | 55 +++++++++++++++++++++++++ 4 files changed, 146 insertions(+), 4 deletions(-) create mode 100644 docs/adr/0010-development-workflow.md diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index e69de29..625dceb 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -0,0 +1,38 @@ +# Contributing + +1. Branch Strategy +2. Development Workflow +3. Commit Messages +4. Definition of Done +5. Coding Standards +6. Documentation Requirements +7. Testing +8. Creating ADRs + +## Development Workflow + +See ADR 0010 for the overall development workflow. + +--- + +## Commit Messages + +Commit messages should be: + +- short +- descriptive +- written in the imperative mood + +Examples: + +- Add meeting import +- Implement transcript persistence +- Refactor storage layer +- Fix timestamp serialization + +Avoid messages such as: + +- Update +- Changes +- Fix +- Miscellaneous diff --git a/docs/adr/0002-use-faster-whisper.md b/docs/adr/0002-use-faster-whisper.md index 30dab4d..7d5f30f 100644 --- a/docs/adr/0002-use-faster-whisper.md +++ b/docs/adr/0002-use-faster-whisper.md @@ -10,16 +10,60 @@ The project requires high-quality speech-to-text transcription for German and En The transcription component should support local execution where practical and should remain replaceable in the future. The project should not depend on a meeting bot or on a proprietary meeting platform. +Practical testing showed that long continuous recordings may suffer from error propagation when the decoder continuously conditions on previously generated text. + +Independent transcription of smaller audio chunks produced significantly more stable results. + ## Decision The project will use `faster-whisper` as the initial transcription engine. -The preferred first model is `large-v3-turbo`, with `large-v3` as quality-oriented fallback if required. +The preferred first model is `large-v3-turbo`, with `large-v3` as a quality-oriented fallback if required. + +The transcription pipeline shall process recordings as independent audio chunks rather than as a single continuous recording. + +The default processing strategy is: + +```text +Recording (WAV) + │ + ▼ +Split into 5-minute chunks + │ + ▼ +Independent Whisper transcription +(condition_on_previous_text = False) + │ + ▼ +JSON transcript per chunk + │ + ▼ +Merge chunk transcripts + │ + ▼ +Canonical transcript +``` + +Chunk duration shall be configurable, with **5 minutes** as the default value. + +Each chunk shall be processed independently. + +The decoder shall not use text generated from previous chunks as context. + +Each chunk shall produce its own intermediate JSON transcript before merging. + +The merge step shall preserve timestamps and chunk ordering. ## Consequences +- Long recordings become more robust. +- Error propagation across chunk boundaries is prevented. +- Failed chunks can be retranscribed independently. +- Parallel processing of multiple chunks becomes possible. +- Intermediate JSON files simplify debugging and future reprocessing. +- Chunk size can be tuned later without changing the overall architecture. - Transcription can run locally on suitable hardware. - NVIDIA CUDA acceleration can be used later on an office AI PC. - The transcription module must hide the concrete engine behind an internal interface. -- Model name, language setting, timestamp and engine version should be stored with every transcript. -- The decision can be revisited if another engine provides clearly better quality, speed or deployment characteristics. +- Model name, language setting, timestamp mode, engine version, chunk duration and relevant decoding parameters shall be stored with every transcript. +- The decision can be revisited if another engine provides clearly better quality, speed or deployment characteristics. \ No newline at end of file diff --git a/docs/adr/0008-meeting-storage-layout.md b/docs/adr/0008-meeting-storage-layout.md index 44535fd..edf61db 100644 --- a/docs/adr/0008-meeting-storage-layout.md +++ b/docs/adr/0008-meeting-storage-layout.md @@ -54,7 +54,12 @@ meeting/ ├── transcript/ │ ├── canonical.json │ ├── working_v1.json -│ └── working_v2.json +│ +├── chunks/ +│ ├── chunk_000.json +│ ├── chunk_001.json +│ ├── chunk_002.json +│ ├── ... │ ├── artifacts/ │ ├── summary.md diff --git a/docs/adr/0010-development-workflow.md b/docs/adr/0010-development-workflow.md new file mode 100644 index 0000000..97e65b3 --- /dev/null +++ b/docs/adr/0010-development-workflow.md @@ -0,0 +1,55 @@ +# ADR 0010: Development Workflow + +## Status + +Accepted + +## Context + +The project is expected to evolve over an extended period and will be developed with the assistance of AI coding tools. + +To maintain code quality, architectural consistency and a comprehensible project history, a common development workflow is required. + +## Decision + +The project follows a feature-branch based development workflow. + +The `main` branch shall always represent a stable and working state. + +Development of new functionality shall take place on dedicated feature branches. + +Examples: + +- feature/domain-model +- feature/storage +- feature/import +- feature/transcription +- feature/diarization +- feature/export +- feature/ui + +Feature branches shall be merged into `main` only after: + +- implementation is complete +- documentation has been updated where required +- all tests pass +- static analysis completes without errors + +Architectural changes shall be documented using Architecture Decision Records (ADRs). + +Project versioning follows Semantic Versioning. + +Version numbers shall be maintained in: + +- `pyproject.toml` +- `CHANGELOG.md` + +Each released version shall be tagged in Git. + + +## Consequences + +- Development history remains easy to understand. +- Experimental work is isolated from the stable branch. +- Architectural decisions remain documented. +- Releases are reproducible through Git tags.