Refine chunked transcription strategy
This commit is contained in:
@@ -0,0 +1,38 @@
|
|||||||
|
# Contributing
|
||||||
|
|
||||||
|
1. Branch Strategy
|
||||||
|
2. Development Workflow
|
||||||
|
3. Commit Messages
|
||||||
|
4. Definition of Done
|
||||||
|
5. Coding Standards
|
||||||
|
6. Documentation Requirements
|
||||||
|
7. Testing
|
||||||
|
8. Creating ADRs
|
||||||
|
|
||||||
|
## Development Workflow
|
||||||
|
|
||||||
|
See ADR 0010 for the overall development workflow.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Commit Messages
|
||||||
|
|
||||||
|
Commit messages should be:
|
||||||
|
|
||||||
|
- short
|
||||||
|
- descriptive
|
||||||
|
- written in the imperative mood
|
||||||
|
|
||||||
|
Examples:
|
||||||
|
|
||||||
|
- Add meeting import
|
||||||
|
- Implement transcript persistence
|
||||||
|
- Refactor storage layer
|
||||||
|
- Fix timestamp serialization
|
||||||
|
|
||||||
|
Avoid messages such as:
|
||||||
|
|
||||||
|
- Update
|
||||||
|
- Changes
|
||||||
|
- Fix
|
||||||
|
- Miscellaneous
|
||||||
|
|||||||
@@ -10,16 +10,60 @@ The project requires high-quality speech-to-text transcription for German and En
|
|||||||
|
|
||||||
The transcription component should support local execution where practical and should remain replaceable in the future. The project should not depend on a meeting bot or on a proprietary meeting platform.
|
The transcription component should support local execution where practical and should remain replaceable in the future. The project should not depend on a meeting bot or on a proprietary meeting platform.
|
||||||
|
|
||||||
|
Practical testing showed that long continuous recordings may suffer from error propagation when the decoder continuously conditions on previously generated text.
|
||||||
|
|
||||||
|
Independent transcription of smaller audio chunks produced significantly more stable results.
|
||||||
|
|
||||||
## Decision
|
## Decision
|
||||||
|
|
||||||
The project will use `faster-whisper` as the initial transcription engine.
|
The project will use `faster-whisper` as the initial transcription engine.
|
||||||
|
|
||||||
The preferred first model is `large-v3-turbo`, with `large-v3` as quality-oriented fallback if required.
|
The preferred first model is `large-v3-turbo`, with `large-v3` as a quality-oriented fallback if required.
|
||||||
|
|
||||||
|
The transcription pipeline shall process recordings as independent audio chunks rather than as a single continuous recording.
|
||||||
|
|
||||||
|
The default processing strategy is:
|
||||||
|
|
||||||
|
```text
|
||||||
|
Recording (WAV)
|
||||||
|
│
|
||||||
|
▼
|
||||||
|
Split into 5-minute chunks
|
||||||
|
│
|
||||||
|
▼
|
||||||
|
Independent Whisper transcription
|
||||||
|
(condition_on_previous_text = False)
|
||||||
|
│
|
||||||
|
▼
|
||||||
|
JSON transcript per chunk
|
||||||
|
│
|
||||||
|
▼
|
||||||
|
Merge chunk transcripts
|
||||||
|
│
|
||||||
|
▼
|
||||||
|
Canonical transcript
|
||||||
|
```
|
||||||
|
|
||||||
|
Chunk duration shall be configurable, with **5 minutes** as the default value.
|
||||||
|
|
||||||
|
Each chunk shall be processed independently.
|
||||||
|
|
||||||
|
The decoder shall not use text generated from previous chunks as context.
|
||||||
|
|
||||||
|
Each chunk shall produce its own intermediate JSON transcript before merging.
|
||||||
|
|
||||||
|
The merge step shall preserve timestamps and chunk ordering.
|
||||||
|
|
||||||
## Consequences
|
## Consequences
|
||||||
|
|
||||||
|
- Long recordings become more robust.
|
||||||
|
- Error propagation across chunk boundaries is prevented.
|
||||||
|
- Failed chunks can be retranscribed independently.
|
||||||
|
- Parallel processing of multiple chunks becomes possible.
|
||||||
|
- Intermediate JSON files simplify debugging and future reprocessing.
|
||||||
|
- Chunk size can be tuned later without changing the overall architecture.
|
||||||
- Transcription can run locally on suitable hardware.
|
- Transcription can run locally on suitable hardware.
|
||||||
- NVIDIA CUDA acceleration can be used later on an office AI PC.
|
- NVIDIA CUDA acceleration can be used later on an office AI PC.
|
||||||
- The transcription module must hide the concrete engine behind an internal interface.
|
- The transcription module must hide the concrete engine behind an internal interface.
|
||||||
- Model name, language setting, timestamp and engine version should be stored with every transcript.
|
- Model name, language setting, timestamp mode, engine version, chunk duration and relevant decoding parameters shall be stored with every transcript.
|
||||||
- The decision can be revisited if another engine provides clearly better quality, speed or deployment characteristics.
|
- The decision can be revisited if another engine provides clearly better quality, speed or deployment characteristics.
|
||||||
@@ -54,7 +54,12 @@ meeting/
|
|||||||
├── transcript/
|
├── transcript/
|
||||||
│ ├── canonical.json
|
│ ├── canonical.json
|
||||||
│ ├── working_v1.json
|
│ ├── working_v1.json
|
||||||
│ └── working_v2.json
|
│
|
||||||
|
├── chunks/
|
||||||
|
│ ├── chunk_000.json
|
||||||
|
│ ├── chunk_001.json
|
||||||
|
│ ├── chunk_002.json
|
||||||
|
│ ├── ...
|
||||||
│
|
│
|
||||||
├── artifacts/
|
├── artifacts/
|
||||||
│ ├── summary.md
|
│ ├── summary.md
|
||||||
|
|||||||
@@ -0,0 +1,55 @@
|
|||||||
|
# ADR 0010: Development Workflow
|
||||||
|
|
||||||
|
## Status
|
||||||
|
|
||||||
|
Accepted
|
||||||
|
|
||||||
|
## Context
|
||||||
|
|
||||||
|
The project is expected to evolve over an extended period and will be developed with the assistance of AI coding tools.
|
||||||
|
|
||||||
|
To maintain code quality, architectural consistency and a comprehensible project history, a common development workflow is required.
|
||||||
|
|
||||||
|
## Decision
|
||||||
|
|
||||||
|
The project follows a feature-branch based development workflow.
|
||||||
|
|
||||||
|
The `main` branch shall always represent a stable and working state.
|
||||||
|
|
||||||
|
Development of new functionality shall take place on dedicated feature branches.
|
||||||
|
|
||||||
|
Examples:
|
||||||
|
|
||||||
|
- feature/domain-model
|
||||||
|
- feature/storage
|
||||||
|
- feature/import
|
||||||
|
- feature/transcription
|
||||||
|
- feature/diarization
|
||||||
|
- feature/export
|
||||||
|
- feature/ui
|
||||||
|
|
||||||
|
Feature branches shall be merged into `main` only after:
|
||||||
|
|
||||||
|
- implementation is complete
|
||||||
|
- documentation has been updated where required
|
||||||
|
- all tests pass
|
||||||
|
- static analysis completes without errors
|
||||||
|
|
||||||
|
Architectural changes shall be documented using Architecture Decision Records (ADRs).
|
||||||
|
|
||||||
|
Project versioning follows Semantic Versioning.
|
||||||
|
|
||||||
|
Version numbers shall be maintained in:
|
||||||
|
|
||||||
|
- `pyproject.toml`
|
||||||
|
- `CHANGELOG.md`
|
||||||
|
|
||||||
|
Each released version shall be tagged in Git.
|
||||||
|
|
||||||
|
|
||||||
|
## Consequences
|
||||||
|
|
||||||
|
- Development history remains easy to understand.
|
||||||
|
- Experimental work is isolated from the stable branch.
|
||||||
|
- Architectural decisions remain documented.
|
||||||
|
- Releases are reproducible through Git tags.
|
||||||
Reference in New Issue
Block a user