Refine chunked transcription strategy
This commit is contained in:
@@ -0,0 +1,38 @@
|
||||
# Contributing
|
||||
|
||||
1. Branch Strategy
|
||||
2. Development Workflow
|
||||
3. Commit Messages
|
||||
4. Definition of Done
|
||||
5. Coding Standards
|
||||
6. Documentation Requirements
|
||||
7. Testing
|
||||
8. Creating ADRs
|
||||
|
||||
## Development Workflow
|
||||
|
||||
See ADR 0010 for the overall development workflow.
|
||||
|
||||
---
|
||||
|
||||
## Commit Messages
|
||||
|
||||
Commit messages should be:
|
||||
|
||||
- short
|
||||
- descriptive
|
||||
- written in the imperative mood
|
||||
|
||||
Examples:
|
||||
|
||||
- Add meeting import
|
||||
- Implement transcript persistence
|
||||
- Refactor storage layer
|
||||
- Fix timestamp serialization
|
||||
|
||||
Avoid messages such as:
|
||||
|
||||
- Update
|
||||
- Changes
|
||||
- Fix
|
||||
- Miscellaneous
|
||||
|
||||
@@ -10,16 +10,60 @@ The project requires high-quality speech-to-text transcription for German and En
|
||||
|
||||
The transcription component should support local execution where practical and should remain replaceable in the future. The project should not depend on a meeting bot or on a proprietary meeting platform.
|
||||
|
||||
Practical testing showed that long continuous recordings may suffer from error propagation when the decoder continuously conditions on previously generated text.
|
||||
|
||||
Independent transcription of smaller audio chunks produced significantly more stable results.
|
||||
|
||||
## Decision
|
||||
|
||||
The project will use `faster-whisper` as the initial transcription engine.
|
||||
|
||||
The preferred first model is `large-v3-turbo`, with `large-v3` as quality-oriented fallback if required.
|
||||
The preferred first model is `large-v3-turbo`, with `large-v3` as a quality-oriented fallback if required.
|
||||
|
||||
The transcription pipeline shall process recordings as independent audio chunks rather than as a single continuous recording.
|
||||
|
||||
The default processing strategy is:
|
||||
|
||||
```text
|
||||
Recording (WAV)
|
||||
│
|
||||
▼
|
||||
Split into 5-minute chunks
|
||||
│
|
||||
▼
|
||||
Independent Whisper transcription
|
||||
(condition_on_previous_text = False)
|
||||
│
|
||||
▼
|
||||
JSON transcript per chunk
|
||||
│
|
||||
▼
|
||||
Merge chunk transcripts
|
||||
│
|
||||
▼
|
||||
Canonical transcript
|
||||
```
|
||||
|
||||
Chunk duration shall be configurable, with **5 minutes** as the default value.
|
||||
|
||||
Each chunk shall be processed independently.
|
||||
|
||||
The decoder shall not use text generated from previous chunks as context.
|
||||
|
||||
Each chunk shall produce its own intermediate JSON transcript before merging.
|
||||
|
||||
The merge step shall preserve timestamps and chunk ordering.
|
||||
|
||||
## Consequences
|
||||
|
||||
- Long recordings become more robust.
|
||||
- Error propagation across chunk boundaries is prevented.
|
||||
- Failed chunks can be retranscribed independently.
|
||||
- Parallel processing of multiple chunks becomes possible.
|
||||
- Intermediate JSON files simplify debugging and future reprocessing.
|
||||
- Chunk size can be tuned later without changing the overall architecture.
|
||||
- Transcription can run locally on suitable hardware.
|
||||
- NVIDIA CUDA acceleration can be used later on an office AI PC.
|
||||
- The transcription module must hide the concrete engine behind an internal interface.
|
||||
- Model name, language setting, timestamp and engine version should be stored with every transcript.
|
||||
- Model name, language setting, timestamp mode, engine version, chunk duration and relevant decoding parameters shall be stored with every transcript.
|
||||
- The decision can be revisited if another engine provides clearly better quality, speed or deployment characteristics.
|
||||
@@ -54,7 +54,12 @@ meeting/
|
||||
├── transcript/
|
||||
│ ├── canonical.json
|
||||
│ ├── working_v1.json
|
||||
│ └── working_v2.json
|
||||
│
|
||||
├── chunks/
|
||||
│ ├── chunk_000.json
|
||||
│ ├── chunk_001.json
|
||||
│ ├── chunk_002.json
|
||||
│ ├── ...
|
||||
│
|
||||
├── artifacts/
|
||||
│ ├── summary.md
|
||||
|
||||
@@ -0,0 +1,55 @@
|
||||
# ADR 0010: Development Workflow
|
||||
|
||||
## Status
|
||||
|
||||
Accepted
|
||||
|
||||
## Context
|
||||
|
||||
The project is expected to evolve over an extended period and will be developed with the assistance of AI coding tools.
|
||||
|
||||
To maintain code quality, architectural consistency and a comprehensible project history, a common development workflow is required.
|
||||
|
||||
## Decision
|
||||
|
||||
The project follows a feature-branch based development workflow.
|
||||
|
||||
The `main` branch shall always represent a stable and working state.
|
||||
|
||||
Development of new functionality shall take place on dedicated feature branches.
|
||||
|
||||
Examples:
|
||||
|
||||
- feature/domain-model
|
||||
- feature/storage
|
||||
- feature/import
|
||||
- feature/transcription
|
||||
- feature/diarization
|
||||
- feature/export
|
||||
- feature/ui
|
||||
|
||||
Feature branches shall be merged into `main` only after:
|
||||
|
||||
- implementation is complete
|
||||
- documentation has been updated where required
|
||||
- all tests pass
|
||||
- static analysis completes without errors
|
||||
|
||||
Architectural changes shall be documented using Architecture Decision Records (ADRs).
|
||||
|
||||
Project versioning follows Semantic Versioning.
|
||||
|
||||
Version numbers shall be maintained in:
|
||||
|
||||
- `pyproject.toml`
|
||||
- `CHANGELOG.md`
|
||||
|
||||
Each released version shall be tagged in Git.
|
||||
|
||||
|
||||
## Consequences
|
||||
|
||||
- Development history remains easy to understand.
|
||||
- Experimental work is isolated from the stable branch.
|
||||
- Architectural decisions remain documented.
|
||||
- Releases are reproducible through Git tags.
|
||||
Reference in New Issue
Block a user