39 lines
1.6 KiB
Markdown
39 lines
1.6 KiB
Markdown
# ADR 0004: Store Original Transcript as Canonical Source
|
|
|
|
## Status
|
|
|
|
Accepted
|
|
|
|
## Context
|
|
|
|
Speech-to-text systems continuously improve over time. Likewise, AI-generated summaries, action items and other derived artifacts may evolve as better models become available.
|
|
|
|
To ensure reproducibility, traceability and future reprocessing, the project requires a single immutable source of truth.
|
|
|
|
## Decision
|
|
|
|
The original transcript generated by the transcription engine is stored as the canonical transcript.
|
|
|
|
The canonical transcript is immutable and must never be modified.
|
|
|
|
A separate working transcript may be created for manual corrections such as:
|
|
|
|
- spelling corrections
|
|
- terminology normalization
|
|
- speaker name assignment
|
|
- punctuation improvements
|
|
|
|
The working transcript shall maintain a reference to the canonical transcript from which it was created.
|
|
|
|
All AI-generated artifacts (summaries, decisions, action items, knowledge objects, etc.) are derived from either the canonical transcript or a specific working transcript. The transcript version used must always be recorded.
|
|
|
|
## Consequences
|
|
|
|
- The original transcript remains permanently available.
|
|
- Future AI models can regenerate improved summaries without requiring a new recording.
|
|
- Manual corrections do not overwrite the canonical transcript.
|
|
- Traceability and reproducibility are preserved throughout the system.
|
|
- A working transcript enables manual improvements without compromising data integrity.
|
|
- Every derived artifact can reference the transcript version from which it was generated.
|
|
- Multiple working transcript revisions may coexist if required.
|