# ADR 0004: Store Original Transcript as Canonical Source ## Status Accepted ## Context Speech-to-text systems continuously improve over time. Likewise, AI-generated summaries, action items and other derived artifacts may evolve as better models become available. To ensure reproducibility, traceability and future reprocessing, the project requires a single immutable source of truth. ## Decision The original transcript generated by the transcription engine is stored as the canonical transcript. The canonical transcript is immutable and must never be modified. A separate working transcript may be created for manual corrections such as: - spelling corrections - terminology normalization - speaker name assignment - punctuation improvements The working transcript shall maintain a reference to the canonical transcript from which it was created. All AI-generated artifacts (summaries, decisions, action items, knowledge objects, etc.) are derived from either the canonical transcript or a specific working transcript. The transcript version used must always be recorded. ## Consequences - The original transcript remains permanently available. - Future AI models can regenerate improved summaries without requiring a new recording. - Manual corrections do not overwrite the canonical transcript. - Traceability and reproducibility are preserved throughout the system. - A working transcript enables manual improvements without compromising data integrity. - Every derived artifact can reference the transcript version from which it was generated. - Multiple working transcript revisions may coexist if required.