Files
meeting-assistant/docs/adr/0004-store-original-transcript.md

1.6 KiB

ADR 0004: Store Original Transcript as Canonical Source

Status

Accepted

Context

Speech-to-text systems continuously improve over time. Likewise, AI-generated summaries, action items and other derived artifacts may evolve as better models become available.

To ensure reproducibility, traceability and future reprocessing, the project requires a single immutable source of truth.

Decision

The original transcript generated by the transcription engine is stored as the canonical transcript.

The canonical transcript is immutable and must never be modified.

A separate working transcript may be created for manual corrections such as:

  • spelling corrections
  • terminology normalization
  • speaker name assignment
  • punctuation improvements

The working transcript shall maintain a reference to the canonical transcript from which it was created.

All AI-generated artifacts (summaries, decisions, action items, knowledge objects, etc.) are derived from either the canonical transcript or a specific working transcript. The transcript version used must always be recorded.

Consequences

  • The original transcript remains permanently available.
  • Future AI models can regenerate improved summaries without requiring a new recording.
  • Manual corrections do not overwrite the canonical transcript.
  • Traceability and reproducibility are preserved throughout the system.
  • A working transcript enables manual improvements without compromising data integrity.
  • Every derived artifact can reference the transcript version from which it was generated.
  • Multiple working transcript revisions may coexist if required.