Establish initial project architecture
This commit is contained in:
@@ -0,0 +1,32 @@
|
||||
# ADR 0001: Project Vision
|
||||
|
||||
## Status
|
||||
|
||||
Accepted
|
||||
|
||||
## Context
|
||||
|
||||
The project aims to create a local-first meeting assistant that transforms recorded meetings into high-quality transcripts and structured organizational knowledge.
|
||||
|
||||
Existing commercial meeting assistants often depend on bots, cloud services, vendor-specific workflows or opaque AI pipelines. This project is intended to provide more control over audio capture, transcription, knowledge extraction and long-term data ownership.
|
||||
|
||||
## Decision
|
||||
|
||||
The project will be developed as a modular Python application with a clear pipeline:
|
||||
|
||||
Recording
|
||||
→ Transcription
|
||||
→ Speaker Diarization
|
||||
→ AI Analysis
|
||||
→ Knowledge Extraction
|
||||
→ Export
|
||||
|
||||
The transcript is treated as the canonical source. AI-generated summaries and knowledge objects are derived artifacts.
|
||||
|
||||
## Consequences
|
||||
|
||||
- Audio and transcript quality have priority over UI features.
|
||||
- The system must preserve original recordings and original transcripts.
|
||||
- AI outputs must be reproducible where practical.
|
||||
- Modules should remain replaceable.
|
||||
- The project should avoid unnecessary vendor lock-in.
|
||||
@@ -0,0 +1,25 @@
|
||||
# ADR 0002: Use faster-whisper for Transcription
|
||||
|
||||
## Status
|
||||
|
||||
Accepted
|
||||
|
||||
## Context
|
||||
|
||||
The project requires high-quality speech-to-text transcription for German and English meetings.
|
||||
|
||||
The transcription component should support local execution where practical and should remain replaceable in the future. The project should not depend on a meeting bot or on a proprietary meeting platform.
|
||||
|
||||
## Decision
|
||||
|
||||
The project will use `faster-whisper` as the initial transcription engine.
|
||||
|
||||
The preferred first model is `large-v3-turbo`, with `large-v3` as quality-oriented fallback if required.
|
||||
|
||||
## Consequences
|
||||
|
||||
- Transcription can run locally on suitable hardware.
|
||||
- NVIDIA CUDA acceleration can be used later on an office AI PC.
|
||||
- The transcription module must hide the concrete engine behind an internal interface.
|
||||
- Model name, language setting, timestamp and engine version should be stored with every transcript.
|
||||
- The decision can be revisited if another engine provides clearly better quality, speed or deployment characteristics.
|
||||
@@ -0,0 +1,25 @@
|
||||
# ADR 0003: Use pyannote for Speaker Diarization
|
||||
|
||||
## Status
|
||||
|
||||
Accepted
|
||||
|
||||
## Context
|
||||
|
||||
The project requires speaker diarization to assign transcript segments to speakers.
|
||||
|
||||
The first implementation should provide reliable separation of speakers in meetings, while keeping the diarization component replaceable.
|
||||
|
||||
## Decision
|
||||
|
||||
The project will use `pyannote.audio` as the initial speaker diarization engine.
|
||||
|
||||
Speaker diarization will be treated as a separate pipeline step after transcription.
|
||||
|
||||
## Consequences
|
||||
|
||||
- Diarization can be improved or replaced independently from transcription.
|
||||
- Speaker labels are derived metadata and must not modify the canonical transcript.
|
||||
- Human correction of speaker names should be supported later.
|
||||
- The diarization module must hide the concrete engine behind an internal interface.
|
||||
- Model name, engine version and timestamp should be stored with every diarization result.
|
||||
@@ -0,0 +1,38 @@
|
||||
# ADR 0004: Store Original Transcript as Canonical Source
|
||||
|
||||
## Status
|
||||
|
||||
Accepted
|
||||
|
||||
## Context
|
||||
|
||||
Speech-to-text systems continuously improve over time. Likewise, AI-generated summaries, action items and other derived artifacts may evolve as better models become available.
|
||||
|
||||
To ensure reproducibility, traceability and future reprocessing, the project requires a single immutable source of truth.
|
||||
|
||||
## Decision
|
||||
|
||||
The original transcript generated by the transcription engine is stored as the canonical transcript.
|
||||
|
||||
The canonical transcript is immutable and must never be modified.
|
||||
|
||||
A separate working transcript may be created for manual corrections such as:
|
||||
|
||||
- spelling corrections
|
||||
- terminology normalization
|
||||
- speaker name assignment
|
||||
- punctuation improvements
|
||||
|
||||
The working transcript shall maintain a reference to the canonical transcript from which it was created.
|
||||
|
||||
All AI-generated artifacts (summaries, decisions, action items, knowledge objects, etc.) are derived from either the canonical transcript or a specific working transcript. The transcript version used must always be recorded.
|
||||
|
||||
## Consequences
|
||||
|
||||
- The original transcript remains permanently available.
|
||||
- Future AI models can regenerate improved summaries without requiring a new recording.
|
||||
- Manual corrections do not overwrite the canonical transcript.
|
||||
- Traceability and reproducibility are preserved throughout the system.
|
||||
- A working transcript enables manual improvements without compromising data integrity.
|
||||
- Every derived artifact can reference the transcript version from which it was generated.
|
||||
- Multiple working transcript revisions may coexist if required.
|
||||
@@ -0,0 +1,42 @@
|
||||
# ADR 0005: Meeting Domain Model
|
||||
|
||||
## Status
|
||||
|
||||
Accepted
|
||||
|
||||
## Context
|
||||
|
||||
The application manages meetings as structured collections of related artifacts rather than as individual files.
|
||||
|
||||
A consistent domain model is required to support future extensions such as semantic search, knowledge extraction, versioning and database storage.
|
||||
|
||||
## Decision
|
||||
|
||||
The central entity of the application is the Meeting.
|
||||
|
||||
A Meeting owns or references all artifacts created during its lifecycle.
|
||||
|
||||
The initial domain model consists of:
|
||||
|
||||
- Meeting
|
||||
- Recording
|
||||
- Transcript
|
||||
- Working Transcript
|
||||
- Speaker
|
||||
- Participant
|
||||
- AI Artifact
|
||||
- Action Item
|
||||
- Decision
|
||||
- Knowledge Object
|
||||
- Attachment
|
||||
|
||||
Each entity has a unique identifier.
|
||||
|
||||
Relationships between entities shall be maintained explicitly.
|
||||
|
||||
## Consequences
|
||||
|
||||
- New artifact types can be added without changing the overall architecture.
|
||||
- Storage implementation (files, SQLite, PostgreSQL, etc.) remains independent from the domain model.
|
||||
- AI outputs become first-class entities rather than generated text files.
|
||||
- Future integrations can operate on domain objects instead of file names.
|
||||
Reference in New Issue
Block a user