Establish initial project architecture

This commit is contained in:
2026-07-08 15:57:35 +02:00
commit bb8a1ccb08
28 changed files with 823 additions and 0 deletions
BIN
View File
Binary file not shown.
+32
View File
@@ -0,0 +1,32 @@
# ADR 0001: Project Vision
## Status
Accepted
## Context
The project aims to create a local-first meeting assistant that transforms recorded meetings into high-quality transcripts and structured organizational knowledge.
Existing commercial meeting assistants often depend on bots, cloud services, vendor-specific workflows or opaque AI pipelines. This project is intended to provide more control over audio capture, transcription, knowledge extraction and long-term data ownership.
## Decision
The project will be developed as a modular Python application with a clear pipeline:
Recording
→ Transcription
→ Speaker Diarization
→ AI Analysis
→ Knowledge Extraction
→ Export
The transcript is treated as the canonical source. AI-generated summaries and knowledge objects are derived artifacts.
## Consequences
- Audio and transcript quality have priority over UI features.
- The system must preserve original recordings and original transcripts.
- AI outputs must be reproducible where practical.
- Modules should remain replaceable.
- The project should avoid unnecessary vendor lock-in.
+25
View File
@@ -0,0 +1,25 @@
# ADR 0002: Use faster-whisper for Transcription
## Status
Accepted
## Context
The project requires high-quality speech-to-text transcription for German and English meetings.
The transcription component should support local execution where practical and should remain replaceable in the future. The project should not depend on a meeting bot or on a proprietary meeting platform.
## Decision
The project will use `faster-whisper` as the initial transcription engine.
The preferred first model is `large-v3-turbo`, with `large-v3` as quality-oriented fallback if required.
## Consequences
- Transcription can run locally on suitable hardware.
- NVIDIA CUDA acceleration can be used later on an office AI PC.
- The transcription module must hide the concrete engine behind an internal interface.
- Model name, language setting, timestamp and engine version should be stored with every transcript.
- The decision can be revisited if another engine provides clearly better quality, speed or deployment characteristics.
+25
View File
@@ -0,0 +1,25 @@
# ADR 0003: Use pyannote for Speaker Diarization
## Status
Accepted
## Context
The project requires speaker diarization to assign transcript segments to speakers.
The first implementation should provide reliable separation of speakers in meetings, while keeping the diarization component replaceable.
## Decision
The project will use `pyannote.audio` as the initial speaker diarization engine.
Speaker diarization will be treated as a separate pipeline step after transcription.
## Consequences
- Diarization can be improved or replaced independently from transcription.
- Speaker labels are derived metadata and must not modify the canonical transcript.
- Human correction of speaker names should be supported later.
- The diarization module must hide the concrete engine behind an internal interface.
- Model name, engine version and timestamp should be stored with every diarization result.
@@ -0,0 +1,38 @@
# ADR 0004: Store Original Transcript as Canonical Source
## Status
Accepted
## Context
Speech-to-text systems continuously improve over time. Likewise, AI-generated summaries, action items and other derived artifacts may evolve as better models become available.
To ensure reproducibility, traceability and future reprocessing, the project requires a single immutable source of truth.
## Decision
The original transcript generated by the transcription engine is stored as the canonical transcript.
The canonical transcript is immutable and must never be modified.
A separate working transcript may be created for manual corrections such as:
- spelling corrections
- terminology normalization
- speaker name assignment
- punctuation improvements
The working transcript shall maintain a reference to the canonical transcript from which it was created.
All AI-generated artifacts (summaries, decisions, action items, knowledge objects, etc.) are derived from either the canonical transcript or a specific working transcript. The transcript version used must always be recorded.
## Consequences
- The original transcript remains permanently available.
- Future AI models can regenerate improved summaries without requiring a new recording.
- Manual corrections do not overwrite the canonical transcript.
- Traceability and reproducibility are preserved throughout the system.
- A working transcript enables manual improvements without compromising data integrity.
- Every derived artifact can reference the transcript version from which it was generated.
- Multiple working transcript revisions may coexist if required.
+42
View File
@@ -0,0 +1,42 @@
# ADR 0005: Meeting Domain Model
## Status
Accepted
## Context
The application manages meetings as structured collections of related artifacts rather than as individual files.
A consistent domain model is required to support future extensions such as semantic search, knowledge extraction, versioning and database storage.
## Decision
The central entity of the application is the Meeting.
A Meeting owns or references all artifacts created during its lifecycle.
The initial domain model consists of:
- Meeting
- Recording
- Transcript
- Working Transcript
- Speaker
- Participant
- AI Artifact
- Action Item
- Decision
- Knowledge Object
- Attachment
Each entity has a unique identifier.
Relationships between entities shall be maintained explicitly.
## Consequences
- New artifact types can be added without changing the overall architecture.
- Storage implementation (files, SQLite, PostgreSQL, etc.) remains independent from the domain model.
- AI outputs become first-class entities rather than generated text files.
- Future integrations can operate on domain objects instead of file names.
+180
View File
@@ -0,0 +1,180 @@
# Architecture
## Purpose
The Meeting Knowledge Assistant transforms recorded meetings into structured organizational knowledge through a modular processing pipeline.
---
# Design Principles
- Architecture first
- Offline first where practical
- Immutable source data
- Replaceable AI components
- Small, focused modules
- Explicit interfaces
- Reproducible AI outputs
---
# High-Level Pipeline
Recording
↓
Transcription
↓
Speaker Diarization
↓
Working Transcript
↓
LLM Analysis
↓
Knowledge Extraction
↓
Export
---
# Domain Model
Meeting
│
├── Recording
├── Transcript
│ ├── Canonical
│ └── Working
├── Speakers
├── Participants
├── AI Artifacts
├── Knowledge Objects
├── Attachments
└── Exports
---
# Module Responsibilities
## Recorder
Responsible only for capturing audio.
Input:
None
Output:
Recording
---
## Transcription
Responsible only for speech-to-text conversion.
Input:
Recording
Output:
Canonical Transcript
---
## Diarization
Responsible only for speaker identification.
Input:
Recording + Canonical Transcript
Output:
Working Transcript
---
## LLM
Responsible for semantic analysis.
Input:
Transcript
Output:
AI Artifacts
---
## Knowledge Extraction
Responsible for creating structured knowledge.
Input:
AI Artifacts
Output:
Knowledge Objects
---
## Export
Responsible for creating user-facing documents.
Input:
Knowledge Objects
Output:
Markdown
PDF
DOCX
---
# Artifact Lifecycle
Recording
↓
Canonical Transcript
↓
Working Transcript
↓
AI Artifacts
↓
Knowledge Objects
↓
Exports
---
# Storage Strategy
The domain model is independent of the storage backend.
Possible implementations:
- File System
- SQLite
- PostgreSQL
- Cloud Storage
---
# Future Extensions
- Live transcription
- Video processing
- OCR
- Semantic search
- Knowledge graph
- Company glossary
- Multi-language meetings
- Local LLM support
+20
View File
@@ -0,0 +1,20 @@
Transcript
Original speech-to-text output.
Diarization
Assignment of transcript segments to speakers.
Knowledge Object
Structured information extracted from meetings.
Action Item
A task assigned during a meeting.
Decision
A formally agreed conclusion.
Meeting Artifact
Any output generated from a meeting.
Canonical Transcript
The immutable original transcript.
View File