Establish initial project architecture
This commit is contained in:
Vendored
BIN
Binary file not shown.
@@ -0,0 +1,32 @@
|
||||
# ADR 0001: Project Vision
|
||||
|
||||
## Status
|
||||
|
||||
Accepted
|
||||
|
||||
## Context
|
||||
|
||||
The project aims to create a local-first meeting assistant that transforms recorded meetings into high-quality transcripts and structured organizational knowledge.
|
||||
|
||||
Existing commercial meeting assistants often depend on bots, cloud services, vendor-specific workflows or opaque AI pipelines. This project is intended to provide more control over audio capture, transcription, knowledge extraction and long-term data ownership.
|
||||
|
||||
## Decision
|
||||
|
||||
The project will be developed as a modular Python application with a clear pipeline:
|
||||
|
||||
Recording
|
||||
→ Transcription
|
||||
→ Speaker Diarization
|
||||
→ AI Analysis
|
||||
→ Knowledge Extraction
|
||||
→ Export
|
||||
|
||||
The transcript is treated as the canonical source. AI-generated summaries and knowledge objects are derived artifacts.
|
||||
|
||||
## Consequences
|
||||
|
||||
- Audio and transcript quality have priority over UI features.
|
||||
- The system must preserve original recordings and original transcripts.
|
||||
- AI outputs must be reproducible where practical.
|
||||
- Modules should remain replaceable.
|
||||
- The project should avoid unnecessary vendor lock-in.
|
||||
@@ -0,0 +1,25 @@
|
||||
# ADR 0002: Use faster-whisper for Transcription
|
||||
|
||||
## Status
|
||||
|
||||
Accepted
|
||||
|
||||
## Context
|
||||
|
||||
The project requires high-quality speech-to-text transcription for German and English meetings.
|
||||
|
||||
The transcription component should support local execution where practical and should remain replaceable in the future. The project should not depend on a meeting bot or on a proprietary meeting platform.
|
||||
|
||||
## Decision
|
||||
|
||||
The project will use `faster-whisper` as the initial transcription engine.
|
||||
|
||||
The preferred first model is `large-v3-turbo`, with `large-v3` as quality-oriented fallback if required.
|
||||
|
||||
## Consequences
|
||||
|
||||
- Transcription can run locally on suitable hardware.
|
||||
- NVIDIA CUDA acceleration can be used later on an office AI PC.
|
||||
- The transcription module must hide the concrete engine behind an internal interface.
|
||||
- Model name, language setting, timestamp and engine version should be stored with every transcript.
|
||||
- The decision can be revisited if another engine provides clearly better quality, speed or deployment characteristics.
|
||||
@@ -0,0 +1,25 @@
|
||||
# ADR 0003: Use pyannote for Speaker Diarization
|
||||
|
||||
## Status
|
||||
|
||||
Accepted
|
||||
|
||||
## Context
|
||||
|
||||
The project requires speaker diarization to assign transcript segments to speakers.
|
||||
|
||||
The first implementation should provide reliable separation of speakers in meetings, while keeping the diarization component replaceable.
|
||||
|
||||
## Decision
|
||||
|
||||
The project will use `pyannote.audio` as the initial speaker diarization engine.
|
||||
|
||||
Speaker diarization will be treated as a separate pipeline step after transcription.
|
||||
|
||||
## Consequences
|
||||
|
||||
- Diarization can be improved or replaced independently from transcription.
|
||||
- Speaker labels are derived metadata and must not modify the canonical transcript.
|
||||
- Human correction of speaker names should be supported later.
|
||||
- The diarization module must hide the concrete engine behind an internal interface.
|
||||
- Model name, engine version and timestamp should be stored with every diarization result.
|
||||
@@ -0,0 +1,38 @@
|
||||
# ADR 0004: Store Original Transcript as Canonical Source
|
||||
|
||||
## Status
|
||||
|
||||
Accepted
|
||||
|
||||
## Context
|
||||
|
||||
Speech-to-text systems continuously improve over time. Likewise, AI-generated summaries, action items and other derived artifacts may evolve as better models become available.
|
||||
|
||||
To ensure reproducibility, traceability and future reprocessing, the project requires a single immutable source of truth.
|
||||
|
||||
## Decision
|
||||
|
||||
The original transcript generated by the transcription engine is stored as the canonical transcript.
|
||||
|
||||
The canonical transcript is immutable and must never be modified.
|
||||
|
||||
A separate working transcript may be created for manual corrections such as:
|
||||
|
||||
- spelling corrections
|
||||
- terminology normalization
|
||||
- speaker name assignment
|
||||
- punctuation improvements
|
||||
|
||||
The working transcript shall maintain a reference to the canonical transcript from which it was created.
|
||||
|
||||
All AI-generated artifacts (summaries, decisions, action items, knowledge objects, etc.) are derived from either the canonical transcript or a specific working transcript. The transcript version used must always be recorded.
|
||||
|
||||
## Consequences
|
||||
|
||||
- The original transcript remains permanently available.
|
||||
- Future AI models can regenerate improved summaries without requiring a new recording.
|
||||
- Manual corrections do not overwrite the canonical transcript.
|
||||
- Traceability and reproducibility are preserved throughout the system.
|
||||
- A working transcript enables manual improvements without compromising data integrity.
|
||||
- Every derived artifact can reference the transcript version from which it was generated.
|
||||
- Multiple working transcript revisions may coexist if required.
|
||||
@@ -0,0 +1,42 @@
|
||||
# ADR 0005: Meeting Domain Model
|
||||
|
||||
## Status
|
||||
|
||||
Accepted
|
||||
|
||||
## Context
|
||||
|
||||
The application manages meetings as structured collections of related artifacts rather than as individual files.
|
||||
|
||||
A consistent domain model is required to support future extensions such as semantic search, knowledge extraction, versioning and database storage.
|
||||
|
||||
## Decision
|
||||
|
||||
The central entity of the application is the Meeting.
|
||||
|
||||
A Meeting owns or references all artifacts created during its lifecycle.
|
||||
|
||||
The initial domain model consists of:
|
||||
|
||||
- Meeting
|
||||
- Recording
|
||||
- Transcript
|
||||
- Working Transcript
|
||||
- Speaker
|
||||
- Participant
|
||||
- AI Artifact
|
||||
- Action Item
|
||||
- Decision
|
||||
- Knowledge Object
|
||||
- Attachment
|
||||
|
||||
Each entity has a unique identifier.
|
||||
|
||||
Relationships between entities shall be maintained explicitly.
|
||||
|
||||
## Consequences
|
||||
|
||||
- New artifact types can be added without changing the overall architecture.
|
||||
- Storage implementation (files, SQLite, PostgreSQL, etc.) remains independent from the domain model.
|
||||
- AI outputs become first-class entities rather than generated text files.
|
||||
- Future integrations can operate on domain objects instead of file names.
|
||||
@@ -0,0 +1,180 @@
|
||||
# Architecture
|
||||
|
||||
## Purpose
|
||||
|
||||
The Meeting Knowledge Assistant transforms recorded meetings into structured organizational knowledge through a modular processing pipeline.
|
||||
|
||||
---
|
||||
|
||||
# Design Principles
|
||||
|
||||
- Architecture first
|
||||
- Offline first where practical
|
||||
- Immutable source data
|
||||
- Replaceable AI components
|
||||
- Small, focused modules
|
||||
- Explicit interfaces
|
||||
- Reproducible AI outputs
|
||||
|
||||
---
|
||||
|
||||
# High-Level Pipeline
|
||||
|
||||
Recording
|
||||
↓
|
||||
Transcription
|
||||
↓
|
||||
Speaker Diarization
|
||||
↓
|
||||
Working Transcript
|
||||
↓
|
||||
LLM Analysis
|
||||
↓
|
||||
Knowledge Extraction
|
||||
↓
|
||||
Export
|
||||
|
||||
---
|
||||
|
||||
# Domain Model
|
||||
|
||||
Meeting
|
||||
│
|
||||
├── Recording
|
||||
├── Transcript
|
||||
│ ├── Canonical
|
||||
│ └── Working
|
||||
├── Speakers
|
||||
├── Participants
|
||||
├── AI Artifacts
|
||||
├── Knowledge Objects
|
||||
├── Attachments
|
||||
└── Exports
|
||||
|
||||
---
|
||||
|
||||
# Module Responsibilities
|
||||
|
||||
## Recorder
|
||||
|
||||
Responsible only for capturing audio.
|
||||
|
||||
Input:
|
||||
None
|
||||
|
||||
Output:
|
||||
Recording
|
||||
|
||||
---
|
||||
|
||||
## Transcription
|
||||
|
||||
Responsible only for speech-to-text conversion.
|
||||
|
||||
Input:
|
||||
Recording
|
||||
|
||||
Output:
|
||||
Canonical Transcript
|
||||
|
||||
---
|
||||
|
||||
## Diarization
|
||||
|
||||
Responsible only for speaker identification.
|
||||
|
||||
Input:
|
||||
Recording + Canonical Transcript
|
||||
|
||||
Output:
|
||||
Working Transcript
|
||||
|
||||
---
|
||||
|
||||
## LLM
|
||||
|
||||
Responsible for semantic analysis.
|
||||
|
||||
Input:
|
||||
Transcript
|
||||
|
||||
Output:
|
||||
AI Artifacts
|
||||
|
||||
---
|
||||
|
||||
## Knowledge Extraction
|
||||
|
||||
Responsible for creating structured knowledge.
|
||||
|
||||
Input:
|
||||
AI Artifacts
|
||||
|
||||
Output:
|
||||
Knowledge Objects
|
||||
|
||||
---
|
||||
|
||||
## Export
|
||||
|
||||
Responsible for creating user-facing documents.
|
||||
|
||||
Input:
|
||||
Knowledge Objects
|
||||
|
||||
Output:
|
||||
Markdown
|
||||
PDF
|
||||
DOCX
|
||||
|
||||
---
|
||||
|
||||
# Artifact Lifecycle
|
||||
|
||||
Recording
|
||||
|
||||
↓
|
||||
|
||||
Canonical Transcript
|
||||
|
||||
↓
|
||||
|
||||
Working Transcript
|
||||
|
||||
↓
|
||||
|
||||
AI Artifacts
|
||||
|
||||
↓
|
||||
|
||||
Knowledge Objects
|
||||
|
||||
↓
|
||||
|
||||
Exports
|
||||
|
||||
---
|
||||
|
||||
# Storage Strategy
|
||||
|
||||
The domain model is independent of the storage backend.
|
||||
|
||||
Possible implementations:
|
||||
|
||||
- File System
|
||||
- SQLite
|
||||
- PostgreSQL
|
||||
- Cloud Storage
|
||||
|
||||
---
|
||||
|
||||
# Future Extensions
|
||||
|
||||
- Live transcription
|
||||
- Video processing
|
||||
- OCR
|
||||
- Semantic search
|
||||
- Knowledge graph
|
||||
- Company glossary
|
||||
- Multi-language meetings
|
||||
- Local LLM support
|
||||
@@ -0,0 +1,20 @@
|
||||
Transcript
|
||||
Original speech-to-text output.
|
||||
|
||||
Diarization
|
||||
Assignment of transcript segments to speakers.
|
||||
|
||||
Knowledge Object
|
||||
Structured information extracted from meetings.
|
||||
|
||||
Action Item
|
||||
A task assigned during a meeting.
|
||||
|
||||
Decision
|
||||
A formally agreed conclusion.
|
||||
|
||||
Meeting Artifact
|
||||
Any output generated from a meeting.
|
||||
|
||||
Canonical Transcript
|
||||
The immutable original transcript.
|
||||
Reference in New Issue
Block a user