Establish initial project architecture

This commit is contained in:
2026-07-08 15:57:35 +02:00
commit bb8a1ccb08
28 changed files with 823 additions and 0 deletions
Vendored
BIN
View File
Binary file not shown.
View File
View File
+111
View File
@@ -0,0 +1,111 @@
# AGENTS.md
## Purpose
This document contains instructions for AI coding agents working on this repository.
---
## General Principles
Prefer clean architecture over quick fixes.
Prefer readability over cleverness.
Keep modules small.
Avoid unnecessary abstractions.
Minimize dependencies.
Document architectural decisions.
---
## Development Workflow
Never work directly on main.
Implement one feature per branch.
Prefer many small commits.
Each commit should compile.
Each feature should include:
- implementation
- tests
- documentation updates
---
## Coding Style
Follow PEP8.
Use type hints.
Prefer dataclasses where appropriate.
Avoid global state.
Prefer dependency injection.
No hidden side effects.
---
## AI Behaviour
If requirements are ambiguous:
STOP.
Ask questions.
Never invent requirements.
Never silently change behaviour.
---
## Refactoring
Refactor only inside the requested scope.
Avoid unrelated changes.
Do not rename files without justification.
---
## Documentation
Update README if user-visible behaviour changes.
Update PROJECT_KNOWLEDGE.md when technical knowledge changes.
Update ROADMAP.md only for future work.
---
## Priority Order
Correctness
↓
Maintainability
↓
Readability
↓
Performance
↓
Premature optimization
View File
View File
+163
View File
@@ -0,0 +1,163 @@
# Project Knowledge
## Vision
The project creates a complete pipeline that transforms spoken meetings into structured organizational knowledge.
The transcript is the canonical source.
AI-generated summaries are reproducible artefacts.
---
## Design Principles
Offline-first where practical.
Replaceable AI components.
Vendor independence.
Modular architecture.
Small focused modules.
Simple interfaces.
---
## Core Pipeline
Recorder
↓
Transcription
↓
Speaker Identification
↓
LLM Analysis
↓
Knowledge Extraction
↓
Storage
↓
Export
---
## AI Components
Recorder
Responsible only for recording.
No AI.
---
Transcription
Responsible only for speech-to-text.
No summarization.
---
Speaker Identification
Responsible only for identifying speakers.
Must not modify transcript text.
---
LLM Analysis
Responsible for:
- Summary
- Decisions
- Action Items
- Risks
- Questions
- Knowledge extraction
---
Storage
Stores:
Audio
Transcript
Metadata
AI Results
Knowledge Objects
---
## Architectural Rule
Every module has exactly one responsibility.
---
## Transcript Policy
Never modify the original transcript.
Corrections must create a new derived version.
---
## AI Output Policy
Every AI output should be reproducible.
Prompt version should be stored.
Model should be stored.
Timestamp should be stored.
---
## Long-Term Goal
Every meeting becomes searchable organizational knowledge.
No information should be lost after the meeting.
---
## Engineering Principles
The project follows an architecture-first development approach.
Before implementing a feature:
- define the domain model
- define module boundaries
- document architectural decisions
Implementation is intentionally delayed until the architecture is considered sufficiently stable.
+86
View File
@@ -0,0 +1,86 @@
# Meeting Knowledge Assistant
## Vision
The Meeting Knowledge Assistant (MKA) is an AI-assisted desktop application that records meetings, creates high-quality transcripts and transforms them into structured knowledge.
Unlike traditional meeting assistants, MKA does not aim to replace human interaction during meetings.
Its goal is to create an accurate, searchable and long-term knowledge base from spoken communication.
The project follows an "AI-first" architecture:
Audio
→ Transcription
→ Speaker Identification
→ AI Analysis
→ Structured Knowledge
→ Search
---
## Core Features
- Local audio recording
- System audio + microphone recording
- High quality transcription
- Speaker diarization
- AI-generated meeting minutes
- Executive Summary
- Action Items
- Decision Tracking
- Knowledge extraction
- Local database
- Full text search
- Export (Markdown, PDF, DOCX)
---
## Design Goals
- High transcription quality
- Modular architecture
- Offline-first whenever practical
- Replaceable AI components
- Vendor independence
- Long-term maintainability
---
## Philosophy
The transcript is the primary asset.
Everything else (summaries, reports, action items, knowledge extraction)
can always be regenerated using better AI models in the future.
Therefore:
Audio
→ Transcript
is considered immutable.
AI output is considered reproducible.
---
## Planned Architecture
Recorder
↓
Transcription
↓
Speaker Identification
↓
LLM Processing
↓
Knowledge Database
↓
Export
---
## Current Status
Project planning.
No implementation has started yet.
+101
View File
@@ -0,0 +1,101 @@
# Roadmap
## Phase 1
Minimum Viable Product
- Audio recording
- Whisper transcription
- Save transcript
- Markdown export
---
## Phase 2
Speaker diarization
- Speaker detection
- Speaker naming
- Timeline view
---
## Phase 3
Meeting intelligence
- Executive Summary
- Action Items
- Decisions
- Risks
- Open Questions
---
## Phase 4
Knowledge system
- SQLite database
- Full text search
- Semantic search
- Tags
- Projects
- Participants
---
## Phase 5
Desktop application
- Recording UI
- Transcript viewer
- Search
- Export
---
## Phase 6
Integrations
- Outlook
- Teams
- Google Calendar
- Local LLM
- OpenAI
- Ollama
---
## Future Ideas
- Live transcription
- Real-time summaries
- Company glossary
- Custom vocabulary
- Automatic project detection
- Meeting templates
- Voice identification
- Audio cleanup
- OCR for shared screens
- Timeline with bookmarks
- RAG knowledge integration
- Multi-language meetings
- Translation
- Automatic follow-up generation
BIN
View File
Binary file not shown.
+32
View File
@@ -0,0 +1,32 @@
# ADR 0001: Project Vision
## Status
Accepted
## Context
The project aims to create a local-first meeting assistant that transforms recorded meetings into high-quality transcripts and structured organizational knowledge.
Existing commercial meeting assistants often depend on bots, cloud services, vendor-specific workflows or opaque AI pipelines. This project is intended to provide more control over audio capture, transcription, knowledge extraction and long-term data ownership.
## Decision
The project will be developed as a modular Python application with a clear pipeline:
Recording
→ Transcription
→ Speaker Diarization
→ AI Analysis
→ Knowledge Extraction
→ Export
The transcript is treated as the canonical source. AI-generated summaries and knowledge objects are derived artifacts.
## Consequences
- Audio and transcript quality have priority over UI features.
- The system must preserve original recordings and original transcripts.
- AI outputs must be reproducible where practical.
- Modules should remain replaceable.
- The project should avoid unnecessary vendor lock-in.
+25
View File
@@ -0,0 +1,25 @@
# ADR 0002: Use faster-whisper for Transcription
## Status
Accepted
## Context
The project requires high-quality speech-to-text transcription for German and English meetings.
The transcription component should support local execution where practical and should remain replaceable in the future. The project should not depend on a meeting bot or on a proprietary meeting platform.
## Decision
The project will use `faster-whisper` as the initial transcription engine.
The preferred first model is `large-v3-turbo`, with `large-v3` as quality-oriented fallback if required.
## Consequences
- Transcription can run locally on suitable hardware.
- NVIDIA CUDA acceleration can be used later on an office AI PC.
- The transcription module must hide the concrete engine behind an internal interface.
- Model name, language setting, timestamp and engine version should be stored with every transcript.
- The decision can be revisited if another engine provides clearly better quality, speed or deployment characteristics.
+25
View File
@@ -0,0 +1,25 @@
# ADR 0003: Use pyannote for Speaker Diarization
## Status
Accepted
## Context
The project requires speaker diarization to assign transcript segments to speakers.
The first implementation should provide reliable separation of speakers in meetings, while keeping the diarization component replaceable.
## Decision
The project will use `pyannote.audio` as the initial speaker diarization engine.
Speaker diarization will be treated as a separate pipeline step after transcription.
## Consequences
- Diarization can be improved or replaced independently from transcription.
- Speaker labels are derived metadata and must not modify the canonical transcript.
- Human correction of speaker names should be supported later.
- The diarization module must hide the concrete engine behind an internal interface.
- Model name, engine version and timestamp should be stored with every diarization result.
@@ -0,0 +1,38 @@
# ADR 0004: Store Original Transcript as Canonical Source
## Status
Accepted
## Context
Speech-to-text systems continuously improve over time. Likewise, AI-generated summaries, action items and other derived artifacts may evolve as better models become available.
To ensure reproducibility, traceability and future reprocessing, the project requires a single immutable source of truth.
## Decision
The original transcript generated by the transcription engine is stored as the canonical transcript.
The canonical transcript is immutable and must never be modified.
A separate working transcript may be created for manual corrections such as:
- spelling corrections
- terminology normalization
- speaker name assignment
- punctuation improvements
The working transcript shall maintain a reference to the canonical transcript from which it was created.
All AI-generated artifacts (summaries, decisions, action items, knowledge objects, etc.) are derived from either the canonical transcript or a specific working transcript. The transcript version used must always be recorded.
## Consequences
- The original transcript remains permanently available.
- Future AI models can regenerate improved summaries without requiring a new recording.
- Manual corrections do not overwrite the canonical transcript.
- Traceability and reproducibility are preserved throughout the system.
- A working transcript enables manual improvements without compromising data integrity.
- Every derived artifact can reference the transcript version from which it was generated.
- Multiple working transcript revisions may coexist if required.
+42
View File
@@ -0,0 +1,42 @@
# ADR 0005: Meeting Domain Model
## Status
Accepted
## Context
The application manages meetings as structured collections of related artifacts rather than as individual files.
A consistent domain model is required to support future extensions such as semantic search, knowledge extraction, versioning and database storage.
## Decision
The central entity of the application is the Meeting.
A Meeting owns or references all artifacts created during its lifecycle.
The initial domain model consists of:
- Meeting
- Recording
- Transcript
- Working Transcript
- Speaker
- Participant
- AI Artifact
- Action Item
- Decision
- Knowledge Object
- Attachment
Each entity has a unique identifier.
Relationships between entities shall be maintained explicitly.
## Consequences
- New artifact types can be added without changing the overall architecture.
- Storage implementation (files, SQLite, PostgreSQL, etc.) remains independent from the domain model.
- AI outputs become first-class entities rather than generated text files.
- Future integrations can operate on domain objects instead of file names.
+180
View File
@@ -0,0 +1,180 @@
# Architecture
## Purpose
The Meeting Knowledge Assistant transforms recorded meetings into structured organizational knowledge through a modular processing pipeline.
---
# Design Principles
- Architecture first
- Offline first where practical
- Immutable source data
- Replaceable AI components
- Small, focused modules
- Explicit interfaces
- Reproducible AI outputs
---
# High-Level Pipeline
Recording
↓
Transcription
↓
Speaker Diarization
↓
Working Transcript
↓
LLM Analysis
↓
Knowledge Extraction
↓
Export
---
# Domain Model
Meeting
│
├── Recording
├── Transcript
│ ├── Canonical
│ └── Working
├── Speakers
├── Participants
├── AI Artifacts
├── Knowledge Objects
├── Attachments
└── Exports
---
# Module Responsibilities
## Recorder
Responsible only for capturing audio.
Input:
None
Output:
Recording
---
## Transcription
Responsible only for speech-to-text conversion.
Input:
Recording
Output:
Canonical Transcript
---
## Diarization
Responsible only for speaker identification.
Input:
Recording + Canonical Transcript
Output:
Working Transcript
---
## LLM
Responsible for semantic analysis.
Input:
Transcript
Output:
AI Artifacts
---
## Knowledge Extraction
Responsible for creating structured knowledge.
Input:
AI Artifacts
Output:
Knowledge Objects
---
## Export
Responsible for creating user-facing documents.
Input:
Knowledge Objects
Output:
Markdown
PDF
DOCX
---
# Artifact Lifecycle
Recording
↓
Canonical Transcript
↓
Working Transcript
↓
AI Artifacts
↓
Knowledge Objects
↓
Exports
---
# Storage Strategy
The domain model is independent of the storage backend.
Possible implementations:
- File System
- SQLite
- PostgreSQL
- Cloud Storage
---
# Future Extensions
- Live transcription
- Video processing
- OCR
- Semantic search
- Knowledge graph
- Company glossary
- Multi-language meetings
- Local LLM support
+20
View File
@@ -0,0 +1,20 @@
Transcript
Original speech-to-text output.
Diarization
Assignment of transcript segments to speakers.
Knowledge Object
Structured information extracted from meetings.
Action Item
A task assigned during a meeting.
Decision
A formally agreed conclusion.
Meeting Artifact
Any output generated from a meeting.
Canonical Transcript
The immutable original transcript.
View File
View File
View File
View File
View File
View File
View File
View File
View File
View File
View File