commit bb8a1ccb081cb878a73420d4b955430ea26de507 Author: Martin Tazl Date: Wed Jul 8 15:57:35 2026 +0200 Establish initial project architecture diff --git a/.DS_Store b/.DS_Store new file mode 100644 index 0000000..232a029 Binary files /dev/null and b/.DS_Store differ diff --git a/.env.example b/.env.example new file mode 100644 index 0000000..e69de29 diff --git a/.gitignore b/.gitignore new file mode 100644 index 0000000..e69de29 diff --git a/AGENTS.md b/AGENTS.md new file mode 100644 index 0000000..c8853a9 --- /dev/null +++ b/AGENTS.md @@ -0,0 +1,111 @@ +# AGENTS.md + +## Purpose + +This document contains instructions for AI coding agents working on this repository. + +--- + +## General Principles + +Prefer clean architecture over quick fixes. + +Prefer readability over cleverness. + +Keep modules small. + +Avoid unnecessary abstractions. + +Minimize dependencies. + +Document architectural decisions. + +--- + +## Development Workflow + +Never work directly on main. + +Implement one feature per branch. + +Prefer many small commits. + +Each commit should compile. + +Each feature should include: + +- implementation +- tests +- documentation updates + +--- + +## Coding Style + +Follow PEP8. + +Use type hints. + +Prefer dataclasses where appropriate. + +Avoid global state. + +Prefer dependency injection. + +No hidden side effects. + +--- + +## AI Behaviour + +If requirements are ambiguous: + +STOP. + +Ask questions. + +Never invent requirements. + +Never silently change behaviour. + +--- + +## Refactoring + +Refactor only inside the requested scope. + +Avoid unrelated changes. + +Do not rename files without justification. + +--- + +## Documentation + +Update README if user-visible behaviour changes. + +Update PROJECT_KNOWLEDGE.md when technical knowledge changes. + +Update ROADMAP.md only for future work. + +--- + +## Priority Order + +Correctness + +↓ + +Maintainability + +↓ + +Readability + +↓ + +Performance + +↓ + +Premature optimization diff --git a/CHANGELOG.md b/CHANGELOG.md new file mode 100644 index 0000000..e69de29 diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md new file mode 100644 index 0000000..e69de29 diff --git a/PROJECT_KNOWLEDGE.md b/PROJECT_KNOWLEDGE.md new file mode 100644 index 0000000..dddb205 --- /dev/null +++ b/PROJECT_KNOWLEDGE.md @@ -0,0 +1,163 @@ +# Project Knowledge + +## Vision + +The project creates a complete pipeline that transforms spoken meetings into structured organizational knowledge. + +The transcript is the canonical source. + +AI-generated summaries are reproducible artefacts. + +--- + +## Design Principles + +Offline-first where practical. + +Replaceable AI components. + +Vendor independence. + +Modular architecture. + +Small focused modules. + +Simple interfaces. + +--- + +## Core Pipeline + +Recorder + +↓ + +Transcription + +↓ + +Speaker Identification + +↓ + +LLM Analysis + +↓ + +Knowledge Extraction + +↓ + +Storage + +↓ + +Export + +--- + +## AI Components + +Recorder + +Responsible only for recording. + +No AI. + +--- + +Transcription + +Responsible only for speech-to-text. + +No summarization. + +--- + +Speaker Identification + +Responsible only for identifying speakers. + +Must not modify transcript text. + +--- + +LLM Analysis + +Responsible for: + +- Summary + +- Decisions + +- Action Items + +- Risks + +- Questions + +- Knowledge extraction + +--- + +Storage + +Stores: + +Audio + +Transcript + +Metadata + +AI Results + +Knowledge Objects + +--- + +## Architectural Rule + +Every module has exactly one responsibility. + +--- + +## Transcript Policy + +Never modify the original transcript. + +Corrections must create a new derived version. + +--- + +## AI Output Policy + +Every AI output should be reproducible. + +Prompt version should be stored. + +Model should be stored. + +Timestamp should be stored. + +--- + +## Long-Term Goal + +Every meeting becomes searchable organizational knowledge. + +No information should be lost after the meeting. + +--- + +## Engineering Principles + +The project follows an architecture-first development approach. + +Before implementing a feature: + +- define the domain model +- define module boundaries +- document architectural decisions + +Implementation is intentionally delayed until the architecture is considered sufficiently stable. \ No newline at end of file diff --git a/README.md b/README.md new file mode 100644 index 0000000..0476f23 --- /dev/null +++ b/README.md @@ -0,0 +1,86 @@ +# Meeting Knowledge Assistant + +## Vision + +The Meeting Knowledge Assistant (MKA) is an AI-assisted desktop application that records meetings, creates high-quality transcripts and transforms them into structured knowledge. + +Unlike traditional meeting assistants, MKA does not aim to replace human interaction during meetings. +Its goal is to create an accurate, searchable and long-term knowledge base from spoken communication. + +The project follows an "AI-first" architecture: + +Audio +→ Transcription +→ Speaker Identification +→ AI Analysis +→ Structured Knowledge +→ Search + +--- + +## Core Features + +- Local audio recording +- System audio + microphone recording +- High quality transcription +- Speaker diarization +- AI-generated meeting minutes +- Executive Summary +- Action Items +- Decision Tracking +- Knowledge extraction +- Local database +- Full text search +- Export (Markdown, PDF, DOCX) + +--- + +## Design Goals + +- High transcription quality +- Modular architecture +- Offline-first whenever practical +- Replaceable AI components +- Vendor independence +- Long-term maintainability + +--- + +## Philosophy + +The transcript is the primary asset. + +Everything else (summaries, reports, action items, knowledge extraction) +can always be regenerated using better AI models in the future. + +Therefore: + +Audio +→ Transcript +is considered immutable. + +AI output is considered reproducible. + +--- + +## Planned Architecture + +Recorder +↓ +Transcription +↓ +Speaker Identification +↓ +LLM Processing +↓ +Knowledge Database +↓ +Export + +--- + +## Current Status + +Project planning. + +No implementation has started yet. diff --git a/ROADMAP.md b/ROADMAP.md new file mode 100644 index 0000000..882ce5f --- /dev/null +++ b/ROADMAP.md @@ -0,0 +1,101 @@ +# Roadmap + +## Phase 1 + +Minimum Viable Product + +- Audio recording +- Whisper transcription +- Save transcript +- Markdown export + +--- + +## Phase 2 + +Speaker diarization + +- Speaker detection +- Speaker naming +- Timeline view + +--- + +## Phase 3 + +Meeting intelligence + +- Executive Summary +- Action Items +- Decisions +- Risks +- Open Questions + +--- + +## Phase 4 + +Knowledge system + +- SQLite database +- Full text search +- Semantic search +- Tags +- Projects +- Participants + +--- + +## Phase 5 + +Desktop application + +- Recording UI +- Transcript viewer +- Search +- Export + +--- + +## Phase 6 + +Integrations + +- Outlook +- Teams +- Google Calendar +- Local LLM +- OpenAI +- Ollama + +--- + +## Future Ideas + +- Live transcription + +- Real-time summaries + +- Company glossary + +- Custom vocabulary + +- Automatic project detection + +- Meeting templates + +- Voice identification + +- Audio cleanup + +- OCR for shared screens + +- Timeline with bookmarks + +- RAG knowledge integration + +- Multi-language meetings + +- Translation + +- Automatic follow-up generation diff --git a/docs/.DS_Store b/docs/.DS_Store new file mode 100644 index 0000000..adaae90 Binary files /dev/null and b/docs/.DS_Store differ diff --git a/docs/adr/0001-project-vision.md b/docs/adr/0001-project-vision.md new file mode 100644 index 0000000..cdc6d81 --- /dev/null +++ b/docs/adr/0001-project-vision.md @@ -0,0 +1,32 @@ +# ADR 0001: Project Vision + +## Status + +Accepted + +## Context + +The project aims to create a local-first meeting assistant that transforms recorded meetings into high-quality transcripts and structured organizational knowledge. + +Existing commercial meeting assistants often depend on bots, cloud services, vendor-specific workflows or opaque AI pipelines. This project is intended to provide more control over audio capture, transcription, knowledge extraction and long-term data ownership. + +## Decision + +The project will be developed as a modular Python application with a clear pipeline: + +Recording +→ Transcription +→ Speaker Diarization +→ AI Analysis +→ Knowledge Extraction +→ Export + +The transcript is treated as the canonical source. AI-generated summaries and knowledge objects are derived artifacts. + +## Consequences + +- Audio and transcript quality have priority over UI features. +- The system must preserve original recordings and original transcripts. +- AI outputs must be reproducible where practical. +- Modules should remain replaceable. +- The project should avoid unnecessary vendor lock-in. \ No newline at end of file diff --git a/docs/adr/0002-use-faster-whisper.md b/docs/adr/0002-use-faster-whisper.md new file mode 100644 index 0000000..30dab4d --- /dev/null +++ b/docs/adr/0002-use-faster-whisper.md @@ -0,0 +1,25 @@ +# ADR 0002: Use faster-whisper for Transcription + +## Status + +Accepted + +## Context + +The project requires high-quality speech-to-text transcription for German and English meetings. + +The transcription component should support local execution where practical and should remain replaceable in the future. The project should not depend on a meeting bot or on a proprietary meeting platform. + +## Decision + +The project will use `faster-whisper` as the initial transcription engine. + +The preferred first model is `large-v3-turbo`, with `large-v3` as quality-oriented fallback if required. + +## Consequences + +- Transcription can run locally on suitable hardware. +- NVIDIA CUDA acceleration can be used later on an office AI PC. +- The transcription module must hide the concrete engine behind an internal interface. +- Model name, language setting, timestamp and engine version should be stored with every transcript. +- The decision can be revisited if another engine provides clearly better quality, speed or deployment characteristics. diff --git a/docs/adr/0003-use-pyannote.md b/docs/adr/0003-use-pyannote.md new file mode 100644 index 0000000..f72e62d --- /dev/null +++ b/docs/adr/0003-use-pyannote.md @@ -0,0 +1,25 @@ +# ADR 0003: Use pyannote for Speaker Diarization + +## Status + +Accepted + +## Context + +The project requires speaker diarization to assign transcript segments to speakers. + +The first implementation should provide reliable separation of speakers in meetings, while keeping the diarization component replaceable. + +## Decision + +The project will use `pyannote.audio` as the initial speaker diarization engine. + +Speaker diarization will be treated as a separate pipeline step after transcription. + +## Consequences + +- Diarization can be improved or replaced independently from transcription. +- Speaker labels are derived metadata and must not modify the canonical transcript. +- Human correction of speaker names should be supported later. +- The diarization module must hide the concrete engine behind an internal interface. +- Model name, engine version and timestamp should be stored with every diarization result. diff --git a/docs/adr/0004-store-original-transcript.md b/docs/adr/0004-store-original-transcript.md new file mode 100644 index 0000000..c1698c9 --- /dev/null +++ b/docs/adr/0004-store-original-transcript.md @@ -0,0 +1,38 @@ +# ADR 0004: Store Original Transcript as Canonical Source + +## Status + +Accepted + +## Context + +Speech-to-text systems continuously improve over time. Likewise, AI-generated summaries, action items and other derived artifacts may evolve as better models become available. + +To ensure reproducibility, traceability and future reprocessing, the project requires a single immutable source of truth. + +## Decision + +The original transcript generated by the transcription engine is stored as the canonical transcript. + +The canonical transcript is immutable and must never be modified. + +A separate working transcript may be created for manual corrections such as: + +- spelling corrections +- terminology normalization +- speaker name assignment +- punctuation improvements + +The working transcript shall maintain a reference to the canonical transcript from which it was created. + +All AI-generated artifacts (summaries, decisions, action items, knowledge objects, etc.) are derived from either the canonical transcript or a specific working transcript. The transcript version used must always be recorded. + +## Consequences + +- The original transcript remains permanently available. +- Future AI models can regenerate improved summaries without requiring a new recording. +- Manual corrections do not overwrite the canonical transcript. +- Traceability and reproducibility are preserved throughout the system. +- A working transcript enables manual improvements without compromising data integrity. +- Every derived artifact can reference the transcript version from which it was generated. +- Multiple working transcript revisions may coexist if required. diff --git a/docs/adr/0005-knowledge-model.md b/docs/adr/0005-knowledge-model.md new file mode 100644 index 0000000..d354f0b --- /dev/null +++ b/docs/adr/0005-knowledge-model.md @@ -0,0 +1,42 @@ +# ADR 0005: Meeting Domain Model + +## Status + +Accepted + +## Context + +The application manages meetings as structured collections of related artifacts rather than as individual files. + +A consistent domain model is required to support future extensions such as semantic search, knowledge extraction, versioning and database storage. + +## Decision + +The central entity of the application is the Meeting. + +A Meeting owns or references all artifacts created during its lifecycle. + +The initial domain model consists of: + +- Meeting +- Recording +- Transcript +- Working Transcript +- Speaker +- Participant +- AI Artifact +- Action Item +- Decision +- Knowledge Object +- Attachment + +Each entity has a unique identifier. + +Relationships between entities shall be maintained explicitly. + +## Consequences + +- New artifact types can be added without changing the overall architecture. +- Storage implementation (files, SQLite, PostgreSQL, etc.) remains independent from the domain model. +- AI outputs become first-class entities rather than generated text files. +- Future integrations can operate on domain objects instead of file names. \ No newline at end of file diff --git a/docs/architecture.md b/docs/architecture.md new file mode 100644 index 0000000..5779ed2 --- /dev/null +++ b/docs/architecture.md @@ -0,0 +1,180 @@ +# Architecture + +## Purpose + +The Meeting Knowledge Assistant transforms recorded meetings into structured organizational knowledge through a modular processing pipeline. + +--- + +# Design Principles + +- Architecture first +- Offline first where practical +- Immutable source data +- Replaceable AI components +- Small, focused modules +- Explicit interfaces +- Reproducible AI outputs + +--- + +# High-Level Pipeline + +Recording + ↓ +Transcription + ↓ +Speaker Diarization + ↓ +Working Transcript + ↓ +LLM Analysis + ↓ +Knowledge Extraction + ↓ +Export + +--- + +# Domain Model + +Meeting +│ +├── Recording +├── Transcript +│ ├── Canonical +│ └── Working +├── Speakers +├── Participants +├── AI Artifacts +├── Knowledge Objects +├── Attachments +└── Exports + +--- + +# Module Responsibilities + +## Recorder + +Responsible only for capturing audio. + +Input: +None + +Output: +Recording + +--- + +## Transcription + +Responsible only for speech-to-text conversion. + +Input: +Recording + +Output: +Canonical Transcript + +--- + +## Diarization + +Responsible only for speaker identification. + +Input: +Recording + Canonical Transcript + +Output: +Working Transcript + +--- + +## LLM + +Responsible for semantic analysis. + +Input: +Transcript + +Output: +AI Artifacts + +--- + +## Knowledge Extraction + +Responsible for creating structured knowledge. + +Input: +AI Artifacts + +Output: +Knowledge Objects + +--- + +## Export + +Responsible for creating user-facing documents. + +Input: +Knowledge Objects + +Output: +Markdown +PDF +DOCX + +--- + +# Artifact Lifecycle + +Recording + +↓ + +Canonical Transcript + +↓ + +Working Transcript + +↓ + +AI Artifacts + +↓ + +Knowledge Objects + +↓ + +Exports + +--- + +# Storage Strategy + +The domain model is independent of the storage backend. + +Possible implementations: + +- File System +- SQLite +- PostgreSQL +- Cloud Storage + +--- + +# Future Extensions + +- Live transcription +- Video processing +- OCR +- Semantic search +- Knowledge graph +- Company glossary +- Multi-language meetings +- Local LLM support \ No newline at end of file diff --git a/docs/glossary.md b/docs/glossary.md new file mode 100644 index 0000000..1ba9bef --- /dev/null +++ b/docs/glossary.md @@ -0,0 +1,20 @@ +Transcript +Original speech-to-text output. + +Diarization +Assignment of transcript segments to speakers. + +Knowledge Object +Structured information extracted from meetings. + +Action Item +A task assigned during a meeting. + +Decision +A formally agreed conclusion. + +Meeting Artifact +Any output generated from a meeting. + +Canonical Transcript +The immutable original transcript. diff --git a/docs/prompts.md b/docs/prompts.md new file mode 100644 index 0000000..e69de29 diff --git a/pyproject.toml b/pyproject.toml new file mode 100644 index 0000000..e69de29 diff --git a/src/mka/__init__.py b/src/mka/__init__.py new file mode 100644 index 0000000..e69de29 diff --git a/src/mka/diarization/__init__.py b/src/mka/diarization/__init__.py new file mode 100644 index 0000000..e69de29 diff --git a/src/mka/export/__init__.py b/src/mka/export/__init__.py new file mode 100644 index 0000000..e69de29 diff --git a/src/mka/llm/__init__.py b/src/mka/llm/__init__.py new file mode 100644 index 0000000..e69de29 diff --git a/src/mka/plugins/__init__.py b/src/mka/plugins/__init__.py new file mode 100644 index 0000000..e69de29 diff --git a/src/mka/recorder/__init__.py b/src/mka/recorder/__init__.py new file mode 100644 index 0000000..e69de29 diff --git a/src/mka/storage/__init__.py b/src/mka/storage/__init__.py new file mode 100644 index 0000000..e69de29 diff --git a/src/mka/transcription/__init__.py b/src/mka/transcription/__init__.py new file mode 100644 index 0000000..e69de29 diff --git a/src/mka/ui/__init__.py b/src/mka/ui/__init__.py new file mode 100644 index 0000000..e69de29