Establish initial project architecture
This commit is contained in:
@@ -0,0 +1,111 @@
|
||||
# AGENTS.md
|
||||
|
||||
## Purpose
|
||||
|
||||
This document contains instructions for AI coding agents working on this repository.
|
||||
|
||||
---
|
||||
|
||||
## General Principles
|
||||
|
||||
Prefer clean architecture over quick fixes.
|
||||
|
||||
Prefer readability over cleverness.
|
||||
|
||||
Keep modules small.
|
||||
|
||||
Avoid unnecessary abstractions.
|
||||
|
||||
Minimize dependencies.
|
||||
|
||||
Document architectural decisions.
|
||||
|
||||
---
|
||||
|
||||
## Development Workflow
|
||||
|
||||
Never work directly on main.
|
||||
|
||||
Implement one feature per branch.
|
||||
|
||||
Prefer many small commits.
|
||||
|
||||
Each commit should compile.
|
||||
|
||||
Each feature should include:
|
||||
|
||||
- implementation
|
||||
- tests
|
||||
- documentation updates
|
||||
|
||||
---
|
||||
|
||||
## Coding Style
|
||||
|
||||
Follow PEP8.
|
||||
|
||||
Use type hints.
|
||||
|
||||
Prefer dataclasses where appropriate.
|
||||
|
||||
Avoid global state.
|
||||
|
||||
Prefer dependency injection.
|
||||
|
||||
No hidden side effects.
|
||||
|
||||
---
|
||||
|
||||
## AI Behaviour
|
||||
|
||||
If requirements are ambiguous:
|
||||
|
||||
STOP.
|
||||
|
||||
Ask questions.
|
||||
|
||||
Never invent requirements.
|
||||
|
||||
Never silently change behaviour.
|
||||
|
||||
---
|
||||
|
||||
## Refactoring
|
||||
|
||||
Refactor only inside the requested scope.
|
||||
|
||||
Avoid unrelated changes.
|
||||
|
||||
Do not rename files without justification.
|
||||
|
||||
---
|
||||
|
||||
## Documentation
|
||||
|
||||
Update README if user-visible behaviour changes.
|
||||
|
||||
Update PROJECT_KNOWLEDGE.md when technical knowledge changes.
|
||||
|
||||
Update ROADMAP.md only for future work.
|
||||
|
||||
---
|
||||
|
||||
## Priority Order
|
||||
|
||||
Correctness
|
||||
|
||||
↓
|
||||
|
||||
Maintainability
|
||||
|
||||
↓
|
||||
|
||||
Readability
|
||||
|
||||
↓
|
||||
|
||||
Performance
|
||||
|
||||
↓
|
||||
|
||||
Premature optimization
|
||||
@@ -0,0 +1,163 @@
|
||||
# Project Knowledge
|
||||
|
||||
## Vision
|
||||
|
||||
The project creates a complete pipeline that transforms spoken meetings into structured organizational knowledge.
|
||||
|
||||
The transcript is the canonical source.
|
||||
|
||||
AI-generated summaries are reproducible artefacts.
|
||||
|
||||
---
|
||||
|
||||
## Design Principles
|
||||
|
||||
Offline-first where practical.
|
||||
|
||||
Replaceable AI components.
|
||||
|
||||
Vendor independence.
|
||||
|
||||
Modular architecture.
|
||||
|
||||
Small focused modules.
|
||||
|
||||
Simple interfaces.
|
||||
|
||||
---
|
||||
|
||||
## Core Pipeline
|
||||
|
||||
Recorder
|
||||
|
||||
↓
|
||||
|
||||
Transcription
|
||||
|
||||
↓
|
||||
|
||||
Speaker Identification
|
||||
|
||||
↓
|
||||
|
||||
LLM Analysis
|
||||
|
||||
↓
|
||||
|
||||
Knowledge Extraction
|
||||
|
||||
↓
|
||||
|
||||
Storage
|
||||
|
||||
↓
|
||||
|
||||
Export
|
||||
|
||||
---
|
||||
|
||||
## AI Components
|
||||
|
||||
Recorder
|
||||
|
||||
Responsible only for recording.
|
||||
|
||||
No AI.
|
||||
|
||||
---
|
||||
|
||||
Transcription
|
||||
|
||||
Responsible only for speech-to-text.
|
||||
|
||||
No summarization.
|
||||
|
||||
---
|
||||
|
||||
Speaker Identification
|
||||
|
||||
Responsible only for identifying speakers.
|
||||
|
||||
Must not modify transcript text.
|
||||
|
||||
---
|
||||
|
||||
LLM Analysis
|
||||
|
||||
Responsible for:
|
||||
|
||||
- Summary
|
||||
|
||||
- Decisions
|
||||
|
||||
- Action Items
|
||||
|
||||
- Risks
|
||||
|
||||
- Questions
|
||||
|
||||
- Knowledge extraction
|
||||
|
||||
---
|
||||
|
||||
Storage
|
||||
|
||||
Stores:
|
||||
|
||||
Audio
|
||||
|
||||
Transcript
|
||||
|
||||
Metadata
|
||||
|
||||
AI Results
|
||||
|
||||
Knowledge Objects
|
||||
|
||||
---
|
||||
|
||||
## Architectural Rule
|
||||
|
||||
Every module has exactly one responsibility.
|
||||
|
||||
---
|
||||
|
||||
## Transcript Policy
|
||||
|
||||
Never modify the original transcript.
|
||||
|
||||
Corrections must create a new derived version.
|
||||
|
||||
---
|
||||
|
||||
## AI Output Policy
|
||||
|
||||
Every AI output should be reproducible.
|
||||
|
||||
Prompt version should be stored.
|
||||
|
||||
Model should be stored.
|
||||
|
||||
Timestamp should be stored.
|
||||
|
||||
---
|
||||
|
||||
## Long-Term Goal
|
||||
|
||||
Every meeting becomes searchable organizational knowledge.
|
||||
|
||||
No information should be lost after the meeting.
|
||||
|
||||
---
|
||||
|
||||
## Engineering Principles
|
||||
|
||||
The project follows an architecture-first development approach.
|
||||
|
||||
Before implementing a feature:
|
||||
|
||||
- define the domain model
|
||||
- define module boundaries
|
||||
- document architectural decisions
|
||||
|
||||
Implementation is intentionally delayed until the architecture is considered sufficiently stable.
|
||||
@@ -0,0 +1,86 @@
|
||||
# Meeting Knowledge Assistant
|
||||
|
||||
## Vision
|
||||
|
||||
The Meeting Knowledge Assistant (MKA) is an AI-assisted desktop application that records meetings, creates high-quality transcripts and transforms them into structured knowledge.
|
||||
|
||||
Unlike traditional meeting assistants, MKA does not aim to replace human interaction during meetings.
|
||||
Its goal is to create an accurate, searchable and long-term knowledge base from spoken communication.
|
||||
|
||||
The project follows an "AI-first" architecture:
|
||||
|
||||
Audio
|
||||
→ Transcription
|
||||
→ Speaker Identification
|
||||
→ AI Analysis
|
||||
→ Structured Knowledge
|
||||
→ Search
|
||||
|
||||
---
|
||||
|
||||
## Core Features
|
||||
|
||||
- Local audio recording
|
||||
- System audio + microphone recording
|
||||
- High quality transcription
|
||||
- Speaker diarization
|
||||
- AI-generated meeting minutes
|
||||
- Executive Summary
|
||||
- Action Items
|
||||
- Decision Tracking
|
||||
- Knowledge extraction
|
||||
- Local database
|
||||
- Full text search
|
||||
- Export (Markdown, PDF, DOCX)
|
||||
|
||||
---
|
||||
|
||||
## Design Goals
|
||||
|
||||
- High transcription quality
|
||||
- Modular architecture
|
||||
- Offline-first whenever practical
|
||||
- Replaceable AI components
|
||||
- Vendor independence
|
||||
- Long-term maintainability
|
||||
|
||||
---
|
||||
|
||||
## Philosophy
|
||||
|
||||
The transcript is the primary asset.
|
||||
|
||||
Everything else (summaries, reports, action items, knowledge extraction)
|
||||
can always be regenerated using better AI models in the future.
|
||||
|
||||
Therefore:
|
||||
|
||||
Audio
|
||||
→ Transcript
|
||||
is considered immutable.
|
||||
|
||||
AI output is considered reproducible.
|
||||
|
||||
---
|
||||
|
||||
## Planned Architecture
|
||||
|
||||
Recorder
|
||||
↓
|
||||
Transcription
|
||||
↓
|
||||
Speaker Identification
|
||||
↓
|
||||
LLM Processing
|
||||
↓
|
||||
Knowledge Database
|
||||
↓
|
||||
Export
|
||||
|
||||
---
|
||||
|
||||
## Current Status
|
||||
|
||||
Project planning.
|
||||
|
||||
No implementation has started yet.
|
||||
+101
@@ -0,0 +1,101 @@
|
||||
# Roadmap
|
||||
|
||||
## Phase 1
|
||||
|
||||
Minimum Viable Product
|
||||
|
||||
- Audio recording
|
||||
- Whisper transcription
|
||||
- Save transcript
|
||||
- Markdown export
|
||||
|
||||
---
|
||||
|
||||
## Phase 2
|
||||
|
||||
Speaker diarization
|
||||
|
||||
- Speaker detection
|
||||
- Speaker naming
|
||||
- Timeline view
|
||||
|
||||
---
|
||||
|
||||
## Phase 3
|
||||
|
||||
Meeting intelligence
|
||||
|
||||
- Executive Summary
|
||||
- Action Items
|
||||
- Decisions
|
||||
- Risks
|
||||
- Open Questions
|
||||
|
||||
---
|
||||
|
||||
## Phase 4
|
||||
|
||||
Knowledge system
|
||||
|
||||
- SQLite database
|
||||
- Full text search
|
||||
- Semantic search
|
||||
- Tags
|
||||
- Projects
|
||||
- Participants
|
||||
|
||||
---
|
||||
|
||||
## Phase 5
|
||||
|
||||
Desktop application
|
||||
|
||||
- Recording UI
|
||||
- Transcript viewer
|
||||
- Search
|
||||
- Export
|
||||
|
||||
---
|
||||
|
||||
## Phase 6
|
||||
|
||||
Integrations
|
||||
|
||||
- Outlook
|
||||
- Teams
|
||||
- Google Calendar
|
||||
- Local LLM
|
||||
- OpenAI
|
||||
- Ollama
|
||||
|
||||
---
|
||||
|
||||
## Future Ideas
|
||||
|
||||
- Live transcription
|
||||
|
||||
- Real-time summaries
|
||||
|
||||
- Company glossary
|
||||
|
||||
- Custom vocabulary
|
||||
|
||||
- Automatic project detection
|
||||
|
||||
- Meeting templates
|
||||
|
||||
- Voice identification
|
||||
|
||||
- Audio cleanup
|
||||
|
||||
- OCR for shared screens
|
||||
|
||||
- Timeline with bookmarks
|
||||
|
||||
- RAG knowledge integration
|
||||
|
||||
- Multi-language meetings
|
||||
|
||||
- Translation
|
||||
|
||||
- Automatic follow-up generation
|
||||
Vendored
BIN
Binary file not shown.
@@ -0,0 +1,32 @@
|
||||
# ADR 0001: Project Vision
|
||||
|
||||
## Status
|
||||
|
||||
Accepted
|
||||
|
||||
## Context
|
||||
|
||||
The project aims to create a local-first meeting assistant that transforms recorded meetings into high-quality transcripts and structured organizational knowledge.
|
||||
|
||||
Existing commercial meeting assistants often depend on bots, cloud services, vendor-specific workflows or opaque AI pipelines. This project is intended to provide more control over audio capture, transcription, knowledge extraction and long-term data ownership.
|
||||
|
||||
## Decision
|
||||
|
||||
The project will be developed as a modular Python application with a clear pipeline:
|
||||
|
||||
Recording
|
||||
→ Transcription
|
||||
→ Speaker Diarization
|
||||
→ AI Analysis
|
||||
→ Knowledge Extraction
|
||||
→ Export
|
||||
|
||||
The transcript is treated as the canonical source. AI-generated summaries and knowledge objects are derived artifacts.
|
||||
|
||||
## Consequences
|
||||
|
||||
- Audio and transcript quality have priority over UI features.
|
||||
- The system must preserve original recordings and original transcripts.
|
||||
- AI outputs must be reproducible where practical.
|
||||
- Modules should remain replaceable.
|
||||
- The project should avoid unnecessary vendor lock-in.
|
||||
@@ -0,0 +1,25 @@
|
||||
# ADR 0002: Use faster-whisper for Transcription
|
||||
|
||||
## Status
|
||||
|
||||
Accepted
|
||||
|
||||
## Context
|
||||
|
||||
The project requires high-quality speech-to-text transcription for German and English meetings.
|
||||
|
||||
The transcription component should support local execution where practical and should remain replaceable in the future. The project should not depend on a meeting bot or on a proprietary meeting platform.
|
||||
|
||||
## Decision
|
||||
|
||||
The project will use `faster-whisper` as the initial transcription engine.
|
||||
|
||||
The preferred first model is `large-v3-turbo`, with `large-v3` as quality-oriented fallback if required.
|
||||
|
||||
## Consequences
|
||||
|
||||
- Transcription can run locally on suitable hardware.
|
||||
- NVIDIA CUDA acceleration can be used later on an office AI PC.
|
||||
- The transcription module must hide the concrete engine behind an internal interface.
|
||||
- Model name, language setting, timestamp and engine version should be stored with every transcript.
|
||||
- The decision can be revisited if another engine provides clearly better quality, speed or deployment characteristics.
|
||||
@@ -0,0 +1,25 @@
|
||||
# ADR 0003: Use pyannote for Speaker Diarization
|
||||
|
||||
## Status
|
||||
|
||||
Accepted
|
||||
|
||||
## Context
|
||||
|
||||
The project requires speaker diarization to assign transcript segments to speakers.
|
||||
|
||||
The first implementation should provide reliable separation of speakers in meetings, while keeping the diarization component replaceable.
|
||||
|
||||
## Decision
|
||||
|
||||
The project will use `pyannote.audio` as the initial speaker diarization engine.
|
||||
|
||||
Speaker diarization will be treated as a separate pipeline step after transcription.
|
||||
|
||||
## Consequences
|
||||
|
||||
- Diarization can be improved or replaced independently from transcription.
|
||||
- Speaker labels are derived metadata and must not modify the canonical transcript.
|
||||
- Human correction of speaker names should be supported later.
|
||||
- The diarization module must hide the concrete engine behind an internal interface.
|
||||
- Model name, engine version and timestamp should be stored with every diarization result.
|
||||
@@ -0,0 +1,38 @@
|
||||
# ADR 0004: Store Original Transcript as Canonical Source
|
||||
|
||||
## Status
|
||||
|
||||
Accepted
|
||||
|
||||
## Context
|
||||
|
||||
Speech-to-text systems continuously improve over time. Likewise, AI-generated summaries, action items and other derived artifacts may evolve as better models become available.
|
||||
|
||||
To ensure reproducibility, traceability and future reprocessing, the project requires a single immutable source of truth.
|
||||
|
||||
## Decision
|
||||
|
||||
The original transcript generated by the transcription engine is stored as the canonical transcript.
|
||||
|
||||
The canonical transcript is immutable and must never be modified.
|
||||
|
||||
A separate working transcript may be created for manual corrections such as:
|
||||
|
||||
- spelling corrections
|
||||
- terminology normalization
|
||||
- speaker name assignment
|
||||
- punctuation improvements
|
||||
|
||||
The working transcript shall maintain a reference to the canonical transcript from which it was created.
|
||||
|
||||
All AI-generated artifacts (summaries, decisions, action items, knowledge objects, etc.) are derived from either the canonical transcript or a specific working transcript. The transcript version used must always be recorded.
|
||||
|
||||
## Consequences
|
||||
|
||||
- The original transcript remains permanently available.
|
||||
- Future AI models can regenerate improved summaries without requiring a new recording.
|
||||
- Manual corrections do not overwrite the canonical transcript.
|
||||
- Traceability and reproducibility are preserved throughout the system.
|
||||
- A working transcript enables manual improvements without compromising data integrity.
|
||||
- Every derived artifact can reference the transcript version from which it was generated.
|
||||
- Multiple working transcript revisions may coexist if required.
|
||||
@@ -0,0 +1,42 @@
|
||||
# ADR 0005: Meeting Domain Model
|
||||
|
||||
## Status
|
||||
|
||||
Accepted
|
||||
|
||||
## Context
|
||||
|
||||
The application manages meetings as structured collections of related artifacts rather than as individual files.
|
||||
|
||||
A consistent domain model is required to support future extensions such as semantic search, knowledge extraction, versioning and database storage.
|
||||
|
||||
## Decision
|
||||
|
||||
The central entity of the application is the Meeting.
|
||||
|
||||
A Meeting owns or references all artifacts created during its lifecycle.
|
||||
|
||||
The initial domain model consists of:
|
||||
|
||||
- Meeting
|
||||
- Recording
|
||||
- Transcript
|
||||
- Working Transcript
|
||||
- Speaker
|
||||
- Participant
|
||||
- AI Artifact
|
||||
- Action Item
|
||||
- Decision
|
||||
- Knowledge Object
|
||||
- Attachment
|
||||
|
||||
Each entity has a unique identifier.
|
||||
|
||||
Relationships between entities shall be maintained explicitly.
|
||||
|
||||
## Consequences
|
||||
|
||||
- New artifact types can be added without changing the overall architecture.
|
||||
- Storage implementation (files, SQLite, PostgreSQL, etc.) remains independent from the domain model.
|
||||
- AI outputs become first-class entities rather than generated text files.
|
||||
- Future integrations can operate on domain objects instead of file names.
|
||||
@@ -0,0 +1,180 @@
|
||||
# Architecture
|
||||
|
||||
## Purpose
|
||||
|
||||
The Meeting Knowledge Assistant transforms recorded meetings into structured organizational knowledge through a modular processing pipeline.
|
||||
|
||||
---
|
||||
|
||||
# Design Principles
|
||||
|
||||
- Architecture first
|
||||
- Offline first where practical
|
||||
- Immutable source data
|
||||
- Replaceable AI components
|
||||
- Small, focused modules
|
||||
- Explicit interfaces
|
||||
- Reproducible AI outputs
|
||||
|
||||
---
|
||||
|
||||
# High-Level Pipeline
|
||||
|
||||
Recording
|
||||
↓
|
||||
Transcription
|
||||
↓
|
||||
Speaker Diarization
|
||||
↓
|
||||
Working Transcript
|
||||
↓
|
||||
LLM Analysis
|
||||
↓
|
||||
Knowledge Extraction
|
||||
↓
|
||||
Export
|
||||
|
||||
---
|
||||
|
||||
# Domain Model
|
||||
|
||||
Meeting
|
||||
│
|
||||
├── Recording
|
||||
├── Transcript
|
||||
│ ├── Canonical
|
||||
│ └── Working
|
||||
├── Speakers
|
||||
├── Participants
|
||||
├── AI Artifacts
|
||||
├── Knowledge Objects
|
||||
├── Attachments
|
||||
└── Exports
|
||||
|
||||
---
|
||||
|
||||
# Module Responsibilities
|
||||
|
||||
## Recorder
|
||||
|
||||
Responsible only for capturing audio.
|
||||
|
||||
Input:
|
||||
None
|
||||
|
||||
Output:
|
||||
Recording
|
||||
|
||||
---
|
||||
|
||||
## Transcription
|
||||
|
||||
Responsible only for speech-to-text conversion.
|
||||
|
||||
Input:
|
||||
Recording
|
||||
|
||||
Output:
|
||||
Canonical Transcript
|
||||
|
||||
---
|
||||
|
||||
## Diarization
|
||||
|
||||
Responsible only for speaker identification.
|
||||
|
||||
Input:
|
||||
Recording + Canonical Transcript
|
||||
|
||||
Output:
|
||||
Working Transcript
|
||||
|
||||
---
|
||||
|
||||
## LLM
|
||||
|
||||
Responsible for semantic analysis.
|
||||
|
||||
Input:
|
||||
Transcript
|
||||
|
||||
Output:
|
||||
AI Artifacts
|
||||
|
||||
---
|
||||
|
||||
## Knowledge Extraction
|
||||
|
||||
Responsible for creating structured knowledge.
|
||||
|
||||
Input:
|
||||
AI Artifacts
|
||||
|
||||
Output:
|
||||
Knowledge Objects
|
||||
|
||||
---
|
||||
|
||||
## Export
|
||||
|
||||
Responsible for creating user-facing documents.
|
||||
|
||||
Input:
|
||||
Knowledge Objects
|
||||
|
||||
Output:
|
||||
Markdown
|
||||
PDF
|
||||
DOCX
|
||||
|
||||
---
|
||||
|
||||
# Artifact Lifecycle
|
||||
|
||||
Recording
|
||||
|
||||
↓
|
||||
|
||||
Canonical Transcript
|
||||
|
||||
↓
|
||||
|
||||
Working Transcript
|
||||
|
||||
↓
|
||||
|
||||
AI Artifacts
|
||||
|
||||
↓
|
||||
|
||||
Knowledge Objects
|
||||
|
||||
↓
|
||||
|
||||
Exports
|
||||
|
||||
---
|
||||
|
||||
# Storage Strategy
|
||||
|
||||
The domain model is independent of the storage backend.
|
||||
|
||||
Possible implementations:
|
||||
|
||||
- File System
|
||||
- SQLite
|
||||
- PostgreSQL
|
||||
- Cloud Storage
|
||||
|
||||
---
|
||||
|
||||
# Future Extensions
|
||||
|
||||
- Live transcription
|
||||
- Video processing
|
||||
- OCR
|
||||
- Semantic search
|
||||
- Knowledge graph
|
||||
- Company glossary
|
||||
- Multi-language meetings
|
||||
- Local LLM support
|
||||
@@ -0,0 +1,20 @@
|
||||
Transcript
|
||||
Original speech-to-text output.
|
||||
|
||||
Diarization
|
||||
Assignment of transcript segments to speakers.
|
||||
|
||||
Knowledge Object
|
||||
Structured information extracted from meetings.
|
||||
|
||||
Action Item
|
||||
A task assigned during a meeting.
|
||||
|
||||
Decision
|
||||
A formally agreed conclusion.
|
||||
|
||||
Meeting Artifact
|
||||
Any output generated from a meeting.
|
||||
|
||||
Canonical Transcript
|
||||
The immutable original transcript.
|
||||
Reference in New Issue
Block a user