# Architecture ## Purpose The Meeting Knowledge Assistant transforms recorded meetings into structured organizational knowledge through a modular processing pipeline. --- # Design Principles - Architecture first - Offline first where practical - Immutable source data - Replaceable AI components - Small, focused modules - Explicit interfaces - Reproducible AI outputs - External recording and internal processing are separate responsibilities. - The processing pipeline must not depend on OBS-specific metadata or behavior. - Imported recordings must be treated like recordings from any other supported source. --- ## High-Level Pipeline External Recorder ↓ Audio Import ↓ Recording Validation ↓ Meeting Storage ↓ Transcription ↓ Speaker Diarization ↓ Working Transcript ↓ LLM Analysis ↓ Knowledge Extraction ↓ Export --- ## Core Pipeline External Recording ↓ Audio Import and Validation ↓ Transcription ↓ Speaker Diarization ↓ Working Transcript ↓ LLM Analysis ↓ Knowledge Extraction ↓ Storage and Export # Domain Model Meeting │ ├── Recording ├── Transcript │ ├── Canonical │ └── Working ├── Speakers ├── Participants ├── AI Artifacts ├── Knowledge Objects ├── Attachments └── Exports --- # Module Responsibilities ## Input and Ingest Responsible for importing existing recordings into the application. Initial reference source: - OBS Studio Initial supported formats: - WAV - FLAC Responsibilities: - validate the input file - collect technical metadata - calculate checksums - copy or move the recording into the meeting directory - create the initial Recording domain object The ingest component must not perform transcription or modify the audio content. ## Recorder The recorder module is reserved for a future integrated recording implementation. It is not required for the MVP. The initial application workflow uses externally created recordings, with OBS Studio as the recommended reference recorder. --- ## Transcription Responsible only for speech-to-text conversion. Input: Recording Output: Canonical Transcript --- ## Diarization Responsible only for speaker identification. Input: Recording + Canonical Transcript Output: Working Transcript --- ## LLM Responsible for semantic analysis. Input: Transcript Output: AI Artifacts --- ## Knowledge Extraction Responsible for creating structured knowledge. Input: AI Artifacts Output: Knowledge Objects --- ## Export Responsible for creating user-facing documents. Input: Knowledge Objects Output: Markdown PDF DOCX --- ## Artifact Lifecycle ```text Imported Recording ↓ Validated Source Recording ↓ Canonical Transcript ↓ Working Transcript ↓ AI Artifacts ↓ Knowledge Objects ↓ Exports --- ## Initial Recording Strategy The MVP does not implement platform-specific audio capture. OBS Studio is the recommended reference recorder for online meetings. The application initially processes existing WAV or FLAC recordings. Integrated recording remains a future extension and must not be required by transcription, diarization or analysis modules. --- # Storage Strategy The domain model is independent of the storage backend. Possible implementations: - File System - SQLite - PostgreSQL - Cloud Storage --- # Future Extensions - Live transcription - Video processing - OCR - Semantic search - Knowledge graph - Company glossary - Multi-language meetings - Local LLM support