Reconcile Meeting Assistant architecture with Meeting Lab

This commit is contained in:
2026-08-23 20:19:47 +02:00
parent e028f1c82c
commit 490d557cb3
12 changed files with 462 additions and 605 deletions
+47 -108
View File
@@ -1,126 +1,65 @@
# Meeting Knowledge Assistant
# Meeting Assistant
## Vision
Meeting Assistant is the user-facing application for turning recorded meetings
into reviewed, distributable protocols. It uses Meeting Lab as its reusable
local processing backend.
The Meeting Knowledge Assistant (MKA) is an AI-assisted desktop application that records meetings, creates high-quality transcripts and transforms them into structured knowledge.
## Current MVP Direction
Unlike traditional meeting assistants, MKA does not aim to replace human interaction during meetings.
Its goal is to create an accurate, searchable and long-term knowledge base from spoken communication.
The first product milestone is a desktop GUI that lets a user:
The project follows an "AI-first" architecture:
- select an existing audio recording
- create and edit structured meeting metadata and participants
- optionally map anonymous `SPEAKER_XX` labels to known participants
- start the Meeting Lab processing pipeline
- follow stage-based progress
- review and edit the generated protocol
- export the reviewed result
Audio
→ Transcription
→ Speaker Identification
→ AI Analysis
→ Structured Knowledge
→ Search
The MVP does not include integrated recording. OBS Studio remains a recommended
reference recorder, but imported audio is not tied to OBS-specific behavior.
---
## Core Features
- Import of existing meeting recordings
- OBS Studio as the recommended reference recorder
- High-quality local transcription
- Speaker diarization
- Editable working transcripts
- AI-generated meeting minutes
- Executive summaries
- Action items
- Decision tracking
- Knowledge extraction
- Local file-based storage
- Export to Markdown, PDF and DOCX
---
## Recording Workflow
The MVP does not include its own audio recorder.
OBS Studio is the recommended reference recorder for online meetings. It can
capture system audio and microphone audio without joining the meeting as a bot.
Initial workflow:
## Processing Pipeline
```text
Audio file
-> FFmpeg preparation (mono, 16 kHz PCM WAV)
-> whisper.cpp transcription (large-v3-turbo)
-> optional pyannote.audio Community-1 diarization
-> direct full-transcript protocol generation
-> human review and export
```
Online Meeting
Meeting Assistant calls Meeting Lab's reusable Python API
(`MvpMeetingConfig`, `MvpRunResult` and `run_mvp_meeting(...)`). It does not run
the Meeting Lab CLI as a subprocess. Fixed five-minute audio chunking is not a
required part of this workflow.
↓
GPU acceleration is supported where available, but CPU execution remains a
compatibility path. Diarization is optional and produces anonymous speaker
labels. A label identifies a participant only when the user explicitly
confirms the mapping; automatic speaker-name inference is not allowed.
OBS Studio
## Product Outputs
↓
The product direction includes:
WAV Recording
- a detailed, contextual protocol for review and continued work
- a shorter version suitable for participants and distribution
↓
The detailed direct-protocol flow is the practical MVP direction. The shorter
version may be implemented after the initial GUI. Human review remains part of
the workflow for all generated protocols.
Meeting Assistant Import
## Project Boundary
↓
Meeting Lab owns reusable audio preparation, transcription, diarization,
orchestration and protocol-generation logic. Meeting Assistant owns the GUI,
context and participant editing, explicit speaker mapping, progress display,
protocol editing and export.
Transcription
The code currently contains the initial Meeting Assistant project and domain
foundation. The GUI has not yet been implemented.
↓
Speaker Diarization
↓
Analysis and Export
---
## Design Goals
- High transcription quality
- Modular architecture
- Offline-first whenever practical
- Replaceable AI components
- Vendor independence
- Long-term maintainability
---
## Philosophy
The transcript is the primary asset.
Everything else (summaries, reports, action items, knowledge extraction)
can always be regenerated using better AI models in the future.
Therefore:
Audio
→ Transcript
is considered immutable.
AI output is considered reproducible.
---
## Planned Architecture
Recorder
↓
Transcription
↓
Speaker Identification
↓
LLM Processing
↓
Knowledge Database
↓
Export
---
## Current Status
Project planning.
No implementation has started yet.
See [Architecture](docs/architecture.md), [Project Knowledge](PROJECT_KNOWLEDGE.md),
[Roadmap](ROADMAP.md) and [ADR 0011](docs/adr/0011-use-meeting-lab-mvp-backend.md).