Reconcile Meeting Assistant architecture with Meeting Lab

This commit is contained in:
2026-08-23 20:19:47 +02:00
parent e028f1c82c
commit 490d557cb3
12 changed files with 462 additions and 605 deletions
+103 -153
View File
@@ -2,162 +2,112 @@
## Vision
The project creates a complete pipeline that transforms spoken meetings into structured organizational knowledge.
Meeting Assistant turns recorded meetings into reviewed, user-facing protocols
and, over time, searchable organizational knowledge. It is the application
layer over the reusable Meeting Lab processing backend.
The transcript is the canonical source.
The source audio and canonical transcript are preserved. Generated protocols
are derived, reproducible artifacts and require human review.
AI-generated summaries are reproducible artefacts.
## System Boundary
---
### Meeting Lab
Meeting Lab owns reusable and experimental processing:
- FFmpeg audio preparation
- `whisper.cpp` transcription
- optional `pyannote.audio` diarization
- protocol generation
- pipeline orchestration and progress events
Its reusable integration surface is the Python API:
- `MvpMeetingConfig`
- `MvpRunResult`
- `run_mvp_meeting(...)`
The CLI is a thin adapter, not the Meeting Assistant integration boundary.
### Meeting Assistant
Meeting Assistant owns user interaction and product workflow:
- audio-file selection
- structured Meeting Context editing
- participant management
- optional explicit speaker mapping
- pipeline launch and progress display
- protocol review and editing
- export and presentation of protocol versions
Processing logic must not be duplicated in the application.
## Validated MVP Pipeline
```text
Imported audio
-> prepare as mono, 16 kHz PCM WAV with FFmpeg
-> transcribe with whisper.cpp and large-v3-turbo
-> optionally diarize with pyannote.audio Community-1
-> generate a protocol directly from the full transcript
-> review and edit by a human
```
Mandatory fixed five-minute chunking is not part of the productive MVP path.
Chunking or segmentation used internally by an engine does not shape the
Meeting Assistant storage or domain model.
On the North reference machine, `whisper.cpp` with Vulkan on an AMD RX 9070
transcribed approximately 93.5 minutes in about 100 seconds. Pyannote through
PyTorch/ROCm processed the same duration in approximately 98.6 seconds. These
are validation observations, not performance guarantees or hardware
requirements. CPU execution remains supported and may be substantially slower.
## Meeting Context and Speakers
`MeetingContext` is a structured domain object containing meeting metadata and
participants. It can also contain optional explicit mappings from
`SPEAKER_XX` labels to `participant_id` values.
A speaker mapping is authoritative only when a user explicitly confirms it.
Speakers otherwise remain anonymous. Automatic speaker-name inference is not
allowed. The GUI must create and edit Meeting Context; hand-written YAML is not
a product requirement.
## Progress Contract
Meeting Lab emits stage-based progress events for:
- `preparing`
- `transcription`
- `diarization`
- `protocol_generation`
- `completed`
- `failed`
The GUI consumes this event interface. Percentages are shown only when real
measurable progress is available.
## Protocol Policy
Direct full-transcript protocol generation is the practical MVP direction.
`qwen3.8:27B` has shown strong readability and contextual synthesis;
`qwen3.6:35B-A3B` has shown somewhat more conservative behavior in some areas.
Experimental dual-model and diarization-assisted hard-fact extraction has not
demonstrated reliably better strict attribution accuracy and is not mandatory.
The product should ultimately offer both a detailed contextual protocol and a
shorter participant/distribution version. The short version remains follow-up
work if it is not available for the first GUI milestone.
## Design Principles
Offline-first where practical.
Replaceable AI components.
Vendor independence.
Modular architecture.
Small focused modules.
Simple interfaces.
---
## Core Pipeline
Recorder
↓
Transcription
↓
Speaker Identification
↓
LLM Analysis
↓
Knowledge Extraction
↓
Storage
↓
Export
---
## AI Components
Recorder
Responsible only for recording.
No AI.
---
Transcription
Responsible only for speech-to-text.
No summarization.
---
Speaker Identification
Responsible only for identifying speakers.
Must not modify transcript text.
---
LLM Analysis
Responsible for:
- Summary
- Decisions
- Action Items
- Risks
- Questions
- Knowledge extraction
---
Storage
Stores:
Audio
Transcript
Metadata
AI Results
Knowledge Objects
---
## Architectural Rule
Every module has exactly one responsibility.
---
## Transcript Policy
Never modify the original transcript.
Corrections must create a new derived version.
---
## AI Output Policy
Every AI output should be reproducible.
Prompt version should be stored.
Model should be stored.
Timestamp should be stored.
---
## Long-Term Goal
Every meeting becomes searchable organizational knowledge.
No information should be lost after the meeting.
---
## Engineering Principles
The project follows an architecture-first development approach.
Before implementing a feature:
- define the domain model
- define module boundaries
- document architectural decisions
Implementation is intentionally delayed until the architecture is considered sufficiently stable.
- offline-first where practical
- explicit application/backend boundaries
- replaceable AI components
- CPU-compatible operation with optional acceleration
- immutable source artifacts and traceable derived artifacts
- no automatic identity claims
- human review of generated protocols
- small modules and simple interfaces