Reconcile Meeting Assistant architecture with Meeting Lab
This commit is contained in:
+103
-153
@@ -2,162 +2,112 @@
|
||||
|
||||
## Vision
|
||||
|
||||
The project creates a complete pipeline that transforms spoken meetings into structured organizational knowledge.
|
||||
Meeting Assistant turns recorded meetings into reviewed, user-facing protocols
|
||||
and, over time, searchable organizational knowledge. It is the application
|
||||
layer over the reusable Meeting Lab processing backend.
|
||||
|
||||
The transcript is the canonical source.
|
||||
The source audio and canonical transcript are preserved. Generated protocols
|
||||
are derived, reproducible artifacts and require human review.
|
||||
|
||||
AI-generated summaries are reproducible artefacts.
|
||||
## System Boundary
|
||||
|
||||
---
|
||||
### Meeting Lab
|
||||
|
||||
Meeting Lab owns reusable and experimental processing:
|
||||
|
||||
- FFmpeg audio preparation
|
||||
- `whisper.cpp` transcription
|
||||
- optional `pyannote.audio` diarization
|
||||
- protocol generation
|
||||
- pipeline orchestration and progress events
|
||||
|
||||
Its reusable integration surface is the Python API:
|
||||
|
||||
- `MvpMeetingConfig`
|
||||
- `MvpRunResult`
|
||||
- `run_mvp_meeting(...)`
|
||||
|
||||
The CLI is a thin adapter, not the Meeting Assistant integration boundary.
|
||||
|
||||
### Meeting Assistant
|
||||
|
||||
Meeting Assistant owns user interaction and product workflow:
|
||||
|
||||
- audio-file selection
|
||||
- structured Meeting Context editing
|
||||
- participant management
|
||||
- optional explicit speaker mapping
|
||||
- pipeline launch and progress display
|
||||
- protocol review and editing
|
||||
- export and presentation of protocol versions
|
||||
|
||||
Processing logic must not be duplicated in the application.
|
||||
|
||||
## Validated MVP Pipeline
|
||||
|
||||
```text
|
||||
Imported audio
|
||||
-> prepare as mono, 16 kHz PCM WAV with FFmpeg
|
||||
-> transcribe with whisper.cpp and large-v3-turbo
|
||||
-> optionally diarize with pyannote.audio Community-1
|
||||
-> generate a protocol directly from the full transcript
|
||||
-> review and edit by a human
|
||||
```
|
||||
|
||||
Mandatory fixed five-minute chunking is not part of the productive MVP path.
|
||||
Chunking or segmentation used internally by an engine does not shape the
|
||||
Meeting Assistant storage or domain model.
|
||||
|
||||
On the North reference machine, `whisper.cpp` with Vulkan on an AMD RX 9070
|
||||
transcribed approximately 93.5 minutes in about 100 seconds. Pyannote through
|
||||
PyTorch/ROCm processed the same duration in approximately 98.6 seconds. These
|
||||
are validation observations, not performance guarantees or hardware
|
||||
requirements. CPU execution remains supported and may be substantially slower.
|
||||
|
||||
## Meeting Context and Speakers
|
||||
|
||||
`MeetingContext` is a structured domain object containing meeting metadata and
|
||||
participants. It can also contain optional explicit mappings from
|
||||
`SPEAKER_XX` labels to `participant_id` values.
|
||||
|
||||
A speaker mapping is authoritative only when a user explicitly confirms it.
|
||||
Speakers otherwise remain anonymous. Automatic speaker-name inference is not
|
||||
allowed. The GUI must create and edit Meeting Context; hand-written YAML is not
|
||||
a product requirement.
|
||||
|
||||
## Progress Contract
|
||||
|
||||
Meeting Lab emits stage-based progress events for:
|
||||
|
||||
- `preparing`
|
||||
- `transcription`
|
||||
- `diarization`
|
||||
- `protocol_generation`
|
||||
- `completed`
|
||||
- `failed`
|
||||
|
||||
The GUI consumes this event interface. Percentages are shown only when real
|
||||
measurable progress is available.
|
||||
|
||||
## Protocol Policy
|
||||
|
||||
Direct full-transcript protocol generation is the practical MVP direction.
|
||||
`qwen3.8:27B` has shown strong readability and contextual synthesis;
|
||||
`qwen3.6:35B-A3B` has shown somewhat more conservative behavior in some areas.
|
||||
Experimental dual-model and diarization-assisted hard-fact extraction has not
|
||||
demonstrated reliably better strict attribution accuracy and is not mandatory.
|
||||
|
||||
The product should ultimately offer both a detailed contextual protocol and a
|
||||
shorter participant/distribution version. The short version remains follow-up
|
||||
work if it is not available for the first GUI milestone.
|
||||
|
||||
## Design Principles
|
||||
|
||||
Offline-first where practical.
|
||||
|
||||
Replaceable AI components.
|
||||
|
||||
Vendor independence.
|
||||
|
||||
Modular architecture.
|
||||
|
||||
Small focused modules.
|
||||
|
||||
Simple interfaces.
|
||||
|
||||
---
|
||||
|
||||
## Core Pipeline
|
||||
|
||||
Recorder
|
||||
|
||||
↓
|
||||
|
||||
Transcription
|
||||
|
||||
↓
|
||||
|
||||
Speaker Identification
|
||||
|
||||
↓
|
||||
|
||||
LLM Analysis
|
||||
|
||||
↓
|
||||
|
||||
Knowledge Extraction
|
||||
|
||||
↓
|
||||
|
||||
Storage
|
||||
|
||||
↓
|
||||
|
||||
Export
|
||||
|
||||
---
|
||||
|
||||
## AI Components
|
||||
|
||||
Recorder
|
||||
|
||||
Responsible only for recording.
|
||||
|
||||
No AI.
|
||||
|
||||
---
|
||||
|
||||
Transcription
|
||||
|
||||
Responsible only for speech-to-text.
|
||||
|
||||
No summarization.
|
||||
|
||||
---
|
||||
|
||||
Speaker Identification
|
||||
|
||||
Responsible only for identifying speakers.
|
||||
|
||||
Must not modify transcript text.
|
||||
|
||||
---
|
||||
|
||||
LLM Analysis
|
||||
|
||||
Responsible for:
|
||||
|
||||
- Summary
|
||||
|
||||
- Decisions
|
||||
|
||||
- Action Items
|
||||
|
||||
- Risks
|
||||
|
||||
- Questions
|
||||
|
||||
- Knowledge extraction
|
||||
|
||||
---
|
||||
|
||||
Storage
|
||||
|
||||
Stores:
|
||||
|
||||
Audio
|
||||
|
||||
Transcript
|
||||
|
||||
Metadata
|
||||
|
||||
AI Results
|
||||
|
||||
Knowledge Objects
|
||||
|
||||
---
|
||||
|
||||
## Architectural Rule
|
||||
|
||||
Every module has exactly one responsibility.
|
||||
|
||||
---
|
||||
|
||||
## Transcript Policy
|
||||
|
||||
Never modify the original transcript.
|
||||
|
||||
Corrections must create a new derived version.
|
||||
|
||||
---
|
||||
|
||||
## AI Output Policy
|
||||
|
||||
Every AI output should be reproducible.
|
||||
|
||||
Prompt version should be stored.
|
||||
|
||||
Model should be stored.
|
||||
|
||||
Timestamp should be stored.
|
||||
|
||||
---
|
||||
|
||||
## Long-Term Goal
|
||||
|
||||
Every meeting becomes searchable organizational knowledge.
|
||||
|
||||
No information should be lost after the meeting.
|
||||
|
||||
---
|
||||
|
||||
## Engineering Principles
|
||||
|
||||
The project follows an architecture-first development approach.
|
||||
|
||||
Before implementing a feature:
|
||||
|
||||
- define the domain model
|
||||
- define module boundaries
|
||||
- document architectural decisions
|
||||
|
||||
Implementation is intentionally delayed until the architecture is considered sufficiently stable.
|
||||
- offline-first where practical
|
||||
- explicit application/backend boundaries
|
||||
- replaceable AI components
|
||||
- CPU-compatible operation with optional acceleration
|
||||
- immutable source artifacts and traceable derived artifacts
|
||||
- no automatic identity claims
|
||||
- human review of generated protocols
|
||||
- small modules and simple interfaces
|
||||
|
||||
Reference in New Issue
Block a user