Reconcile Meeting Assistant architecture with Meeting Lab
This commit is contained in:
+125
-233
@@ -2,254 +2,146 @@
|
||||
|
||||
## Purpose
|
||||
|
||||
The Meeting Knowledge Assistant transforms recorded meetings into structured organizational knowledge through a modular processing pipeline.
|
||||
Meeting Assistant is the user-facing application for preparing meeting
|
||||
context, running the Meeting Lab pipeline and reviewing its results. Meeting
|
||||
Lab is the reusable processing backend and experimental engine.
|
||||
|
||||
---
|
||||
|
||||
# Design Principles
|
||||
|
||||
- Architecture first
|
||||
- Offline first where practical
|
||||
- Immutable source data
|
||||
- Replaceable AI components
|
||||
- Small, focused modules
|
||||
- Explicit interfaces
|
||||
- Reproducible AI outputs
|
||||
- External recording and internal processing are separate responsibilities.
|
||||
- The processing pipeline must not depend on OBS-specific metadata or behavior.
|
||||
- Imported recordings must be treated like recordings from any other supported source.
|
||||
|
||||
---
|
||||
|
||||
## High-Level Pipeline
|
||||
|
||||
External Recorder
|
||||
↓
|
||||
Audio Import
|
||||
↓
|
||||
Recording Validation
|
||||
↓
|
||||
Meeting Storage
|
||||
↓
|
||||
Transcription
|
||||
↓
|
||||
Speaker Diarization
|
||||
↓
|
||||
Working Transcript
|
||||
↓
|
||||
LLM Analysis
|
||||
↓
|
||||
Knowledge Extraction
|
||||
↓
|
||||
Export
|
||||
|
||||
---
|
||||
|
||||
## Core Pipeline
|
||||
|
||||
External Recording
|
||||
|
||||
↓
|
||||
|
||||
Audio Import and Validation
|
||||
|
||||
↓
|
||||
|
||||
Transcription
|
||||
|
||||
↓
|
||||
|
||||
Speaker Diarization
|
||||
|
||||
↓
|
||||
|
||||
Working Transcript
|
||||
|
||||
↓
|
||||
|
||||
LLM Analysis
|
||||
|
||||
↓
|
||||
|
||||
Knowledge Extraction
|
||||
|
||||
↓
|
||||
|
||||
Storage and Export
|
||||
|
||||
# Domain Model
|
||||
|
||||
Meeting
|
||||
│
|
||||
├── Recording
|
||||
├── Transcript
|
||||
│ ├── Canonical
|
||||
│ └── Working
|
||||
├── Speakers
|
||||
├── Participants
|
||||
├── AI Artifacts
|
||||
├── Knowledge Objects
|
||||
├── Attachments
|
||||
└── Exports
|
||||
|
||||
---
|
||||
|
||||
# Module Responsibilities
|
||||
|
||||
## Input and Ingest
|
||||
|
||||
Responsible for importing existing recordings into the application.
|
||||
|
||||
Initial reference source:
|
||||
|
||||
- OBS Studio
|
||||
|
||||
Initial supported formats:
|
||||
|
||||
- WAV
|
||||
|
||||
- FLAC
|
||||
|
||||
Responsibilities:
|
||||
|
||||
- validate the input file
|
||||
|
||||
- collect technical metadata
|
||||
|
||||
- calculate checksums
|
||||
|
||||
- copy or move the recording into the meeting directory
|
||||
|
||||
- create the initial Recording domain object
|
||||
|
||||
The ingest component must not perform transcription or modify the audio content.
|
||||
|
||||
## Recorder
|
||||
|
||||
The recorder module is reserved for a future integrated recording implementation.
|
||||
|
||||
It is not required for the MVP.
|
||||
|
||||
The initial application workflow uses externally created recordings, with OBS
|
||||
Studio as the recommended reference recorder.
|
||||
|
||||
---
|
||||
|
||||
## Transcription
|
||||
|
||||
Responsible only for speech-to-text conversion.
|
||||
|
||||
Input:
|
||||
Recording
|
||||
|
||||
Output:
|
||||
Canonical Transcript
|
||||
|
||||
---
|
||||
|
||||
## Diarization
|
||||
|
||||
Responsible only for speaker identification.
|
||||
|
||||
Input:
|
||||
Recording + Canonical Transcript
|
||||
|
||||
Output:
|
||||
Working Transcript
|
||||
|
||||
---
|
||||
|
||||
## LLM
|
||||
|
||||
Responsible for semantic analysis.
|
||||
|
||||
Input:
|
||||
Transcript
|
||||
|
||||
Output:
|
||||
AI Artifacts
|
||||
|
||||
---
|
||||
|
||||
## Knowledge Extraction
|
||||
|
||||
Responsible for creating structured knowledge.
|
||||
|
||||
Input:
|
||||
AI Artifacts
|
||||
|
||||
Output:
|
||||
Knowledge Objects
|
||||
|
||||
---
|
||||
|
||||
## Export
|
||||
|
||||
Responsible for creating user-facing documents.
|
||||
|
||||
Input:
|
||||
Knowledge Objects
|
||||
|
||||
Output:
|
||||
Markdown
|
||||
PDF
|
||||
DOCX
|
||||
|
||||
---
|
||||
|
||||
## Artifact Lifecycle
|
||||
## System Boundary
|
||||
|
||||
```text
|
||||
Imported Recording
|
||||
↓
|
||||
Validated Source Recording
|
||||
↓
|
||||
Canonical Transcript
|
||||
↓
|
||||
Working Transcript
|
||||
↓
|
||||
AI Artifacts
|
||||
↓
|
||||
Knowledge Objects
|
||||
↓
|
||||
Exports
|
||||
Meeting Assistant Meeting Lab
|
||||
----------------- -----------
|
||||
Audio selection ------> Audio preparation
|
||||
Meeting Context editor ------> Transcription
|
||||
Participant management ------> Optional diarization
|
||||
Explicit speaker mapping ------> Protocol generation
|
||||
Progress presentation <------ Progress events
|
||||
Protocol editor and export <------ Run result and artifacts
|
||||
```
|
||||
|
||||
---
|
||||
Meeting Assistant calls the Meeting Lab Python API directly. The Meeting Lab
|
||||
CLI is a thin adapter over that API and must not be launched as an application
|
||||
subprocess.
|
||||
|
||||
## Initial Recording Strategy
|
||||
## MVP Processing Flow
|
||||
|
||||
The MVP does not implement platform-specific audio capture.
|
||||
```text
|
||||
Source audio
|
||||
-> FFmpeg normalization/preparation
|
||||
-> mono, 16 kHz PCM WAV
|
||||
-> whisper.cpp transcription with large-v3-turbo
|
||||
-> optional pyannote.audio Community-1 diarization
|
||||
-> direct full-transcript protocol generation
|
||||
-> human review and editing
|
||||
-> export
|
||||
```
|
||||
|
||||
OBS Studio is the recommended reference recorder for online meetings.
|
||||
The MVP does not require fixed five-minute audio chunks, semantic chunking, a
|
||||
separate evidence-extraction pipeline or multiple LLMs. Any internal
|
||||
segmentation remains a Meeting Lab implementation detail.
|
||||
|
||||
The application initially processes existing WAV or FLAC recordings. Integrated
|
||||
## Backend Integration
|
||||
|
||||
recording remains a future extension and must not be required by transcription,
|
||||
The reusable Meeting Lab interface consists of:
|
||||
|
||||
diarization or analysis modules.
|
||||
- `MvpMeetingConfig`, the run configuration
|
||||
- `run_mvp_meeting(...)`, the orchestration entry point
|
||||
- `MvpRunResult`, the completed run result
|
||||
- stage-based progress events
|
||||
|
||||
---
|
||||
The known stages are:
|
||||
|
||||
# Storage Strategy
|
||||
```text
|
||||
preparing
|
||||
transcription
|
||||
diarization
|
||||
protocol_generation
|
||||
completed
|
||||
failed
|
||||
```
|
||||
|
||||
The domain model is independent of the storage backend.
|
||||
The GUI shows the current stage. It shows a percentage only when the event
|
||||
contains real measurable progress; stage changes must not be presented as
|
||||
invented percentages.
|
||||
|
||||
Possible implementations:
|
||||
## Meeting Context
|
||||
|
||||
- File System
|
||||
- SQLite
|
||||
- PostgreSQL
|
||||
- Cloud Storage
|
||||
`MeetingContext` is structured domain input rather than an informal prompt or a
|
||||
file users must author manually. It contains meeting metadata and participants
|
||||
and can include optional mappings:
|
||||
|
||||
---
|
||||
```text
|
||||
SPEAKER_XX -> participant_id
|
||||
```
|
||||
|
||||
# Future Extensions
|
||||
Mappings are authoritative only after explicit user confirmation. Diarization
|
||||
labels otherwise remain anonymous, and the application must not infer speaker
|
||||
names automatically. The Meeting Assistant GUI owns creation and editing of
|
||||
this context and its mappings.
|
||||
|
||||
- Live transcription
|
||||
- Video processing
|
||||
- OCR
|
||||
- Semantic search
|
||||
- Knowledge graph
|
||||
- Company glossary
|
||||
- Multi-language meetings
|
||||
- Local LLM support
|
||||
## Processing Components
|
||||
|
||||
### Audio Preparation
|
||||
|
||||
Meeting Lab uses FFmpeg to prepare a consistent local-processing input. The
|
||||
current practical target is mono, 16 kHz PCM WAV. The imported source remains a
|
||||
separate source artifact.
|
||||
|
||||
### Transcription
|
||||
|
||||
`whisper.cpp` with `large-v3-turbo` is the preferred local backend. Vulkan is a
|
||||
validated acceleration path on North's AMD RX 9070, while CPU execution remains
|
||||
a compatibility fallback. Engine and model details should be retained with the
|
||||
result for traceability.
|
||||
|
||||
### Diarization
|
||||
|
||||
`pyannote.audio` Community-1 is preferred when diarization is enabled. It can
|
||||
run on CPU and may use GPU acceleration such as PyTorch/ROCm where available.
|
||||
Diarization is optional and its anonymous labels do not modify the canonical
|
||||
transcript or assert participant identity.
|
||||
|
||||
### Protocol Generation
|
||||
|
||||
The direct full-transcript path is the practical MVP. Current model experiments
|
||||
favor `qwen3.8:27B` for readability and contextual synthesis, while
|
||||
`qwen3.6:35B-A3B` has shown more conservative behavior in some areas. Neither a
|
||||
specific dual-model arrangement nor an evidence pipeline is an application
|
||||
requirement. Generated protocols require human review.
|
||||
|
||||
## Application Responsibilities
|
||||
|
||||
The first GUI milestone provides:
|
||||
|
||||
- audio-file selection
|
||||
- structured Meeting Context and participant editing
|
||||
- optional explicit speaker mapping
|
||||
- pipeline start and configuration
|
||||
- progress and failure presentation
|
||||
- detailed protocol display and editing
|
||||
- export of the reviewed result
|
||||
|
||||
The product direction also includes a shorter participant/distribution
|
||||
protocol. It may be delivered after the first GUI milestone.
|
||||
|
||||
## Artifact and Storage Principles
|
||||
|
||||
The existing Meeting domain remains the application-level container for source
|
||||
recordings, canonical and working transcripts, context, generated artifacts and
|
||||
exports. Storage remains independent of processing segmentation. In particular,
|
||||
chunks are not primary Meeting Assistant domain objects or the required unit of
|
||||
MVP persistence.
|
||||
|
||||
Source audio and the canonical transcription output are preserved. Corrections
|
||||
and reviewed protocols are derived versions. Generated artifacts should retain
|
||||
their input version, backend/model configuration, prompt version and timestamp
|
||||
where practical.
|
||||
|
||||
## Future Extensions
|
||||
|
||||
- shorter distribution protocols
|
||||
- integrated recording
|
||||
- searchable meeting history and optional database indexes
|
||||
- live transcription
|
||||
- OCR and video processing
|
||||
- semantic search and organizational knowledge features
|
||||
|
||||
Reference in New Issue
Block a user