161 lines
5.7 KiB
Markdown
161 lines
5.7 KiB
Markdown
# Architecture
|
|
|
|
## Purpose
|
|
|
|
Meeting Assistant is the user-facing application for preparing meeting
|
|
context, running the Meeting Lab pipeline and reviewing its results. Meeting
|
|
Lab is the reusable processing backend and experimental engine.
|
|
|
|
## System Boundary
|
|
|
|
```text
|
|
Meeting Assistant Meeting Lab
|
|
----------------- -----------
|
|
Audio selection ------> Audio preparation
|
|
Meeting Context editor ------> Transcription
|
|
Participant management ------> Optional diarization
|
|
Explicit speaker mapping ------> Protocol generation
|
|
Progress presentation <------ Progress events
|
|
Protocol editor and export <------ Run result and artifacts
|
|
```
|
|
|
|
Meeting Assistant calls the Meeting Lab Python API directly. The Meeting Lab
|
|
CLI is a thin adapter over that API and must not be launched as an application
|
|
subprocess.
|
|
|
|
## MVP Processing Flow
|
|
|
|
```text
|
|
Source audio
|
|
-> FFmpeg preparation (normalization optional, default on)
|
|
-> mono, 16 kHz PCM WAV
|
|
-> whisper.cpp transcription with large-v3-turbo
|
|
-> optional pyannote.audio Community-1 diarization
|
|
-> direct full-transcript protocol generation
|
|
-> human review and editing
|
|
-> export
|
|
```
|
|
|
|
The MVP does not require fixed five-minute audio chunks, semantic chunking, a
|
|
separate evidence-extraction pipeline or multiple LLMs. Any internal
|
|
segmentation remains a Meeting Lab implementation detail.
|
|
|
|
## Backend Integration
|
|
|
|
The reusable Meeting Lab interface consists of:
|
|
|
|
- `MvpMeetingConfig`, the run configuration
|
|
- `run_mvp_meeting(...)`, the orchestration entry point
|
|
- `MvpRunResult`, the completed run result
|
|
- stage-based progress events
|
|
|
|
The known stages are:
|
|
|
|
```text
|
|
preparing
|
|
transcription
|
|
diarization
|
|
protocol_generation
|
|
completed
|
|
failed
|
|
```
|
|
|
|
The GUI shows the current stage. It shows a percentage only when the event
|
|
contains real measurable progress; stage changes must not be presented as
|
|
invented percentages.
|
|
|
|
## Meeting Context
|
|
|
|
`MeetingContext` is structured domain input rather than an informal prompt or a
|
|
file users must author manually. It contains meeting metadata and participants
|
|
and can include optional mappings:
|
|
|
|
```text
|
|
SPEAKER_XX -> participant_id
|
|
```
|
|
|
|
Mappings are authoritative only after explicit user confirmation. Diarization
|
|
labels otherwise remain anonymous, and the application must not infer speaker
|
|
names automatically. The Meeting Assistant GUI owns creation and editing of
|
|
this context and its mappings.
|
|
|
|
## Processing Components
|
|
|
|
### Audio Preparation
|
|
|
|
Meeting Lab uses FFmpeg to prepare a consistent local-processing input. The
|
|
current practical target is mono, 16 kHz PCM WAV. The imported source remains a
|
|
separate source artifact. Preparation always runs for WAV, FLAC and M4A,
|
|
regardless of the normalization switch. When enabled, Meeting Lab currently
|
|
uses `loudnorm=I=-16:LRA=11:TP=-1.5`, an isolated conservative default for
|
|
speech recordings that may be revisited after empirical comparison. Meeting
|
|
Assistant passes only an on/off choice and does not own filter parameters.
|
|
|
|
### Transcription
|
|
|
|
`whisper.cpp` with `large-v3-turbo` is the preferred local backend. Vulkan is a
|
|
validated acceleration path on North's AMD RX 9070, while CPU execution remains
|
|
a compatibility fallback. Engine and model details should be retained with the
|
|
result for traceability.
|
|
|
|
### Diarization
|
|
|
|
`pyannote.audio` Community-1 is preferred when diarization is enabled. It can
|
|
run on CPU and may use GPU acceleration such as PyTorch/ROCm where available.
|
|
Diarization is optional and its anonymous labels do not modify the canonical
|
|
transcript or assert participant identity.
|
|
|
|
### Protocol Generation
|
|
|
|
The direct full-transcript path is the practical MVP. Current model experiments
|
|
favor `qwen3.8:27B` for readability and contextual synthesis, while
|
|
`qwen3.6:35B-A3B` has shown more conservative behavior in some areas. Neither a
|
|
specific dual-model arrangement nor an evidence pipeline is an application
|
|
requirement. Generated protocols require human review.
|
|
|
|
## Application Responsibilities
|
|
|
|
The first GUI milestone provides:
|
|
|
|
- audio-file selection
|
|
- structured Meeting Context and participant editing
|
|
- optional explicit speaker mapping
|
|
- pipeline start and configuration
|
|
- progress and failure presentation
|
|
- detailed protocol display and editing
|
|
- export of the reviewed result
|
|
|
|
The product direction also includes a shorter participant/distribution
|
|
protocol. It may be delivered after the first GUI milestone.
|
|
|
|
## Artifact and Storage Principles
|
|
|
|
The existing Meeting domain remains the application-level container for source
|
|
recordings, canonical and working transcripts, context, generated artifacts and
|
|
exports. Storage remains independent of processing segmentation. In particular,
|
|
chunks are not primary Meeting Assistant domain objects or the required unit of
|
|
MVP persistence.
|
|
|
|
Source audio and the canonical transcription output are preserved. Corrections
|
|
and reviewed protocols are derived versions. Generated artifacts should retain
|
|
their input version, backend/model configuration, prompt version and timestamp
|
|
where practical.
|
|
|
|
## Later Post-Run Correction Flow
|
|
|
|
Post-run corrections are planned as a separate, non-mandatory workflow after
|
|
initial protocol generation. As detailed in [ADR 0012](adr/0012-post-run-corrections.md),
|
|
confirmed name, person and anonymous-speaker corrections should become
|
|
meeting-specific knowledge. They may then be applied deterministically to safe
|
|
derived artifacts or used for protocol-only regeneration without needlessly
|
|
rerunning transcription or diarization.
|
|
|
|
## Future Extensions
|
|
|
|
- shorter distribution protocols
|
|
- integrated recording
|
|
- searchable meeting history and optional database indexes
|
|
- live transcription
|
|
- OCR and video processing
|
|
- semantic search and organizational knowledge features
|