Files
meeting-assistant/docs/architecture.md

5.7 KiB

Architecture

Purpose

Meeting Assistant is the user-facing application for preparing meeting context, running the Meeting Lab pipeline and reviewing its results. Meeting Lab is the reusable processing backend and experimental engine.

System Boundary

Meeting Assistant                         Meeting Lab
-----------------                         -----------
Audio selection                  ------>  Audio preparation
Meeting Context editor           ------>  Transcription
Participant management           ------>  Optional diarization
Explicit speaker mapping         ------>  Protocol generation
Progress presentation            <------  Progress events
Protocol editor and export       <------  Run result and artifacts

Meeting Assistant calls the Meeting Lab Python API directly. The Meeting Lab CLI is a thin adapter over that API and must not be launched as an application subprocess.

MVP Processing Flow

Source audio
    -> FFmpeg preparation (normalization optional, default on)
    -> mono, 16 kHz PCM WAV
    -> whisper.cpp transcription with large-v3-turbo
    -> optional pyannote.audio Community-1 diarization
    -> direct full-transcript protocol generation
    -> human review and editing
    -> export

The MVP does not require fixed five-minute audio chunks, semantic chunking, a separate evidence-extraction pipeline or multiple LLMs. Any internal segmentation remains a Meeting Lab implementation detail.

Backend Integration

The reusable Meeting Lab interface consists of:

  • MvpMeetingConfig, the run configuration
  • run_mvp_meeting(...), the orchestration entry point
  • MvpRunResult, the completed run result
  • stage-based progress events

The known stages are:

preparing
transcription
diarization
protocol_generation
completed
failed

The GUI shows the current stage. It shows a percentage only when the event contains real measurable progress; stage changes must not be presented as invented percentages.

Meeting Context

MeetingContext is structured domain input rather than an informal prompt or a file users must author manually. It contains meeting metadata and participants and can include optional mappings:

SPEAKER_XX -> participant_id

Mappings are authoritative only after explicit user confirmation. Diarization labels otherwise remain anonymous, and the application must not infer speaker names automatically. The Meeting Assistant GUI owns creation and editing of this context and its mappings.

Processing Components

Audio Preparation

Meeting Lab uses FFmpeg to prepare a consistent local-processing input. The current practical target is mono, 16 kHz PCM WAV. The imported source remains a separate source artifact. Preparation always runs for WAV, FLAC and M4A, regardless of the normalization switch. When enabled, Meeting Lab currently uses loudnorm=I=-16:LRA=11:TP=-1.5, an isolated conservative default for speech recordings that may be revisited after empirical comparison. Meeting Assistant passes only an on/off choice and does not own filter parameters.

Transcription

whisper.cpp with large-v3-turbo is the preferred local backend. Vulkan is a validated acceleration path on North's AMD RX 9070, while CPU execution remains a compatibility fallback. Engine and model details should be retained with the result for traceability.

Diarization

pyannote.audio Community-1 is preferred when diarization is enabled. It can run on CPU and may use GPU acceleration such as PyTorch/ROCm where available. Diarization is optional and its anonymous labels do not modify the canonical transcript or assert participant identity.

Protocol Generation

The direct full-transcript path is the practical MVP. Current model experiments favor qwen3.8:27B for readability and contextual synthesis, while qwen3.6:35B-A3B has shown more conservative behavior in some areas. Neither a specific dual-model arrangement nor an evidence pipeline is an application requirement. Generated protocols require human review.

Application Responsibilities

The first GUI milestone provides:

  • audio-file selection
  • structured Meeting Context and participant editing
  • optional explicit speaker mapping
  • pipeline start and configuration
  • progress and failure presentation
  • detailed protocol display and editing
  • export of the reviewed result

The product direction also includes a shorter participant/distribution protocol. It may be delivered after the first GUI milestone.

Artifact and Storage Principles

The existing Meeting domain remains the application-level container for source recordings, canonical and working transcripts, context, generated artifacts and exports. Storage remains independent of processing segmentation. In particular, chunks are not primary Meeting Assistant domain objects or the required unit of MVP persistence.

Source audio and the canonical transcription output are preserved. Corrections and reviewed protocols are derived versions. Generated artifacts should retain their input version, backend/model configuration, prompt version and timestamp where practical.

Later Post-Run Correction Flow

Post-run corrections are planned as a separate, non-mandatory workflow after initial protocol generation. As detailed in ADR 0012, confirmed name, person and anonymous-speaker corrections should become meeting-specific knowledge. They may then be applied deterministically to safe derived artifacts or used for protocol-only regeneration without needlessly rerunning transcription or diarization.

Future Extensions

  • shorter distribution protocols
  • integrated recording
  • searchable meeting history and optional database indexes
  • live transcription
  • OCR and video processing
  • semantic search and organizational knowledge features