Reconcile Meeting Assistant architecture with Meeting Lab

This commit is contained in:
2026-08-23 20:19:47 +02:00
parent e028f1c82c
commit 490d557cb3
12 changed files with 462 additions and 605 deletions
+7 -2
View File
@@ -2,7 +2,7 @@
## Status
Accepted
Accepted, amended by [ADR 0011](0011-use-meeting-lab-mvp-backend.md)
## Context
@@ -29,4 +29,9 @@ The transcript is treated as the canonical source. AI-generated summaries and kn
- The system must preserve original recordings and original transcripts.
- AI outputs must be reproducible where practical.
- Modules should remain replaceable.
- The project should avoid unnecessary vendor lock-in.
- The project should avoid unnecessary vendor lock-in.
ADR 0011 narrows the practical MVP to a user-facing application over the
reusable Meeting Lab backend. Direct protocol generation is the validated MVP
path; a separate knowledge-extraction stage remains a possible future
extension rather than an MVP requirement.
+1 -1
View File
@@ -2,7 +2,7 @@
## Status
Accepted
Superseded by [ADR 0011](0011-use-meeting-lab-mvp-backend.md)
## Context
+9 -1
View File
@@ -2,7 +2,7 @@
## Status
Accepted
Accepted, amended by [ADR 0011](0011-use-meeting-lab-mvp-backend.md)
## Context
@@ -23,3 +23,11 @@ Speaker diarization will be treated as a separate pipeline step after transcript
- Human correction of speaker names should be supported later.
- The diarization module must hide the concrete engine behind an internal interface.
- Model name, engine version and timestamp should be stored with every diarization result.
## Amendment
ADR 0011 makes diarization optional for the MVP and selects `pyannote.audio`
Community-1 as the current preferred backend. CPU execution remains supported,
with GPU acceleration used when available. Speaker labels remain anonymous
unless a user explicitly confirms a `SPEAKER_XX` to participant mapping in the
Meeting Context. Automatic speaker-name inference is not allowed.
+3 -3
View File
@@ -2,7 +2,7 @@
## Status
Accepted
Accepted, amended by [ADR 0011](0011-use-meeting-lab-mvp-backend.md)
## Context
@@ -18,8 +18,8 @@ removing temporary source files without risking data loss.
Recordings shall initially be captured in a processing-friendly lossless
format, normally WAV.
After transcription and diarization have completed, the recording may be
converted to an archive format.
After transcription and, when enabled, diarization have completed, the
recording may be converted to an archive format.
The preferred archive formats are:
@@ -2,7 +2,7 @@
## Status
Accepted
Accepted, amended by [ADR 0011](0011-use-meeting-lab-mvp-backend.md)
## Context
@@ -13,8 +13,9 @@ OBS Studio already provides stable, configurable and cross-platform audio
capture. Reimplementing audio capture during the initial project phase would
add substantial complexity without directly improving transcription quality.
The primary value of the Meeting Assistant lies in transcription, speaker
diarization, analysis, structured knowledge extraction and export.
The primary value of the Meeting Assistant lies in transcription, optional
speaker diarization, protocol generation, review and export. ADR 0011 defines
the validated processing boundary and current MVP protocol path.
## Decision
@@ -0,0 +1,94 @@
# ADR 0011: Use Meeting Lab as the MVP Processing Backend
## Status
Accepted
Supersedes ADR 0002 and amends ADRs 0001 and 0003.
## Context
Meeting Lab experiments have validated a practical local pipeline for turning
a recorded meeting into an editable protocol. Earlier Meeting Assistant plans
selected `faster-whisper`, treated diarization as required and described a
separate knowledge-extraction stage as part of the primary pipeline. Work on an
old feature branch also proposed mandatory fixed five-minute audio chunks and
chunk-oriented storage.
The validated backend now has a reusable Python API, optional diarization and a
direct full-transcript protocol path. The product boundary must keep backend
processing in Meeting Lab and user interaction in Meeting Assistant.
## Decision
Meeting Lab is the reusable processing backend and experimental engine.
Meeting Assistant is the user-facing application and calls Meeting Lab through
its Python API rather than launching its CLI as a subprocess.
The MVP processing flow is:
```text
Source audio
-> FFmpeg preparation (mono, 16 kHz PCM WAV)
-> whisper.cpp transcription (large-v3-turbo)
-> optional pyannote.audio Community-1 diarization
-> direct full-transcript protocol generation
-> human review and editing
```
The productive path does not require fixed five-minute audio chunks. Internal
streaming or segmentation remains an implementation detail of Meeting Lab and
must not define Meeting Assistant's domain or storage model.
Meeting Assistant integrates with these public Meeting Lab interfaces:
- `MvpMeetingConfig`
- `MvpRunResult`
- `run_mvp_meeting(...)`
- stage-based progress events: `preparing`, `transcription`, `diarization`,
`protocol_generation`, `completed` and `failed`
Progress percentages are displayed only when the backend reports real
measurable progress.
`MeetingContext` is the structured input containing meeting metadata and
participants. It may contain explicitly confirmed `SPEAKER_XX -> participant_id`
mappings. Such mappings are authoritative only when confirmed by the user.
The system must not infer speaker names automatically.
The preferred local engines are:
- `whisper.cpp` with `large-v3-turbo` for transcription
- `pyannote.audio` Community-1 for optional diarization
GPU acceleration is optional. Vulkan transcription and PyTorch/ROCm
diarization are validated acceleration paths on North's AMD RX 9070. CPU
execution remains the compatibility fallback.
The direct full-transcript protocol path is the practical MVP direction.
`qwen3.8:27B` has produced strong readability and contextual synthesis, while
`qwen3.6:35B-A3B` has sometimes behaved more conservatively. Dual-model and
diarization-assisted hard-fact experiments have not established sufficiently
reliable attribution gains to justify mandatory MVP complexity. Model choice
remains backend configuration, and human review is required.
Meeting Assistant owns the GUI, structured context editor, participant and
speaker-mapping controls, progress presentation, protocol editor and export.
Meeting Lab owns audio preparation, transcription, diarization, orchestration
and protocol-generation logic.
The product should ultimately provide a detailed contextual protocol and a
shorter participant/distribution version. The shorter version may follow the
first GUI milestone.
## Consequences
- ADR 0002's `faster-whisper` selection is no longer current.
- Fixed-duration chunking and chunk-centric storage are not MVP requirements.
- Diarization and GPU acceleration remain optional.
- The Meeting Lab CLI remains useful for command-line operation but is not the
application integration boundary.
- Meeting Context is created through the future GUI; users are not expected to
hand-write YAML.
- Backend experiments can evolve without moving processing logic into the GUI.
- Protocols remain AI-assisted outputs that require human review.
+125 -233
View File
@@ -2,254 +2,146 @@
## Purpose
The Meeting Knowledge Assistant transforms recorded meetings into structured organizational knowledge through a modular processing pipeline.
Meeting Assistant is the user-facing application for preparing meeting
context, running the Meeting Lab pipeline and reviewing its results. Meeting
Lab is the reusable processing backend and experimental engine.
---
# Design Principles
- Architecture first
- Offline first where practical
- Immutable source data
- Replaceable AI components
- Small, focused modules
- Explicit interfaces
- Reproducible AI outputs
- External recording and internal processing are separate responsibilities.
- The processing pipeline must not depend on OBS-specific metadata or behavior.
- Imported recordings must be treated like recordings from any other supported source.
---
## High-Level Pipeline
External Recorder
↓
Audio Import
↓
Recording Validation
↓
Meeting Storage
↓
Transcription
↓
Speaker Diarization
↓
Working Transcript
↓
LLM Analysis
↓
Knowledge Extraction
↓
Export
---
## Core Pipeline
External Recording
↓
Audio Import and Validation
↓
Transcription
↓
Speaker Diarization
↓
Working Transcript
↓
LLM Analysis
↓
Knowledge Extraction
↓
Storage and Export
# Domain Model
Meeting
│
├── Recording
├── Transcript
│ ├── Canonical
│ └── Working
├── Speakers
├── Participants
├── AI Artifacts
├── Knowledge Objects
├── Attachments
└── Exports
---
# Module Responsibilities
## Input and Ingest
Responsible for importing existing recordings into the application.
Initial reference source:
- OBS Studio
Initial supported formats:
- WAV
- FLAC
Responsibilities:
- validate the input file
- collect technical metadata
- calculate checksums
- copy or move the recording into the meeting directory
- create the initial Recording domain object
The ingest component must not perform transcription or modify the audio content.
## Recorder
The recorder module is reserved for a future integrated recording implementation.
It is not required for the MVP.
The initial application workflow uses externally created recordings, with OBS
Studio as the recommended reference recorder.
---
## Transcription
Responsible only for speech-to-text conversion.
Input:
Recording
Output:
Canonical Transcript
---
## Diarization
Responsible only for speaker identification.
Input:
Recording + Canonical Transcript
Output:
Working Transcript
---
## LLM
Responsible for semantic analysis.
Input:
Transcript
Output:
AI Artifacts
---
## Knowledge Extraction
Responsible for creating structured knowledge.
Input:
AI Artifacts
Output:
Knowledge Objects
---
## Export
Responsible for creating user-facing documents.
Input:
Knowledge Objects
Output:
Markdown
PDF
DOCX
---
## Artifact Lifecycle
## System Boundary
```text
Imported Recording
↓
Validated Source Recording
↓
Canonical Transcript
↓
Working Transcript
↓
AI Artifacts
↓
Knowledge Objects
↓
Exports
Meeting Assistant Meeting Lab
----------------- -----------
Audio selection ------> Audio preparation
Meeting Context editor ------> Transcription
Participant management ------> Optional diarization
Explicit speaker mapping ------> Protocol generation
Progress presentation <------ Progress events
Protocol editor and export <------ Run result and artifacts
```
---
Meeting Assistant calls the Meeting Lab Python API directly. The Meeting Lab
CLI is a thin adapter over that API and must not be launched as an application
subprocess.
## Initial Recording Strategy
## MVP Processing Flow
The MVP does not implement platform-specific audio capture.
```text
Source audio
-> FFmpeg normalization/preparation
-> mono, 16 kHz PCM WAV
-> whisper.cpp transcription with large-v3-turbo
-> optional pyannote.audio Community-1 diarization
-> direct full-transcript protocol generation
-> human review and editing
-> export
```
OBS Studio is the recommended reference recorder for online meetings.
The MVP does not require fixed five-minute audio chunks, semantic chunking, a
separate evidence-extraction pipeline or multiple LLMs. Any internal
segmentation remains a Meeting Lab implementation detail.
The application initially processes existing WAV or FLAC recordings. Integrated
## Backend Integration
recording remains a future extension and must not be required by transcription,
The reusable Meeting Lab interface consists of:
diarization or analysis modules.
- `MvpMeetingConfig`, the run configuration
- `run_mvp_meeting(...)`, the orchestration entry point
- `MvpRunResult`, the completed run result
- stage-based progress events
---
The known stages are:
# Storage Strategy
```text
preparing
transcription
diarization
protocol_generation
completed
failed
```
The domain model is independent of the storage backend.
The GUI shows the current stage. It shows a percentage only when the event
contains real measurable progress; stage changes must not be presented as
invented percentages.
Possible implementations:
## Meeting Context
- File System
- SQLite
- PostgreSQL
- Cloud Storage
`MeetingContext` is structured domain input rather than an informal prompt or a
file users must author manually. It contains meeting metadata and participants
and can include optional mappings:
---
```text
SPEAKER_XX -> participant_id
```
# Future Extensions
Mappings are authoritative only after explicit user confirmation. Diarization
labels otherwise remain anonymous, and the application must not infer speaker
names automatically. The Meeting Assistant GUI owns creation and editing of
this context and its mappings.
- Live transcription
- Video processing
- OCR
- Semantic search
- Knowledge graph
- Company glossary
- Multi-language meetings
- Local LLM support
## Processing Components
### Audio Preparation
Meeting Lab uses FFmpeg to prepare a consistent local-processing input. The
current practical target is mono, 16 kHz PCM WAV. The imported source remains a
separate source artifact.
### Transcription
`whisper.cpp` with `large-v3-turbo` is the preferred local backend. Vulkan is a
validated acceleration path on North's AMD RX 9070, while CPU execution remains
a compatibility fallback. Engine and model details should be retained with the
result for traceability.
### Diarization
`pyannote.audio` Community-1 is preferred when diarization is enabled. It can
run on CPU and may use GPU acceleration such as PyTorch/ROCm where available.
Diarization is optional and its anonymous labels do not modify the canonical
transcript or assert participant identity.
### Protocol Generation
The direct full-transcript path is the practical MVP. Current model experiments
favor `qwen3.8:27B` for readability and contextual synthesis, while
`qwen3.6:35B-A3B` has shown more conservative behavior in some areas. Neither a
specific dual-model arrangement nor an evidence pipeline is an application
requirement. Generated protocols require human review.
## Application Responsibilities
The first GUI milestone provides:
- audio-file selection
- structured Meeting Context and participant editing
- optional explicit speaker mapping
- pipeline start and configuration
- progress and failure presentation
- detailed protocol display and editing
- export of the reviewed result
The product direction also includes a shorter participant/distribution
protocol. It may be delivered after the first GUI milestone.
## Artifact and Storage Principles
The existing Meeting domain remains the application-level container for source
recordings, canonical and working transcripts, context, generated artifacts and
exports. Storage remains independent of processing segmentation. In particular,
chunks are not primary Meeting Assistant domain objects or the required unit of
MVP persistence.
Source audio and the canonical transcription output are preserved. Corrections
and reviewed protocols are derived versions. Generated artifacts should retain
their input version, backend/model configuration, prompt version and timestamp
where practical.
## Future Extensions
- shorter distribution protocols
- integrated recording
- searchable meeting history and optional database indexes
- live transcription
- OCR and video processing
- semantic search and organizational knowledge features
+17 -1
View File
@@ -2,7 +2,23 @@ Transcript
Original speech-to-text output.
Diarization
Assignment of transcript segments to speakers.
Assignment of transcript segments to anonymous speaker labels. It does not by
itself identify participants.
Meeting Context
Structured meeting metadata, participants and optional explicitly confirmed
speaker-to-participant mappings supplied to the processing pipeline.
Speaker Mapping
An explicitly confirmed association from a diarization label such as
`SPEAKER_00` to a participant ID. It must not be inferred automatically.
Detailed Protocol
Contextual protocol generated from the full transcript for human review and
continued work.
Distribution Protocol
Shorter, participant-facing version of a reviewed meeting protocol.
Knowledge Object
Structured information extracted from meetings.