172 lines
7.2 KiB
Markdown
172 lines
7.2 KiB
Markdown
# Project Knowledge
|
|
|
|
## Meeting language
|
|
|
|
The GUI's `meeting_language` is passed through `MeetingDetails.language` to both
|
|
Meeting Context `meeting.language` and `MvpMeetingConfig.language` (Whisper).
|
|
Meeting Lab derives explicit protocol output language from the persisted context,
|
|
also during regeneration, and records derived `output_language` provenance in
|
|
protocol runtime metadata. Missing context language defaults to German. There is
|
|
no separate protocol-language setting or automatic context/transcript translation.
|
|
|
|
## Vision
|
|
|
|
Meeting Assistant turns recorded meetings into reviewed, user-facing protocols
|
|
and, over time, searchable organizational knowledge. It is the application
|
|
layer over the reusable Meeting Lab processing backend.
|
|
|
|
The source audio and canonical transcript are preserved. Generated protocols
|
|
are derived, reproducible artifacts and require human review.
|
|
|
|
## System Boundary
|
|
|
|
### Meeting Lab
|
|
|
|
Meeting Lab owns reusable and experimental processing:
|
|
|
|
- FFmpeg audio preparation, with optional normalization enabled by default
|
|
- `whisper.cpp` transcription
|
|
- optional `pyannote.audio` diarization
|
|
- protocol generation
|
|
- pipeline orchestration and progress events
|
|
|
|
Its reusable integration surface is the Python API:
|
|
|
|
- `MvpMeetingConfig`
|
|
- `MvpRunResult`
|
|
- `run_mvp_meeting(...)`
|
|
|
|
The CLI is a thin adapter, not the Meeting Assistant integration boundary.
|
|
|
|
### Meeting Assistant
|
|
|
|
Meeting Assistant owns user interaction and product workflow:
|
|
|
|
- audio-file selection
|
|
- structured Meeting Context editing
|
|
- participant management
|
|
- versioned YAML import/export for reusable People lists, using replace semantics
|
|
- versioned JSON import/export for the complete user-configurable run-input
|
|
form; source media is represented only by optional filename metadata and must
|
|
be selected again after import
|
|
- optional explicit speaker mapping
|
|
- post-diarization speaker review and protocol-only regeneration from existing
|
|
run artifacts
|
|
- pipeline launch and progress display
|
|
- protocol review and editing
|
|
- a global SQLite terminology glossary whose active canonical core terms and
|
|
recognition aliases are rendered through Meeting Context into direct
|
|
protocol prompts
|
|
- versioned UTF-8 YAML glossary interchange (`version: 1`) that preserves entry
|
|
IDs and atomically replaces SQLite state only after complete validation
|
|
- export and presentation of protocol versions
|
|
|
|
Processing logic must not be duplicated in the application.
|
|
|
|
Diarization runtime, container image and ordered container arguments are
|
|
machine configuration. `MKA_DIARIZATION_CONTAINER_ARGS` is a JSON array of
|
|
strings forwarded unchanged to Meeting Lab; these details are not normal UI
|
|
controls.
|
|
|
|
People-list YAML is a Meeting Assistant application concern and contains only
|
|
stable IDs, display names, roles, organizations and attendance states. It does
|
|
not contain meeting metadata or processing settings. Named team or meeting
|
|
templates may build on this later, but are not part of the current mechanism.
|
|
|
|
## Validated MVP Pipeline
|
|
|
|
```text
|
|
Imported audio
|
|
-> always prepare as mono, 16 kHz PCM WAV with FFmpeg
|
|
(optionally normalize loudness; default on)
|
|
-> transcribe with whisper.cpp and large-v3-turbo
|
|
-> optionally diarize with pyannote.audio Community-1
|
|
-> generate a protocol directly from the full transcript
|
|
-> review and edit by a human
|
|
```
|
|
|
|
Mandatory fixed five-minute chunking is not part of the productive MVP path.
|
|
Chunking or segmentation used internally by an engine does not shape the
|
|
Meeting Assistant storage or domain model.
|
|
|
|
On the North reference machine, `whisper.cpp` with Vulkan on an AMD RX 9070
|
|
transcribed approximately 93.5 minutes in about 100 seconds. Pyannote through
|
|
PyTorch/ROCm processed the same duration in approximately 98.6 seconds. These
|
|
are validation observations, not performance guarantees or hardware
|
|
requirements. CPU execution remains supported and may be substantially slower.
|
|
|
|
Protocol-generation performance is selected through intentionally abstract
|
|
profiles rather than hardware controls in the GUI. `auto` is the default and
|
|
delegates thread selection to the inference backend; with Ollama this means
|
|
omitting `num_thread` entirely. The explicit resource profiles currently map
|
|
`fast` to 16 Ollama CPU threads, `efficient` to 10, and `powersave` to 4. These
|
|
concrete mappings belong to the backend/configuration boundary and may evolve
|
|
independently of the user-facing profile semantics.
|
|
|
|
## Meeting Context and Speakers
|
|
|
|
`MeetingContext` is a structured domain object containing meeting metadata,
|
|
present participants, and people who were mentioned without attending. The GUI
|
|
records exactly `present` or `mentioned_only`; missing legacy participant status
|
|
defaults to `present` in Meeting Lab. It can also contain optional explicit mappings from
|
|
`SPEAKER_XX` labels to `participant_id` values.
|
|
|
|
A speaker mapping is authoritative only when a user explicitly confirms it and
|
|
may reference only a present participant.
|
|
Speakers otherwise remain anonymous. Automatic speaker-name inference is not
|
|
allowed. The GUI must create and edit Meeting Context; hand-written YAML is not
|
|
a product requirement.
|
|
|
|
Speaker selectors filter participants using a snapshot of all current widget
|
|
selections, falling back to saved mappings before first interaction. Filtering
|
|
retains existing assignments; clearing a selection releases the participant.
|
|
Completeness counts and warnings cover detected speakers only and do not gate
|
|
protocol regeneration or change mapping persistence.
|
|
|
|
## Progress Contract
|
|
|
|
Meeting Lab emits stage-based progress events for:
|
|
|
|
- `preparing`
|
|
- `transcription`
|
|
- `diarization`
|
|
- `protocol_generation`
|
|
- `completed`
|
|
- `failed`
|
|
|
|
The GUI consumes this event interface. Percentages are shown only when real
|
|
measurable progress is available.
|
|
|
|
## Protocol Policy
|
|
|
|
Direct full-transcript protocol generation is the practical MVP direction.
|
|
`qwen3.8:27b` has shown strong readability and contextual synthesis;
|
|
`qwen3.6:35B-A3B` has shown somewhat more conservative behavior in some areas.
|
|
Experimental dual-model and diarization-assisted hard-fact extraction has not
|
|
demonstrated reliably better strict attribution accuracy and is not mandatory.
|
|
|
|
The product should ultimately offer both a detailed contextual protocol and a
|
|
shorter participant/distribution version. The short version remains follow-up
|
|
work if it is not available for the first GUI milestone.
|
|
|
|
Confirmed meeting-specific corrections should ultimately be reusable by both
|
|
protocol views. The detailed correction and selective-regeneration decision is
|
|
recorded in [ADR 0012](docs/adr/0012-post-run-corrections.md); it is later
|
|
product work, not a current MVP requirement.
|
|
|
|
## Design Principles
|
|
|
|
- offline-first where practical
|
|
- explicit application/backend boundaries
|
|
- replaceable AI components
|
|
- CPU-compatible operation with optional acceleration
|
|
- immutable source artifacts and traceable derived artifacts
|
|
- no automatic identity claims
|
|
- human review of generated protocols
|
|
- small modules and simple interfaces
|
|
|
|
Active glossary aliases are forwarded to Meeting Lab as provenance metadata.
|
|
Canonical terminology remains Meeting Context guidance; the exact protocol
|
|
transcript input and the raw Whisper/diarization artifacts are not rewritten.
|
|
`glossary_replacements` remains empty.
|