Files

265 lines
12 KiB
Markdown

# Meeting Assistant
Meeting Assistant is the user-facing application for turning recorded meetings
into reviewed, distributable protocols. It uses Meeting Lab as its reusable
local processing backend.
## Current MVP Direction
The first product milestone is a desktop GUI that lets a user:
- select an existing WAV, FLAC, or M4A recording
- create and edit structured meeting metadata and relevant people, including whether
they were present or only mentioned
- import and export the reusable People list as versioned YAML
- optionally map anonymous `SPEAKER_XX` labels to known participants
- start the Meeting Lab processing pipeline
- follow stage-based progress
- review and edit the generated protocol
- export the reviewed result
The MVP does not include integrated recording. OBS Studio remains a recommended
reference recorder, but imported audio is not tied to OBS-specific behavior.
## Processing Pipeline
The existing **Meeting language** selector controls both transcription and protocol
output: `de` means German for both, and `en` means English for both. The selection
is saved in Meeting Context (`meeting.language`) and reused for protocol-only
regeneration without retranscription. Older contexts without a language retain
German protocol output. Names, speaker mappings and authored Meeting Context are
preserved; no transcript translation stage is added.
```text
Audio file
-> FFmpeg preparation (mono, 16 kHz PCM WAV; normalization optional)
-> whisper.cpp transcription (large-v3-turbo)
-> optional pyannote.audio Community-1 diarization
-> optional speaker review and explicit protocol generation (when diarized)
or direct full-transcript protocol generation (without diarization)
-> human review and export
```
Meeting Assistant calls Meeting Lab's reusable Python API
(`MvpMeetingConfig`, `MvpRunResult` and `run_mvp_meeting(...)`). It does not run
the Meeting Lab CLI as a subprocess. Fixed five-minute audio chunking is not a
required part of this workflow.
GPU acceleration is supported where available, but CPU execution remains a
compatibility path. Diarization is optional and produces anonymous speaker
labels. A label identifies a participant only when the user explicitly
confirms the mapping; automatic speaker-name inference is not allowed.
Speaker selectors hide participants already assigned to other speakers while keeping
the current assignment available. The UI shows detected, assigned and unassigned
counts, marks unassigned speakers, and warns when mappings remain incomplete.
Anonymous-speaker protocol generation remains available.
After a diarized run, the result view lists detected `SPEAKER_XX` labels with
short transcript excerpts. Confirmed mappings regenerate only the protocol
from the existing diarized transcript; audio preparation, Whisper and Pyannote
are not rerun. `Unmapped / Unknown` remains valid, and the original anonymous
diarized transcript is preserved.
## Product Outputs
The product direction includes:
- a detailed, contextual protocol for review and continued work
- a shorter version suitable for participants and distribution
The detailed direct-protocol flow is the practical MVP direction. The shorter
version may be implemented after the initial GUI. Human review remains part of
the workflow for all generated protocols.
## Project Boundary
Meeting Lab owns reusable audio preparation, transcription, diarization,
orchestration and protocol-generation logic. Meeting Assistant owns the GUI,
context and participant editing, explicit speaker mapping, progress display,
protocol editing and export.
Audio normalization is enabled by default and can be disabled in the processing
options. This controls loudness normalization only: Meeting Lab still prepares
every WAV, FLAC or M4A source as canonical audio before transcription.
The **Performance profile** selector controls protocol-generation runtime using
the abstract Auto (default), Fast, Efficient, and Powersave profiles. Auto lets
the inference backend select its own thread configuration; for Ollama, Meeting
Assistant intentionally sends no `num_thread` option. Fast, Efficient, and
Powersave are explicit resource profiles currently mapped to 16, 10, and 4
Ollama CPU threads. These concrete mappings may evolve independently of the UI
semantics.
The People section can export its current entries to a UTF-8 `people.yaml` file
and replace them from a previous `.yaml` or `.yml` export. Stable person IDs,
names, roles, organizations and attendance states are retained. This is a small
reuse mechanism, not a server-side participant library or named meeting-template
system.
## Run the Streamlit MVP
The development setup expects `meeting-assistant` and `meeting-lab` to be
sibling repositories. Meeting Lab currently imports its API through the
`src.meeting_lab` package path, so its repository root must be supplied on
`PYTHONPATH`. Meeting Assistant does not modify `sys.path` at runtime.
Create a virtual environment and install Meeting Assistant, including
Streamlit:
```bash
python3 -m venv .venv
.venv/bin/pip install -e '.[dev]'
.venv/bin/pip install -e ../meeting-lab
```
Configure at least the local whisper.cpp model. All supported values are shown
in `.env.example`; export them in the shell because the application does not
load `.env` files implicitly:
```bash
export MKA_WHISPER_MODEL=/path/to/ggml-large-v3-turbo.bin
export MKA_WHISPER_EXECUTABLE=whisper-cli
export MKA_FFMPEG_EXECUTABLE=ffmpeg
export MKA_PROTOCOL_MODEL=qwen3.8:27b
```
Optional machine-specific settings include:
- `MKA_DATA_ROOT` (default: `data/meetings`)
- `MKA_OLLAMA_ENDPOINT` (default: `http://127.0.0.1:11434`)
- `MKA_WHISPER_THREADS` (default: `auto`)
- `MKA_FFMPEG_EXECUTABLE` (default: `ffmpeg` found on `PATH`)
- `MKA_DIARIZATION_MODE` (`auto`, `cpu`, or `gpu`; default: `auto`)
- `MKA_DIARIZATION_RUNTIME` (`native` or `container`; default: `native`)
- `MKA_DIARIZATION_CONTAINER_IMAGE` (required for container diarization)
- `MKA_DIARIZATION_CONTAINER_ARGS` (default: no extra arguments), encoded as a
JSON array of strings so ordering and leading dashes are preserved exactly
- `MKA_GLOSSARY_DATABASE` (default: `data/database/glossary.sqlite3`), the local
SQLite file used by the global terminology glossary
## Terminology glossary
The Streamlit **Terminology glossary** section manages recurring product,
material, organization, acronym, and technical names. Store canonical core
terms such as `Secugrid HS`, not every compound such as `Secugrid HS Düse`.
Aliases help the protocol model recognize transcript variants while retaining
the surrounding wording. Only active entries are added to Meeting Context as
authoritative terminology for direct protocol generation; inactive entries
remain stored but are omitted. The production database starts empty.
The database defaults to `data/database/glossary.sqlite3`. It is created and
bootstrapped automatically and can be moved with `MKA_GLOSSARY_DATABASE`.
Glossary integration does not rewrite raw or diarized transcript artifacts.
Use **Export glossary** to download the complete SQLite-backed glossary as
UTF-8 YAML, and **Import glossary** to replace it from a previously exported
file. Imports are fully validated before a single SQLite transaction replaces
the current glossary, so malformed or conflicting files leave existing data
unchanged. Schema version 1 is:
```yaml
version: 1
glossary:
- id: 1
canonical_term: Secugrid HS
aliases:
- Sikirgut
category: product
description: Canonical product spelling
active: true
```
`id` is the stable SQLite glossary-entry identifier. Categories are `product`,
`material`, `organization`, `technical_term`, `acronym`, or `other`.
## Input configuration import and export
Use **Export inputs** to save the current Meeting Assistant run-input form as a
small, versioned JSON file and **Import inputs** to restore it later. The file
contains meeting metadata, participant records, language, date, audio
normalization, and diarization choices. It contains form configuration only:
generated prompts, protocols, run artifacts, and source-media contents are not
included.
The original source filename may be retained as a reminder, but importing does
not restore an upload. Select the audio file explicitly before starting the new
run.
Protocol generation with `qwen3.8:27b` explicitly requests a 32,768-token
Ollama context with thinking disabled; Ollama's machine default may otherwise
be only 4,096. Keep normal prompt input at approximately 29,000 tokens or less.
Although a 31,038-token synthetic prompt passed, larger input is not assumed
safe merely because the model advertises a 262,144-token native context.
For example, a compatible AMD ROCm workstation can configure validated
container access without adding controls to the Streamlit UI:
```bash
export MKA_DIARIZATION_MODE=gpu
export MKA_DIARIZATION_RUNTIME=container
export MKA_DIARIZATION_CONTAINER_IMAGE=rocm/pytorch:rocm7.2.1_ubuntu24.04_py3.12_pytorch_release_2.9.1
export MKA_DIARIZATION_CONTAINER_ARGS='["--device=/dev/kfd","--device=/dev/dri","--group-add","video"]'
```
The JSON elements are forwarded as four separate, ordered Meeting Lab
container arguments. Runtime and hardware details remain machine-specific
environment configuration; the UI continues to expose only the diarization
on/off choice.
From the Meeting Assistant repository, start the UI with:
```bash
PYTHONPATH=src:../meeting-lab .venv/bin/streamlit run src/mka/ui/streamlit_app.py
```
### North launcher
In the post-diarization Assistant worktree on North, use the checked-in
launcher instead of creating another virtual environment:
```bash
./run_north.sh
```
It reuses North's validated Assistant virtual environment and external
Whisper/Ollama/Docker-ROCm runtime while loading Meeting Assistant from this
worktree and Meeting Lab from
`/opt/git-projekts/meeting-lab`. `HF_TOKEN`, when present, is
passed through to the disposable diarization container without being stored by
the launcher.
Streamlit opens `http://localhost:8501` by default. Installing the sibling
Meeting Lab project supplies its runtime requirements such as PyYAML and
Requests. Optional diarization dependencies are needed only when diarization
is enabled.
Uploaded source files are stored under
`data/meetings/<meeting-id>/uploads/`. Meeting Lab run artifacts are stored
under `data/meetings/<meeting-id>/runs/`. The reviewed protocol is saved as
`protocol_edited.md` inside its run directory; the generated `protocol.md`
remains unchanged.
The upload remains in its original format and is passed unchanged to Meeting Lab.
Meeting Lab creates the canonical mono 16 kHz signed PCM16 WAV run artifact used by
transcription; Meeting Assistant does not duplicate audio conversion.
See [Architecture](docs/architecture.md), [Project Knowledge](PROJECT_KNOWLEDGE.md),
[Roadmap](ROADMAP.md) and [ADR 0011](docs/adr/0011-use-meeting-lab-mvp-backend.md).
Active glossary aliases are forwarded to Meeting Lab as provenance metadata.
Canonical terminology remains Meeting Context guidance; the exact protocol
transcript input and the raw Whisper/diarization artifacts are not rewritten.
`glossary_replacements` remains empty.
Processing and protocol regeneration show measured monotonic stage/total timings,
including frozen failure durations. The latest timing display survives ordinary
Streamlit reruns. Initial-run durations are saved in `run_metadata.json`;
regeneration display timings remain session-local. Worker callbacks capture UI
configuration before dispatch and never read Streamlit session state.
Meeting Lab retains immutable generation records including prompt, transcript
input, model response/metadata, glossary configuration, mappings and Meeting
Context. Latest protocol/diagnostic/context paths resolve through an atomic
`protocol/current` link. Copy whole runs with relative symlinks preserved.