12 KiB
Meeting Assistant
Meeting Assistant is the user-facing application for turning recorded meetings into reviewed, distributable protocols. It uses Meeting Lab as its reusable local processing backend.
Current MVP Direction
The first product milestone is a desktop GUI that lets a user:
- select an existing WAV, FLAC, or M4A recording
- create and edit structured meeting metadata and relevant people, including whether they were present or only mentioned
- import and export the reusable People list as versioned YAML
- optionally map anonymous
SPEAKER_XXlabels to known participants - start the Meeting Lab processing pipeline
- follow stage-based progress
- review and edit the generated protocol
- export the reviewed result
The MVP does not include integrated recording. OBS Studio remains a recommended reference recorder, but imported audio is not tied to OBS-specific behavior.
Processing Pipeline
The existing Meeting language selector controls both transcription and protocol
output: de means German for both, and en means English for both. The selection
is saved in Meeting Context (meeting.language) and reused for protocol-only
regeneration without retranscription. Older contexts without a language retain
German protocol output. Names, speaker mappings and authored Meeting Context are
preserved; no transcript translation stage is added.
Audio file
-> FFmpeg preparation (mono, 16 kHz PCM WAV; normalization optional)
-> whisper.cpp transcription (large-v3-turbo)
-> optional pyannote.audio Community-1 diarization
-> optional speaker review and explicit protocol generation (when diarized)
or direct full-transcript protocol generation (without diarization)
-> human review and export
Meeting Assistant calls Meeting Lab's reusable Python API
(MvpMeetingConfig, MvpRunResult and run_mvp_meeting(...)). It does not run
the Meeting Lab CLI as a subprocess. Fixed five-minute audio chunking is not a
required part of this workflow.
GPU acceleration is supported where available, but CPU execution remains a compatibility path. Diarization is optional and produces anonymous speaker labels. A label identifies a participant only when the user explicitly confirms the mapping; automatic speaker-name inference is not allowed.
Speaker selectors hide participants already assigned to other speakers while keeping the current assignment available. The UI shows detected, assigned and unassigned counts, marks unassigned speakers, and warns when mappings remain incomplete. Anonymous-speaker protocol generation remains available.
After a diarized run, the result view lists detected SPEAKER_XX labels with
short transcript excerpts. Confirmed mappings regenerate only the protocol
from the existing diarized transcript; audio preparation, Whisper and Pyannote
are not rerun. Unmapped / Unknown remains valid, and the original anonymous
diarized transcript is preserved.
Product Outputs
The product direction includes:
- a detailed, contextual protocol for review and continued work
- a shorter version suitable for participants and distribution
The detailed direct-protocol flow is the practical MVP direction. The shorter version may be implemented after the initial GUI. Human review remains part of the workflow for all generated protocols.
Project Boundary
Meeting Lab owns reusable audio preparation, transcription, diarization, orchestration and protocol-generation logic. Meeting Assistant owns the GUI, context and participant editing, explicit speaker mapping, progress display, protocol editing and export.
Audio normalization is enabled by default and can be disabled in the processing options. This controls loudness normalization only: Meeting Lab still prepares every WAV, FLAC or M4A source as canonical audio before transcription.
The Performance profile selector controls protocol-generation runtime using
the abstract Auto (default), Fast, Efficient, and Powersave profiles. Auto lets
the inference backend select its own thread configuration; for Ollama, Meeting
Assistant intentionally sends no num_thread option. Fast, Efficient, and
Powersave are explicit resource profiles currently mapped to 16, 10, and 4
Ollama CPU threads. These concrete mappings may evolve independently of the UI
semantics.
The People section can export its current entries to a UTF-8 people.yaml file
and replace them from a previous .yaml or .yml export. Stable person IDs,
names, roles, organizations and attendance states are retained. This is a small
reuse mechanism, not a server-side participant library or named meeting-template
system.
Run the Streamlit MVP
The development setup expects meeting-assistant and meeting-lab to be
sibling repositories. Meeting Lab currently imports its API through the
src.meeting_lab package path, so its repository root must be supplied on
PYTHONPATH. Meeting Assistant does not modify sys.path at runtime.
Create a virtual environment and install Meeting Assistant, including Streamlit:
python3 -m venv .venv
.venv/bin/pip install -e '.[dev]'
.venv/bin/pip install -e ../meeting-lab
Configure at least the local whisper.cpp model. All supported values are shown
in .env.example; export them in the shell because the application does not
load .env files implicitly:
export MKA_WHISPER_MODEL=/path/to/ggml-large-v3-turbo.bin
export MKA_WHISPER_EXECUTABLE=whisper-cli
export MKA_FFMPEG_EXECUTABLE=ffmpeg
export MKA_PROTOCOL_MODEL=qwen3.8:27b
Optional machine-specific settings include:
MKA_DATA_ROOT(default:data/meetings)MKA_OLLAMA_ENDPOINT(default:http://127.0.0.1:11434)MKA_WHISPER_THREADS(default:auto)MKA_FFMPEG_EXECUTABLE(default:ffmpegfound onPATH)MKA_DIARIZATION_MODE(auto,cpu, orgpu; default:auto)MKA_DIARIZATION_RUNTIME(nativeorcontainer; default:native)MKA_DIARIZATION_CONTAINER_IMAGE(required for container diarization)MKA_DIARIZATION_CONTAINER_ARGS(default: no extra arguments), encoded as a JSON array of strings so ordering and leading dashes are preserved exactlyMKA_GLOSSARY_DATABASE(default:data/database/glossary.sqlite3), the local SQLite file used by the global terminology glossary
Terminology glossary
The Streamlit Terminology glossary section manages recurring product,
material, organization, acronym, and technical names. Store canonical core
terms such as Secugrid HS, not every compound such as Secugrid HS Düse.
Aliases help the protocol model recognize transcript variants while retaining
the surrounding wording. Only active entries are added to Meeting Context as
authoritative terminology for direct protocol generation; inactive entries
remain stored but are omitted. The production database starts empty.
The database defaults to data/database/glossary.sqlite3. It is created and
bootstrapped automatically and can be moved with MKA_GLOSSARY_DATABASE.
Glossary integration does not rewrite raw or diarized transcript artifacts.
Use Export glossary to download the complete SQLite-backed glossary as UTF-8 YAML, and Import glossary to replace it from a previously exported file. Imports are fully validated before a single SQLite transaction replaces the current glossary, so malformed or conflicting files leave existing data unchanged. Schema version 1 is:
version: 1
glossary:
- id: 1
canonical_term: Secugrid HS
aliases:
- Sikirgut
category: product
description: Canonical product spelling
active: true
id is the stable SQLite glossary-entry identifier. Categories are product,
material, organization, technical_term, acronym, or other.
Input configuration import and export
Use Export inputs to save the current Meeting Assistant run-input form as a small, versioned JSON file and Import inputs to restore it later. The file contains meeting metadata, participant records, language, date, audio normalization, and diarization choices. It contains form configuration only: generated prompts, protocols, run artifacts, and source-media contents are not included.
The original source filename may be retained as a reminder, but importing does not restore an upload. Select the audio file explicitly before starting the new run.
Protocol generation with qwen3.8:27b explicitly requests a 32,768-token
Ollama context with thinking disabled; Ollama's machine default may otherwise
be only 4,096. Keep normal prompt input at approximately 29,000 tokens or less.
Although a 31,038-token synthetic prompt passed, larger input is not assumed
safe merely because the model advertises a 262,144-token native context.
For example, a compatible AMD ROCm workstation can configure validated container access without adding controls to the Streamlit UI:
export MKA_DIARIZATION_MODE=gpu
export MKA_DIARIZATION_RUNTIME=container
export MKA_DIARIZATION_CONTAINER_IMAGE=rocm/pytorch:rocm7.2.1_ubuntu24.04_py3.12_pytorch_release_2.9.1
export MKA_DIARIZATION_CONTAINER_ARGS='["--device=/dev/kfd","--device=/dev/dri","--group-add","video"]'
The JSON elements are forwarded as four separate, ordered Meeting Lab container arguments. Runtime and hardware details remain machine-specific environment configuration; the UI continues to expose only the diarization on/off choice.
From the Meeting Assistant repository, start the UI with:
PYTHONPATH=src:../meeting-lab .venv/bin/streamlit run src/mka/ui/streamlit_app.py
North launcher
In the post-diarization Assistant worktree on North, use the checked-in launcher instead of creating another virtual environment:
./run_north.sh
It reuses North's validated Assistant virtual environment and external
Whisper/Ollama/Docker-ROCm runtime while loading Meeting Assistant from this
worktree and Meeting Lab from
/opt/git-projekts/meeting-lab. HF_TOKEN, when present, is
passed through to the disposable diarization container without being stored by
the launcher.
Streamlit opens http://localhost:8501 by default. Installing the sibling
Meeting Lab project supplies its runtime requirements such as PyYAML and
Requests. Optional diarization dependencies are needed only when diarization
is enabled.
Uploaded source files are stored under
data/meetings/<meeting-id>/uploads/. Meeting Lab run artifacts are stored
under data/meetings/<meeting-id>/runs/. The reviewed protocol is saved as
protocol_edited.md inside its run directory; the generated protocol.md
remains unchanged.
The upload remains in its original format and is passed unchanged to Meeting Lab. Meeting Lab creates the canonical mono 16 kHz signed PCM16 WAV run artifact used by transcription; Meeting Assistant does not duplicate audio conversion.
See Architecture, Project Knowledge, Roadmap and ADR 0011.
Active glossary aliases are forwarded to Meeting Lab as provenance metadata.
Canonical terminology remains Meeting Context guidance; the exact protocol
transcript input and the raw Whisper/diarization artifacts are not rewritten.
glossary_replacements remains empty.
Processing and protocol regeneration show measured monotonic stage/total timings,
including frozen failure durations. The latest timing display survives ordinary
Streamlit reruns. Initial-run durations are saved in run_metadata.json;
regeneration display timings remain session-local. Worker callbacks capture UI
configuration before dispatch and never read Streamlit session state.
Meeting Lab retains immutable generation records including prompt, transcript
input, model response/metadata, glossary configuration, mappings and Meeting
Context. Latest protocol/diagnostic/context paths resolve through an atomic
protocol/current link. Copy whole runs with relative symlinks preserved.