203 lines
9.0 KiB
Markdown
203 lines
9.0 KiB
Markdown
# Meeting Assistant
|
|
|
|
Meeting Assistant is the user-facing application for turning recorded meetings
|
|
into reviewed, distributable protocols. It uses Meeting Lab as its reusable
|
|
local processing backend.
|
|
|
|
## Current MVP Direction
|
|
|
|
The first product milestone is a desktop GUI that lets a user:
|
|
|
|
- select an existing WAV, FLAC, or M4A recording
|
|
- create and edit structured meeting metadata and relevant people, including whether
|
|
they were present or only mentioned
|
|
- import and export the reusable People list as versioned YAML
|
|
- optionally map anonymous `SPEAKER_XX` labels to known participants
|
|
- start the Meeting Lab processing pipeline
|
|
- follow stage-based progress
|
|
- review and edit the generated protocol
|
|
- export the reviewed result
|
|
|
|
The MVP does not include integrated recording. OBS Studio remains a recommended
|
|
reference recorder, but imported audio is not tied to OBS-specific behavior.
|
|
|
|
## Processing Pipeline
|
|
|
|
The existing **Meeting language** selector controls both transcription and protocol
|
|
output: `de` means German for both, and `en` means English for both. The selection
|
|
is saved in Meeting Context (`meeting.language`) and reused for protocol-only
|
|
regeneration without retranscription. Older contexts without a language retain
|
|
German protocol output. Names, speaker mappings and authored Meeting Context are
|
|
preserved; no transcript translation stage is added.
|
|
|
|
```text
|
|
Audio file
|
|
-> FFmpeg preparation (mono, 16 kHz PCM WAV; normalization optional)
|
|
-> whisper.cpp transcription (large-v3-turbo)
|
|
-> optional pyannote.audio Community-1 diarization
|
|
-> direct full-transcript protocol generation
|
|
-> human review and export
|
|
```
|
|
|
|
Meeting Assistant calls Meeting Lab's reusable Python API
|
|
(`MvpMeetingConfig`, `MvpRunResult` and `run_mvp_meeting(...)`). It does not run
|
|
the Meeting Lab CLI as a subprocess. Fixed five-minute audio chunking is not a
|
|
required part of this workflow.
|
|
|
|
GPU acceleration is supported where available, but CPU execution remains a
|
|
compatibility path. Diarization is optional and produces anonymous speaker
|
|
labels. A label identifies a participant only when the user explicitly
|
|
confirms the mapping; automatic speaker-name inference is not allowed.
|
|
|
|
Speaker selectors hide participants already assigned to other speakers while keeping
|
|
the current assignment available. The UI shows detected, assigned and unassigned
|
|
counts, marks unassigned speakers, and warns when mappings remain incomplete.
|
|
Anonymous-speaker protocol generation remains available.
|
|
|
|
After a diarized run, the result view lists detected `SPEAKER_XX` labels with
|
|
short transcript excerpts. Confirmed mappings regenerate only the protocol
|
|
from the existing diarized transcript; audio preparation, Whisper and Pyannote
|
|
are not rerun. `Unmapped / Unknown` remains valid, and the original anonymous
|
|
diarized transcript is preserved.
|
|
|
|
## Product Outputs
|
|
|
|
The product direction includes:
|
|
|
|
- a detailed, contextual protocol for review and continued work
|
|
- a shorter version suitable for participants and distribution
|
|
|
|
The detailed direct-protocol flow is the practical MVP direction. The shorter
|
|
version may be implemented after the initial GUI. Human review remains part of
|
|
the workflow for all generated protocols.
|
|
|
|
## Project Boundary
|
|
|
|
Meeting Lab owns reusable audio preparation, transcription, diarization,
|
|
orchestration and protocol-generation logic. Meeting Assistant owns the GUI,
|
|
context and participant editing, explicit speaker mapping, progress display,
|
|
protocol editing and export.
|
|
|
|
Audio normalization is enabled by default and can be disabled in the processing
|
|
options. This controls loudness normalization only: Meeting Lab still prepares
|
|
every WAV, FLAC or M4A source as canonical audio before transcription.
|
|
|
|
The People section can export its current entries to a UTF-8 `people.yaml` file
|
|
and replace them from a previous `.yaml` or `.yml` export. Stable person IDs,
|
|
names, roles, organizations and attendance states are retained. This is a small
|
|
reuse mechanism, not a server-side participant library or named meeting-template
|
|
system.
|
|
|
|
## Run the Streamlit MVP
|
|
|
|
The development setup expects `meeting-assistant` and `meeting-lab` to be
|
|
sibling repositories. Meeting Lab currently imports its API through the
|
|
`src.meeting_lab` package path, so its repository root must be supplied on
|
|
`PYTHONPATH`. Meeting Assistant does not modify `sys.path` at runtime.
|
|
|
|
Create a virtual environment and install Meeting Assistant, including
|
|
Streamlit:
|
|
|
|
```bash
|
|
python3 -m venv .venv
|
|
.venv/bin/pip install -e '.[dev]'
|
|
.venv/bin/pip install -e ../meeting-lab
|
|
```
|
|
|
|
Configure at least the local whisper.cpp model. All supported values are shown
|
|
in `.env.example`; export them in the shell because the application does not
|
|
load `.env` files implicitly:
|
|
|
|
```bash
|
|
export MKA_WHISPER_MODEL=/path/to/ggml-large-v3-turbo.bin
|
|
export MKA_WHISPER_EXECUTABLE=whisper-cli
|
|
export MKA_FFMPEG_EXECUTABLE=ffmpeg
|
|
export MKA_PROTOCOL_MODEL=qwen3.8:27b
|
|
```
|
|
|
|
Optional machine-specific settings include:
|
|
|
|
- `MKA_DATA_ROOT` (default: `data/meetings`)
|
|
- `MKA_OLLAMA_ENDPOINT` (default: `http://127.0.0.1:11434`)
|
|
- `MKA_WHISPER_THREADS` (default: `auto`)
|
|
- `MKA_FFMPEG_EXECUTABLE` (default: `ffmpeg` found on `PATH`)
|
|
- `MKA_DIARIZATION_MODE` (`auto`, `cpu`, or `gpu`; default: `auto`)
|
|
- `MKA_DIARIZATION_RUNTIME` (`native` or `container`; default: `native`)
|
|
- `MKA_DIARIZATION_CONTAINER_IMAGE` (required for container diarization)
|
|
- `MKA_DIARIZATION_CONTAINER_ARGS` (default: no extra arguments), encoded as a
|
|
JSON array of strings so ordering and leading dashes are preserved exactly
|
|
- `MKA_GLOSSARY_DATABASE` (default: `data/database/glossary.sqlite3`), the local
|
|
SQLite file used by the global terminology glossary
|
|
|
|
## Terminology glossary
|
|
|
|
The Streamlit **Terminology glossary** section manages recurring product,
|
|
material, organization, acronym, and technical names. Store canonical core
|
|
terms such as `Secugrid HS`, not every compound such as `Secugrid HS Düse`.
|
|
Aliases help the protocol model recognize transcript variants while retaining
|
|
the surrounding wording. Only active entries are added to Meeting Context as
|
|
authoritative terminology for direct protocol generation; inactive entries
|
|
remain stored but are omitted. The production database starts empty.
|
|
|
|
The database defaults to `data/database/glossary.sqlite3`. It is created and
|
|
bootstrapped automatically and can be moved with `MKA_GLOSSARY_DATABASE`.
|
|
Glossary integration does not rewrite raw or diarized transcript artifacts.
|
|
|
|
## Input configuration import and export
|
|
|
|
Use **Export inputs** to save the current Meeting Assistant run-input form as a
|
|
small, versioned JSON file and **Import inputs** to restore it later. The file
|
|
contains meeting metadata, participant records, language, date, audio
|
|
normalization, and diarization choices. It contains form configuration only:
|
|
generated prompts, protocols, run artifacts, and source-media contents are not
|
|
included.
|
|
|
|
The original source filename may be retained as a reminder, but importing does
|
|
not restore an upload. Select the audio file explicitly before starting the new
|
|
run.
|
|
|
|
Protocol generation with `qwen3.8:27b` explicitly requests a 32,768-token
|
|
Ollama context with thinking disabled; Ollama's machine default may otherwise
|
|
be only 4,096. Keep normal prompt input at approximately 29,000 tokens or less.
|
|
Although a 31,038-token synthetic prompt passed, larger input is not assumed
|
|
safe merely because the model advertises a 262,144-token native context.
|
|
|
|
For example, a compatible AMD ROCm workstation can configure validated
|
|
container access without adding controls to the Streamlit UI:
|
|
|
|
```bash
|
|
export MKA_DIARIZATION_MODE=gpu
|
|
export MKA_DIARIZATION_RUNTIME=container
|
|
export MKA_DIARIZATION_CONTAINER_IMAGE=rocm/pytorch:rocm7.2.1_ubuntu24.04_py3.12_pytorch_release_2.9.1
|
|
export MKA_DIARIZATION_CONTAINER_ARGS='["--device=/dev/kfd","--device=/dev/dri","--group-add","video"]'
|
|
```
|
|
|
|
The JSON elements are forwarded as four separate, ordered Meeting Lab
|
|
container arguments. Runtime and hardware details remain machine-specific
|
|
environment configuration; the UI continues to expose only the diarization
|
|
on/off choice.
|
|
|
|
From the Meeting Assistant repository, start the UI with:
|
|
|
|
```bash
|
|
PYTHONPATH=src:../meeting-lab .venv/bin/streamlit run src/mka/ui/streamlit_app.py
|
|
```
|
|
|
|
Streamlit opens `http://localhost:8501` by default. Installing the sibling
|
|
Meeting Lab project supplies its runtime requirements such as PyYAML and
|
|
Requests. Optional diarization dependencies are needed only when diarization
|
|
is enabled.
|
|
|
|
Uploaded source files are stored under
|
|
`data/meetings/<meeting-id>/uploads/`. Meeting Lab run artifacts are stored
|
|
under `data/meetings/<meeting-id>/runs/`. The reviewed protocol is saved as
|
|
`protocol_edited.md` inside its run directory; the generated `protocol.md`
|
|
remains unchanged.
|
|
|
|
The upload remains in its original format and is passed unchanged to Meeting Lab.
|
|
Meeting Lab creates the canonical mono 16 kHz signed PCM16 WAV run artifact used by
|
|
transcription; Meeting Assistant does not duplicate audio conversion.
|
|
|
|
See [Architecture](docs/architecture.md), [Project Knowledge](PROJECT_KNOWLEDGE.md),
|
|
[Roadmap](ROADMAP.md) and [ADR 0011](docs/adr/0011-use-meeting-lab-mvp-backend.md).
|