3.2 KiB
Optional Speaker Diarization
The direct-protocol MVP keeps speaker diarization disabled by default. Enable
anonymous Community-1 speaker labels with --diarization auto, gpu, or
cpu:
python3 scripts/run_mvp_meeting.py meeting.wav \
--whisper-model /path/to/ggml-model.bin \
--diarization auto
Native mode (the default runtime) requires a compatible local PyTorch and
pyannote.audio==4.0.7. For isolated ROCm/CUDA environments, select the
container runtime and provide its image and hardware arguments explicitly:
python3 scripts/run_mvp_meeting.py meeting.wav \
--whisper-model /path/to/ggml-model.bin \
--diarization gpu \
--diarization-runtime container \
--diarization-container-image IMAGE \
--diarization-container-arg=--device=/dev/kfd \
--diarization-container-arg=--device=/dev/dri \
--diarization-container-arg=--group-add \
--diarization-container-arg=video
The container receives HF_TOKEN by environment-variable name only. It mounts
the source audio and repository read-only and writes diarization artifacts into
the current run directory. Meeting Lab loads mono 16 kHz PCM16 WAV with
Python's wave module and sends an in-memory tensor to pyannote, avoiding its
torchcodec file decoder.
Anonymous SPEAKER_XX labels are aligned to Whisper segments by maximum
temporal overlap with Community-1 exclusive diarization. The original Whisper
transcript is preserved; the derived transcript under diarization/ is used as
the direct-protocol generator's source input.
The full diarized JSON and timestamped text remain immutable audit artifacts,
but their per-segment formatting is too verbose for a full-meeting LLM prompt:
timestamps and repeated speaker labels can more than double input size. For
protocol generation, Meeting Lab deterministically groups only adjacent
segments assigned to the same anonymous speaker and omits timestamps. A later
return by the same speaker starts a new block, and unassigned segments remain
under SPEAKER_UNASSIGNED. protocol/transcript_input.txt preserves the exact
derived representation sent to prompt construction.
Before contacting Ollama, Meeting Lab conservatively estimates prompt tokens
from UTF-8 byte count without adding a model tokenizer dependency. The default safe
budget is 29,000 estimated tokens within the explicitly configured 32,768-token
Ollama context. The estimate is calibrated against the currently validated
German BPD input and is configurable through
MvpMeetingConfig.protocol_safe_input_token_budget or
--protocol-safe-input-token-budget.
If compact diarized input exceeds the budget, the generator deterministically uses the complete plain segment transcript and records the fallback. If that also exceeds the budget, generation fails before model lookup or generation; it never truncates, chunks, summarizes, retries, or makes multiple protocol calls implicitly. Full diarization artifacts are never overwritten by this selection.
Glossary aliases are recorded as configuration provenance, never applied as
deterministic replacements to the compact/plain protocol input. Canonical
terminology guides generation through Meeting Context. Raw Whisper and
diarization artifacts remain unchanged. glossary_replacements is always empty.