Files
meeting-lab/docs/diarization.md
T

3.2 KiB

Optional Speaker Diarization

The direct-protocol MVP keeps speaker diarization disabled by default. Enable anonymous Community-1 speaker labels with --diarization auto, gpu, or cpu:

python3 scripts/run_mvp_meeting.py meeting.wav \
  --whisper-model /path/to/ggml-model.bin \
  --diarization auto

Native mode (the default runtime) requires a compatible local PyTorch and pyannote.audio==4.0.7. For isolated ROCm/CUDA environments, select the container runtime and provide its image and hardware arguments explicitly:

python3 scripts/run_mvp_meeting.py meeting.wav \
  --whisper-model /path/to/ggml-model.bin \
  --diarization gpu \
  --diarization-runtime container \
  --diarization-container-image IMAGE \
  --diarization-container-arg=--device=/dev/kfd \
  --diarization-container-arg=--device=/dev/dri \
  --diarization-container-arg=--group-add \
  --diarization-container-arg=video

The container receives HF_TOKEN by environment-variable name only. It mounts the source audio and repository read-only and writes diarization artifacts into the current run directory. Meeting Lab loads mono 16 kHz PCM16 WAV with Python's wave module and sends an in-memory tensor to pyannote, avoiding its torchcodec file decoder.

Anonymous SPEAKER_XX labels are aligned to Whisper segments by maximum temporal overlap with Community-1 exclusive diarization. The original Whisper transcript is preserved; the derived transcript under diarization/ is used as the direct-protocol generator's source input.

The full diarized JSON and timestamped text remain immutable audit artifacts, but their per-segment formatting is too verbose for a full-meeting LLM prompt: timestamps and repeated speaker labels can more than double input size. For protocol generation, Meeting Lab deterministically groups only adjacent segments assigned to the same anonymous speaker and omits timestamps. A later return by the same speaker starts a new block, and unassigned segments remain under SPEAKER_UNASSIGNED. protocol/transcript_input.txt preserves the exact derived representation sent to prompt construction.

Before contacting Ollama, Meeting Lab conservatively estimates prompt tokens from UTF-8 byte count without adding a model tokenizer dependency. The default safe budget is 29,000 estimated tokens within the explicitly configured 32,768-token Ollama context. The estimate is calibrated against the currently validated German BPD input and is configurable through MvpMeetingConfig.protocol_safe_input_token_budget or --protocol-safe-input-token-budget.

If compact diarized input exceeds the budget, the generator deterministically uses the complete plain segment transcript and records the fallback. If that also exceeds the budget, generation fails before model lookup or generation; it never truncates, chunks, summarizes, retries, or makes multiple protocol calls implicitly. Full diarization artifacts are never overwritten by this selection.

Glossary aliases are recorded as configuration provenance, never applied as deterministic replacements to the compact/plain protocol input. Canonical terminology guides generation through Meeting Context. Raw Whisper and diarization artifacts remain unchanged. glossary_replacements is always empty.