1.5 KiB
Optional Speaker Diarization
The direct-protocol MVP keeps speaker diarization disabled by default. Enable
anonymous Community-1 speaker labels with --diarization auto, gpu, or
cpu:
python3 scripts/run_mvp_meeting.py meeting.wav \
--whisper-model /path/to/ggml-model.bin \
--diarization auto
Native mode (the default runtime) requires a compatible local PyTorch and
pyannote.audio==4.0.7. For isolated ROCm/CUDA environments, select the
container runtime and provide its image and hardware arguments explicitly:
python3 scripts/run_mvp_meeting.py meeting.wav \
--whisper-model /path/to/ggml-model.bin \
--diarization gpu \
--diarization-runtime container \
--diarization-container-image IMAGE \
--diarization-container-arg=--device=/dev/kfd \
--diarization-container-arg=--device=/dev/dri \
--diarization-container-arg=--group-add \
--diarization-container-arg=video
The container receives HF_TOKEN by environment-variable name only. It mounts
the source audio and repository read-only and writes diarization artifacts into
the current run directory. Meeting Lab loads mono 16 kHz PCM16 WAV with
Python's wave module and sends an in-memory tensor to pyannote, avoiding its
torchcodec file decoder.
Anonymous SPEAKER_XX labels are aligned to Whisper segments by maximum
temporal overlap with Community-1 exclusive diarization. The original Whisper
transcript is preserved; the derived transcript under diarization/ is used as
the unchanged direct-protocol generator's input.