Add diarization and reusable MVP meeting pipeline
This commit is contained in:
@@ -0,0 +1,38 @@
|
||||
# Optional Speaker Diarization
|
||||
|
||||
The direct-protocol MVP keeps speaker diarization disabled by default. Enable
|
||||
anonymous Community-1 speaker labels with `--diarization auto`, `gpu`, or
|
||||
`cpu`:
|
||||
|
||||
```bash
|
||||
python3 scripts/run_mvp_meeting.py meeting.wav \
|
||||
--whisper-model /path/to/ggml-model.bin \
|
||||
--diarization auto
|
||||
```
|
||||
|
||||
Native mode (the default runtime) requires a compatible local PyTorch and
|
||||
`pyannote.audio==4.0.7`. For isolated ROCm/CUDA environments, select the
|
||||
container runtime and provide its image and hardware arguments explicitly:
|
||||
|
||||
```bash
|
||||
python3 scripts/run_mvp_meeting.py meeting.wav \
|
||||
--whisper-model /path/to/ggml-model.bin \
|
||||
--diarization gpu \
|
||||
--diarization-runtime container \
|
||||
--diarization-container-image IMAGE \
|
||||
--diarization-container-arg=--device=/dev/kfd \
|
||||
--diarization-container-arg=--device=/dev/dri \
|
||||
--diarization-container-arg=--group-add \
|
||||
--diarization-container-arg=video
|
||||
```
|
||||
|
||||
The container receives `HF_TOKEN` by environment-variable name only. It mounts
|
||||
the source audio and repository read-only and writes diarization artifacts into
|
||||
the current run directory. Meeting Lab loads mono 16 kHz PCM16 WAV with
|
||||
Python's `wave` module and sends an in-memory tensor to pyannote, avoiding its
|
||||
torchcodec file decoder.
|
||||
|
||||
Anonymous `SPEAKER_XX` labels are aligned to Whisper segments by maximum
|
||||
temporal overlap with Community-1 exclusive diarization. The original Whisper
|
||||
transcript is preserved; the derived transcript under `diarization/` is used as
|
||||
the unchanged direct-protocol generator's input.
|
||||
Reference in New Issue
Block a user