68 lines
3.2 KiB
Markdown
68 lines
3.2 KiB
Markdown
# Optional Speaker Diarization
|
|
|
|
The direct-protocol MVP keeps speaker diarization disabled by default. Enable
|
|
anonymous Community-1 speaker labels with `--diarization auto`, `gpu`, or
|
|
`cpu`:
|
|
|
|
```bash
|
|
python3 scripts/run_mvp_meeting.py meeting.wav \
|
|
--whisper-model /path/to/ggml-model.bin \
|
|
--diarization auto
|
|
```
|
|
|
|
Native mode (the default runtime) requires a compatible local PyTorch and
|
|
`pyannote.audio==4.0.7`. For isolated ROCm/CUDA environments, select the
|
|
container runtime and provide its image and hardware arguments explicitly:
|
|
|
|
```bash
|
|
python3 scripts/run_mvp_meeting.py meeting.wav \
|
|
--whisper-model /path/to/ggml-model.bin \
|
|
--diarization gpu \
|
|
--diarization-runtime container \
|
|
--diarization-container-image IMAGE \
|
|
--diarization-container-arg=--device=/dev/kfd \
|
|
--diarization-container-arg=--device=/dev/dri \
|
|
--diarization-container-arg=--group-add \
|
|
--diarization-container-arg=video
|
|
```
|
|
|
|
The container receives `HF_TOKEN` by environment-variable name only. It mounts
|
|
the source audio and repository read-only and writes diarization artifacts into
|
|
the current run directory. Meeting Lab loads mono 16 kHz PCM16 WAV with
|
|
Python's `wave` module and sends an in-memory tensor to pyannote, avoiding its
|
|
torchcodec file decoder.
|
|
|
|
Anonymous `SPEAKER_XX` labels are aligned to Whisper segments by maximum
|
|
temporal overlap with Community-1 exclusive diarization. The original Whisper
|
|
transcript is preserved; the derived transcript under `diarization/` is used as
|
|
the direct-protocol generator's source input.
|
|
|
|
The full diarized JSON and timestamped text remain immutable audit artifacts,
|
|
but their per-segment formatting is too verbose for a full-meeting LLM prompt:
|
|
timestamps and repeated speaker labels can more than double input size. For
|
|
protocol generation, Meeting Lab deterministically groups only adjacent
|
|
segments assigned to the same anonymous speaker and omits timestamps. A later
|
|
return by the same speaker starts a new block, and unassigned segments remain
|
|
under `SPEAKER_UNASSIGNED`. `protocol/transcript_input.txt` preserves the exact
|
|
derived representation sent to prompt construction.
|
|
|
|
Before contacting Ollama, Meeting Lab conservatively estimates prompt tokens
|
|
from UTF-8 byte count without adding a model tokenizer dependency. The default safe
|
|
budget is 29,000 estimated tokens within the explicitly configured 32,768-token
|
|
Ollama context. The estimate is calibrated against the currently validated
|
|
German BPD input and is configurable through
|
|
`MvpMeetingConfig.protocol_safe_input_token_budget` or
|
|
`--protocol-safe-input-token-budget`.
|
|
|
|
If compact diarized input exceeds the budget, the generator deterministically
|
|
uses the complete plain segment transcript and records the fallback. If that
|
|
also exceeds the budget, generation fails before model lookup or generation;
|
|
it never truncates, chunks, summarizes, retries, or makes multiple protocol
|
|
calls implicitly. Full diarization artifacts are never overwritten by this
|
|
selection.
|
|
|
|
Glossary aliases are recorded as configuration provenance, never applied as
|
|
deterministic replacements to the compact/plain protocol input. Canonical
|
|
terminology guides generation through Meeting Context. Raw Whisper and
|
|
diarization artifacts remain unchanged. `glossary_replacements` is always empty.
|