feat: add post-diarization speaker mapping workflow

This commit is contained in:
2026-08-25 15:30:16 +02:00
parent df89a38829
commit 8a0f38fce4
6 changed files with 233 additions and 2 deletions
+8
View File
@@ -39,6 +39,14 @@ Implemented:
configurable safe input budget, falls back to complete plain transcript text
when necessary, and fails before any Ollama request if even that input is too
large. Silent head/tail truncation is prohibited.
- The `qwen3.8:27b` direct-protocol stage explicitly requests `num_ctx=32768`
and `think=false`; the practical prompt target is approximately 29,000 tokens.
A 31,038-token synthetic prompt passed, but larger prompts are not assumed safe
from the model's advertised 262,144-token native context alone.
- `regenerate_mvp_protocol` updates the run's validated Meeting Context and
regenerates protocol artifacts from the existing diarized transcript when
available. It never reruns audio preparation, Whisper or Pyannote, and it
preserves anonymous speaker labels in the source transcript.
- Non-LLM unit tests for chunking, extraction helpers, protocol rendering and
gold-test runner validation.
- Meeting Context V1 scaffold and documentation for manually maintained