feat: add post-diarization speaker mapping workflow
This commit is contained in:
@@ -39,6 +39,14 @@ Implemented:
|
||||
configurable safe input budget, falls back to complete plain transcript text
|
||||
when necessary, and fails before any Ollama request if even that input is too
|
||||
large. Silent head/tail truncation is prohibited.
|
||||
- The `qwen3.8:27b` direct-protocol stage explicitly requests `num_ctx=32768`
|
||||
and `think=false`; the practical prompt target is approximately 29,000 tokens.
|
||||
A 31,038-token synthetic prompt passed, but larger prompts are not assumed safe
|
||||
from the model's advertised 262,144-token native context alone.
|
||||
- `regenerate_mvp_protocol` updates the run's validated Meeting Context and
|
||||
regenerates protocol artifacts from the existing diarized transcript when
|
||||
available. It never reruns audio preparation, Whisper or Pyannote, and it
|
||||
preserves anonymous speaker labels in the source transcript.
|
||||
- Non-LLM unit tests for chunking, extraction helpers, protocol rendering and
|
||||
gold-test runner validation.
|
||||
- Meeting Context V1 scaffold and documentation for manually maintained
|
||||
|
||||
Reference in New Issue
Block a user