Record glossary provenance without transcript mutation

This commit is contained in:
2026-09-12 11:40:57 +02:00
parent 9177a2660b
commit 92247ef43e
7 changed files with 160 additions and 0 deletions
+5
View File
@@ -60,3 +60,8 @@ also exceeds the budget, generation fails before model lookup or generation;
it never truncates, chunks, summarizes, retries, or makes multiple protocol
calls implicitly. Full diarization artifacts are never overwritten by this
selection.
Glossary aliases are recorded as configuration provenance, never applied as
deterministic replacements to the compact/plain protocol input. Canonical
terminology guides generation through Meeting Context. Raw Whisper and
diarization artifacts remain unchanged. `glossary_replacements` is always empty.
+25
View File
@@ -0,0 +1,25 @@
# Direct-protocol regression reproduction
The August 2026 glossary regression was isolated with frozen inputs. Condition
A used the original derived transcript; condition B differed only by these
seven deterministic substitutions:
- `Carbofool` to `Carbofol` (two occurrences)
- `Bento Fix` to `Bentofix` (two occurrences)
- `Sikirgut-Heistlöse` to `Secugrid HS` (one occurrence)
- `Lumini` to `Luminy` (two occurrences)
To repeat the manual comparison, copy the investigated run to a new temporary
directory, retain its Meeting Context and speaker mapping, and invoke
`regenerate_mvp_protocol` through the same parameters used by Meeting
Assistant. Never run the comparison in the historical run directory. Preserve
the model, `num_ctx`, `num_predict`, temperature, think setting, thread setting,
and Meeting Context. Compare the new generation's `exact_prompt.txt` and
`transcript_input.txt` with the frozen A and B artifacts before comparing model
output.
The required fixed condition is A: configured glossary aliases remain visible
in Meeting Context terminology guidance, while `transcript_input.txt` retains
the original seven source spellings. Automated tests mock the model and enforce
that invariant; this full historical experiment remains an explicit manual LLM
validation so normal tests do not depend on a local model.