Document validation architecture and renderer faithfulness findings
- document Entity Registry and Meeting Context V2 architecture - preserve meeting_context.yaml as the authoritative meeting-specific input - define immutable authoritative metadata across all pipeline stages - restrict Constraint Repair to deterministic structured-data operations - record BUG-003 root cause and deferred entity-verification resolution - document BUG-005 attendance-consistency design - add BUG-006 renderer faithfulness root-cause analysis - distinguish Engineering Readiness from Practical Usability - update the persistent regression bug tracker
This commit is contained in:
@@ -0,0 +1,265 @@
|
||||
# BUG-003 Root Cause Analysis
|
||||
|
||||
## Observed Behaviour
|
||||
|
||||
The final working protocol contains two names that were flagged as not
|
||||
belonging to the real meeting:
|
||||
|
||||
- `Guido`
|
||||
- `Noah`
|
||||
|
||||
Observed output:
|
||||
|
||||
- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/working_protocol.md:36`
|
||||
contains `Einbinden der Fachbereiche (stellvertretend durch Guido) zur
|
||||
Definition von Prüfsteinen.`
|
||||
- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/working_protocol.md:44`
|
||||
contains `Sollte ein eigener Prozess für Noah (Business Development)
|
||||
benötigt werden oder reicht der bestehende?`
|
||||
|
||||
## Evidence
|
||||
|
||||
The names are present in the final renderer raw response as well as
|
||||
`working_protocol.md`:
|
||||
|
||||
- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/raw_model_response.txt:36`
|
||||
contains `Guido`
|
||||
- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/raw_model_response.txt:44`
|
||||
contains `Noah`
|
||||
|
||||
They are also present before rendering in the repaired consolidated input:
|
||||
|
||||
- `samples/benchmarks/meeting_context_v1/semantic_consolidator_repair_v1/consolidated_extractions.json:1047`
|
||||
contains the open question text with `Noah`
|
||||
- `samples/benchmarks/meeting_context_v1/semantic_consolidator_repair_v1/consolidated_extractions.json:1048`
|
||||
contains evidence `Oder brauchen wir einen eigenen Prozess bei Noah?`
|
||||
- `samples/benchmarks/meeting_context_v1/semantic_consolidator_repair_v1/consolidated_extractions.json:1537`
|
||||
contains the action item text with `Guido`
|
||||
|
||||
They are present before semantic consolidation in the canonicalized output:
|
||||
|
||||
- `samples/benchmarks/meeting_context_v1/canonicalizer/canonicalized_extractions.json:411`
|
||||
contains the open question text with `Noah`
|
||||
- `samples/benchmarks/meeting_context_v1/canonicalizer/canonicalized_extractions.json:412`
|
||||
contains evidence `Oder brauchen wir einen eigenen Prozess bei Noah?`
|
||||
- `samples/benchmarks/meeting_context_v1/canonicalizer/canonicalized_extractions.json:1401`
|
||||
contains the action item text with `Guido`
|
||||
|
||||
They are present before canonicalization in extraction output and raw model
|
||||
responses:
|
||||
|
||||
- `samples/benchmarks/meeting_context_v1/full_context_run/chunk_02_extraction.json:18`
|
||||
contains the extracted open question with `Noah`
|
||||
- `samples/benchmarks/meeting_context_v1/full_context_run/chunk_02_extraction.raw.txt:76`
|
||||
contains the same raw model output question with `Noah`
|
||||
- `samples/benchmarks/meeting_context_v1/full_context_run/chunk_07_extraction.json:11`
|
||||
contains the extracted todo with `Guido`
|
||||
- `samples/benchmarks/meeting_context_v1/full_context_run/chunk_07_extraction.raw.txt:53`
|
||||
contains the same raw model output todo with `Guido`
|
||||
|
||||
They are present before extraction in the reconstructed normalized chunks:
|
||||
|
||||
- `samples/benchmarks/meeting_context_v1/source_reconstruction/chunk_02_normalized.txt:79`
|
||||
contains `Oder brauchen wir einen eigenen Prozess bei Noah?`
|
||||
- `samples/benchmarks/meeting_context_v1/source_reconstruction/chunk_07_normalized.txt:213`
|
||||
contains `Guido`
|
||||
|
||||
They are present before normalization in reconstructed raw chunks:
|
||||
|
||||
- `samples/benchmarks/meeting_context_v1/source_reconstruction/chunk_02.txt:79`
|
||||
contains `Oder brauchen wir einen eigenen Prozess bei Noah?`
|
||||
- `samples/benchmarks/meeting_context_v1/source_reconstruction/chunk_07.txt:213`
|
||||
contains `Guido`
|
||||
|
||||
They are present in the source cleaned Whisper transcript used to reconstruct
|
||||
the chunks:
|
||||
|
||||
- `samples/real_live/project_process_meeting/transcript/meeting_speech_cleaned.json:5783`
|
||||
has segment text `Oder brauchen wir einen eigenen Prozess bei Noah?`
|
||||
- `samples/real_live/project_process_meeting/transcript/meeting_speech_cleaned.json:26315`
|
||||
has segment text `Guido`
|
||||
|
||||
Segment-level verification:
|
||||
|
||||
- `Noah`: segment `id=182`, `start=899.6400000000001`, `end=901.94`,
|
||||
text `Oder brauchen wir einen eigenen Prozess bei Noah?`
|
||||
- `Guido`: segment `id=776`, `start=4157.86`, `end=4160.14`, text `Guido`
|
||||
|
||||
The original Whisper transcript also contains both names:
|
||||
|
||||
- `samples/real_live/project_process_meeting/transcript/meeting_speech.json:1`
|
||||
contains both `Noah` and `Guido` in the top-level transcript text and segment
|
||||
data.
|
||||
|
||||
Meeting Context does not contain either name:
|
||||
|
||||
- `rg -n "Guido|Noah" samples/real_live/project_process_meeting/meeting_context.yaml`
|
||||
returned no matches.
|
||||
|
||||
Prompt files do not contain either name:
|
||||
|
||||
- `rg -n "Guido|Noah" prompts`
|
||||
returned no matches.
|
||||
|
||||
## Earliest Pipeline Stage Containing the Names
|
||||
|
||||
The earliest verified pipeline artifact containing the names is the Whisper
|
||||
transcript stage:
|
||||
|
||||
- `samples/real_live/project_process_meeting/transcript/meeting_speech.json`
|
||||
- `samples/real_live/project_process_meeting/transcript/meeting_speech_cleaned.json`
|
||||
|
||||
For the reconstructed benchmark specifically, the earliest input used by the
|
||||
TODO 4 pipeline is:
|
||||
|
||||
- `samples/real_live/project_process_meeting/transcript/meeting_speech_cleaned.json`
|
||||
|
||||
The reconstructed chunks preserve the names from that cleaned Whisper JSON.
|
||||
Extraction, canonicalization, consolidation and rendering propagate them.
|
||||
|
||||
## Repository Search Results
|
||||
|
||||
Repository search for `Guido|Noah` found these occurrence groups:
|
||||
|
||||
- `docs/regression-bugs.md`: tracker entry for BUG-003 mentions both names.
|
||||
- `tests/gold/responsibility_attribution_negative/transcript.txt`: contains
|
||||
`Noah` in a synthetic gold scenario.
|
||||
- `tests/gold/responsibility_attribution_negative/expected.json`: contains
|
||||
expected `Noah` entries for that synthetic gold scenario.
|
||||
- `samples/real_live/project_process_meeting/transcript/meeting_speech.json`:
|
||||
contains `Noah` and `Guido`.
|
||||
- `samples/real_live/project_process_meeting/transcript/meeting_speech_cleaned.json`:
|
||||
contains `Noah` and `Guido`.
|
||||
- `samples/benchmarks/meeting_context_v1/source_reconstruction/chunk_02.txt`
|
||||
and `chunk_02_normalized.txt`: contain `Noah`.
|
||||
- `samples/benchmarks/meeting_context_v1/source_reconstruction/chunk_07.txt`
|
||||
and `chunk_07_normalized.txt`: contain `Guido`.
|
||||
- `samples/benchmarks/meeting_context_v1/full_context_run/chunk_02_extraction.json`
|
||||
and `chunk_02_extraction.raw.txt`: contain `Noah`.
|
||||
- `samples/benchmarks/meeting_context_v1/full_context_run/chunk_07_extraction.json`
|
||||
and `chunk_07_extraction.raw.txt`: contain `Guido`.
|
||||
- `samples/benchmarks/meeting_context_v1/canonicalizer/canonicalized_extractions.json`:
|
||||
contains both names.
|
||||
- `samples/benchmarks/meeting_context_v1/semantic_consolidator_repair_v1/consolidated_extractions.json`:
|
||||
contains both names.
|
||||
- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/raw_model_response.txt`
|
||||
and `working_protocol.md`: contain both names.
|
||||
- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/comparison.md`:
|
||||
mentions `Noah` in the evaluation notes.
|
||||
|
||||
Search results did not find `Guido` or `Noah` in:
|
||||
|
||||
- `prompts/`
|
||||
- `samples/real_live/project_process_meeting/meeting_context.yaml`
|
||||
- `src/`
|
||||
|
||||
## Prompt Inspection
|
||||
|
||||
Renderer prompt:
|
||||
|
||||
- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/metadata.json`
|
||||
records `prompt_file: "prompts/working_protocol.md"` and
|
||||
`input_source:
|
||||
"samples/benchmarks/meeting_context_v1/semantic_consolidator_repair_v1/consolidated_extractions.json"`.
|
||||
- `prompts/working_protocol.md` contains no few-shot examples and no `Guido` or
|
||||
`Noah` occurrences.
|
||||
|
||||
Extraction prompt assembly:
|
||||
|
||||
- `src/meeting_lab/llm/prompts.py` builds extraction prompts from
|
||||
`prompts/common.md`, optional Meeting Context, explicitly requested task
|
||||
prompts and the provided transcript.
|
||||
- `src/meeting_lab/extraction/extract_chunks.py` calls that shared builder with
|
||||
`decisions.md` and `todos.md` as task prompts.
|
||||
|
||||
Semantic Consolidator prompt assembly:
|
||||
|
||||
- `src/meeting_lab/consolidation/consolidate_facts.py` builds its prompt from
|
||||
`prompts/consolidate_facts.md` plus the canonicalized fact payload.
|
||||
|
||||
Gold/test prompt inspection:
|
||||
|
||||
- `tests/gold/responsibility_attribution_negative/` contains synthetic `Noah`
|
||||
examples.
|
||||
- `scripts/run_gold_test.py` is a gold-test runner and reads
|
||||
`scenario_dir / "transcript.txt"` when explicitly invoked.
|
||||
- The inspected production metadata for the TODO 4 extraction run records
|
||||
`input_path` values under
|
||||
`samples/benchmarks/meeting_context_v1/source_reconstruction/`, not
|
||||
`tests/gold/`.
|
||||
- The inspected renderer metadata records only the repaired consolidated JSON
|
||||
and `prompts/working_protocol.md`.
|
||||
|
||||
No evidence was found that few-shot examples, embedded examples, test
|
||||
transcripts or gold scenarios became part of the production prompt for this
|
||||
benchmark run.
|
||||
|
||||
## Pipeline Verification
|
||||
|
||||
The persisted metadata verifies the production execution path:
|
||||
|
||||
- `chunk_02_extraction.metadata.json` input:
|
||||
`samples/benchmarks/meeting_context_v1/source_reconstruction/chunk_02_normalized.txt`
|
||||
- `chunk_07_extraction.metadata.json` input:
|
||||
`samples/benchmarks/meeting_context_v1/source_reconstruction/chunk_07_normalized.txt`
|
||||
- both extraction metadata files record Meeting Context source:
|
||||
`samples/real_live/project_process_meeting/meeting_context.yaml`
|
||||
- `canonicalizer/metadata.json` records input directory:
|
||||
`samples/benchmarks/meeting_context_v1/full_context_run`
|
||||
- `semantic_consolidator_repair_v1/report.md` records output:
|
||||
`samples/benchmarks/meeting_context_v1/semantic_consolidator_repair_v1/consolidated_extractions.json`
|
||||
- `final_protocol_with_context/metadata.json` records renderer input:
|
||||
`samples/benchmarks/meeting_context_v1/semantic_consolidator_repair_v1/consolidated_extractions.json`
|
||||
and prompt file `prompts/working_protocol.md`
|
||||
|
||||
No execution metadata points to unrelated examples, templates, tests or gold
|
||||
scenarios.
|
||||
|
||||
## Verified Root Cause
|
||||
|
||||
The verified root cause for BUG-003 in this pipeline run is upstream transcript
|
||||
contamination: `Guido` and `Noah` already exist in the Whisper transcript and
|
||||
the cleaned Whisper transcript used as the benchmark source.
|
||||
|
||||
The later stages did not introduce these names from prompts, Meeting Context,
|
||||
gold tests or renderer examples. They propagated names already present in the
|
||||
pipeline input:
|
||||
|
||||
`meeting_speech.json`
|
||||
-> `meeting_speech_cleaned.json`
|
||||
-> reconstructed chunks
|
||||
-> extraction JSON
|
||||
-> canonicalized JSON
|
||||
-> repaired consolidated JSON
|
||||
-> final working protocol
|
||||
|
||||
What cannot be established from repository artifacts alone:
|
||||
|
||||
- whether Whisper hallucinated these names from audio
|
||||
- whether the source audio actually contains words that sound like these names
|
||||
- what the correct intended tokens should be
|
||||
|
||||
Those questions require audio-level or human-transcript verification and are
|
||||
outside the evidence available in this repository trace.
|
||||
|
||||
## Confidence
|
||||
|
||||
High
|
||||
|
||||
The conclusion that the final protocol names originated before extraction is
|
||||
directly supported by persisted source transcript, chunk, extraction,
|
||||
canonicalization, consolidation and renderer artifacts.
|
||||
|
||||
The confidence does not extend to identifying the correct replacement words.
|
||||
That remains unverified.
|
||||
|
||||
## Recommended Fix
|
||||
|
||||
Conceptually, add a transcript/source-quality validation step before extraction
|
||||
that flags person or organization names not present in Meeting Context or an
|
||||
approved alias list. The step should preserve the original transcript text, but
|
||||
mark suspicious entity mentions for review before they become structured
|
||||
knowledge and final protocol content.
|
||||
|
||||
Do not treat gold scenarios or prompt examples as the root cause for this bug;
|
||||
the evidence does not support that.
|
||||
Reference in New Issue
Block a user