Document validation architecture and renderer faithfulness findings
- document Entity Registry and Meeting Context V2 architecture - preserve meeting_context.yaml as the authoritative meeting-specific input - define immutable authoritative metadata across all pipeline stages - restrict Constraint Repair to deterministic structured-data operations - record BUG-003 root cause and deferred entity-verification resolution - document BUG-005 attendance-consistency design - add BUG-006 renderer faithfulness root-cause analysis - distinguish Engineering Readiness from Practical Usability - update the persistent regression bug tracker
This commit is contained in:
@@ -0,0 +1,153 @@
|
||||
# ADR: Meeting Context V2 and Entity Registry
|
||||
|
||||
Status: Accepted Architecture
|
||||
|
||||
Implementation: Deferred
|
||||
|
||||
Date: 2026-08-03
|
||||
|
||||
## Context
|
||||
|
||||
Meeting Context V1 proved that authoritative meeting context can significantly
|
||||
improve extraction quality. It helps the extractor normalize known aliases,
|
||||
identify participants and avoid treating context metadata as evidence for
|
||||
responsibility or decisions.
|
||||
|
||||
End-to-end evaluation also showed that manually writing Meeting Context is not
|
||||
the right long-term primary workflow. The pipeline needs an earlier,
|
||||
interactive entity confirmation step after Whisper transcription. That step
|
||||
should identify candidate entities, ask the user to confirm them and then build
|
||||
the meeting-specific context from confirmed data.
|
||||
|
||||
## Decision
|
||||
|
||||
Meeting Context should evolve toward V2 as an authoritative,
|
||||
meeting-specific Point of Truth generated or assisted from a persistent Entity
|
||||
Registry, user confirmations and meeting metadata.
|
||||
|
||||
Preferred future pipeline:
|
||||
|
||||
```text
|
||||
Whisper
|
||||
↓
|
||||
Entity Detection
|
||||
↓
|
||||
User Confirmation
|
||||
↓
|
||||
Entity Registry Update
|
||||
↓
|
||||
Meeting Context Builder
|
||||
↓
|
||||
meeting_context.yaml
|
||||
↓
|
||||
Extraction Pipeline
|
||||
```
|
||||
|
||||
The Entity Registry is the persistent cross-meeting knowledge source for
|
||||
confirmed entities, aliases and organizational metadata.
|
||||
|
||||
For each individual meeting, `meeting_context.yaml` remains the authoritative
|
||||
meeting-specific Point of Truth and reproducible input artifact consumed by
|
||||
the extraction pipeline. Meeting Context V2 changes how this YAML is prepared,
|
||||
not its authority for a meeting run.
|
||||
|
||||
The generated `meeting_context.yaml` is a meeting-specific snapshot. The
|
||||
Registry must not override explicit meeting-specific confirmations. Changes to
|
||||
the Registry after a meeting run must not silently change the historical
|
||||
Meeting Context used for that run.
|
||||
|
||||
## Entity Registry
|
||||
|
||||
The Entity Registry is the persistent cross-meeting knowledge source and is
|
||||
independent from individual meetings. It stores confirmed entities such as:
|
||||
|
||||
- people
|
||||
- organizations
|
||||
- departments
|
||||
- products
|
||||
- projects
|
||||
- locations
|
||||
- abbreviations
|
||||
|
||||
Each entity receives a stable internal identifier. The displayed name may
|
||||
change over time, but the identifier must remain stable.
|
||||
|
||||
## Learning Principle
|
||||
|
||||
The registry never learns automatically.
|
||||
|
||||
It may propose matches, but only confirmed user actions update the registry.
|
||||
No autonomous learning is allowed.
|
||||
|
||||
## Alias Handling
|
||||
|
||||
Aliases are first-class data.
|
||||
|
||||
Examples:
|
||||
|
||||
```text
|
||||
Jovana
|
||||
Giovanna
|
||||
Jovanna
|
||||
Giovana
|
||||
```
|
||||
|
||||
These variants may all refer to one confirmed entity. Future runs should
|
||||
automatically suggest previously confirmed aliases, but those suggestions still
|
||||
require explicit confirmation when they would update registry data.
|
||||
|
||||
## Unknown Entities
|
||||
|
||||
Previously unseen names are presented to the user for classification.
|
||||
|
||||
Possible classifications:
|
||||
|
||||
- meeting participant
|
||||
- mentioned person
|
||||
- external person
|
||||
- transcription error
|
||||
- ignore
|
||||
|
||||
Nothing is automatically accepted.
|
||||
|
||||
## Similarity Search
|
||||
|
||||
Similarity search is a future extension. It can propose likely matches for:
|
||||
|
||||
- spelling variants
|
||||
- Whisper transcription variants
|
||||
- umlaut handling
|
||||
- OCR-like mistakes
|
||||
|
||||
Similarity suggestions require explicit confirmation.
|
||||
|
||||
## Rationale
|
||||
|
||||
Expected advantages:
|
||||
|
||||
- significantly less manual work
|
||||
- earlier detection of transcription errors
|
||||
- robust alias handling
|
||||
- reusable organizational knowledge
|
||||
- improved Meeting Context quality
|
||||
- easier GUI workflow
|
||||
- better scalability across many meetings
|
||||
|
||||
## Consequences
|
||||
|
||||
Meeting Context V1 remains the current implemented interface.
|
||||
|
||||
Meeting Context V2 should preserve the YAML interface for extraction.
|
||||
`meeting_context.yaml` remains the authoritative meeting-specific Point of
|
||||
Truth and reproducible input artifact for a meeting run. The Entity Registry is
|
||||
the persistent cross-meeting knowledge source used to prepare that artifact.
|
||||
|
||||
The Registry must not override explicit meeting-specific confirmations. Changes
|
||||
to the Registry after a meeting run must not silently change the historical
|
||||
Meeting Context used for that run.
|
||||
|
||||
The registry must not infer responsibility, decisions, attendance or ownership.
|
||||
Those still require meeting evidence and remain governed by the existing
|
||||
responsibility attribution invariant.
|
||||
|
||||
No implementation is part of this ADR.
|
||||
@@ -99,6 +99,8 @@ The analyzer gradually transforms an unstructured discussion into structured kno
|
||||
|
||||
# High-Level Pipeline
|
||||
|
||||
Current implemented and intended analysis flow:
|
||||
|
||||
```text
|
||||
Transcript
|
||||
↓
|
||||
@@ -123,6 +125,24 @@ Output View Rendering
|
||||
Working Protocol / Distribution Protocol / Knowledge Objects
|
||||
```
|
||||
|
||||
Accepted future Meeting Context V2 preparation flow:
|
||||
|
||||
```text
|
||||
Whisper
|
||||
↓
|
||||
Entity Detection
|
||||
↓
|
||||
User Confirmation
|
||||
↓
|
||||
Entity Registry Update
|
||||
↓
|
||||
Meeting Context Builder
|
||||
↓
|
||||
meeting_context.yaml
|
||||
↓
|
||||
Extraction Pipeline
|
||||
```
|
||||
|
||||
Each stage solves one clearly defined problem.
|
||||
|
||||
No module should perform multiple semantic tasks simultaneously.
|
||||
@@ -204,6 +224,32 @@ The context can help prevent non-participants from being interpreted as
|
||||
attendees and can normalize known aliases for extraction. It must not infer
|
||||
roles, departments, responsibilities or decisions.
|
||||
|
||||
Accepted future direction:
|
||||
|
||||
Meeting Context V2 should be generated or assisted from an interactive entity
|
||||
confirmation workflow and a persistent Entity Registry. The Entity Registry is
|
||||
the persistent cross-meeting knowledge source for confirmed people,
|
||||
organizations, departments, products, projects, locations, aliases and
|
||||
organizational metadata under stable internal IDs. Display names may change,
|
||||
but internal IDs remain stable. Aliases are first-class data.
|
||||
|
||||
The registry never learns automatically. It may propose matches and aliases,
|
||||
including spelling variants, Whisper transcription variants, umlaut variants
|
||||
and OCR-like mistakes, but only explicit user confirmation updates registry
|
||||
state. Unknown names should be presented to the user as meeting participant,
|
||||
mentioned person, external person, transcription error or ignore.
|
||||
|
||||
In this architecture, `meeting_context.yaml` remains the authoritative
|
||||
meeting-specific Point of Truth and reproducible input artifact consumed by the
|
||||
extraction pipeline. It is a meeting-specific snapshot built from the Entity
|
||||
Registry, user confirmations and meeting metadata. The Registry must not
|
||||
override explicit meeting-specific confirmations, and Registry changes after a
|
||||
meeting run must not silently change the historical Meeting Context used for
|
||||
that run.
|
||||
|
||||
Status: Accepted Architecture; implementation deferred. See
|
||||
`docs/adr-meeting-context-v2-entity-registry.md`.
|
||||
|
||||
---
|
||||
|
||||
## consolidation/
|
||||
|
||||
@@ -0,0 +1,265 @@
|
||||
# BUG-003 Root Cause Analysis
|
||||
|
||||
## Observed Behaviour
|
||||
|
||||
The final working protocol contains two names that were flagged as not
|
||||
belonging to the real meeting:
|
||||
|
||||
- `Guido`
|
||||
- `Noah`
|
||||
|
||||
Observed output:
|
||||
|
||||
- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/working_protocol.md:36`
|
||||
contains `Einbinden der Fachbereiche (stellvertretend durch Guido) zur
|
||||
Definition von Prüfsteinen.`
|
||||
- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/working_protocol.md:44`
|
||||
contains `Sollte ein eigener Prozess für Noah (Business Development)
|
||||
benötigt werden oder reicht der bestehende?`
|
||||
|
||||
## Evidence
|
||||
|
||||
The names are present in the final renderer raw response as well as
|
||||
`working_protocol.md`:
|
||||
|
||||
- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/raw_model_response.txt:36`
|
||||
contains `Guido`
|
||||
- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/raw_model_response.txt:44`
|
||||
contains `Noah`
|
||||
|
||||
They are also present before rendering in the repaired consolidated input:
|
||||
|
||||
- `samples/benchmarks/meeting_context_v1/semantic_consolidator_repair_v1/consolidated_extractions.json:1047`
|
||||
contains the open question text with `Noah`
|
||||
- `samples/benchmarks/meeting_context_v1/semantic_consolidator_repair_v1/consolidated_extractions.json:1048`
|
||||
contains evidence `Oder brauchen wir einen eigenen Prozess bei Noah?`
|
||||
- `samples/benchmarks/meeting_context_v1/semantic_consolidator_repair_v1/consolidated_extractions.json:1537`
|
||||
contains the action item text with `Guido`
|
||||
|
||||
They are present before semantic consolidation in the canonicalized output:
|
||||
|
||||
- `samples/benchmarks/meeting_context_v1/canonicalizer/canonicalized_extractions.json:411`
|
||||
contains the open question text with `Noah`
|
||||
- `samples/benchmarks/meeting_context_v1/canonicalizer/canonicalized_extractions.json:412`
|
||||
contains evidence `Oder brauchen wir einen eigenen Prozess bei Noah?`
|
||||
- `samples/benchmarks/meeting_context_v1/canonicalizer/canonicalized_extractions.json:1401`
|
||||
contains the action item text with `Guido`
|
||||
|
||||
They are present before canonicalization in extraction output and raw model
|
||||
responses:
|
||||
|
||||
- `samples/benchmarks/meeting_context_v1/full_context_run/chunk_02_extraction.json:18`
|
||||
contains the extracted open question with `Noah`
|
||||
- `samples/benchmarks/meeting_context_v1/full_context_run/chunk_02_extraction.raw.txt:76`
|
||||
contains the same raw model output question with `Noah`
|
||||
- `samples/benchmarks/meeting_context_v1/full_context_run/chunk_07_extraction.json:11`
|
||||
contains the extracted todo with `Guido`
|
||||
- `samples/benchmarks/meeting_context_v1/full_context_run/chunk_07_extraction.raw.txt:53`
|
||||
contains the same raw model output todo with `Guido`
|
||||
|
||||
They are present before extraction in the reconstructed normalized chunks:
|
||||
|
||||
- `samples/benchmarks/meeting_context_v1/source_reconstruction/chunk_02_normalized.txt:79`
|
||||
contains `Oder brauchen wir einen eigenen Prozess bei Noah?`
|
||||
- `samples/benchmarks/meeting_context_v1/source_reconstruction/chunk_07_normalized.txt:213`
|
||||
contains `Guido`
|
||||
|
||||
They are present before normalization in reconstructed raw chunks:
|
||||
|
||||
- `samples/benchmarks/meeting_context_v1/source_reconstruction/chunk_02.txt:79`
|
||||
contains `Oder brauchen wir einen eigenen Prozess bei Noah?`
|
||||
- `samples/benchmarks/meeting_context_v1/source_reconstruction/chunk_07.txt:213`
|
||||
contains `Guido`
|
||||
|
||||
They are present in the source cleaned Whisper transcript used to reconstruct
|
||||
the chunks:
|
||||
|
||||
- `samples/real_live/project_process_meeting/transcript/meeting_speech_cleaned.json:5783`
|
||||
has segment text `Oder brauchen wir einen eigenen Prozess bei Noah?`
|
||||
- `samples/real_live/project_process_meeting/transcript/meeting_speech_cleaned.json:26315`
|
||||
has segment text `Guido`
|
||||
|
||||
Segment-level verification:
|
||||
|
||||
- `Noah`: segment `id=182`, `start=899.6400000000001`, `end=901.94`,
|
||||
text `Oder brauchen wir einen eigenen Prozess bei Noah?`
|
||||
- `Guido`: segment `id=776`, `start=4157.86`, `end=4160.14`, text `Guido`
|
||||
|
||||
The original Whisper transcript also contains both names:
|
||||
|
||||
- `samples/real_live/project_process_meeting/transcript/meeting_speech.json:1`
|
||||
contains both `Noah` and `Guido` in the top-level transcript text and segment
|
||||
data.
|
||||
|
||||
Meeting Context does not contain either name:
|
||||
|
||||
- `rg -n "Guido|Noah" samples/real_live/project_process_meeting/meeting_context.yaml`
|
||||
returned no matches.
|
||||
|
||||
Prompt files do not contain either name:
|
||||
|
||||
- `rg -n "Guido|Noah" prompts`
|
||||
returned no matches.
|
||||
|
||||
## Earliest Pipeline Stage Containing the Names
|
||||
|
||||
The earliest verified pipeline artifact containing the names is the Whisper
|
||||
transcript stage:
|
||||
|
||||
- `samples/real_live/project_process_meeting/transcript/meeting_speech.json`
|
||||
- `samples/real_live/project_process_meeting/transcript/meeting_speech_cleaned.json`
|
||||
|
||||
For the reconstructed benchmark specifically, the earliest input used by the
|
||||
TODO 4 pipeline is:
|
||||
|
||||
- `samples/real_live/project_process_meeting/transcript/meeting_speech_cleaned.json`
|
||||
|
||||
The reconstructed chunks preserve the names from that cleaned Whisper JSON.
|
||||
Extraction, canonicalization, consolidation and rendering propagate them.
|
||||
|
||||
## Repository Search Results
|
||||
|
||||
Repository search for `Guido|Noah` found these occurrence groups:
|
||||
|
||||
- `docs/regression-bugs.md`: tracker entry for BUG-003 mentions both names.
|
||||
- `tests/gold/responsibility_attribution_negative/transcript.txt`: contains
|
||||
`Noah` in a synthetic gold scenario.
|
||||
- `tests/gold/responsibility_attribution_negative/expected.json`: contains
|
||||
expected `Noah` entries for that synthetic gold scenario.
|
||||
- `samples/real_live/project_process_meeting/transcript/meeting_speech.json`:
|
||||
contains `Noah` and `Guido`.
|
||||
- `samples/real_live/project_process_meeting/transcript/meeting_speech_cleaned.json`:
|
||||
contains `Noah` and `Guido`.
|
||||
- `samples/benchmarks/meeting_context_v1/source_reconstruction/chunk_02.txt`
|
||||
and `chunk_02_normalized.txt`: contain `Noah`.
|
||||
- `samples/benchmarks/meeting_context_v1/source_reconstruction/chunk_07.txt`
|
||||
and `chunk_07_normalized.txt`: contain `Guido`.
|
||||
- `samples/benchmarks/meeting_context_v1/full_context_run/chunk_02_extraction.json`
|
||||
and `chunk_02_extraction.raw.txt`: contain `Noah`.
|
||||
- `samples/benchmarks/meeting_context_v1/full_context_run/chunk_07_extraction.json`
|
||||
and `chunk_07_extraction.raw.txt`: contain `Guido`.
|
||||
- `samples/benchmarks/meeting_context_v1/canonicalizer/canonicalized_extractions.json`:
|
||||
contains both names.
|
||||
- `samples/benchmarks/meeting_context_v1/semantic_consolidator_repair_v1/consolidated_extractions.json`:
|
||||
contains both names.
|
||||
- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/raw_model_response.txt`
|
||||
and `working_protocol.md`: contain both names.
|
||||
- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/comparison.md`:
|
||||
mentions `Noah` in the evaluation notes.
|
||||
|
||||
Search results did not find `Guido` or `Noah` in:
|
||||
|
||||
- `prompts/`
|
||||
- `samples/real_live/project_process_meeting/meeting_context.yaml`
|
||||
- `src/`
|
||||
|
||||
## Prompt Inspection
|
||||
|
||||
Renderer prompt:
|
||||
|
||||
- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/metadata.json`
|
||||
records `prompt_file: "prompts/working_protocol.md"` and
|
||||
`input_source:
|
||||
"samples/benchmarks/meeting_context_v1/semantic_consolidator_repair_v1/consolidated_extractions.json"`.
|
||||
- `prompts/working_protocol.md` contains no few-shot examples and no `Guido` or
|
||||
`Noah` occurrences.
|
||||
|
||||
Extraction prompt assembly:
|
||||
|
||||
- `src/meeting_lab/llm/prompts.py` builds extraction prompts from
|
||||
`prompts/common.md`, optional Meeting Context, explicitly requested task
|
||||
prompts and the provided transcript.
|
||||
- `src/meeting_lab/extraction/extract_chunks.py` calls that shared builder with
|
||||
`decisions.md` and `todos.md` as task prompts.
|
||||
|
||||
Semantic Consolidator prompt assembly:
|
||||
|
||||
- `src/meeting_lab/consolidation/consolidate_facts.py` builds its prompt from
|
||||
`prompts/consolidate_facts.md` plus the canonicalized fact payload.
|
||||
|
||||
Gold/test prompt inspection:
|
||||
|
||||
- `tests/gold/responsibility_attribution_negative/` contains synthetic `Noah`
|
||||
examples.
|
||||
- `scripts/run_gold_test.py` is a gold-test runner and reads
|
||||
`scenario_dir / "transcript.txt"` when explicitly invoked.
|
||||
- The inspected production metadata for the TODO 4 extraction run records
|
||||
`input_path` values under
|
||||
`samples/benchmarks/meeting_context_v1/source_reconstruction/`, not
|
||||
`tests/gold/`.
|
||||
- The inspected renderer metadata records only the repaired consolidated JSON
|
||||
and `prompts/working_protocol.md`.
|
||||
|
||||
No evidence was found that few-shot examples, embedded examples, test
|
||||
transcripts or gold scenarios became part of the production prompt for this
|
||||
benchmark run.
|
||||
|
||||
## Pipeline Verification
|
||||
|
||||
The persisted metadata verifies the production execution path:
|
||||
|
||||
- `chunk_02_extraction.metadata.json` input:
|
||||
`samples/benchmarks/meeting_context_v1/source_reconstruction/chunk_02_normalized.txt`
|
||||
- `chunk_07_extraction.metadata.json` input:
|
||||
`samples/benchmarks/meeting_context_v1/source_reconstruction/chunk_07_normalized.txt`
|
||||
- both extraction metadata files record Meeting Context source:
|
||||
`samples/real_live/project_process_meeting/meeting_context.yaml`
|
||||
- `canonicalizer/metadata.json` records input directory:
|
||||
`samples/benchmarks/meeting_context_v1/full_context_run`
|
||||
- `semantic_consolidator_repair_v1/report.md` records output:
|
||||
`samples/benchmarks/meeting_context_v1/semantic_consolidator_repair_v1/consolidated_extractions.json`
|
||||
- `final_protocol_with_context/metadata.json` records renderer input:
|
||||
`samples/benchmarks/meeting_context_v1/semantic_consolidator_repair_v1/consolidated_extractions.json`
|
||||
and prompt file `prompts/working_protocol.md`
|
||||
|
||||
No execution metadata points to unrelated examples, templates, tests or gold
|
||||
scenarios.
|
||||
|
||||
## Verified Root Cause
|
||||
|
||||
The verified root cause for BUG-003 in this pipeline run is upstream transcript
|
||||
contamination: `Guido` and `Noah` already exist in the Whisper transcript and
|
||||
the cleaned Whisper transcript used as the benchmark source.
|
||||
|
||||
The later stages did not introduce these names from prompts, Meeting Context,
|
||||
gold tests or renderer examples. They propagated names already present in the
|
||||
pipeline input:
|
||||
|
||||
`meeting_speech.json`
|
||||
-> `meeting_speech_cleaned.json`
|
||||
-> reconstructed chunks
|
||||
-> extraction JSON
|
||||
-> canonicalized JSON
|
||||
-> repaired consolidated JSON
|
||||
-> final working protocol
|
||||
|
||||
What cannot be established from repository artifacts alone:
|
||||
|
||||
- whether Whisper hallucinated these names from audio
|
||||
- whether the source audio actually contains words that sound like these names
|
||||
- what the correct intended tokens should be
|
||||
|
||||
Those questions require audio-level or human-transcript verification and are
|
||||
outside the evidence available in this repository trace.
|
||||
|
||||
## Confidence
|
||||
|
||||
High
|
||||
|
||||
The conclusion that the final protocol names originated before extraction is
|
||||
directly supported by persisted source transcript, chunk, extraction,
|
||||
canonicalization, consolidation and renderer artifacts.
|
||||
|
||||
The confidence does not extend to identifying the correct replacement words.
|
||||
That remains unverified.
|
||||
|
||||
## Recommended Fix
|
||||
|
||||
Conceptually, add a transcript/source-quality validation step before extraction
|
||||
that flags person or organization names not present in Meeting Context or an
|
||||
approved alias list. The step should preserve the original transcript text, but
|
||||
mark suspicious entity mentions for review before they become structured
|
||||
knowledge and final protocol content.
|
||||
|
||||
Do not treat gold scenarios or prompt examples as the root cause for this bug;
|
||||
the evidence does not support that.
|
||||
@@ -0,0 +1,198 @@
|
||||
# BUG-005 Root Cause Analysis
|
||||
|
||||
## Observed Behaviour
|
||||
|
||||
The final working protocol treats Björn as one of the "fehlende Teilnehmer":
|
||||
|
||||
- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/working_protocol.md:53`
|
||||
|
||||
This contradicts Meeting Context, which lists Björn as an actual participant
|
||||
with `attendance_status: "present"`:
|
||||
|
||||
- `samples/real_live/project_process_meeting/meeting_context.yaml:45`
|
||||
- `samples/real_live/project_process_meeting/meeting_context.yaml:52`
|
||||
|
||||
## Evidence
|
||||
|
||||
Meeting Context marks Björn present:
|
||||
|
||||
```text
|
||||
participant_id: "bjoern"
|
||||
display_name: "Björn"
|
||||
role: "Leiter Marketing"
|
||||
attendance_status: "present"
|
||||
```
|
||||
|
||||
The extraction prompt path does include Meeting Context:
|
||||
|
||||
- `src/meeting_lab/extraction/extract_chunks.py:161` renders Meeting Context
|
||||
for the prompt when it is supplied.
|
||||
- `src/meeting_lab/models/meeting_context.py:108` renders the heading
|
||||
`MEETING CONTEXT V1 (AUTHORITATIVE METADATA)`.
|
||||
- `src/meeting_lab/models/meeting_context.py:112` renders the rule
|
||||
`The participant list is authoritative.`
|
||||
- `src/meeting_lab/models/meeting_context.py:128` to
|
||||
`src/meeting_lab/models/meeting_context.py:132` renders actual
|
||||
participants.
|
||||
|
||||
Execution metadata confirms chunk 08 used the Meeting Context file:
|
||||
|
||||
- `samples/benchmarks/meeting_context_v1/full_context_run/chunk_08_extraction.metadata.json:29`
|
||||
to `:32`
|
||||
- `samples/benchmarks/meeting_context_v1/full_context_run/chunk_08_extraction.metadata.json:34`
|
||||
to `:38`
|
||||
|
||||
The source chunk does not say Björn was absent from the meeting. It says
|
||||
Giovanna and Björn had not yet understood or reviewed the discussed mail or
|
||||
process state:
|
||||
|
||||
- `samples/benchmarks/meeting_context_v1/source_reconstruction/chunk_08_normalized.txt:211`
|
||||
to `:215`
|
||||
- `samples/benchmarks/meeting_context_v1/source_reconstruction/chunk_08_normalized.txt:229`
|
||||
to `:245`
|
||||
- `samples/benchmarks/meeting_context_v1/source_reconstruction/chunk_08_normalized.txt:257`
|
||||
to `:267`
|
||||
|
||||
The earliest generated artifact that treats Björn as absent is the raw chunk
|
||||
08 extraction response:
|
||||
|
||||
- `samples/benchmarks/meeting_context_v1/full_context_run/chunk_08_extraction.raw.txt:4`
|
||||
summarizes "Einbeziehung fehlender Teilnehmer wie Jovana und Björn".
|
||||
- `samples/benchmarks/meeting_context_v1/full_context_run/chunk_08_extraction.raw.txt:6`
|
||||
to `:10` lists participants as only Martin, Lars and Malte.
|
||||
- `samples/benchmarks/meeting_context_v1/full_context_run/chunk_08_extraction.raw.txt:64`
|
||||
to `:68` emits the open question
|
||||
`Wie sollen fehlende Teilnehmer (Jovana, Björn) in den Prozess einbezogen werden?`
|
||||
|
||||
The persisted extraction JSON keeps that open question:
|
||||
|
||||
- `samples/benchmarks/meeting_context_v1/full_context_run/chunk_08_extraction.json:14`
|
||||
to `:16`
|
||||
|
||||
The persisted extraction JSON records only Meeting Context provenance, not the
|
||||
participant list or attendance status:
|
||||
|
||||
- `samples/benchmarks/meeting_context_v1/full_context_run/chunk_08_extraction.json:23`
|
||||
to `:27`
|
||||
|
||||
The normalizer drops raw response fields such as `participants` and `topics`:
|
||||
|
||||
- `src/meeting_lab/extraction/extract_chunks.py:318` to `:365` returns only
|
||||
normalized `facts`, `decisions`, `todos`, `questions`, `positions` and
|
||||
`technical`.
|
||||
- `src/meeting_lab/extraction/extract_chunks.py:432` to `:434` adds only
|
||||
provenance from Meeting Context after normalization.
|
||||
|
||||
Canonicalizer preserves the bad open question:
|
||||
|
||||
- `samples/benchmarks/meeting_context_v1/canonicalizer/canonicalized_extractions.json:1633`
|
||||
to `:1645`
|
||||
|
||||
Semantic Consolidator output also preserves the bad open question:
|
||||
|
||||
- `samples/benchmarks/meeting_context_v1/semantic_consolidator_repair_v1/consolidated_extractions.json:1669`
|
||||
to `:1681`
|
||||
|
||||
The renderer receives the repaired consolidated JSON as its only recorded
|
||||
input:
|
||||
|
||||
- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/metadata.json:2`
|
||||
to `:4`
|
||||
|
||||
The renderer prompt requires using only the provided input:
|
||||
|
||||
- `prompts/working_protocol.md:7` to `:12`
|
||||
|
||||
Therefore the renderer did not receive Meeting Context attendance metadata
|
||||
that would let it distinguish a present participant from a mentioned-only or
|
||||
absent person.
|
||||
|
||||
## Stage-by-Stage Verification
|
||||
|
||||
1. Does extraction receive Björn as a present participant?
|
||||
|
||||
Yes. The extraction code renders Meeting Context into the prompt when supplied,
|
||||
and the chunk 08 metadata records
|
||||
`samples/real_live/project_process_meeting/meeting_context.yaml` as the Meeting
|
||||
Context source.
|
||||
|
||||
2. Does extraction output preserve that information?
|
||||
|
||||
No. The raw chunk 08 response lists participants as Martin, Lars and Malte
|
||||
only, despite Björn being present in Meeting Context. The official persisted
|
||||
extraction JSON contains only a minimal `context` provenance object and does
|
||||
not preserve the authoritative participant list or attendance status.
|
||||
|
||||
3. Does Canonicalizer preserve it?
|
||||
|
||||
No. Canonicalizer input lacks the participant list and attendance status. The
|
||||
Canonicalizer preserves the already bad open question from chunk 08.
|
||||
|
||||
4. Does Semantic Consolidator preserve it?
|
||||
|
||||
No. The repaired consolidated JSON lacks Meeting Context participant metadata
|
||||
and preserves the bad open question.
|
||||
|
||||
5. Does the Renderer receive enough information to distinguish present from
|
||||
mentioned-only participants?
|
||||
|
||||
No. Renderer metadata records only
|
||||
`semantic_consolidator_repair_v1/consolidated_extractions.json` as input, and
|
||||
that file contains no Meeting Context participant list or attendance status.
|
||||
|
||||
## Earliest Failing Pipeline Stage
|
||||
|
||||
The earliest failing stage is chunk extraction for
|
||||
`chunk_08_normalized.txt`.
|
||||
|
||||
More precisely, the raw model response for chunk 08 is the first artifact that
|
||||
both omits Björn from the participant list and groups Björn with absent or
|
||||
"fehlende" participants.
|
||||
|
||||
## Verified Root Cause
|
||||
|
||||
The immediate root cause is ignored Meeting Context metadata during extraction
|
||||
for chunk 08. The extraction model received authoritative Meeting Context, but
|
||||
interpreted transcript evidence about not having reviewed or understood an
|
||||
email/process state as meeting absence.
|
||||
|
||||
The downstream cause is missing propagation of participant attendance metadata.
|
||||
After extraction, only Meeting Context provenance is persisted. Canonicalizer,
|
||||
Semantic Consolidator and Renderer do not receive the authoritative participant
|
||||
list or attendance status, so they cannot detect or repair the contradiction.
|
||||
|
||||
This is not primarily a renderer inference bug. The renderer preserved an open
|
||||
question already present in its input, and its prompt explicitly restricts it
|
||||
to the supplied consolidated input.
|
||||
|
||||
## Cause Classification
|
||||
|
||||
- Missing propagation: verified.
|
||||
- Ignored metadata: verified at extraction.
|
||||
- Prompt wording: not verified as the direct cause.
|
||||
- Renderer inference: disproved as the earliest cause.
|
||||
- Lost provenance: partially verified; provenance remains, but the substantive
|
||||
participant metadata is not propagated.
|
||||
- Another cause: not identified.
|
||||
|
||||
## Confidence
|
||||
|
||||
High.
|
||||
|
||||
The conclusion is supported by Meeting Context, extraction metadata, raw chunk
|
||||
08 model output, persisted extraction JSON, Canonicalizer output, Semantic
|
||||
Consolidator output and renderer metadata.
|
||||
|
||||
## Recommended Fix
|
||||
|
||||
Concept only:
|
||||
|
||||
Propagate authoritative Meeting Context participant metadata, including
|
||||
attendance status, beyond extraction as structured data. Add validation that
|
||||
flags contradictions where generated items classify an actual participant as
|
||||
absent or mentioned-only. The validation should run before rendering, and
|
||||
ideally immediately after extraction so the contradiction is caught at the
|
||||
earliest stage.
|
||||
|
||||
Do not rely on the Working Protocol Renderer to correct attendance semantics
|
||||
from prose-only consolidated items.
|
||||
@@ -0,0 +1,426 @@
|
||||
# BUG-006 Root Cause Analysis
|
||||
|
||||
## Observed Behaviour
|
||||
|
||||
In benchmark `meeting_context_v1/e2e_current_20260803_151000`, the Working
|
||||
Protocol Renderer output is not a faithful view of the matching consolidated
|
||||
representation.
|
||||
|
||||
Compared files:
|
||||
|
||||
- Consolidated input:
|
||||
`samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/semantic_consolidator/consolidated_extractions.json`
|
||||
- Rendered output:
|
||||
`samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/working_protocol/working_protocol.md`
|
||||
- Raw renderer response:
|
||||
`samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/working_protocol/raw_model_response.txt`
|
||||
|
||||
Observed problems:
|
||||
|
||||
- consolidated decisions are omitted from the rendered `Decisions` sections
|
||||
- consolidated open questions are omitted from rendered `Open Questions`
|
||||
sections
|
||||
- consolidated facts are rendered as decisions
|
||||
- some consolidated action items are omitted
|
||||
- some action-item and decision meanings are merged into other rendered
|
||||
sections
|
||||
|
||||
The raw renderer response and `working_protocol.md` are byte-identical
|
||||
(`cmp` exit code `0`), so the mismatch is already present in the LLM renderer
|
||||
response. It is not introduced by Markdown file writing.
|
||||
|
||||
## Renderer Input And Prompt Evidence
|
||||
|
||||
Renderer metadata records the exact input and prompt:
|
||||
|
||||
- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/working_protocol/metadata.json:2`
|
||||
to `:4`
|
||||
|
||||
The renderer prompt explicitly says:
|
||||
|
||||
- use only the provided input:
|
||||
`prompts/working_protocol.md:9` to `:10`
|
||||
- preserve all decisions, action items and open questions:
|
||||
`prompts/working_protocol.md:24` to `:26`
|
||||
- remove presentation-level redundancy:
|
||||
`prompts/working_protocol.md:23`
|
||||
- do not repeat information merely because it appears in multiple categories:
|
||||
`prompts/working_protocol.md:65`
|
||||
- condense everything else:
|
||||
`prompts/working_protocol.md:75`
|
||||
|
||||
The prompt is therefore present and relevant, but the direct observed failure
|
||||
occurs in the model-generated renderer response.
|
||||
|
||||
## Faithfulness Analysis
|
||||
|
||||
### Rendered Section: Prozessrahmen und Projektideen-Eingang / Background
|
||||
|
||||
Status: modified.
|
||||
|
||||
Rendered lines:
|
||||
|
||||
- `working_protocol.md:7`
|
||||
- `working_protocol.md:9`
|
||||
|
||||
The section condenses several consolidated facts and decisions into background
|
||||
paragraphs. Examples:
|
||||
|
||||
- `decision_0001` is restated in background and again as a decision.
|
||||
- `fact`: `Der Prozess ist ein F&E-Prozess gewesen, einfach historisch.`
|
||||
- `fact`: `Die Bearbeitung der Projekte muss und kann nur in den
|
||||
Fachabteilungen passieren.`
|
||||
- `fact`: `Bevor F&E ein Projekt durchdenkt, müssen im Vorfeld Kriterien wie
|
||||
Marktexistenz und Zulassungen geprüft werden.`
|
||||
- `decision_0003` and `decision_0005` are partially restated as background.
|
||||
|
||||
This is a condensation, not a renderer-only invention.
|
||||
|
||||
### Rendered Section: Prozessrahmen und Projektideen-Eingang / Decisions
|
||||
|
||||
Status: partially faithful, partially omitted.
|
||||
|
||||
Faithfully or near-faithfully preserved:
|
||||
|
||||
- `decision_0001` -> `working_protocol.md:13`
|
||||
- `decision_0002` -> `working_protocol.md:14`
|
||||
- `decision_0003` -> `working_protocol.md:15`
|
||||
- `decision_0005` -> `working_protocol.md:16`
|
||||
- `decision_0006` -> `working_protocol.md:17`
|
||||
- `decision_0009` -> `working_protocol.md:18`
|
||||
- `decision_0011` -> `working_protocol.md:19`
|
||||
|
||||
Omitted as decisions:
|
||||
|
||||
- `decision_0004`: `Festlegung eines Reporting-Zyklus für abgelehnte
|
||||
Projekte.`
|
||||
- Consolidated evidence: `consolidated_extractions.json:1217` to `:1227`
|
||||
- Rendered only as background: `working_protocol.md:72`
|
||||
- `decision_0007`: `Zusammengetragen und beantwortete Fragen führen zur
|
||||
Sitzung mit fünf Leuten...`
|
||||
- Consolidated evidence: `consolidated_extractions.json:1329` to `:1339`
|
||||
- No matching rendered decision.
|
||||
- `decision_0008`: `Es wird vereinbart, den bestehenden Prozessrahmen zu
|
||||
nutzen und auf die Bedürfnisse der F&E abzustimmen...`
|
||||
- Consolidated evidence present in the input.
|
||||
- No matching rendered decision.
|
||||
- `decision_0010`: `Es wird vereinbart, dass betriebliche
|
||||
Verbesserungsvorschläge... ausgeschleust werden.`
|
||||
- Consolidated evidence present in the input.
|
||||
- No matching rendered decision.
|
||||
- `decision_0012`: `Die Diskussion wird beendet und die Änderungen werden
|
||||
verschickt; Abwarten auf Rückmeldung von Johanna/Björn.`
|
||||
- Consolidated evidence: `consolidated_extractions.json:1747` to `:1757`
|
||||
- No matching rendered decision.
|
||||
|
||||
### Rendered Section: Prozessrahmen und Projektideen-Eingang / Action Items
|
||||
|
||||
Status: partially faithful, partially omitted, with upstream semantic problems
|
||||
preserved.
|
||||
|
||||
Rendered lines:
|
||||
|
||||
- `working_protocol.md:23` to `:32`
|
||||
|
||||
Preserved action items include:
|
||||
|
||||
- `action_item_0001` -> `working_protocol.md:23`
|
||||
- `action_item_0002` -> `working_protocol.md:24`
|
||||
- `action_item_0003` -> `working_protocol.md:25`
|
||||
- `action_item_0004` -> `working_protocol.md:26`
|
||||
- `action_item_0005` -> `working_protocol.md:27`
|
||||
- `action_item_0006` -> `working_protocol.md:28`
|
||||
- `action_item_0007` -> `working_protocol.md:29`
|
||||
- `action_item_0009` -> `working_protocol.md:30`
|
||||
- `action_item_0011` -> `working_protocol.md:31`
|
||||
- `action_item_0012` -> `working_protocol.md:32`
|
||||
|
||||
Omitted or merged action items:
|
||||
|
||||
- `action_item_0008`: `Diskussion über die Relevanz eines speziellen Marktes
|
||||
(z.B. Turkmenistan) führen.`
|
||||
- Present in consolidated input.
|
||||
- No matching rendered action item.
|
||||
- `action_item_0010`: `Filterkriterien für verschiedene Fälle... definieren.`
|
||||
- Present in consolidated input.
|
||||
- Its meaning overlaps with rendered decision `working_protocol.md:18` and
|
||||
action item `working_protocol.md:30`, but it is not preserved as its own
|
||||
action item.
|
||||
|
||||
Important boundary:
|
||||
|
||||
Items such as `Feedback zu den Kriterien...`, `Einbinden der Fachbereiche...`
|
||||
and `Mail an Jovana und Björn...` are questionable todos, but they are already
|
||||
`action_item` entries in the consolidated input:
|
||||
|
||||
- `consolidated_extractions.json:1007` to `:1019`
|
||||
- `consolidated_extractions.json:1537` to `:1549`
|
||||
- `consolidated_extractions.json:1651` to `:1663`
|
||||
|
||||
The renderer preserves those upstream action-item categories. It does not
|
||||
create those todos from facts or prose.
|
||||
|
||||
### Rendered Section: Prozessrahmen und Projektideen-Eingang / Open Questions
|
||||
|
||||
Status: partially faithful, partially omitted.
|
||||
|
||||
Rendered lines:
|
||||
|
||||
- `working_protocol.md:36` to `:47`
|
||||
|
||||
Preserved open questions:
|
||||
|
||||
- `open_question_0001` -> `working_protocol.md:36`
|
||||
- `open_question_0002` -> `working_protocol.md:37`
|
||||
- `open_question_0003` -> `working_protocol.md:38`
|
||||
- `open_question_0004` -> `working_protocol.md:39`
|
||||
- `open_question_0005` -> `working_protocol.md:40`
|
||||
- `open_question_0006` -> `working_protocol.md:41`
|
||||
- `open_question_0007` -> `working_protocol.md:42`
|
||||
- `open_question_0009` -> `working_protocol.md:43`
|
||||
- `open_question_0010` -> `working_protocol.md:44`
|
||||
- `open_question_0011` -> `working_protocol.md:45`
|
||||
- `open_question_0012` -> `working_protocol.md:46`
|
||||
- `open_question_0013` -> `working_protocol.md:47`
|
||||
|
||||
Omitted open questions:
|
||||
|
||||
- `open_question_0008`: `Wer soll die Rolle des Gatekeepers übernehmen und wie
|
||||
wird der Prozess konkret implementiert?`
|
||||
- Consolidated evidence: `consolidated_extractions.json:1387` to `:1397`
|
||||
- No matching rendered open question.
|
||||
- `open_question_0014`: `Welche Prüfsteine sind relevant für die
|
||||
Fachabteilung?`
|
||||
- Consolidated evidence: `consolidated_extractions.json:1765` to `:1775`
|
||||
- No matching rendered open question.
|
||||
|
||||
### Rendered Section: Projektkategorisierung und Status / Background
|
||||
|
||||
Status: modified.
|
||||
|
||||
Rendered line:
|
||||
|
||||
- `working_protocol.md:53`
|
||||
|
||||
This condenses:
|
||||
|
||||
- fact: `Der Leiter F&E führt die Projektliste...`
|
||||
- `technical_detail_0010`: project type definition
|
||||
- fact: `Ein Projekt wandert direkt wieder in Business Development...`
|
||||
|
||||
No renderer-only invention was verified in this section.
|
||||
|
||||
### Rendered Section: Projektkategorisierung und Status / Decisions
|
||||
|
||||
Status: category-changed.
|
||||
|
||||
Rendered lines:
|
||||
|
||||
- `working_protocol.md:57`
|
||||
- `working_protocol.md:58`
|
||||
|
||||
Both rendered decisions are facts in the consolidated input:
|
||||
|
||||
- fact: `Im Zweifel gehen Projekte durch.`
|
||||
- Consolidated evidence: `consolidated_extractions.json:775` to `:784`
|
||||
- Rendered as decision: `working_protocol.md:57`
|
||||
- fact: `Ein Projekt kann abgelehnt werden, aber es muss sichergestellt sein,
|
||||
dass relevante Projekte nicht weggeschmissen werden.`
|
||||
- Consolidated evidence: `consolidated_extractions.json:755` to `:764`
|
||||
- Rendered as decision: `working_protocol.md:58`
|
||||
|
||||
This is the clearest category change in the renderer output.
|
||||
|
||||
### Rendered Section: Projektkategorisierung und Status / Action Items
|
||||
|
||||
Status: omitted.
|
||||
|
||||
Rendered line:
|
||||
|
||||
- `working_protocol.md:62`: `*Keine.*`
|
||||
|
||||
The consolidated input contains at least one action item that belongs to this
|
||||
topic area:
|
||||
|
||||
- `action_item_0008`: `Diskussion über die Relevanz eines speziellen Marktes
|
||||
(z.B. Turkmenistan) führen.`
|
||||
|
||||
The renderer omitted it.
|
||||
|
||||
### Rendered Section: Projektkategorisierung und Status / Open Questions
|
||||
|
||||
Status: faithful to rendered grouping, but unfaithful to complete input.
|
||||
|
||||
Rendered line:
|
||||
|
||||
- `working_protocol.md:66`: `*Keine.*`
|
||||
|
||||
The complete consolidated input still contains project-process open questions,
|
||||
including omitted `open_question_0008` and `open_question_0014`. They were not
|
||||
rendered elsewhere.
|
||||
|
||||
### Rendered Section: Technische Details und Infrastruktur / Background
|
||||
|
||||
Status: modified.
|
||||
|
||||
Rendered line:
|
||||
|
||||
- `working_protocol.md:72`
|
||||
|
||||
This condenses facts and technical details:
|
||||
|
||||
- fact: `Das Netzwerk bei Gemeinden ist momentan langsam.`
|
||||
- `technical_detail_0006`: `Das Netzwerk ist langsam (wie zu ISDN-Zeiten).`
|
||||
- `technical_detail_0007`: `Der Rechner schmiert ab...`
|
||||
- fact: `Das Dokument ist das aktuelle Projektdeckblatt...`
|
||||
- fact: `Es gibt einen Reporting-Zyklus.`
|
||||
- fact: `Sekretariate sind sehr abweisend...`
|
||||
|
||||
The problem is not invention; the problem is that `decision_0004`
|
||||
(`Festlegung eines Reporting-Zyklus...`) is downgraded from decision to
|
||||
background.
|
||||
|
||||
### Rendered Section: Technische Details und Infrastruktur / Decisions
|
||||
|
||||
Status: omitted.
|
||||
|
||||
Rendered line:
|
||||
|
||||
- `working_protocol.md:76`: `*Keine.*`
|
||||
|
||||
The consolidated input contains `decision_0004`, which is about a reporting
|
||||
cycle and is rendered only as background.
|
||||
|
||||
### Rendered Section: Technische Details und Infrastruktur / Action Items
|
||||
|
||||
Status: faithful to rendered items.
|
||||
|
||||
Rendered lines:
|
||||
|
||||
- `working_protocol.md:80`
|
||||
- `working_protocol.md:81`
|
||||
|
||||
These correspond to:
|
||||
|
||||
- `action_item_0013`
|
||||
- `action_item_0014`
|
||||
|
||||
The semantic quality of `action_item_0014` is questionable, but that issue is
|
||||
upstream because the consolidated input already categorizes it as an action
|
||||
item.
|
||||
|
||||
### Rendered Section: Technische Details und Infrastruktur / Open Questions
|
||||
|
||||
Status: omitted relative to complete input.
|
||||
|
||||
Rendered line:
|
||||
|
||||
- `working_protocol.md:85`: `*Keine.*`
|
||||
|
||||
No technical open questions are clearly required here, but the overall rendered
|
||||
protocol still omits consolidated open questions elsewhere.
|
||||
|
||||
## Exact Mismatches
|
||||
|
||||
| Consolidated item | Category before | Rendered text | Category after | Result |
|
||||
| --- | --- | --- | --- | --- |
|
||||
| `decision_0004`: `Festlegung eines Reporting-Zyklus...` | decision | `Es gibt einen Reporting-Zyklus...` | background | category weakened |
|
||||
| `decision_0007`: `...Sitzung mit fünf Leuten...` | decision | none | omitted | omitted |
|
||||
| `decision_0008`: `...Prozessrahmen nutzen...` | decision | none | omitted | omitted |
|
||||
| `decision_0010`: `...Verbesserungsvorschläge... ausgeschleust...` | decision | none | omitted | omitted |
|
||||
| `decision_0012`: `Die Diskussion wird beendet...` | decision | none | omitted | omitted |
|
||||
| `open_question_0008`: `Wer soll die Rolle des Gatekeepers übernehmen...` | open_question | none | omitted | omitted |
|
||||
| `open_question_0014`: `Welche Prüfsteine sind relevant...` | open_question | none | omitted | omitted |
|
||||
| fact: `Im Zweifel gehen Projekte durch.` | fact | `Im Zweifel gehen Projekte durch.` | decision | promoted |
|
||||
| fact: `Ein Projekt kann abgelehnt werden...` | fact | same meaning | decision | promoted |
|
||||
| `action_item_0008`: `Diskussion über die Relevanz... Turkmenistan...` | action_item | none | omitted | omitted |
|
||||
| `action_item_0010`: `Filterkriterien für verschiedene Fälle... definieren.` | action_item | overlapped by decision/action item text | merged / not preserved as action item | modified |
|
||||
|
||||
## Invented Content Check
|
||||
|
||||
No renderer-only invented people were verified in this comparison. `Guido` and
|
||||
`Noah` appear in the rendered protocol, but they are already present in the
|
||||
consolidated input. This belongs to BUG-003, not BUG-006.
|
||||
|
||||
No clear renderer-only invented responsibility was verified in this comparison.
|
||||
Questionable todos in the rendered protocol are already categorized as action
|
||||
items in the consolidated input.
|
||||
|
||||
## Merged And Split Items
|
||||
|
||||
Verified merges:
|
||||
|
||||
- Multiple facts and decisions are merged into background paragraphs in
|
||||
`working_protocol.md:7`, `:9`, `:53` and `:72`.
|
||||
- `action_item_0010` is not preserved as a separate action item and overlaps
|
||||
with rendered filter-criteria decision/action-item wording.
|
||||
|
||||
Verified splits:
|
||||
|
||||
- No clear split from one consolidated item into multiple materially different
|
||||
rendered items was identified.
|
||||
|
||||
## Earliest Point Where The Mismatch Occurs
|
||||
|
||||
The earliest point is the Working Protocol Renderer LLM response.
|
||||
|
||||
Evidence:
|
||||
|
||||
- `metadata.json` records renderer input as
|
||||
`semantic_consolidator/consolidated_extractions.json` and prompt file
|
||||
`prompts/working_protocol.md`.
|
||||
- `raw_model_response.txt` already contains the omissions and category
|
||||
changes.
|
||||
- `raw_model_response.txt` and `working_protocol.md` are byte-identical.
|
||||
|
||||
Therefore the mismatch is not introduced by renderer post-processing, output
|
||||
normalization or Markdown generation.
|
||||
|
||||
## Verified Root Cause
|
||||
|
||||
Verified cause:
|
||||
|
||||
The single-pass LLM Working Protocol Renderer does not reliably preserve the
|
||||
category and coverage constraints of the consolidated input. It generates a
|
||||
Markdown view that condenses and reorganizes the input, but the generated raw
|
||||
response omits required items and changes some categories.
|
||||
|
||||
Cause classification:
|
||||
|
||||
- Renderer prompt: not proven as the sole root cause. The prompt includes
|
||||
explicit preservation rules, but also includes condensation and redundancy
|
||||
instructions. The evidence proves the model output violates the preservation
|
||||
rules; it does not prove which prompt sentence caused the violation.
|
||||
- Renderer post-processing: disproved. Raw response and final Markdown are
|
||||
identical.
|
||||
- Renderer output normalization: disproved. No separate normalization changed
|
||||
the raw response.
|
||||
- Markdown generation: disproved. The final file is the raw Markdown response.
|
||||
- Another verified cause: single-pass LLM renderer generation violates
|
||||
structural faithfulness requirements.
|
||||
|
||||
## Confidence
|
||||
|
||||
High for the location of the failure and for disproving post-processing,
|
||||
normalization and Markdown generation as causes.
|
||||
|
||||
Medium for root-cause classification beyond that. The evidence identifies the
|
||||
renderer LLM response as the failing point, but does not isolate one prompt
|
||||
sentence as the cause.
|
||||
|
||||
## Conceptual Repair Strategy
|
||||
|
||||
Do not attempt to repair this with free-form post-processing.
|
||||
|
||||
Conceptually, renderer faithfulness needs a structural coverage contract:
|
||||
|
||||
- every consolidated decision must be accounted for as a decision
|
||||
- every consolidated action item must be accounted for as an action item
|
||||
- every consolidated open question must be accounted for as an open question
|
||||
- facts may appear in background, but must not be promoted to decisions
|
||||
- omissions and category changes should be validator-detectable before the
|
||||
protocol is accepted
|
||||
|
||||
The renderer output should be validated against the consolidated input by item
|
||||
ID or another stable structured reference, rather than relying on prose-only
|
||||
Markdown to preserve category semantics implicitly.
|
||||
@@ -0,0 +1,348 @@
|
||||
# Constraint Repair Engine V1
|
||||
|
||||
Constraint Repair Engine V1 is a reusable structural repair stage for JSON
|
||||
outputs produced by local LLM pipeline steps.
|
||||
|
||||
It exists for cases where a model produced semantically usable JSON, but a
|
||||
strict validator rejected the document because a structural invariant was
|
||||
violated. Typical examples are repeated identifiers, missing identifiers, empty
|
||||
containers or unstable item order.
|
||||
|
||||
The repair stage is intentionally not integrated into the production pipeline
|
||||
yet. It is a standalone module that future stages can opt into explicitly.
|
||||
|
||||
## Architecture
|
||||
|
||||
Package:
|
||||
|
||||
```text
|
||||
src/meeting_lab/constraint_repair/
|
||||
```
|
||||
|
||||
The package contains:
|
||||
|
||||
- `engine.py`: generic JSON repair logic and the generic repair prompt
|
||||
- `adapters/`: thin adapters from stage-specific validator results to the
|
||||
generic validator report format
|
||||
|
||||
The engine receives exactly two logical inputs:
|
||||
|
||||
1. the complete original JSON output
|
||||
2. a machine-generated validator report
|
||||
|
||||
It never reads the original transcript and never receives Meeting Context or
|
||||
other source material. The engine is domain-neutral: it operates on JSON
|
||||
pointers, list keys and validator-reported identifiers.
|
||||
|
||||
## Validator Interface
|
||||
|
||||
The generic validator report has this shape:
|
||||
|
||||
```json
|
||||
{
|
||||
"valid": false,
|
||||
"violations": [
|
||||
{
|
||||
"type": "duplicate_id",
|
||||
"collection_pointer": "/groups",
|
||||
"id_list_key": "source_item_ids",
|
||||
"id": "item_0001",
|
||||
"occurrences": [
|
||||
{"item_index": 0, "id_index": 1},
|
||||
{"item_index": 3, "id_index": 0}
|
||||
],
|
||||
"keep_occurrence": 0
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
Supported V1 violation types:
|
||||
|
||||
- `duplicate_id`: remove repeated identifier occurrences from a list field
|
||||
- `missing_id`: restore an identifier by appending a validator-provided item
|
||||
template or adding it to a validator-specified existing item
|
||||
- `empty_group`: remove an item whose identifier list is empty
|
||||
- `reorder_items`: reorder a collection by validator-provided keys
|
||||
|
||||
The report must provide enough structural information for the engine to repair
|
||||
the document without interpreting content.
|
||||
|
||||
## Meeting Context Constraint Validation Extension
|
||||
|
||||
Meeting Context constraints can be added without making the repair engine
|
||||
Meeting Context aware.
|
||||
|
||||
The validator may inspect a stage output together with authoritative
|
||||
`meeting_context.yaml`. It then converts violations into the generic validator
|
||||
report format. The Constraint Repair Engine still receives only the original
|
||||
JSON document and the validator report.
|
||||
|
||||
Authoritative metadata is configuration, not meeting content. Examples include:
|
||||
|
||||
- participant attendance
|
||||
- canonical participant identity
|
||||
- aliases
|
||||
- departments
|
||||
- roles
|
||||
|
||||
Authoritative metadata may be consumed by LLMs and downstream pipeline stages.
|
||||
It must never be redefined, overwritten or inferred by any pipeline component.
|
||||
|
||||
This includes, but is not limited to:
|
||||
|
||||
- LLM extraction
|
||||
- Canonicalizer
|
||||
- Semantic Consolidator
|
||||
- Constraint Repair
|
||||
- Renderer
|
||||
|
||||
These components may consume authoritative metadata, but they must treat it as
|
||||
immutable configuration.
|
||||
|
||||
### Constraint Types
|
||||
|
||||
`participant_attendance_conflict`
|
||||
|
||||
A known participant from Meeting Context is represented in structured output
|
||||
as absent, mentioned-only, external or otherwise not attending, contradicting
|
||||
the authoritative attendance status.
|
||||
|
||||
Example:
|
||||
|
||||
```json
|
||||
{
|
||||
"valid": false,
|
||||
"constraint_source": {
|
||||
"type": "meeting_context",
|
||||
"meeting_id": "2026-07-27-projektprozess",
|
||||
"source_file": "samples/real_live/project_process_meeting/meeting_context.yaml",
|
||||
"schema_version": "1"
|
||||
},
|
||||
"violations": [
|
||||
{
|
||||
"type": "participant_attendance_conflict",
|
||||
"severity": "error",
|
||||
"collection_pointer": "/participants",
|
||||
"item_index": 3,
|
||||
"entity_id": "bjoern",
|
||||
"display_name": "Björn",
|
||||
"matched_alias": "Björn",
|
||||
"field": "attendance_status",
|
||||
"actual_value": "absent",
|
||||
"expected_value": "present",
|
||||
"repair": {
|
||||
"operation": "set_field",
|
||||
"field": "attendance_status",
|
||||
"value": "present"
|
||||
}
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
`participant_unknown`
|
||||
|
||||
A person-like entity appears in a structured participant or mentioned-person
|
||||
field but is not known in Meeting Context as a participant, mentioned person or
|
||||
alias.
|
||||
|
||||
Example:
|
||||
|
||||
```json
|
||||
{
|
||||
"type": "participant_unknown",
|
||||
"severity": "error",
|
||||
"collection_pointer": "/participants",
|
||||
"item_index": 4,
|
||||
"matched_text": "Guido",
|
||||
"repair": {
|
||||
"operation": "remove_item"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
Removal is allowed only when the unknown entity is a standalone structured
|
||||
participant-like item and removing it does not remove unrelated content. If the
|
||||
unknown name appears only inside generated prose, the repair must fail closed.
|
||||
|
||||
`participant_alias_conflict`
|
||||
|
||||
A structured entity reference uses an alias that maps to a different
|
||||
authoritative entity, or uses an ambiguous alias that cannot be resolved to
|
||||
exactly one Meeting Context entity.
|
||||
|
||||
Example:
|
||||
|
||||
```json
|
||||
{
|
||||
"type": "participant_alias_conflict",
|
||||
"severity": "error",
|
||||
"json_pointer": "/participants/2/entity_id",
|
||||
"matched_alias": "Johanna",
|
||||
"expected_entity_id": "jovana",
|
||||
"expected_display_name": "Jovana",
|
||||
"actual_entity_id": "unknown",
|
||||
"repair": {
|
||||
"operation": "set_field",
|
||||
"field": "entity_id",
|
||||
"value": "jovana"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
Alias repair is allowed only for structured identity fields and only when
|
||||
Meeting Context maps the alias to exactly one entity.
|
||||
|
||||
### Validation Principle
|
||||
|
||||
The validator validates structured data whenever possible. Lexical analysis is
|
||||
a fallback only when no structured representation exists. The long-term
|
||||
objective is to reduce lexical validation over time by preserving Meeting
|
||||
Context metadata as structured data throughout the pipeline.
|
||||
|
||||
Lexical fallback may report a violation, but it must not create repair
|
||||
instructions that require editing generated prose.
|
||||
|
||||
### Repair Workflow
|
||||
|
||||
```text
|
||||
Stage output
|
||||
↓
|
||||
Meeting Context Constraint Validator
|
||||
↓
|
||||
Generic validator report
|
||||
↓
|
||||
Constraint Repair Engine
|
||||
↓
|
||||
Meeting Context Constraint Validator
|
||||
```
|
||||
|
||||
If Validator #1 succeeds, repair is skipped. If Validator #1 fails, exactly one
|
||||
repair pass may run when all violations map to deterministic structural
|
||||
operations. Validator #2 then checks the repaired document. If Validator #2
|
||||
fails, the pipeline stops and reports the remaining violations.
|
||||
|
||||
### Allowed Repairs
|
||||
|
||||
Allowed repairs are deterministic structural operations only:
|
||||
|
||||
- set a structured attendance field to the authoritative value
|
||||
- set a structured entity ID to the authoritative entity ID
|
||||
- move a structured participant item between participant collections
|
||||
- remove a standalone structured unknown participant item
|
||||
- remove empty structured participant containers
|
||||
- reorder structured participant collections deterministically
|
||||
|
||||
### Forbidden Repairs
|
||||
|
||||
The repair engine must never perform free-form text editing.
|
||||
|
||||
It must not:
|
||||
|
||||
- remove names from generated prose
|
||||
- replace text inside prose fields
|
||||
- rewrite sentences
|
||||
- infer attendance from transcript content
|
||||
- invent participants
|
||||
- invent aliases
|
||||
- merge people
|
||||
- split people
|
||||
- change responsibility attribution
|
||||
- change fact, decision, action-item or open-question meaning
|
||||
- use Meeting Context directly
|
||||
|
||||
If a violation exists only inside generated prose and cannot be repaired by a
|
||||
deterministic structural operation, the repair must fail closed and report the
|
||||
violation.
|
||||
|
||||
### Integration Strategy
|
||||
|
||||
Minimum useful integration for the current architecture:
|
||||
|
||||
```text
|
||||
Semantic Consolidator
|
||||
↓
|
||||
Meeting Context Constraint Validator
|
||||
↓
|
||||
Constraint Repair
|
||||
↓
|
||||
Meeting Context Constraint Validator
|
||||
↓
|
||||
Renderer
|
||||
```
|
||||
|
||||
This catches contradictions in the consolidated representation before the
|
||||
Working Protocol Renderer receives it. It does not require the renderer to
|
||||
infer attendance semantics from prose.
|
||||
|
||||
Future architecture:
|
||||
|
||||
Meeting Context metadata should remain structured throughout the pipeline so
|
||||
that the Renderer receives authoritative participant metadata directly instead
|
||||
of having to infer it from generated prose. In that architecture, Meeting
|
||||
Context validation can operate primarily on structured fields and use lexical
|
||||
analysis only as a diagnostic fallback.
|
||||
|
||||
## Allowed Operations
|
||||
|
||||
The engine may:
|
||||
|
||||
- move identifiers
|
||||
- remove duplicate identifiers
|
||||
- restore missing identifiers
|
||||
- remove empty containers
|
||||
- reorder items
|
||||
|
||||
The engine must not:
|
||||
|
||||
- invent information
|
||||
- rewrite extracted text
|
||||
- reinterpret reasons or explanations
|
||||
- create new semantic relationships
|
||||
- split semantic relationships
|
||||
|
||||
Missing identifier repair is intentionally conservative. If the validator does
|
||||
not provide an explicit target item or an item template, the engine refuses the
|
||||
repair.
|
||||
|
||||
## Generic Repair Prompt
|
||||
|
||||
The module defines a generic prompt contract for future model-backed repair
|
||||
backends. The prompt describes the task as repairing a structured JSON document
|
||||
from a validator report and deliberately avoids stage-specific vocabulary.
|
||||
|
||||
V1 unit tests assert that the prompt does not mention the current consolidation
|
||||
stage, Meeting Context, facts, or merge groups.
|
||||
|
||||
## Semantic Consolidator Adapter
|
||||
|
||||
`constraint_repair.adapters.semantic_consolidator` converts the current
|
||||
consolidation validator shape into the generic validator report.
|
||||
|
||||
The adapter knows the current consolidation output field names such as
|
||||
`groups` and `source_item_ids`. The generic repair engine does not. This keeps
|
||||
stage-specific schema knowledge at the edge and preserves the repair engine as
|
||||
a reusable JSON utility.
|
||||
|
||||
The adapter currently reports:
|
||||
|
||||
- repeated source IDs
|
||||
- missing expected source IDs
|
||||
- empty source-ID containers
|
||||
|
||||
It does not call Ollama and does not change the production consolidation path.
|
||||
|
||||
## Future Reuse
|
||||
|
||||
Future pipeline stages can reuse the same engine by writing a small adapter
|
||||
that maps their validator failures to the generic report format. The required
|
||||
contract is that the adapter reports structural locations and does not ask the
|
||||
engine to infer domain meaning.
|
||||
|
||||
Potential future uses:
|
||||
|
||||
- enforcing exact identifier coverage in LLM-generated grouping output
|
||||
- removing empty generated containers before strict parsing
|
||||
- restoring validator-known singleton items
|
||||
- normalizing deterministic order after otherwise valid generation
|
||||
@@ -266,6 +266,90 @@ renderer integration remains planned.
|
||||
|
||||
---
|
||||
|
||||
# Entity Registry
|
||||
|
||||
Accepted Architecture. Implementation deferred.
|
||||
|
||||
The Entity Registry is the persistent cross-meeting knowledge source for
|
||||
confirmed entities, aliases and organizational metadata. It is independent from
|
||||
individual meetings and is the planned long-term source used to prepare Meeting
|
||||
Context V2.
|
||||
|
||||
Entity types include:
|
||||
|
||||
- people
|
||||
- organizations
|
||||
- departments
|
||||
- products
|
||||
- projects
|
||||
- locations
|
||||
- abbreviations
|
||||
|
||||
Each entity has a stable internal identifier. The displayed name may change
|
||||
over time, but the internal identifier must remain stable.
|
||||
|
||||
Conceptual shape:
|
||||
|
||||
```json
|
||||
{
|
||||
"entity_id": "person_0001",
|
||||
"entity_type": "person",
|
||||
"display_name": "Jovana",
|
||||
"aliases": [
|
||||
"Jovana",
|
||||
"Giovanna",
|
||||
"Jovanna",
|
||||
"Giovana"
|
||||
],
|
||||
"status": "confirmed"
|
||||
}
|
||||
```
|
||||
|
||||
The registry never learns automatically. It may propose matches, but only
|
||||
confirmed user actions update it. Similarity search may suggest spelling
|
||||
variants, Whisper transcription variants, umlaut variants or OCR-like mistakes,
|
||||
but suggestions require explicit confirmation.
|
||||
|
||||
Previously unseen names should be classified by the user as one of:
|
||||
|
||||
- meeting participant
|
||||
- mentioned person
|
||||
- external person
|
||||
- transcription error
|
||||
- ignore
|
||||
|
||||
The Entity Registry must not infer responsibility, decisions, attendance or
|
||||
ownership.
|
||||
|
||||
---
|
||||
|
||||
# Meeting Context V2
|
||||
|
||||
Accepted Architecture. Implementation deferred.
|
||||
|
||||
Meeting Context V2 is an authoritative meeting-specific YAML Point of Truth
|
||||
generated or assisted from:
|
||||
|
||||
- Entity Registry
|
||||
- user confirmations
|
||||
- meeting metadata
|
||||
|
||||
The YAML remains the extraction pipeline interface and the authoritative
|
||||
meeting-specific Point of Truth for that meeting run. It is also a reproducible
|
||||
input artifact: changes to the Entity Registry after a meeting run must not
|
||||
silently change the historical Meeting Context used for that run.
|
||||
|
||||
The Entity Registry remains the persistent cross-meeting knowledge source. It
|
||||
must not override explicit meeting-specific confirmations.
|
||||
|
||||
Meeting Context V2 should reduce manual work, improve alias handling, detect
|
||||
transcription errors earlier and make Meeting Context quality scalable across
|
||||
many meetings.
|
||||
|
||||
See `docs/adr-meeting-context-v2-entity-registry.md`.
|
||||
|
||||
---
|
||||
|
||||
# Topic Result
|
||||
|
||||
After extraction, every topic contains the collected information.
|
||||
|
||||
@@ -26,6 +26,48 @@ Supported metadata:
|
||||
- abbreviations
|
||||
- relevant products, projects, systems, locations and technical terms
|
||||
|
||||
## Future Direction: V2
|
||||
|
||||
Accepted Architecture. Implementation deferred.
|
||||
|
||||
Meeting Context V2 should be generated from an interactive entity confirmation
|
||||
workflow after Whisper transcription:
|
||||
|
||||
```text
|
||||
Whisper
|
||||
↓
|
||||
Entity Detection
|
||||
↓
|
||||
User Confirmation
|
||||
↓
|
||||
Entity Registry Update
|
||||
↓
|
||||
Meeting Context Builder
|
||||
↓
|
||||
meeting_context.yaml
|
||||
↓
|
||||
Extraction Pipeline
|
||||
```
|
||||
|
||||
The YAML remains the extraction interface and the authoritative
|
||||
meeting-specific Point of Truth for a meeting run. It should become a
|
||||
meeting-specific snapshot generated or assisted from the Entity Registry, user
|
||||
confirmations and meeting metadata.
|
||||
|
||||
The Entity Registry is the persistent cross-meeting knowledge source for
|
||||
confirmed entities, aliases and organizational metadata. It stores stable
|
||||
internal identifiers and never learns automatically. The Registry must not
|
||||
override explicit meeting-specific confirmations, and Registry changes after a
|
||||
meeting run must not silently change the historical Meeting Context used for
|
||||
that run.
|
||||
|
||||
Unknown names should be explicitly classified by the user as meeting
|
||||
participant, mentioned person, external person, transcription error or ignore.
|
||||
Similarity suggestions for spelling variants, Whisper variants, umlaut
|
||||
handling and OCR-like mistakes require explicit confirmation.
|
||||
|
||||
See `docs/adr-meeting-context-v2-entity-registry.md`.
|
||||
|
||||
## File Locations
|
||||
|
||||
- Generic template: `samples/templates/meeting_context.template.yaml`
|
||||
|
||||
@@ -30,6 +30,8 @@ Canonical Meeting Knowledge and every Output View renderer.
|
||||
|
||||
# Pipeline Overview
|
||||
|
||||
Current analysis pipeline:
|
||||
|
||||
```text
|
||||
Whisper Transcript
|
||||
│
|
||||
@@ -73,6 +75,46 @@ extraction, deterministic prompt injection and minimal extraction JSON
|
||||
provenance. Later Canonicalizer, Semantic Consolidator, Canonical Meeting
|
||||
Knowledge and renderer integration remains future work.
|
||||
|
||||
Accepted future Meeting Context V2 preparation flow:
|
||||
|
||||
```text
|
||||
Whisper
|
||||
│
|
||||
▼
|
||||
Entity Detection
|
||||
│
|
||||
▼
|
||||
User Confirmation
|
||||
│
|
||||
▼
|
||||
Entity Registry Update
|
||||
│
|
||||
▼
|
||||
Meeting Context Builder
|
||||
│
|
||||
▼
|
||||
meeting_context.yaml
|
||||
│
|
||||
▼
|
||||
Extraction Pipeline
|
||||
```
|
||||
|
||||
This preparation flow is not implemented. It is the accepted long-term
|
||||
direction for reducing manual Meeting Context work while preserving explicit
|
||||
user control. The Entity Registry is the persistent cross-meeting knowledge
|
||||
source for confirmed entities, aliases and organizational metadata. It stores
|
||||
stable internal IDs and confirmed aliases, and never updates itself
|
||||
automatically.
|
||||
|
||||
For each meeting run, `meeting_context.yaml` remains the authoritative
|
||||
meeting-specific Point of Truth and reproducible input artifact consumed by the
|
||||
pipeline. V2 changes how that artifact is prepared: it may be generated or
|
||||
assisted from Registry data, user confirmations and meeting metadata. The
|
||||
Registry must not override explicit meeting-specific confirmations, and
|
||||
Registry changes after a meeting run must not silently change the historical
|
||||
Meeting Context used for that run. Similarity suggestions and unknown entity
|
||||
classifications require explicit user confirmation.
|
||||
|
||||
---
|
||||
|
||||
# Stage 1 – Normalization
|
||||
|
||||
@@ -0,0 +1,54 @@
|
||||
# Quality Readiness
|
||||
|
||||
Meeting Lab quality status should distinguish engineering readiness from
|
||||
practical usability.
|
||||
|
||||
## Engineering Readiness
|
||||
|
||||
Status: NOT READY
|
||||
|
||||
The current end-to-end pipeline is not ready to replace the previous
|
||||
extraction pipeline. The latest quality milestone records seven known
|
||||
regression bugs, two documented root-cause investigations and a renderer
|
||||
faithfulness failure where consolidated items are omitted or reclassified.
|
||||
|
||||
Engineering readiness requires at least:
|
||||
|
||||
- faithful rendering of consolidated decisions, action items and open questions
|
||||
- no false responsibility attribution
|
||||
- no participant-attendance contradictions against Meeting Context
|
||||
- no unverified person/entity names entering final protocols unchecked
|
||||
- regression tests or validation coverage for fixed bugs
|
||||
|
||||
## Practical Usability
|
||||
|
||||
Status: READY FOR MANUAL EDIT
|
||||
|
||||
The generated working protocol can still be useful as a draft when reviewed by
|
||||
a human editor against the known meeting content. This status does not imply
|
||||
engineering readiness and must not be used as evidence that the pipeline is
|
||||
faithful or production-ready.
|
||||
|
||||
Practical usability means:
|
||||
|
||||
- the broad meeting structure is recognizable
|
||||
- many relevant topics and process points are present
|
||||
- manual correction is still required before the protocol can be trusted
|
||||
|
||||
Known manual correction areas include:
|
||||
|
||||
- false or over-strong responsibility attribution
|
||||
- incorrect open questions
|
||||
- synthetic or unverified names from transcript artifacts
|
||||
- participant-attendance contradictions
|
||||
- reversed or over-broad process meaning
|
||||
|
||||
## Current Benchmark Reference
|
||||
|
||||
Latest benchmark reviewed for this distinction:
|
||||
|
||||
- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/benchmark_report.md`
|
||||
|
||||
The benchmark report concludes `NOT READY` for engineering replacement. This
|
||||
document adds the separate practical-usability classification for manual-edit
|
||||
workflows only.
|
||||
@@ -0,0 +1,478 @@
|
||||
# Regression Bug Tracker
|
||||
|
||||
Living tracker for real bugs discovered during end-to-end Meeting Lab
|
||||
evaluation.
|
||||
|
||||
Purpose:
|
||||
|
||||
- prevent forgotten regressions
|
||||
- document known root causes
|
||||
- record fixes when they happen
|
||||
- verify that bugs never silently return
|
||||
|
||||
## BUG-001
|
||||
|
||||
ID: BUG-001
|
||||
|
||||
Title: Renderer reverses the intended process direction
|
||||
|
||||
Pipeline stage: Working Protocol Renderer
|
||||
|
||||
Severity: High
|
||||
|
||||
Status: Open
|
||||
|
||||
Date discovered: 2026-08-03
|
||||
|
||||
Version first observed: `meeting_context_v1/final_protocol_with_context`
|
||||
|
||||
Description:
|
||||
|
||||
The generated working protocol phrases the process direction as if the
|
||||
existing F&E process is broadly taken over for all project types. The intended
|
||||
meaning from the evaluation is more constrained: the existing F&E process is a
|
||||
framework that may be adapted or extended for broader project handling.
|
||||
|
||||
Expected behaviour:
|
||||
|
||||
The protocol should preserve the direction and uncertainty of the source
|
||||
discussion: F&E process/framework adapted for other project types where
|
||||
appropriate, without implying unconditional adoption for all project types.
|
||||
|
||||
Actual behaviour:
|
||||
|
||||
The protocol states that the existing F&E process is generally adopted for all
|
||||
project types.
|
||||
|
||||
Likely root cause:
|
||||
|
||||
Unknown.
|
||||
|
||||
Related files:
|
||||
|
||||
- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/working_protocol.md`
|
||||
- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/comparison.md`
|
||||
- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/working_protocol/working_protocol.md`
|
||||
- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/semantic_consolidator/consolidated_extractions.json`
|
||||
- `prompts/working_protocol.md`
|
||||
|
||||
Regression test available (yes/no): no
|
||||
|
||||
Current status:
|
||||
|
||||
Open. Documented from end-to-end evaluation output.
|
||||
|
||||
Notes:
|
||||
|
||||
Do not treat this entry as a prompt-change instruction. It records the observed
|
||||
bug only.
|
||||
|
||||
Current benchmark `meeting_context_v1/e2e_current_20260803_151000` confirms
|
||||
the issue remains. The new protocol still says the existing F&E process is
|
||||
basically taken over for all project types.
|
||||
|
||||
## BUG-002
|
||||
|
||||
ID: BUG-002
|
||||
|
||||
Title: Jovana todo is strengthened beyond the meeting content
|
||||
|
||||
Pipeline stage: Extraction / Working Protocol Renderer
|
||||
|
||||
Severity: High
|
||||
|
||||
Status: Open
|
||||
|
||||
Date discovered: 2026-08-03
|
||||
|
||||
Version first observed: `meeting_context_v1/final_protocol_with_context`
|
||||
|
||||
Description:
|
||||
|
||||
The generated output strengthens a proposal or discussion about Jovana
|
||||
assembling criteria into an assigned todo. The evaluation identified this as
|
||||
overstating the meeting content.
|
||||
|
||||
Expected behaviour:
|
||||
|
||||
The pipeline should preserve the weaker source meaning unless the transcript
|
||||
explicitly assigns, accepts or confirms responsibility.
|
||||
|
||||
Actual behaviour:
|
||||
|
||||
The consolidated input and final protocol include an action item assigning
|
||||
criterion compilation to Jovana.
|
||||
|
||||
Likely root cause:
|
||||
|
||||
Unknown.
|
||||
|
||||
Related files:
|
||||
|
||||
- `samples/benchmarks/meeting_context_v1/semantic_consolidator_repair_v1/consolidated_extractions.json`
|
||||
- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/working_protocol.md`
|
||||
- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/comparison.md`
|
||||
- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/semantic_consolidator/consolidated_extractions.json`
|
||||
- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/working_protocol/working_protocol.md`
|
||||
- `prompts/todos.md`
|
||||
- `prompts/working_protocol.md`
|
||||
|
||||
Regression test available (yes/no): no
|
||||
|
||||
Current status:
|
||||
|
||||
Open. Documented from end-to-end evaluation output.
|
||||
|
||||
Notes:
|
||||
|
||||
This bug is subject to the responsibility attribution invariant in
|
||||
`AGENTS.md`.
|
||||
|
||||
Current benchmark `meeting_context_v1/e2e_current_20260803_151000` confirms
|
||||
the issue remains. The consolidated output still contains `Zusammenstellen der
|
||||
Kriterien durch Jovana` with `responsible: "Jovana"`, and the final protocol
|
||||
renders it as an action item.
|
||||
|
||||
## BUG-003
|
||||
|
||||
ID: BUG-003
|
||||
|
||||
Title: Unexpected synthetic person names appear in generated protocols
|
||||
|
||||
Pipeline stage: Extraction / Working Protocol Renderer
|
||||
|
||||
Severity: Medium
|
||||
|
||||
Status: Root Cause Identified
|
||||
|
||||
Date discovered: 2026-08-03
|
||||
|
||||
Version first observed: `meeting_context_v1/final_protocol_with_context`
|
||||
|
||||
Description:
|
||||
|
||||
The generated protocol contains unexpected or risky names/aliases such as
|
||||
`Guido` and `Noah`. These names were flagged during evaluation as synthetic or
|
||||
inconsistent in the generated protocol context.
|
||||
|
||||
Expected behaviour:
|
||||
|
||||
The protocol should use only names supported by the transcript, consolidated
|
||||
input or Meeting Context, and should preserve uncertainty when a name or alias
|
||||
is unclear.
|
||||
|
||||
Actual behaviour:
|
||||
|
||||
The final protocol includes names or name variants that were flagged as
|
||||
unexpected in evaluation.
|
||||
|
||||
Likely root cause:
|
||||
|
||||
The unexpected names already exist in the Whisper transcript before extraction.
|
||||
They are propagated through the pipeline and are not introduced by prompts,
|
||||
gold tests or the renderer.
|
||||
|
||||
Related files:
|
||||
|
||||
- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/working_protocol.md`
|
||||
- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/comparison.md`
|
||||
- `samples/benchmarks/meeting_context_v1/semantic_consolidator_repair_v1/consolidated_extractions.json`
|
||||
- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/working_protocol/working_protocol.md`
|
||||
- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/semantic_consolidator/consolidated_extractions.json`
|
||||
- `samples/real_live/project_process_meeting/meeting_context.yaml`
|
||||
|
||||
Regression test available (yes/no): no
|
||||
|
||||
Current status:
|
||||
|
||||
Root Cause Identified. Planned resolution is deferred until Meeting Context V2
|
||||
/ Entity Registry implementation.
|
||||
|
||||
Notes:
|
||||
|
||||
Resolution strategy:
|
||||
|
||||
Introduce an interactive entity verification step after Whisper transcription:
|
||||
|
||||
```text
|
||||
Whisper
|
||||
↓
|
||||
Entity Detection
|
||||
↓
|
||||
User Verification
|
||||
↓
|
||||
Meeting Context Builder
|
||||
↓
|
||||
Extraction
|
||||
```
|
||||
|
||||
Expected effect:
|
||||
|
||||
Unknown or suspicious person names are detected before extraction begins and
|
||||
can be classified by the user as participant, mentioned person, external
|
||||
person, transcription error or ignore.
|
||||
|
||||
Implementation status:
|
||||
|
||||
Architecture accepted. Implementation deferred.
|
||||
|
||||
Regression test:
|
||||
|
||||
Not yet possible before Meeting Context V2 exists.
|
||||
|
||||
Current benchmark `meeting_context_v1/e2e_current_20260803_151000` confirms
|
||||
the issue remains. `Guido` and `Noah` both appear in the generated working
|
||||
protocol.
|
||||
|
||||
## BUG-004
|
||||
|
||||
ID: BUG-004
|
||||
|
||||
Title: Rhetorical question extracted as an open question
|
||||
|
||||
Pipeline stage: Extraction
|
||||
|
||||
Severity: Medium
|
||||
|
||||
Status: Open
|
||||
|
||||
Date discovered: 2026-08-03
|
||||
|
||||
Version first observed: `meeting_context_v1/final_protocol_with_context`
|
||||
|
||||
Description:
|
||||
|
||||
A rhetorical or discussion-framing question is extracted and later rendered as
|
||||
an open question. This makes the protocol imply that the meeting left a real
|
||||
follow-up question unresolved.
|
||||
|
||||
Expected behaviour:
|
||||
|
||||
Only genuine unresolved questions should be extracted as open questions.
|
||||
Rhetorical questions or conversational framing should not become protocol open
|
||||
questions.
|
||||
|
||||
Actual behaviour:
|
||||
|
||||
The final protocol includes at least one open question identified during
|
||||
evaluation as rhetorical rather than genuinely open.
|
||||
|
||||
Likely root cause:
|
||||
|
||||
Unknown.
|
||||
|
||||
Related files:
|
||||
|
||||
- `samples/benchmarks/meeting_context_v1/semantic_consolidator_repair_v1/consolidated_extractions.json`
|
||||
- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/working_protocol.md`
|
||||
- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/comparison.md`
|
||||
- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/semantic_consolidator/consolidated_extractions.json`
|
||||
- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/working_protocol/working_protocol.md`
|
||||
- `prompts/questions.md`
|
||||
- `prompts/working_protocol.md`
|
||||
|
||||
Regression test available (yes/no): no
|
||||
|
||||
Current status:
|
||||
|
||||
Open. Documented from end-to-end evaluation output.
|
||||
|
||||
Notes:
|
||||
|
||||
No specific fix has been proposed.
|
||||
|
||||
Current benchmark `meeting_context_v1/e2e_current_20260803_151000` confirms
|
||||
the issue remains. The final protocol still contains discussion-framing or
|
||||
underspecified questions as open protocol questions, including the question
|
||||
about exactly who is responsible for defining filter criteria.
|
||||
|
||||
## BUG-005
|
||||
|
||||
ID: BUG-005
|
||||
|
||||
Title: Renderer treats Björn as absent although Meeting Context marks him present
|
||||
|
||||
Pipeline stage: Working Protocol Renderer
|
||||
|
||||
Severity: Medium
|
||||
|
||||
Status: Investigating
|
||||
|
||||
Date discovered: 2026-08-03
|
||||
|
||||
Version first observed: `meeting_context_v1/final_protocol_with_context`
|
||||
|
||||
Description:
|
||||
|
||||
The generated protocol asks how missing participants such as Jovana and Björn
|
||||
should be included in the process, even though Meeting Context marks Björn as
|
||||
present. This creates a misleading participant-status implication.
|
||||
|
||||
Expected behaviour:
|
||||
|
||||
The protocol should not mark or imply Björn as absent when Meeting Context
|
||||
records him as present. If the source discussion concerns whether a person had
|
||||
reviewed materials or been included in a process, that should not be converted
|
||||
into meeting absence.
|
||||
|
||||
Actual behaviour:
|
||||
|
||||
The final protocol renders an open question about including missing
|
||||
participants, including Björn.
|
||||
|
||||
Likely root cause:
|
||||
|
||||
The extraction model receives Meeting Context but still classifies Björn with
|
||||
missing participants in chunk 08. Downstream artifacts preserve only minimal
|
||||
Meeting Context provenance, not structured participant attendance metadata, so
|
||||
Canonicalizer, Semantic Consolidator and Renderer cannot correct the conflict.
|
||||
|
||||
Related files:
|
||||
|
||||
- `samples/real_live/project_process_meeting/meeting_context.yaml`
|
||||
- `samples/benchmarks/meeting_context_v1/semantic_consolidator_repair_v1/consolidated_extractions.json`
|
||||
- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/working_protocol.md`
|
||||
- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/comparison.md`
|
||||
- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/semantic_consolidator/consolidated_extractions.json`
|
||||
- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/working_protocol/working_protocol.md`
|
||||
- `docs/bug-005-root-cause.md`
|
||||
- `prompts/working_protocol.md`
|
||||
|
||||
Regression test available (yes/no): no
|
||||
|
||||
Current status:
|
||||
|
||||
Investigating. Root-cause evidence has been documented in
|
||||
`docs/bug-005-root-cause.md`, but no fix has been implemented or verified.
|
||||
Current benchmark `meeting_context_v1/e2e_current_20260803_151000` confirms
|
||||
the issue remains.
|
||||
|
||||
Notes:
|
||||
|
||||
This bug is distinct from responsibility attribution. It concerns participant
|
||||
presence/status handling.
|
||||
|
||||
Recommended resolution remains a Meeting Context constraint validator over
|
||||
structured participant metadata before rendering.
|
||||
|
||||
## BUG-006
|
||||
|
||||
ID: BUG-006
|
||||
|
||||
Title: Renderer drops required consolidated items and changes item categories
|
||||
|
||||
Pipeline stage: Working Protocol Renderer
|
||||
|
||||
Severity: High
|
||||
|
||||
Status: Open
|
||||
|
||||
Date discovered: 2026-08-03
|
||||
|
||||
Version first observed: `meeting_context_v1/e2e_current_20260803_151000`
|
||||
|
||||
Description:
|
||||
|
||||
The Working Protocol Renderer does not preserve every consolidated decision and
|
||||
open question even though the renderer prompt requires preserving all
|
||||
decisions, action items and open questions. It also renders some consolidated
|
||||
facts as decisions.
|
||||
|
||||
Expected behaviour:
|
||||
|
||||
Every consolidated decision and open question should be represented in the
|
||||
final protocol unless there is an explicit, validated reason to omit it.
|
||||
Background facts must not be promoted into decisions.
|
||||
|
||||
Actual behaviour:
|
||||
|
||||
The consolidated input contains decisions such as `Festlegung eines
|
||||
Reporting-Zyklus für abgelehnte Projekte`, `Zusammengetragen und beantwortete
|
||||
Fragen führen zur Sitzung mit fünf Leuten...`, and `Die Diskussion wird
|
||||
beendet und die Änderungen werden verschickt...`; these are not preserved as
|
||||
decisions in the final protocol. The consolidated input also contains open
|
||||
questions such as `Wer soll die Rolle des Gatekeepers übernehmen...` and
|
||||
`Welche Prüfsteine sind relevant für die Fachabteilung?`; these are omitted
|
||||
from the final protocol. Conversely, consolidated facts such as `Im Zweifel
|
||||
gehen Projekte durch` and `Ein Projekt kann abgelehnt werden...` are rendered
|
||||
as decisions.
|
||||
|
||||
Likely root cause:
|
||||
|
||||
Unknown.
|
||||
|
||||
Related files:
|
||||
|
||||
- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/semantic_consolidator/consolidated_extractions.json`
|
||||
- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/working_protocol/working_protocol.md`
|
||||
- `prompts/working_protocol.md`
|
||||
|
||||
Regression test available (yes/no): no
|
||||
|
||||
Current status:
|
||||
|
||||
Open. Documented from current end-to-end benchmark output.
|
||||
|
||||
Notes:
|
||||
|
||||
This is a renderer faithfulness problem independent from whether the upstream
|
||||
extracted items are themselves correct.
|
||||
|
||||
## BUG-007
|
||||
|
||||
ID: BUG-007
|
||||
|
||||
Title: Non-committed discussion items are extracted and rendered as todos
|
||||
|
||||
Pipeline stage: Extraction / Working Protocol Renderer
|
||||
|
||||
Severity: High
|
||||
|
||||
Status: Open
|
||||
|
||||
Date discovered: 2026-08-03
|
||||
|
||||
Version first observed: `meeting_context_v1/e2e_current_20260803_151000`
|
||||
|
||||
Description:
|
||||
|
||||
The pipeline extracts and renders several action items that are not clearly
|
||||
assigned, accepted or confirmed commitments in the meeting evidence.
|
||||
|
||||
Expected behaviour:
|
||||
|
||||
Only explicit assignments, accepted responsibilities or confirmed follow-up
|
||||
actions should become todos. Discussion, examples, proposals, broad process
|
||||
needs or unclear alternatives should remain facts/questions or be marked
|
||||
unclear where supported.
|
||||
|
||||
Actual behaviour:
|
||||
|
||||
The current consolidated output and final protocol contain todos such as
|
||||
`Feedback zu den Kriterien und Änderungen am Prozess abwarten`, `Einbinden der
|
||||
Fachbereiche (stellvertretend durch Guido) zur Definition von Prüfsteinen`,
|
||||
and `Mail an Jovana und Björn erneut senden oder Prozessanpassung vornehmen`.
|
||||
The evidence for these items is discussion-level or ambiguous rather than a
|
||||
clear todo commitment.
|
||||
|
||||
Likely root cause:
|
||||
|
||||
Unknown.
|
||||
|
||||
Related files:
|
||||
|
||||
- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/semantic_consolidator/consolidated_extractions.json`
|
||||
- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/working_protocol/working_protocol.md`
|
||||
- `prompts/todos.md`
|
||||
- `prompts/working_protocol.md`
|
||||
|
||||
Regression test available (yes/no): no
|
||||
|
||||
Current status:
|
||||
|
||||
Open. Documented from current end-to-end benchmark output.
|
||||
|
||||
Notes:
|
||||
|
||||
This generalizes BUG-002 beyond the specific Jovana assignment case and should
|
||||
be evaluated against the responsibility attribution invariant.
|
||||
Reference in New Issue
Block a user