Document validation architecture and renderer faithfulness findings

- document Entity Registry and Meeting Context V2 architecture
- preserve meeting_context.yaml as the authoritative meeting-specific input
- define immutable authoritative metadata across all pipeline stages
- restrict Constraint Repair to deterministic structured-data operations
- record BUG-003 root cause and deferred entity-verification resolution
- document BUG-005 attendance-consistency design
- add BUG-006 renderer faithfulness root-cause analysis
- distinguish Engineering Readiness from Practical Usability
- update the persistent regression bug tracker
This commit is contained in:
2026-08-03 16:05:47 +02:00
parent 06f0e7e651
commit 9446c6e0be
13 changed files with 2208 additions and 0 deletions
@@ -0,0 +1,153 @@
# ADR: Meeting Context V2 and Entity Registry
Status: Accepted Architecture
Implementation: Deferred
Date: 2026-08-03
## Context
Meeting Context V1 proved that authoritative meeting context can significantly
improve extraction quality. It helps the extractor normalize known aliases,
identify participants and avoid treating context metadata as evidence for
responsibility or decisions.
End-to-end evaluation also showed that manually writing Meeting Context is not
the right long-term primary workflow. The pipeline needs an earlier,
interactive entity confirmation step after Whisper transcription. That step
should identify candidate entities, ask the user to confirm them and then build
the meeting-specific context from confirmed data.
## Decision
Meeting Context should evolve toward V2 as an authoritative,
meeting-specific Point of Truth generated or assisted from a persistent Entity
Registry, user confirmations and meeting metadata.
Preferred future pipeline:
```text
Whisper
↓
Entity Detection
↓
User Confirmation
↓
Entity Registry Update
↓
Meeting Context Builder
↓
meeting_context.yaml
↓
Extraction Pipeline
```
The Entity Registry is the persistent cross-meeting knowledge source for
confirmed entities, aliases and organizational metadata.
For each individual meeting, `meeting_context.yaml` remains the authoritative
meeting-specific Point of Truth and reproducible input artifact consumed by
the extraction pipeline. Meeting Context V2 changes how this YAML is prepared,
not its authority for a meeting run.
The generated `meeting_context.yaml` is a meeting-specific snapshot. The
Registry must not override explicit meeting-specific confirmations. Changes to
the Registry after a meeting run must not silently change the historical
Meeting Context used for that run.
## Entity Registry
The Entity Registry is the persistent cross-meeting knowledge source and is
independent from individual meetings. It stores confirmed entities such as:
- people
- organizations
- departments
- products
- projects
- locations
- abbreviations
Each entity receives a stable internal identifier. The displayed name may
change over time, but the identifier must remain stable.
## Learning Principle
The registry never learns automatically.
It may propose matches, but only confirmed user actions update the registry.
No autonomous learning is allowed.
## Alias Handling
Aliases are first-class data.
Examples:
```text
Jovana
Giovanna
Jovanna
Giovana
```
These variants may all refer to one confirmed entity. Future runs should
automatically suggest previously confirmed aliases, but those suggestions still
require explicit confirmation when they would update registry data.
## Unknown Entities
Previously unseen names are presented to the user for classification.
Possible classifications:
- meeting participant
- mentioned person
- external person
- transcription error
- ignore
Nothing is automatically accepted.
## Similarity Search
Similarity search is a future extension. It can propose likely matches for:
- spelling variants
- Whisper transcription variants
- umlaut handling
- OCR-like mistakes
Similarity suggestions require explicit confirmation.
## Rationale
Expected advantages:
- significantly less manual work
- earlier detection of transcription errors
- robust alias handling
- reusable organizational knowledge
- improved Meeting Context quality
- easier GUI workflow
- better scalability across many meetings
## Consequences
Meeting Context V1 remains the current implemented interface.
Meeting Context V2 should preserve the YAML interface for extraction.
`meeting_context.yaml` remains the authoritative meeting-specific Point of
Truth and reproducible input artifact for a meeting run. The Entity Registry is
the persistent cross-meeting knowledge source used to prepare that artifact.
The Registry must not override explicit meeting-specific confirmations. Changes
to the Registry after a meeting run must not silently change the historical
Meeting Context used for that run.
The registry must not infer responsibility, decisions, attendance or ownership.
Those still require meeting evidence and remain governed by the existing
responsibility attribution invariant.
No implementation is part of this ADR.
+46
View File
@@ -99,6 +99,8 @@ The analyzer gradually transforms an unstructured discussion into structured kno
# High-Level Pipeline
Current implemented and intended analysis flow:
```text
Transcript
↓
@@ -123,6 +125,24 @@ Output View Rendering
Working Protocol / Distribution Protocol / Knowledge Objects
```
Accepted future Meeting Context V2 preparation flow:
```text
Whisper
↓
Entity Detection
↓
User Confirmation
↓
Entity Registry Update
↓
Meeting Context Builder
↓
meeting_context.yaml
↓
Extraction Pipeline
```
Each stage solves one clearly defined problem.
No module should perform multiple semantic tasks simultaneously.
@@ -204,6 +224,32 @@ The context can help prevent non-participants from being interpreted as
attendees and can normalize known aliases for extraction. It must not infer
roles, departments, responsibilities or decisions.
Accepted future direction:
Meeting Context V2 should be generated or assisted from an interactive entity
confirmation workflow and a persistent Entity Registry. The Entity Registry is
the persistent cross-meeting knowledge source for confirmed people,
organizations, departments, products, projects, locations, aliases and
organizational metadata under stable internal IDs. Display names may change,
but internal IDs remain stable. Aliases are first-class data.
The registry never learns automatically. It may propose matches and aliases,
including spelling variants, Whisper transcription variants, umlaut variants
and OCR-like mistakes, but only explicit user confirmation updates registry
state. Unknown names should be presented to the user as meeting participant,
mentioned person, external person, transcription error or ignore.
In this architecture, `meeting_context.yaml` remains the authoritative
meeting-specific Point of Truth and reproducible input artifact consumed by the
extraction pipeline. It is a meeting-specific snapshot built from the Entity
Registry, user confirmations and meeting metadata. The Registry must not
override explicit meeting-specific confirmations, and Registry changes after a
meeting run must not silently change the historical Meeting Context used for
that run.
Status: Accepted Architecture; implementation deferred. See
`docs/adr-meeting-context-v2-entity-registry.md`.
---
## consolidation/
+265
View File
@@ -0,0 +1,265 @@
# BUG-003 Root Cause Analysis
## Observed Behaviour
The final working protocol contains two names that were flagged as not
belonging to the real meeting:
- `Guido`
- `Noah`
Observed output:
- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/working_protocol.md:36`
contains `Einbinden der Fachbereiche (stellvertretend durch Guido) zur
Definition von Prüfsteinen.`
- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/working_protocol.md:44`
contains `Sollte ein eigener Prozess für Noah (Business Development)
benötigt werden oder reicht der bestehende?`
## Evidence
The names are present in the final renderer raw response as well as
`working_protocol.md`:
- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/raw_model_response.txt:36`
contains `Guido`
- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/raw_model_response.txt:44`
contains `Noah`
They are also present before rendering in the repaired consolidated input:
- `samples/benchmarks/meeting_context_v1/semantic_consolidator_repair_v1/consolidated_extractions.json:1047`
contains the open question text with `Noah`
- `samples/benchmarks/meeting_context_v1/semantic_consolidator_repair_v1/consolidated_extractions.json:1048`
contains evidence `Oder brauchen wir einen eigenen Prozess bei Noah?`
- `samples/benchmarks/meeting_context_v1/semantic_consolidator_repair_v1/consolidated_extractions.json:1537`
contains the action item text with `Guido`
They are present before semantic consolidation in the canonicalized output:
- `samples/benchmarks/meeting_context_v1/canonicalizer/canonicalized_extractions.json:411`
contains the open question text with `Noah`
- `samples/benchmarks/meeting_context_v1/canonicalizer/canonicalized_extractions.json:412`
contains evidence `Oder brauchen wir einen eigenen Prozess bei Noah?`
- `samples/benchmarks/meeting_context_v1/canonicalizer/canonicalized_extractions.json:1401`
contains the action item text with `Guido`
They are present before canonicalization in extraction output and raw model
responses:
- `samples/benchmarks/meeting_context_v1/full_context_run/chunk_02_extraction.json:18`
contains the extracted open question with `Noah`
- `samples/benchmarks/meeting_context_v1/full_context_run/chunk_02_extraction.raw.txt:76`
contains the same raw model output question with `Noah`
- `samples/benchmarks/meeting_context_v1/full_context_run/chunk_07_extraction.json:11`
contains the extracted todo with `Guido`
- `samples/benchmarks/meeting_context_v1/full_context_run/chunk_07_extraction.raw.txt:53`
contains the same raw model output todo with `Guido`
They are present before extraction in the reconstructed normalized chunks:
- `samples/benchmarks/meeting_context_v1/source_reconstruction/chunk_02_normalized.txt:79`
contains `Oder brauchen wir einen eigenen Prozess bei Noah?`
- `samples/benchmarks/meeting_context_v1/source_reconstruction/chunk_07_normalized.txt:213`
contains `Guido`
They are present before normalization in reconstructed raw chunks:
- `samples/benchmarks/meeting_context_v1/source_reconstruction/chunk_02.txt:79`
contains `Oder brauchen wir einen eigenen Prozess bei Noah?`
- `samples/benchmarks/meeting_context_v1/source_reconstruction/chunk_07.txt:213`
contains `Guido`
They are present in the source cleaned Whisper transcript used to reconstruct
the chunks:
- `samples/real_live/project_process_meeting/transcript/meeting_speech_cleaned.json:5783`
has segment text `Oder brauchen wir einen eigenen Prozess bei Noah?`
- `samples/real_live/project_process_meeting/transcript/meeting_speech_cleaned.json:26315`
has segment text `Guido`
Segment-level verification:
- `Noah`: segment `id=182`, `start=899.6400000000001`, `end=901.94`,
text `Oder brauchen wir einen eigenen Prozess bei Noah?`
- `Guido`: segment `id=776`, `start=4157.86`, `end=4160.14`, text `Guido`
The original Whisper transcript also contains both names:
- `samples/real_live/project_process_meeting/transcript/meeting_speech.json:1`
contains both `Noah` and `Guido` in the top-level transcript text and segment
data.
Meeting Context does not contain either name:
- `rg -n "Guido|Noah" samples/real_live/project_process_meeting/meeting_context.yaml`
returned no matches.
Prompt files do not contain either name:
- `rg -n "Guido|Noah" prompts`
returned no matches.
## Earliest Pipeline Stage Containing the Names
The earliest verified pipeline artifact containing the names is the Whisper
transcript stage:
- `samples/real_live/project_process_meeting/transcript/meeting_speech.json`
- `samples/real_live/project_process_meeting/transcript/meeting_speech_cleaned.json`
For the reconstructed benchmark specifically, the earliest input used by the
TODO 4 pipeline is:
- `samples/real_live/project_process_meeting/transcript/meeting_speech_cleaned.json`
The reconstructed chunks preserve the names from that cleaned Whisper JSON.
Extraction, canonicalization, consolidation and rendering propagate them.
## Repository Search Results
Repository search for `Guido|Noah` found these occurrence groups:
- `docs/regression-bugs.md`: tracker entry for BUG-003 mentions both names.
- `tests/gold/responsibility_attribution_negative/transcript.txt`: contains
`Noah` in a synthetic gold scenario.
- `tests/gold/responsibility_attribution_negative/expected.json`: contains
expected `Noah` entries for that synthetic gold scenario.
- `samples/real_live/project_process_meeting/transcript/meeting_speech.json`:
contains `Noah` and `Guido`.
- `samples/real_live/project_process_meeting/transcript/meeting_speech_cleaned.json`:
contains `Noah` and `Guido`.
- `samples/benchmarks/meeting_context_v1/source_reconstruction/chunk_02.txt`
and `chunk_02_normalized.txt`: contain `Noah`.
- `samples/benchmarks/meeting_context_v1/source_reconstruction/chunk_07.txt`
and `chunk_07_normalized.txt`: contain `Guido`.
- `samples/benchmarks/meeting_context_v1/full_context_run/chunk_02_extraction.json`
and `chunk_02_extraction.raw.txt`: contain `Noah`.
- `samples/benchmarks/meeting_context_v1/full_context_run/chunk_07_extraction.json`
and `chunk_07_extraction.raw.txt`: contain `Guido`.
- `samples/benchmarks/meeting_context_v1/canonicalizer/canonicalized_extractions.json`:
contains both names.
- `samples/benchmarks/meeting_context_v1/semantic_consolidator_repair_v1/consolidated_extractions.json`:
contains both names.
- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/raw_model_response.txt`
and `working_protocol.md`: contain both names.
- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/comparison.md`:
mentions `Noah` in the evaluation notes.
Search results did not find `Guido` or `Noah` in:
- `prompts/`
- `samples/real_live/project_process_meeting/meeting_context.yaml`
- `src/`
## Prompt Inspection
Renderer prompt:
- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/metadata.json`
records `prompt_file: "prompts/working_protocol.md"` and
`input_source:
"samples/benchmarks/meeting_context_v1/semantic_consolidator_repair_v1/consolidated_extractions.json"`.
- `prompts/working_protocol.md` contains no few-shot examples and no `Guido` or
`Noah` occurrences.
Extraction prompt assembly:
- `src/meeting_lab/llm/prompts.py` builds extraction prompts from
`prompts/common.md`, optional Meeting Context, explicitly requested task
prompts and the provided transcript.
- `src/meeting_lab/extraction/extract_chunks.py` calls that shared builder with
`decisions.md` and `todos.md` as task prompts.
Semantic Consolidator prompt assembly:
- `src/meeting_lab/consolidation/consolidate_facts.py` builds its prompt from
`prompts/consolidate_facts.md` plus the canonicalized fact payload.
Gold/test prompt inspection:
- `tests/gold/responsibility_attribution_negative/` contains synthetic `Noah`
examples.
- `scripts/run_gold_test.py` is a gold-test runner and reads
`scenario_dir / "transcript.txt"` when explicitly invoked.
- The inspected production metadata for the TODO 4 extraction run records
`input_path` values under
`samples/benchmarks/meeting_context_v1/source_reconstruction/`, not
`tests/gold/`.
- The inspected renderer metadata records only the repaired consolidated JSON
and `prompts/working_protocol.md`.
No evidence was found that few-shot examples, embedded examples, test
transcripts or gold scenarios became part of the production prompt for this
benchmark run.
## Pipeline Verification
The persisted metadata verifies the production execution path:
- `chunk_02_extraction.metadata.json` input:
`samples/benchmarks/meeting_context_v1/source_reconstruction/chunk_02_normalized.txt`
- `chunk_07_extraction.metadata.json` input:
`samples/benchmarks/meeting_context_v1/source_reconstruction/chunk_07_normalized.txt`
- both extraction metadata files record Meeting Context source:
`samples/real_live/project_process_meeting/meeting_context.yaml`
- `canonicalizer/metadata.json` records input directory:
`samples/benchmarks/meeting_context_v1/full_context_run`
- `semantic_consolidator_repair_v1/report.md` records output:
`samples/benchmarks/meeting_context_v1/semantic_consolidator_repair_v1/consolidated_extractions.json`
- `final_protocol_with_context/metadata.json` records renderer input:
`samples/benchmarks/meeting_context_v1/semantic_consolidator_repair_v1/consolidated_extractions.json`
and prompt file `prompts/working_protocol.md`
No execution metadata points to unrelated examples, templates, tests or gold
scenarios.
## Verified Root Cause
The verified root cause for BUG-003 in this pipeline run is upstream transcript
contamination: `Guido` and `Noah` already exist in the Whisper transcript and
the cleaned Whisper transcript used as the benchmark source.
The later stages did not introduce these names from prompts, Meeting Context,
gold tests or renderer examples. They propagated names already present in the
pipeline input:
`meeting_speech.json`
-> `meeting_speech_cleaned.json`
-> reconstructed chunks
-> extraction JSON
-> canonicalized JSON
-> repaired consolidated JSON
-> final working protocol
What cannot be established from repository artifacts alone:
- whether Whisper hallucinated these names from audio
- whether the source audio actually contains words that sound like these names
- what the correct intended tokens should be
Those questions require audio-level or human-transcript verification and are
outside the evidence available in this repository trace.
## Confidence
High
The conclusion that the final protocol names originated before extraction is
directly supported by persisted source transcript, chunk, extraction,
canonicalization, consolidation and renderer artifacts.
The confidence does not extend to identifying the correct replacement words.
That remains unverified.
## Recommended Fix
Conceptually, add a transcript/source-quality validation step before extraction
that flags person or organization names not present in Meeting Context or an
approved alias list. The step should preserve the original transcript text, but
mark suspicious entity mentions for review before they become structured
knowledge and final protocol content.
Do not treat gold scenarios or prompt examples as the root cause for this bug;
the evidence does not support that.
+198
View File
@@ -0,0 +1,198 @@
# BUG-005 Root Cause Analysis
## Observed Behaviour
The final working protocol treats Björn as one of the "fehlende Teilnehmer":
- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/working_protocol.md:53`
This contradicts Meeting Context, which lists Björn as an actual participant
with `attendance_status: "present"`:
- `samples/real_live/project_process_meeting/meeting_context.yaml:45`
- `samples/real_live/project_process_meeting/meeting_context.yaml:52`
## Evidence
Meeting Context marks Björn present:
```text
participant_id: "bjoern"
display_name: "Björn"
role: "Leiter Marketing"
attendance_status: "present"
```
The extraction prompt path does include Meeting Context:
- `src/meeting_lab/extraction/extract_chunks.py:161` renders Meeting Context
for the prompt when it is supplied.
- `src/meeting_lab/models/meeting_context.py:108` renders the heading
`MEETING CONTEXT V1 (AUTHORITATIVE METADATA)`.
- `src/meeting_lab/models/meeting_context.py:112` renders the rule
`The participant list is authoritative.`
- `src/meeting_lab/models/meeting_context.py:128` to
`src/meeting_lab/models/meeting_context.py:132` renders actual
participants.
Execution metadata confirms chunk 08 used the Meeting Context file:
- `samples/benchmarks/meeting_context_v1/full_context_run/chunk_08_extraction.metadata.json:29`
to `:32`
- `samples/benchmarks/meeting_context_v1/full_context_run/chunk_08_extraction.metadata.json:34`
to `:38`
The source chunk does not say Björn was absent from the meeting. It says
Giovanna and Björn had not yet understood or reviewed the discussed mail or
process state:
- `samples/benchmarks/meeting_context_v1/source_reconstruction/chunk_08_normalized.txt:211`
to `:215`
- `samples/benchmarks/meeting_context_v1/source_reconstruction/chunk_08_normalized.txt:229`
to `:245`
- `samples/benchmarks/meeting_context_v1/source_reconstruction/chunk_08_normalized.txt:257`
to `:267`
The earliest generated artifact that treats Björn as absent is the raw chunk
08 extraction response:
- `samples/benchmarks/meeting_context_v1/full_context_run/chunk_08_extraction.raw.txt:4`
summarizes "Einbeziehung fehlender Teilnehmer wie Jovana und Björn".
- `samples/benchmarks/meeting_context_v1/full_context_run/chunk_08_extraction.raw.txt:6`
to `:10` lists participants as only Martin, Lars and Malte.
- `samples/benchmarks/meeting_context_v1/full_context_run/chunk_08_extraction.raw.txt:64`
to `:68` emits the open question
`Wie sollen fehlende Teilnehmer (Jovana, Björn) in den Prozess einbezogen werden?`
The persisted extraction JSON keeps that open question:
- `samples/benchmarks/meeting_context_v1/full_context_run/chunk_08_extraction.json:14`
to `:16`
The persisted extraction JSON records only Meeting Context provenance, not the
participant list or attendance status:
- `samples/benchmarks/meeting_context_v1/full_context_run/chunk_08_extraction.json:23`
to `:27`
The normalizer drops raw response fields such as `participants` and `topics`:
- `src/meeting_lab/extraction/extract_chunks.py:318` to `:365` returns only
normalized `facts`, `decisions`, `todos`, `questions`, `positions` and
`technical`.
- `src/meeting_lab/extraction/extract_chunks.py:432` to `:434` adds only
provenance from Meeting Context after normalization.
Canonicalizer preserves the bad open question:
- `samples/benchmarks/meeting_context_v1/canonicalizer/canonicalized_extractions.json:1633`
to `:1645`
Semantic Consolidator output also preserves the bad open question:
- `samples/benchmarks/meeting_context_v1/semantic_consolidator_repair_v1/consolidated_extractions.json:1669`
to `:1681`
The renderer receives the repaired consolidated JSON as its only recorded
input:
- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/metadata.json:2`
to `:4`
The renderer prompt requires using only the provided input:
- `prompts/working_protocol.md:7` to `:12`
Therefore the renderer did not receive Meeting Context attendance metadata
that would let it distinguish a present participant from a mentioned-only or
absent person.
## Stage-by-Stage Verification
1. Does extraction receive Björn as a present participant?
Yes. The extraction code renders Meeting Context into the prompt when supplied,
and the chunk 08 metadata records
`samples/real_live/project_process_meeting/meeting_context.yaml` as the Meeting
Context source.
2. Does extraction output preserve that information?
No. The raw chunk 08 response lists participants as Martin, Lars and Malte
only, despite Björn being present in Meeting Context. The official persisted
extraction JSON contains only a minimal `context` provenance object and does
not preserve the authoritative participant list or attendance status.
3. Does Canonicalizer preserve it?
No. Canonicalizer input lacks the participant list and attendance status. The
Canonicalizer preserves the already bad open question from chunk 08.
4. Does Semantic Consolidator preserve it?
No. The repaired consolidated JSON lacks Meeting Context participant metadata
and preserves the bad open question.
5. Does the Renderer receive enough information to distinguish present from
mentioned-only participants?
No. Renderer metadata records only
`semantic_consolidator_repair_v1/consolidated_extractions.json` as input, and
that file contains no Meeting Context participant list or attendance status.
## Earliest Failing Pipeline Stage
The earliest failing stage is chunk extraction for
`chunk_08_normalized.txt`.
More precisely, the raw model response for chunk 08 is the first artifact that
both omits Björn from the participant list and groups Björn with absent or
"fehlende" participants.
## Verified Root Cause
The immediate root cause is ignored Meeting Context metadata during extraction
for chunk 08. The extraction model received authoritative Meeting Context, but
interpreted transcript evidence about not having reviewed or understood an
email/process state as meeting absence.
The downstream cause is missing propagation of participant attendance metadata.
After extraction, only Meeting Context provenance is persisted. Canonicalizer,
Semantic Consolidator and Renderer do not receive the authoritative participant
list or attendance status, so they cannot detect or repair the contradiction.
This is not primarily a renderer inference bug. The renderer preserved an open
question already present in its input, and its prompt explicitly restricts it
to the supplied consolidated input.
## Cause Classification
- Missing propagation: verified.
- Ignored metadata: verified at extraction.
- Prompt wording: not verified as the direct cause.
- Renderer inference: disproved as the earliest cause.
- Lost provenance: partially verified; provenance remains, but the substantive
participant metadata is not propagated.
- Another cause: not identified.
## Confidence
High.
The conclusion is supported by Meeting Context, extraction metadata, raw chunk
08 model output, persisted extraction JSON, Canonicalizer output, Semantic
Consolidator output and renderer metadata.
## Recommended Fix
Concept only:
Propagate authoritative Meeting Context participant metadata, including
attendance status, beyond extraction as structured data. Add validation that
flags contradictions where generated items classify an actual participant as
absent or mentioned-only. The validation should run before rendering, and
ideally immediately after extraction so the contradiction is caught at the
earliest stage.
Do not rely on the Working Protocol Renderer to correct attendance semantics
from prose-only consolidated items.
+426
View File
@@ -0,0 +1,426 @@
# BUG-006 Root Cause Analysis
## Observed Behaviour
In benchmark `meeting_context_v1/e2e_current_20260803_151000`, the Working
Protocol Renderer output is not a faithful view of the matching consolidated
representation.
Compared files:
- Consolidated input:
`samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/semantic_consolidator/consolidated_extractions.json`
- Rendered output:
`samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/working_protocol/working_protocol.md`
- Raw renderer response:
`samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/working_protocol/raw_model_response.txt`
Observed problems:
- consolidated decisions are omitted from the rendered `Decisions` sections
- consolidated open questions are omitted from rendered `Open Questions`
sections
- consolidated facts are rendered as decisions
- some consolidated action items are omitted
- some action-item and decision meanings are merged into other rendered
sections
The raw renderer response and `working_protocol.md` are byte-identical
(`cmp` exit code `0`), so the mismatch is already present in the LLM renderer
response. It is not introduced by Markdown file writing.
## Renderer Input And Prompt Evidence
Renderer metadata records the exact input and prompt:
- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/working_protocol/metadata.json:2`
to `:4`
The renderer prompt explicitly says:
- use only the provided input:
`prompts/working_protocol.md:9` to `:10`
- preserve all decisions, action items and open questions:
`prompts/working_protocol.md:24` to `:26`
- remove presentation-level redundancy:
`prompts/working_protocol.md:23`
- do not repeat information merely because it appears in multiple categories:
`prompts/working_protocol.md:65`
- condense everything else:
`prompts/working_protocol.md:75`
The prompt is therefore present and relevant, but the direct observed failure
occurs in the model-generated renderer response.
## Faithfulness Analysis
### Rendered Section: Prozessrahmen und Projektideen-Eingang / Background
Status: modified.
Rendered lines:
- `working_protocol.md:7`
- `working_protocol.md:9`
The section condenses several consolidated facts and decisions into background
paragraphs. Examples:
- `decision_0001` is restated in background and again as a decision.
- `fact`: `Der Prozess ist ein F&E-Prozess gewesen, einfach historisch.`
- `fact`: `Die Bearbeitung der Projekte muss und kann nur in den
Fachabteilungen passieren.`
- `fact`: `Bevor F&E ein Projekt durchdenkt, müssen im Vorfeld Kriterien wie
Marktexistenz und Zulassungen geprüft werden.`
- `decision_0003` and `decision_0005` are partially restated as background.
This is a condensation, not a renderer-only invention.
### Rendered Section: Prozessrahmen und Projektideen-Eingang / Decisions
Status: partially faithful, partially omitted.
Faithfully or near-faithfully preserved:
- `decision_0001` -> `working_protocol.md:13`
- `decision_0002` -> `working_protocol.md:14`
- `decision_0003` -> `working_protocol.md:15`
- `decision_0005` -> `working_protocol.md:16`
- `decision_0006` -> `working_protocol.md:17`
- `decision_0009` -> `working_protocol.md:18`
- `decision_0011` -> `working_protocol.md:19`
Omitted as decisions:
- `decision_0004`: `Festlegung eines Reporting-Zyklus für abgelehnte
Projekte.`
- Consolidated evidence: `consolidated_extractions.json:1217` to `:1227`
- Rendered only as background: `working_protocol.md:72`
- `decision_0007`: `Zusammengetragen und beantwortete Fragen führen zur
Sitzung mit fünf Leuten...`
- Consolidated evidence: `consolidated_extractions.json:1329` to `:1339`
- No matching rendered decision.
- `decision_0008`: `Es wird vereinbart, den bestehenden Prozessrahmen zu
nutzen und auf die Bedürfnisse der F&E abzustimmen...`
- Consolidated evidence present in the input.
- No matching rendered decision.
- `decision_0010`: `Es wird vereinbart, dass betriebliche
Verbesserungsvorschläge... ausgeschleust werden.`
- Consolidated evidence present in the input.
- No matching rendered decision.
- `decision_0012`: `Die Diskussion wird beendet und die Änderungen werden
verschickt; Abwarten auf Rückmeldung von Johanna/Björn.`
- Consolidated evidence: `consolidated_extractions.json:1747` to `:1757`
- No matching rendered decision.
### Rendered Section: Prozessrahmen und Projektideen-Eingang / Action Items
Status: partially faithful, partially omitted, with upstream semantic problems
preserved.
Rendered lines:
- `working_protocol.md:23` to `:32`
Preserved action items include:
- `action_item_0001` -> `working_protocol.md:23`
- `action_item_0002` -> `working_protocol.md:24`
- `action_item_0003` -> `working_protocol.md:25`
- `action_item_0004` -> `working_protocol.md:26`
- `action_item_0005` -> `working_protocol.md:27`
- `action_item_0006` -> `working_protocol.md:28`
- `action_item_0007` -> `working_protocol.md:29`
- `action_item_0009` -> `working_protocol.md:30`
- `action_item_0011` -> `working_protocol.md:31`
- `action_item_0012` -> `working_protocol.md:32`
Omitted or merged action items:
- `action_item_0008`: `Diskussion über die Relevanz eines speziellen Marktes
(z.B. Turkmenistan) führen.`
- Present in consolidated input.
- No matching rendered action item.
- `action_item_0010`: `Filterkriterien für verschiedene Fälle... definieren.`
- Present in consolidated input.
- Its meaning overlaps with rendered decision `working_protocol.md:18` and
action item `working_protocol.md:30`, but it is not preserved as its own
action item.
Important boundary:
Items such as `Feedback zu den Kriterien...`, `Einbinden der Fachbereiche...`
and `Mail an Jovana und Björn...` are questionable todos, but they are already
`action_item` entries in the consolidated input:
- `consolidated_extractions.json:1007` to `:1019`
- `consolidated_extractions.json:1537` to `:1549`
- `consolidated_extractions.json:1651` to `:1663`
The renderer preserves those upstream action-item categories. It does not
create those todos from facts or prose.
### Rendered Section: Prozessrahmen und Projektideen-Eingang / Open Questions
Status: partially faithful, partially omitted.
Rendered lines:
- `working_protocol.md:36` to `:47`
Preserved open questions:
- `open_question_0001` -> `working_protocol.md:36`
- `open_question_0002` -> `working_protocol.md:37`
- `open_question_0003` -> `working_protocol.md:38`
- `open_question_0004` -> `working_protocol.md:39`
- `open_question_0005` -> `working_protocol.md:40`
- `open_question_0006` -> `working_protocol.md:41`
- `open_question_0007` -> `working_protocol.md:42`
- `open_question_0009` -> `working_protocol.md:43`
- `open_question_0010` -> `working_protocol.md:44`
- `open_question_0011` -> `working_protocol.md:45`
- `open_question_0012` -> `working_protocol.md:46`
- `open_question_0013` -> `working_protocol.md:47`
Omitted open questions:
- `open_question_0008`: `Wer soll die Rolle des Gatekeepers übernehmen und wie
wird der Prozess konkret implementiert?`
- Consolidated evidence: `consolidated_extractions.json:1387` to `:1397`
- No matching rendered open question.
- `open_question_0014`: `Welche Prüfsteine sind relevant für die
Fachabteilung?`
- Consolidated evidence: `consolidated_extractions.json:1765` to `:1775`
- No matching rendered open question.
### Rendered Section: Projektkategorisierung und Status / Background
Status: modified.
Rendered line:
- `working_protocol.md:53`
This condenses:
- fact: `Der Leiter F&E führt die Projektliste...`
- `technical_detail_0010`: project type definition
- fact: `Ein Projekt wandert direkt wieder in Business Development...`
No renderer-only invention was verified in this section.
### Rendered Section: Projektkategorisierung und Status / Decisions
Status: category-changed.
Rendered lines:
- `working_protocol.md:57`
- `working_protocol.md:58`
Both rendered decisions are facts in the consolidated input:
- fact: `Im Zweifel gehen Projekte durch.`
- Consolidated evidence: `consolidated_extractions.json:775` to `:784`
- Rendered as decision: `working_protocol.md:57`
- fact: `Ein Projekt kann abgelehnt werden, aber es muss sichergestellt sein,
dass relevante Projekte nicht weggeschmissen werden.`
- Consolidated evidence: `consolidated_extractions.json:755` to `:764`
- Rendered as decision: `working_protocol.md:58`
This is the clearest category change in the renderer output.
### Rendered Section: Projektkategorisierung und Status / Action Items
Status: omitted.
Rendered line:
- `working_protocol.md:62`: `*Keine.*`
The consolidated input contains at least one action item that belongs to this
topic area:
- `action_item_0008`: `Diskussion über die Relevanz eines speziellen Marktes
(z.B. Turkmenistan) führen.`
The renderer omitted it.
### Rendered Section: Projektkategorisierung und Status / Open Questions
Status: faithful to rendered grouping, but unfaithful to complete input.
Rendered line:
- `working_protocol.md:66`: `*Keine.*`
The complete consolidated input still contains project-process open questions,
including omitted `open_question_0008` and `open_question_0014`. They were not
rendered elsewhere.
### Rendered Section: Technische Details und Infrastruktur / Background
Status: modified.
Rendered line:
- `working_protocol.md:72`
This condenses facts and technical details:
- fact: `Das Netzwerk bei Gemeinden ist momentan langsam.`
- `technical_detail_0006`: `Das Netzwerk ist langsam (wie zu ISDN-Zeiten).`
- `technical_detail_0007`: `Der Rechner schmiert ab...`
- fact: `Das Dokument ist das aktuelle Projektdeckblatt...`
- fact: `Es gibt einen Reporting-Zyklus.`
- fact: `Sekretariate sind sehr abweisend...`
The problem is not invention; the problem is that `decision_0004`
(`Festlegung eines Reporting-Zyklus...`) is downgraded from decision to
background.
### Rendered Section: Technische Details und Infrastruktur / Decisions
Status: omitted.
Rendered line:
- `working_protocol.md:76`: `*Keine.*`
The consolidated input contains `decision_0004`, which is about a reporting
cycle and is rendered only as background.
### Rendered Section: Technische Details und Infrastruktur / Action Items
Status: faithful to rendered items.
Rendered lines:
- `working_protocol.md:80`
- `working_protocol.md:81`
These correspond to:
- `action_item_0013`
- `action_item_0014`
The semantic quality of `action_item_0014` is questionable, but that issue is
upstream because the consolidated input already categorizes it as an action
item.
### Rendered Section: Technische Details und Infrastruktur / Open Questions
Status: omitted relative to complete input.
Rendered line:
- `working_protocol.md:85`: `*Keine.*`
No technical open questions are clearly required here, but the overall rendered
protocol still omits consolidated open questions elsewhere.
## Exact Mismatches
| Consolidated item | Category before | Rendered text | Category after | Result |
| --- | --- | --- | --- | --- |
| `decision_0004`: `Festlegung eines Reporting-Zyklus...` | decision | `Es gibt einen Reporting-Zyklus...` | background | category weakened |
| `decision_0007`: `...Sitzung mit fünf Leuten...` | decision | none | omitted | omitted |
| `decision_0008`: `...Prozessrahmen nutzen...` | decision | none | omitted | omitted |
| `decision_0010`: `...Verbesserungsvorschläge... ausgeschleust...` | decision | none | omitted | omitted |
| `decision_0012`: `Die Diskussion wird beendet...` | decision | none | omitted | omitted |
| `open_question_0008`: `Wer soll die Rolle des Gatekeepers übernehmen...` | open_question | none | omitted | omitted |
| `open_question_0014`: `Welche Prüfsteine sind relevant...` | open_question | none | omitted | omitted |
| fact: `Im Zweifel gehen Projekte durch.` | fact | `Im Zweifel gehen Projekte durch.` | decision | promoted |
| fact: `Ein Projekt kann abgelehnt werden...` | fact | same meaning | decision | promoted |
| `action_item_0008`: `Diskussion über die Relevanz... Turkmenistan...` | action_item | none | omitted | omitted |
| `action_item_0010`: `Filterkriterien für verschiedene Fälle... definieren.` | action_item | overlapped by decision/action item text | merged / not preserved as action item | modified |
## Invented Content Check
No renderer-only invented people were verified in this comparison. `Guido` and
`Noah` appear in the rendered protocol, but they are already present in the
consolidated input. This belongs to BUG-003, not BUG-006.
No clear renderer-only invented responsibility was verified in this comparison.
Questionable todos in the rendered protocol are already categorized as action
items in the consolidated input.
## Merged And Split Items
Verified merges:
- Multiple facts and decisions are merged into background paragraphs in
`working_protocol.md:7`, `:9`, `:53` and `:72`.
- `action_item_0010` is not preserved as a separate action item and overlaps
with rendered filter-criteria decision/action-item wording.
Verified splits:
- No clear split from one consolidated item into multiple materially different
rendered items was identified.
## Earliest Point Where The Mismatch Occurs
The earliest point is the Working Protocol Renderer LLM response.
Evidence:
- `metadata.json` records renderer input as
`semantic_consolidator/consolidated_extractions.json` and prompt file
`prompts/working_protocol.md`.
- `raw_model_response.txt` already contains the omissions and category
changes.
- `raw_model_response.txt` and `working_protocol.md` are byte-identical.
Therefore the mismatch is not introduced by renderer post-processing, output
normalization or Markdown generation.
## Verified Root Cause
Verified cause:
The single-pass LLM Working Protocol Renderer does not reliably preserve the
category and coverage constraints of the consolidated input. It generates a
Markdown view that condenses and reorganizes the input, but the generated raw
response omits required items and changes some categories.
Cause classification:
- Renderer prompt: not proven as the sole root cause. The prompt includes
explicit preservation rules, but also includes condensation and redundancy
instructions. The evidence proves the model output violates the preservation
rules; it does not prove which prompt sentence caused the violation.
- Renderer post-processing: disproved. Raw response and final Markdown are
identical.
- Renderer output normalization: disproved. No separate normalization changed
the raw response.
- Markdown generation: disproved. The final file is the raw Markdown response.
- Another verified cause: single-pass LLM renderer generation violates
structural faithfulness requirements.
## Confidence
High for the location of the failure and for disproving post-processing,
normalization and Markdown generation as causes.
Medium for root-cause classification beyond that. The evidence identifies the
renderer LLM response as the failing point, but does not isolate one prompt
sentence as the cause.
## Conceptual Repair Strategy
Do not attempt to repair this with free-form post-processing.
Conceptually, renderer faithfulness needs a structural coverage contract:
- every consolidated decision must be accounted for as a decision
- every consolidated action item must be accounted for as an action item
- every consolidated open question must be accounted for as an open question
- facts may appear in background, but must not be promoted to decisions
- omissions and category changes should be validator-detectable before the
protocol is accepted
The renderer output should be validated against the consolidated input by item
ID or another stable structured reference, rather than relying on prose-only
Markdown to preserve category semantics implicitly.
+348
View File
@@ -0,0 +1,348 @@
# Constraint Repair Engine V1
Constraint Repair Engine V1 is a reusable structural repair stage for JSON
outputs produced by local LLM pipeline steps.
It exists for cases where a model produced semantically usable JSON, but a
strict validator rejected the document because a structural invariant was
violated. Typical examples are repeated identifiers, missing identifiers, empty
containers or unstable item order.
The repair stage is intentionally not integrated into the production pipeline
yet. It is a standalone module that future stages can opt into explicitly.
## Architecture
Package:
```text
src/meeting_lab/constraint_repair/
```
The package contains:
- `engine.py`: generic JSON repair logic and the generic repair prompt
- `adapters/`: thin adapters from stage-specific validator results to the
generic validator report format
The engine receives exactly two logical inputs:
1. the complete original JSON output
2. a machine-generated validator report
It never reads the original transcript and never receives Meeting Context or
other source material. The engine is domain-neutral: it operates on JSON
pointers, list keys and validator-reported identifiers.
## Validator Interface
The generic validator report has this shape:
```json
{
"valid": false,
"violations": [
{
"type": "duplicate_id",
"collection_pointer": "/groups",
"id_list_key": "source_item_ids",
"id": "item_0001",
"occurrences": [
{"item_index": 0, "id_index": 1},
{"item_index": 3, "id_index": 0}
],
"keep_occurrence": 0
}
]
}
```
Supported V1 violation types:
- `duplicate_id`: remove repeated identifier occurrences from a list field
- `missing_id`: restore an identifier by appending a validator-provided item
template or adding it to a validator-specified existing item
- `empty_group`: remove an item whose identifier list is empty
- `reorder_items`: reorder a collection by validator-provided keys
The report must provide enough structural information for the engine to repair
the document without interpreting content.
## Meeting Context Constraint Validation Extension
Meeting Context constraints can be added without making the repair engine
Meeting Context aware.
The validator may inspect a stage output together with authoritative
`meeting_context.yaml`. It then converts violations into the generic validator
report format. The Constraint Repair Engine still receives only the original
JSON document and the validator report.
Authoritative metadata is configuration, not meeting content. Examples include:
- participant attendance
- canonical participant identity
- aliases
- departments
- roles
Authoritative metadata may be consumed by LLMs and downstream pipeline stages.
It must never be redefined, overwritten or inferred by any pipeline component.
This includes, but is not limited to:
- LLM extraction
- Canonicalizer
- Semantic Consolidator
- Constraint Repair
- Renderer
These components may consume authoritative metadata, but they must treat it as
immutable configuration.
### Constraint Types
`participant_attendance_conflict`
A known participant from Meeting Context is represented in structured output
as absent, mentioned-only, external or otherwise not attending, contradicting
the authoritative attendance status.
Example:
```json
{
"valid": false,
"constraint_source": {
"type": "meeting_context",
"meeting_id": "2026-07-27-projektprozess",
"source_file": "samples/real_live/project_process_meeting/meeting_context.yaml",
"schema_version": "1"
},
"violations": [
{
"type": "participant_attendance_conflict",
"severity": "error",
"collection_pointer": "/participants",
"item_index": 3,
"entity_id": "bjoern",
"display_name": "Björn",
"matched_alias": "Björn",
"field": "attendance_status",
"actual_value": "absent",
"expected_value": "present",
"repair": {
"operation": "set_field",
"field": "attendance_status",
"value": "present"
}
}
]
}
```
`participant_unknown`
A person-like entity appears in a structured participant or mentioned-person
field but is not known in Meeting Context as a participant, mentioned person or
alias.
Example:
```json
{
"type": "participant_unknown",
"severity": "error",
"collection_pointer": "/participants",
"item_index": 4,
"matched_text": "Guido",
"repair": {
"operation": "remove_item"
}
}
```
Removal is allowed only when the unknown entity is a standalone structured
participant-like item and removing it does not remove unrelated content. If the
unknown name appears only inside generated prose, the repair must fail closed.
`participant_alias_conflict`
A structured entity reference uses an alias that maps to a different
authoritative entity, or uses an ambiguous alias that cannot be resolved to
exactly one Meeting Context entity.
Example:
```json
{
"type": "participant_alias_conflict",
"severity": "error",
"json_pointer": "/participants/2/entity_id",
"matched_alias": "Johanna",
"expected_entity_id": "jovana",
"expected_display_name": "Jovana",
"actual_entity_id": "unknown",
"repair": {
"operation": "set_field",
"field": "entity_id",
"value": "jovana"
}
}
```
Alias repair is allowed only for structured identity fields and only when
Meeting Context maps the alias to exactly one entity.
### Validation Principle
The validator validates structured data whenever possible. Lexical analysis is
a fallback only when no structured representation exists. The long-term
objective is to reduce lexical validation over time by preserving Meeting
Context metadata as structured data throughout the pipeline.
Lexical fallback may report a violation, but it must not create repair
instructions that require editing generated prose.
### Repair Workflow
```text
Stage output
↓
Meeting Context Constraint Validator
↓
Generic validator report
↓
Constraint Repair Engine
↓
Meeting Context Constraint Validator
```
If Validator #1 succeeds, repair is skipped. If Validator #1 fails, exactly one
repair pass may run when all violations map to deterministic structural
operations. Validator #2 then checks the repaired document. If Validator #2
fails, the pipeline stops and reports the remaining violations.
### Allowed Repairs
Allowed repairs are deterministic structural operations only:
- set a structured attendance field to the authoritative value
- set a structured entity ID to the authoritative entity ID
- move a structured participant item between participant collections
- remove a standalone structured unknown participant item
- remove empty structured participant containers
- reorder structured participant collections deterministically
### Forbidden Repairs
The repair engine must never perform free-form text editing.
It must not:
- remove names from generated prose
- replace text inside prose fields
- rewrite sentences
- infer attendance from transcript content
- invent participants
- invent aliases
- merge people
- split people
- change responsibility attribution
- change fact, decision, action-item or open-question meaning
- use Meeting Context directly
If a violation exists only inside generated prose and cannot be repaired by a
deterministic structural operation, the repair must fail closed and report the
violation.
### Integration Strategy
Minimum useful integration for the current architecture:
```text
Semantic Consolidator
↓
Meeting Context Constraint Validator
↓
Constraint Repair
↓
Meeting Context Constraint Validator
↓
Renderer
```
This catches contradictions in the consolidated representation before the
Working Protocol Renderer receives it. It does not require the renderer to
infer attendance semantics from prose.
Future architecture:
Meeting Context metadata should remain structured throughout the pipeline so
that the Renderer receives authoritative participant metadata directly instead
of having to infer it from generated prose. In that architecture, Meeting
Context validation can operate primarily on structured fields and use lexical
analysis only as a diagnostic fallback.
## Allowed Operations
The engine may:
- move identifiers
- remove duplicate identifiers
- restore missing identifiers
- remove empty containers
- reorder items
The engine must not:
- invent information
- rewrite extracted text
- reinterpret reasons or explanations
- create new semantic relationships
- split semantic relationships
Missing identifier repair is intentionally conservative. If the validator does
not provide an explicit target item or an item template, the engine refuses the
repair.
## Generic Repair Prompt
The module defines a generic prompt contract for future model-backed repair
backends. The prompt describes the task as repairing a structured JSON document
from a validator report and deliberately avoids stage-specific vocabulary.
V1 unit tests assert that the prompt does not mention the current consolidation
stage, Meeting Context, facts, or merge groups.
## Semantic Consolidator Adapter
`constraint_repair.adapters.semantic_consolidator` converts the current
consolidation validator shape into the generic validator report.
The adapter knows the current consolidation output field names such as
`groups` and `source_item_ids`. The generic repair engine does not. This keeps
stage-specific schema knowledge at the edge and preserves the repair engine as
a reusable JSON utility.
The adapter currently reports:
- repeated source IDs
- missing expected source IDs
- empty source-ID containers
It does not call Ollama and does not change the production consolidation path.
## Future Reuse
Future pipeline stages can reuse the same engine by writing a small adapter
that maps their validator failures to the generic report format. The required
contract is that the adapter reports structural locations and does not ask the
engine to infer domain meaning.
Potential future uses:
- enforcing exact identifier coverage in LLM-generated grouping output
- removing empty generated containers before strict parsing
- restoring validator-known singleton items
- normalizing deterministic order after otherwise valid generation
+84
View File
@@ -266,6 +266,90 @@ renderer integration remains planned.
---
# Entity Registry
Accepted Architecture. Implementation deferred.
The Entity Registry is the persistent cross-meeting knowledge source for
confirmed entities, aliases and organizational metadata. It is independent from
individual meetings and is the planned long-term source used to prepare Meeting
Context V2.
Entity types include:
- people
- organizations
- departments
- products
- projects
- locations
- abbreviations
Each entity has a stable internal identifier. The displayed name may change
over time, but the internal identifier must remain stable.
Conceptual shape:
```json
{
"entity_id": "person_0001",
"entity_type": "person",
"display_name": "Jovana",
"aliases": [
"Jovana",
"Giovanna",
"Jovanna",
"Giovana"
],
"status": "confirmed"
}
```
The registry never learns automatically. It may propose matches, but only
confirmed user actions update it. Similarity search may suggest spelling
variants, Whisper transcription variants, umlaut variants or OCR-like mistakes,
but suggestions require explicit confirmation.
Previously unseen names should be classified by the user as one of:
- meeting participant
- mentioned person
- external person
- transcription error
- ignore
The Entity Registry must not infer responsibility, decisions, attendance or
ownership.
---
# Meeting Context V2
Accepted Architecture. Implementation deferred.
Meeting Context V2 is an authoritative meeting-specific YAML Point of Truth
generated or assisted from:
- Entity Registry
- user confirmations
- meeting metadata
The YAML remains the extraction pipeline interface and the authoritative
meeting-specific Point of Truth for that meeting run. It is also a reproducible
input artifact: changes to the Entity Registry after a meeting run must not
silently change the historical Meeting Context used for that run.
The Entity Registry remains the persistent cross-meeting knowledge source. It
must not override explicit meeting-specific confirmations.
Meeting Context V2 should reduce manual work, improve alias handling, detect
transcription errors earlier and make Meeting Context quality scalable across
many meetings.
See `docs/adr-meeting-context-v2-entity-registry.md`.
---
# Topic Result
After extraction, every topic contains the collected information.
+42
View File
@@ -26,6 +26,48 @@ Supported metadata:
- abbreviations
- relevant products, projects, systems, locations and technical terms
## Future Direction: V2
Accepted Architecture. Implementation deferred.
Meeting Context V2 should be generated from an interactive entity confirmation
workflow after Whisper transcription:
```text
Whisper
↓
Entity Detection
↓
User Confirmation
↓
Entity Registry Update
↓
Meeting Context Builder
↓
meeting_context.yaml
↓
Extraction Pipeline
```
The YAML remains the extraction interface and the authoritative
meeting-specific Point of Truth for a meeting run. It should become a
meeting-specific snapshot generated or assisted from the Entity Registry, user
confirmations and meeting metadata.
The Entity Registry is the persistent cross-meeting knowledge source for
confirmed entities, aliases and organizational metadata. It stores stable
internal identifiers and never learns automatically. The Registry must not
override explicit meeting-specific confirmations, and Registry changes after a
meeting run must not silently change the historical Meeting Context used for
that run.
Unknown names should be explicitly classified by the user as meeting
participant, mentioned person, external person, transcription error or ignore.
Similarity suggestions for spelling variants, Whisper variants, umlaut
handling and OCR-like mistakes require explicit confirmation.
See `docs/adr-meeting-context-v2-entity-registry.md`.
## File Locations
- Generic template: `samples/templates/meeting_context.template.yaml`
+42
View File
@@ -30,6 +30,8 @@ Canonical Meeting Knowledge and every Output View renderer.
# Pipeline Overview
Current analysis pipeline:
```text
Whisper Transcript
│
@@ -73,6 +75,46 @@ extraction, deterministic prompt injection and minimal extraction JSON
provenance. Later Canonicalizer, Semantic Consolidator, Canonical Meeting
Knowledge and renderer integration remains future work.
Accepted future Meeting Context V2 preparation flow:
```text
Whisper
│
▼
Entity Detection
│
▼
User Confirmation
│
▼
Entity Registry Update
│
▼
Meeting Context Builder
│
▼
meeting_context.yaml
│
▼
Extraction Pipeline
```
This preparation flow is not implemented. It is the accepted long-term
direction for reducing manual Meeting Context work while preserving explicit
user control. The Entity Registry is the persistent cross-meeting knowledge
source for confirmed entities, aliases and organizational metadata. It stores
stable internal IDs and confirmed aliases, and never updates itself
automatically.
For each meeting run, `meeting_context.yaml` remains the authoritative
meeting-specific Point of Truth and reproducible input artifact consumed by the
pipeline. V2 changes how that artifact is prepared: it may be generated or
assisted from Registry data, user confirmations and meeting metadata. The
Registry must not override explicit meeting-specific confirmations, and
Registry changes after a meeting run must not silently change the historical
Meeting Context used for that run. Similarity suggestions and unknown entity
classifications require explicit user confirmation.
---
# Stage 1 – Normalization
+54
View File
@@ -0,0 +1,54 @@
# Quality Readiness
Meeting Lab quality status should distinguish engineering readiness from
practical usability.
## Engineering Readiness
Status: NOT READY
The current end-to-end pipeline is not ready to replace the previous
extraction pipeline. The latest quality milestone records seven known
regression bugs, two documented root-cause investigations and a renderer
faithfulness failure where consolidated items are omitted or reclassified.
Engineering readiness requires at least:
- faithful rendering of consolidated decisions, action items and open questions
- no false responsibility attribution
- no participant-attendance contradictions against Meeting Context
- no unverified person/entity names entering final protocols unchecked
- regression tests or validation coverage for fixed bugs
## Practical Usability
Status: READY FOR MANUAL EDIT
The generated working protocol can still be useful as a draft when reviewed by
a human editor against the known meeting content. This status does not imply
engineering readiness and must not be used as evidence that the pipeline is
faithful or production-ready.
Practical usability means:
- the broad meeting structure is recognizable
- many relevant topics and process points are present
- manual correction is still required before the protocol can be trusted
Known manual correction areas include:
- false or over-strong responsibility attribution
- incorrect open questions
- synthetic or unverified names from transcript artifacts
- participant-attendance contradictions
- reversed or over-broad process meaning
## Current Benchmark Reference
Latest benchmark reviewed for this distinction:
- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/benchmark_report.md`
The benchmark report concludes `NOT READY` for engineering replacement. This
document adds the separate practical-usability classification for manual-edit
workflows only.
+478
View File
@@ -0,0 +1,478 @@
# Regression Bug Tracker
Living tracker for real bugs discovered during end-to-end Meeting Lab
evaluation.
Purpose:
- prevent forgotten regressions
- document known root causes
- record fixes when they happen
- verify that bugs never silently return
## BUG-001
ID: BUG-001
Title: Renderer reverses the intended process direction
Pipeline stage: Working Protocol Renderer
Severity: High
Status: Open
Date discovered: 2026-08-03
Version first observed: `meeting_context_v1/final_protocol_with_context`
Description:
The generated working protocol phrases the process direction as if the
existing F&E process is broadly taken over for all project types. The intended
meaning from the evaluation is more constrained: the existing F&E process is a
framework that may be adapted or extended for broader project handling.
Expected behaviour:
The protocol should preserve the direction and uncertainty of the source
discussion: F&E process/framework adapted for other project types where
appropriate, without implying unconditional adoption for all project types.
Actual behaviour:
The protocol states that the existing F&E process is generally adopted for all
project types.
Likely root cause:
Unknown.
Related files:
- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/working_protocol.md`
- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/comparison.md`
- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/working_protocol/working_protocol.md`
- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/semantic_consolidator/consolidated_extractions.json`
- `prompts/working_protocol.md`
Regression test available (yes/no): no
Current status:
Open. Documented from end-to-end evaluation output.
Notes:
Do not treat this entry as a prompt-change instruction. It records the observed
bug only.
Current benchmark `meeting_context_v1/e2e_current_20260803_151000` confirms
the issue remains. The new protocol still says the existing F&E process is
basically taken over for all project types.
## BUG-002
ID: BUG-002
Title: Jovana todo is strengthened beyond the meeting content
Pipeline stage: Extraction / Working Protocol Renderer
Severity: High
Status: Open
Date discovered: 2026-08-03
Version first observed: `meeting_context_v1/final_protocol_with_context`
Description:
The generated output strengthens a proposal or discussion about Jovana
assembling criteria into an assigned todo. The evaluation identified this as
overstating the meeting content.
Expected behaviour:
The pipeline should preserve the weaker source meaning unless the transcript
explicitly assigns, accepts or confirms responsibility.
Actual behaviour:
The consolidated input and final protocol include an action item assigning
criterion compilation to Jovana.
Likely root cause:
Unknown.
Related files:
- `samples/benchmarks/meeting_context_v1/semantic_consolidator_repair_v1/consolidated_extractions.json`
- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/working_protocol.md`
- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/comparison.md`
- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/semantic_consolidator/consolidated_extractions.json`
- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/working_protocol/working_protocol.md`
- `prompts/todos.md`
- `prompts/working_protocol.md`
Regression test available (yes/no): no
Current status:
Open. Documented from end-to-end evaluation output.
Notes:
This bug is subject to the responsibility attribution invariant in
`AGENTS.md`.
Current benchmark `meeting_context_v1/e2e_current_20260803_151000` confirms
the issue remains. The consolidated output still contains `Zusammenstellen der
Kriterien durch Jovana` with `responsible: "Jovana"`, and the final protocol
renders it as an action item.
## BUG-003
ID: BUG-003
Title: Unexpected synthetic person names appear in generated protocols
Pipeline stage: Extraction / Working Protocol Renderer
Severity: Medium
Status: Root Cause Identified
Date discovered: 2026-08-03
Version first observed: `meeting_context_v1/final_protocol_with_context`
Description:
The generated protocol contains unexpected or risky names/aliases such as
`Guido` and `Noah`. These names were flagged during evaluation as synthetic or
inconsistent in the generated protocol context.
Expected behaviour:
The protocol should use only names supported by the transcript, consolidated
input or Meeting Context, and should preserve uncertainty when a name or alias
is unclear.
Actual behaviour:
The final protocol includes names or name variants that were flagged as
unexpected in evaluation.
Likely root cause:
The unexpected names already exist in the Whisper transcript before extraction.
They are propagated through the pipeline and are not introduced by prompts,
gold tests or the renderer.
Related files:
- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/working_protocol.md`
- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/comparison.md`
- `samples/benchmarks/meeting_context_v1/semantic_consolidator_repair_v1/consolidated_extractions.json`
- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/working_protocol/working_protocol.md`
- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/semantic_consolidator/consolidated_extractions.json`
- `samples/real_live/project_process_meeting/meeting_context.yaml`
Regression test available (yes/no): no
Current status:
Root Cause Identified. Planned resolution is deferred until Meeting Context V2
/ Entity Registry implementation.
Notes:
Resolution strategy:
Introduce an interactive entity verification step after Whisper transcription:
```text
Whisper
↓
Entity Detection
↓
User Verification
↓
Meeting Context Builder
↓
Extraction
```
Expected effect:
Unknown or suspicious person names are detected before extraction begins and
can be classified by the user as participant, mentioned person, external
person, transcription error or ignore.
Implementation status:
Architecture accepted. Implementation deferred.
Regression test:
Not yet possible before Meeting Context V2 exists.
Current benchmark `meeting_context_v1/e2e_current_20260803_151000` confirms
the issue remains. `Guido` and `Noah` both appear in the generated working
protocol.
## BUG-004
ID: BUG-004
Title: Rhetorical question extracted as an open question
Pipeline stage: Extraction
Severity: Medium
Status: Open
Date discovered: 2026-08-03
Version first observed: `meeting_context_v1/final_protocol_with_context`
Description:
A rhetorical or discussion-framing question is extracted and later rendered as
an open question. This makes the protocol imply that the meeting left a real
follow-up question unresolved.
Expected behaviour:
Only genuine unresolved questions should be extracted as open questions.
Rhetorical questions or conversational framing should not become protocol open
questions.
Actual behaviour:
The final protocol includes at least one open question identified during
evaluation as rhetorical rather than genuinely open.
Likely root cause:
Unknown.
Related files:
- `samples/benchmarks/meeting_context_v1/semantic_consolidator_repair_v1/consolidated_extractions.json`
- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/working_protocol.md`
- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/comparison.md`
- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/semantic_consolidator/consolidated_extractions.json`
- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/working_protocol/working_protocol.md`
- `prompts/questions.md`
- `prompts/working_protocol.md`
Regression test available (yes/no): no
Current status:
Open. Documented from end-to-end evaluation output.
Notes:
No specific fix has been proposed.
Current benchmark `meeting_context_v1/e2e_current_20260803_151000` confirms
the issue remains. The final protocol still contains discussion-framing or
underspecified questions as open protocol questions, including the question
about exactly who is responsible for defining filter criteria.
## BUG-005
ID: BUG-005
Title: Renderer treats Björn as absent although Meeting Context marks him present
Pipeline stage: Working Protocol Renderer
Severity: Medium
Status: Investigating
Date discovered: 2026-08-03
Version first observed: `meeting_context_v1/final_protocol_with_context`
Description:
The generated protocol asks how missing participants such as Jovana and Björn
should be included in the process, even though Meeting Context marks Björn as
present. This creates a misleading participant-status implication.
Expected behaviour:
The protocol should not mark or imply Björn as absent when Meeting Context
records him as present. If the source discussion concerns whether a person had
reviewed materials or been included in a process, that should not be converted
into meeting absence.
Actual behaviour:
The final protocol renders an open question about including missing
participants, including Björn.
Likely root cause:
The extraction model receives Meeting Context but still classifies Björn with
missing participants in chunk 08. Downstream artifacts preserve only minimal
Meeting Context provenance, not structured participant attendance metadata, so
Canonicalizer, Semantic Consolidator and Renderer cannot correct the conflict.
Related files:
- `samples/real_live/project_process_meeting/meeting_context.yaml`
- `samples/benchmarks/meeting_context_v1/semantic_consolidator_repair_v1/consolidated_extractions.json`
- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/working_protocol.md`
- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/comparison.md`
- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/semantic_consolidator/consolidated_extractions.json`
- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/working_protocol/working_protocol.md`
- `docs/bug-005-root-cause.md`
- `prompts/working_protocol.md`
Regression test available (yes/no): no
Current status:
Investigating. Root-cause evidence has been documented in
`docs/bug-005-root-cause.md`, but no fix has been implemented or verified.
Current benchmark `meeting_context_v1/e2e_current_20260803_151000` confirms
the issue remains.
Notes:
This bug is distinct from responsibility attribution. It concerns participant
presence/status handling.
Recommended resolution remains a Meeting Context constraint validator over
structured participant metadata before rendering.
## BUG-006
ID: BUG-006
Title: Renderer drops required consolidated items and changes item categories
Pipeline stage: Working Protocol Renderer
Severity: High
Status: Open
Date discovered: 2026-08-03
Version first observed: `meeting_context_v1/e2e_current_20260803_151000`
Description:
The Working Protocol Renderer does not preserve every consolidated decision and
open question even though the renderer prompt requires preserving all
decisions, action items and open questions. It also renders some consolidated
facts as decisions.
Expected behaviour:
Every consolidated decision and open question should be represented in the
final protocol unless there is an explicit, validated reason to omit it.
Background facts must not be promoted into decisions.
Actual behaviour:
The consolidated input contains decisions such as `Festlegung eines
Reporting-Zyklus für abgelehnte Projekte`, `Zusammengetragen und beantwortete
Fragen führen zur Sitzung mit fünf Leuten...`, and `Die Diskussion wird
beendet und die Änderungen werden verschickt...`; these are not preserved as
decisions in the final protocol. The consolidated input also contains open
questions such as `Wer soll die Rolle des Gatekeepers übernehmen...` and
`Welche Prüfsteine sind relevant für die Fachabteilung?`; these are omitted
from the final protocol. Conversely, consolidated facts such as `Im Zweifel
gehen Projekte durch` and `Ein Projekt kann abgelehnt werden...` are rendered
as decisions.
Likely root cause:
Unknown.
Related files:
- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/semantic_consolidator/consolidated_extractions.json`
- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/working_protocol/working_protocol.md`
- `prompts/working_protocol.md`
Regression test available (yes/no): no
Current status:
Open. Documented from current end-to-end benchmark output.
Notes:
This is a renderer faithfulness problem independent from whether the upstream
extracted items are themselves correct.
## BUG-007
ID: BUG-007
Title: Non-committed discussion items are extracted and rendered as todos
Pipeline stage: Extraction / Working Protocol Renderer
Severity: High
Status: Open
Date discovered: 2026-08-03
Version first observed: `meeting_context_v1/e2e_current_20260803_151000`
Description:
The pipeline extracts and renders several action items that are not clearly
assigned, accepted or confirmed commitments in the meeting evidence.
Expected behaviour:
Only explicit assignments, accepted responsibilities or confirmed follow-up
actions should become todos. Discussion, examples, proposals, broad process
needs or unclear alternatives should remain facts/questions or be marked
unclear where supported.
Actual behaviour:
The current consolidated output and final protocol contain todos such as
`Feedback zu den Kriterien und Änderungen am Prozess abwarten`, `Einbinden der
Fachbereiche (stellvertretend durch Guido) zur Definition von Prüfsteinen`,
and `Mail an Jovana und Björn erneut senden oder Prozessanpassung vornehmen`.
The evidence for these items is discussion-level or ambiguous rather than a
clear todo commitment.
Likely root cause:
Unknown.
Related files:
- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/semantic_consolidator/consolidated_extractions.json`
- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/working_protocol/working_protocol.md`
- `prompts/todos.md`
- `prompts/working_protocol.md`
Regression test available (yes/no): no
Current status:
Open. Documented from current end-to-end benchmark output.
Notes:
This generalizes BUG-002 beyond the specific Jovana assignment case and should
be evaluated against the responsibility attribution invariant.