diff --git a/README.md b/README.md index 2c0fe6e..aed136b 100644 --- a/README.md +++ b/README.md @@ -114,6 +114,17 @@ optional validieren, als autoritative Metadaten in den Prompt aufnehmen und minimale Kontext-Provenienz im Extraction JSON speichern. Konsolidierung, Canonical Meeting Knowledge und Rendering sind noch nicht daran angeschlossen. +Als akzeptierte, aber noch nicht implementierte Architektur soll Meeting +Context V2 kuenftig nicht mehr primaer manuell geschrieben werden. Nach Whisper +soll ein Entity-Detection- und User-Confirmation-Schritt unbekannte Namen und +Begriffe sichtbar machen. Bestaetigte Entitaeten werden in einer persistenten +Entity Registry als cross-meeting Wissensquelle mit stabilen internen IDs, +Anzeigenamen und Aliasen gepflegt. Aus Registry, Nutzerbestaetigungen und +Meeting-Metadaten erzeugt ein Meeting Context Builder dann das +meeting-spezifische `meeting_context.yaml`. Diese YAML-Datei bleibt der +authoritative meeting-spezifische Point of Truth und ein reproduzierbares +Input-Artefakt fuer den jeweiligen Pipeline-Lauf. + Der gemeinsame Extraktionsprompt besteht aktuell aus `common.md`, `decisions.md` und `todos.md`. Die Todo-Regeln verlangen explizite Zuweisung, Freiwilligenmeldung oder Annahme, bevor eine verantwortliche Person gesetzt diff --git a/ROADMAP.md b/ROADMAP.md index dd67b87..fe3db7e 100644 --- a/ROADMAP.md +++ b/ROADMAP.md @@ -126,6 +126,67 @@ Out of scope: - Export formats beyond those needed to validate the model. - Knowledge-system storage design. +## Phase 4A - Meeting Context V2 and Entity Registry + +Goal: + +- Replace primarily manual Meeting Context authoring with an interactive + entity confirmation workflow backed by a persistent Entity Registry. + +Accepted architecture: + +```text +Whisper + ↓ +Entity Detection + ↓ +User Confirmation + ↓ +Entity Registry Update + ↓ +Meeting Context Builder + ↓ +meeting_context.yaml + ↓ +Extraction Pipeline +``` + +Deliverables: + +- Persistent Entity Registry independent from individual meetings as the + cross-meeting knowledge source for confirmed entities, aliases and + organizational metadata. +- Stable internal entity IDs for people, organizations, departments, products, + projects, locations and abbreviations. +- First-class aliases, including spelling and transcription variants. +- User confirmation UI/workflow for unknown names. +- Meeting Context Builder that produces meeting-specific `meeting_context.yaml` + from registry entries, user confirmations and meeting metadata. +- Generated `meeting_context.yaml` as the authoritative meeting-specific Point + of Truth and reproducible input artifact for each meeting run. +- Similarity suggestions for spelling variants, Whisper variants, umlaut + handling and OCR-like mistakes. + +Prerequisites: + +- Agreement on the Entity Registry data model. +- Review workflow for confirming unknown entities after Whisper transcription. +- Meeting Context V1 remains the extraction interface until V2 is implemented. + +Current status: + +- Accepted Architecture. +- Implementation deferred. + +Out of scope: + +- Autonomous learning. +- Silent registry updates. +- Registry overrides of explicit meeting-specific confirmations. +- Silent changes to historical Meeting Context after Registry updates. +- Inferring responsibility, decisions, ownership or attendance from registry + metadata. + ## Phase 5 - Output Views Goal: diff --git a/docs/adr-meeting-context-v2-entity-registry.md b/docs/adr-meeting-context-v2-entity-registry.md new file mode 100644 index 0000000..71fa20c --- /dev/null +++ b/docs/adr-meeting-context-v2-entity-registry.md @@ -0,0 +1,153 @@ +# ADR: Meeting Context V2 and Entity Registry + +Status: Accepted Architecture + +Implementation: Deferred + +Date: 2026-08-03 + +## Context + +Meeting Context V1 proved that authoritative meeting context can significantly +improve extraction quality. It helps the extractor normalize known aliases, +identify participants and avoid treating context metadata as evidence for +responsibility or decisions. + +End-to-end evaluation also showed that manually writing Meeting Context is not +the right long-term primary workflow. The pipeline needs an earlier, +interactive entity confirmation step after Whisper transcription. That step +should identify candidate entities, ask the user to confirm them and then build +the meeting-specific context from confirmed data. + +## Decision + +Meeting Context should evolve toward V2 as an authoritative, +meeting-specific Point of Truth generated or assisted from a persistent Entity +Registry, user confirmations and meeting metadata. + +Preferred future pipeline: + +```text +Whisper + ↓ +Entity Detection + ↓ +User Confirmation + ↓ +Entity Registry Update + ↓ +Meeting Context Builder + ↓ +meeting_context.yaml + ↓ +Extraction Pipeline +``` + +The Entity Registry is the persistent cross-meeting knowledge source for +confirmed entities, aliases and organizational metadata. + +For each individual meeting, `meeting_context.yaml` remains the authoritative +meeting-specific Point of Truth and reproducible input artifact consumed by +the extraction pipeline. Meeting Context V2 changes how this YAML is prepared, +not its authority for a meeting run. + +The generated `meeting_context.yaml` is a meeting-specific snapshot. The +Registry must not override explicit meeting-specific confirmations. Changes to +the Registry after a meeting run must not silently change the historical +Meeting Context used for that run. + +## Entity Registry + +The Entity Registry is the persistent cross-meeting knowledge source and is +independent from individual meetings. It stores confirmed entities such as: + +- people +- organizations +- departments +- products +- projects +- locations +- abbreviations + +Each entity receives a stable internal identifier. The displayed name may +change over time, but the identifier must remain stable. + +## Learning Principle + +The registry never learns automatically. + +It may propose matches, but only confirmed user actions update the registry. +No autonomous learning is allowed. + +## Alias Handling + +Aliases are first-class data. + +Examples: + +```text +Jovana +Giovanna +Jovanna +Giovana +``` + +These variants may all refer to one confirmed entity. Future runs should +automatically suggest previously confirmed aliases, but those suggestions still +require explicit confirmation when they would update registry data. + +## Unknown Entities + +Previously unseen names are presented to the user for classification. + +Possible classifications: + +- meeting participant +- mentioned person +- external person +- transcription error +- ignore + +Nothing is automatically accepted. + +## Similarity Search + +Similarity search is a future extension. It can propose likely matches for: + +- spelling variants +- Whisper transcription variants +- umlaut handling +- OCR-like mistakes + +Similarity suggestions require explicit confirmation. + +## Rationale + +Expected advantages: + +- significantly less manual work +- earlier detection of transcription errors +- robust alias handling +- reusable organizational knowledge +- improved Meeting Context quality +- easier GUI workflow +- better scalability across many meetings + +## Consequences + +Meeting Context V1 remains the current implemented interface. + +Meeting Context V2 should preserve the YAML interface for extraction. +`meeting_context.yaml` remains the authoritative meeting-specific Point of +Truth and reproducible input artifact for a meeting run. The Entity Registry is +the persistent cross-meeting knowledge source used to prepare that artifact. + +The Registry must not override explicit meeting-specific confirmations. Changes +to the Registry after a meeting run must not silently change the historical +Meeting Context used for that run. + +The registry must not infer responsibility, decisions, attendance or ownership. +Those still require meeting evidence and remain governed by the existing +responsibility attribution invariant. + +No implementation is part of this ADR. diff --git a/docs/architecture.md b/docs/architecture.md index b1d5d2d..ef00cf1 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -99,6 +99,8 @@ The analyzer gradually transforms an unstructured discussion into structured kno # High-Level Pipeline +Current implemented and intended analysis flow: + ```text Transcript ↓ @@ -123,6 +125,24 @@ Output View Rendering Working Protocol / Distribution Protocol / Knowledge Objects ``` +Accepted future Meeting Context V2 preparation flow: + +```text +Whisper + ↓ +Entity Detection + ↓ +User Confirmation + ↓ +Entity Registry Update + ↓ +Meeting Context Builder + ↓ +meeting_context.yaml + ↓ +Extraction Pipeline +``` + Each stage solves one clearly defined problem. No module should perform multiple semantic tasks simultaneously. @@ -204,6 +224,32 @@ The context can help prevent non-participants from being interpreted as attendees and can normalize known aliases for extraction. It must not infer roles, departments, responsibilities or decisions. +Accepted future direction: + +Meeting Context V2 should be generated or assisted from an interactive entity +confirmation workflow and a persistent Entity Registry. The Entity Registry is +the persistent cross-meeting knowledge source for confirmed people, +organizations, departments, products, projects, locations, aliases and +organizational metadata under stable internal IDs. Display names may change, +but internal IDs remain stable. Aliases are first-class data. + +The registry never learns automatically. It may propose matches and aliases, +including spelling variants, Whisper transcription variants, umlaut variants +and OCR-like mistakes, but only explicit user confirmation updates registry +state. Unknown names should be presented to the user as meeting participant, +mentioned person, external person, transcription error or ignore. + +In this architecture, `meeting_context.yaml` remains the authoritative +meeting-specific Point of Truth and reproducible input artifact consumed by the +extraction pipeline. It is a meeting-specific snapshot built from the Entity +Registry, user confirmations and meeting metadata. The Registry must not +override explicit meeting-specific confirmations, and Registry changes after a +meeting run must not silently change the historical Meeting Context used for +that run. + +Status: Accepted Architecture; implementation deferred. See +`docs/adr-meeting-context-v2-entity-registry.md`. + --- ## consolidation/ diff --git a/docs/bug-003-root-cause.md b/docs/bug-003-root-cause.md new file mode 100644 index 0000000..48e3c72 --- /dev/null +++ b/docs/bug-003-root-cause.md @@ -0,0 +1,265 @@ +# BUG-003 Root Cause Analysis + +## Observed Behaviour + +The final working protocol contains two names that were flagged as not +belonging to the real meeting: + +- `Guido` +- `Noah` + +Observed output: + +- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/working_protocol.md:36` + contains `Einbinden der Fachbereiche (stellvertretend durch Guido) zur + Definition von Prüfsteinen.` +- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/working_protocol.md:44` + contains `Sollte ein eigener Prozess für Noah (Business Development) + benötigt werden oder reicht der bestehende?` + +## Evidence + +The names are present in the final renderer raw response as well as +`working_protocol.md`: + +- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/raw_model_response.txt:36` + contains `Guido` +- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/raw_model_response.txt:44` + contains `Noah` + +They are also present before rendering in the repaired consolidated input: + +- `samples/benchmarks/meeting_context_v1/semantic_consolidator_repair_v1/consolidated_extractions.json:1047` + contains the open question text with `Noah` +- `samples/benchmarks/meeting_context_v1/semantic_consolidator_repair_v1/consolidated_extractions.json:1048` + contains evidence `Oder brauchen wir einen eigenen Prozess bei Noah?` +- `samples/benchmarks/meeting_context_v1/semantic_consolidator_repair_v1/consolidated_extractions.json:1537` + contains the action item text with `Guido` + +They are present before semantic consolidation in the canonicalized output: + +- `samples/benchmarks/meeting_context_v1/canonicalizer/canonicalized_extractions.json:411` + contains the open question text with `Noah` +- `samples/benchmarks/meeting_context_v1/canonicalizer/canonicalized_extractions.json:412` + contains evidence `Oder brauchen wir einen eigenen Prozess bei Noah?` +- `samples/benchmarks/meeting_context_v1/canonicalizer/canonicalized_extractions.json:1401` + contains the action item text with `Guido` + +They are present before canonicalization in extraction output and raw model +responses: + +- `samples/benchmarks/meeting_context_v1/full_context_run/chunk_02_extraction.json:18` + contains the extracted open question with `Noah` +- `samples/benchmarks/meeting_context_v1/full_context_run/chunk_02_extraction.raw.txt:76` + contains the same raw model output question with `Noah` +- `samples/benchmarks/meeting_context_v1/full_context_run/chunk_07_extraction.json:11` + contains the extracted todo with `Guido` +- `samples/benchmarks/meeting_context_v1/full_context_run/chunk_07_extraction.raw.txt:53` + contains the same raw model output todo with `Guido` + +They are present before extraction in the reconstructed normalized chunks: + +- `samples/benchmarks/meeting_context_v1/source_reconstruction/chunk_02_normalized.txt:79` + contains `Oder brauchen wir einen eigenen Prozess bei Noah?` +- `samples/benchmarks/meeting_context_v1/source_reconstruction/chunk_07_normalized.txt:213` + contains `Guido` + +They are present before normalization in reconstructed raw chunks: + +- `samples/benchmarks/meeting_context_v1/source_reconstruction/chunk_02.txt:79` + contains `Oder brauchen wir einen eigenen Prozess bei Noah?` +- `samples/benchmarks/meeting_context_v1/source_reconstruction/chunk_07.txt:213` + contains `Guido` + +They are present in the source cleaned Whisper transcript used to reconstruct +the chunks: + +- `samples/real_live/project_process_meeting/transcript/meeting_speech_cleaned.json:5783` + has segment text `Oder brauchen wir einen eigenen Prozess bei Noah?` +- `samples/real_live/project_process_meeting/transcript/meeting_speech_cleaned.json:26315` + has segment text `Guido` + +Segment-level verification: + +- `Noah`: segment `id=182`, `start=899.6400000000001`, `end=901.94`, + text `Oder brauchen wir einen eigenen Prozess bei Noah?` +- `Guido`: segment `id=776`, `start=4157.86`, `end=4160.14`, text `Guido` + +The original Whisper transcript also contains both names: + +- `samples/real_live/project_process_meeting/transcript/meeting_speech.json:1` + contains both `Noah` and `Guido` in the top-level transcript text and segment + data. + +Meeting Context does not contain either name: + +- `rg -n "Guido|Noah" samples/real_live/project_process_meeting/meeting_context.yaml` + returned no matches. + +Prompt files do not contain either name: + +- `rg -n "Guido|Noah" prompts` + returned no matches. + +## Earliest Pipeline Stage Containing the Names + +The earliest verified pipeline artifact containing the names is the Whisper +transcript stage: + +- `samples/real_live/project_process_meeting/transcript/meeting_speech.json` +- `samples/real_live/project_process_meeting/transcript/meeting_speech_cleaned.json` + +For the reconstructed benchmark specifically, the earliest input used by the +TODO 4 pipeline is: + +- `samples/real_live/project_process_meeting/transcript/meeting_speech_cleaned.json` + +The reconstructed chunks preserve the names from that cleaned Whisper JSON. +Extraction, canonicalization, consolidation and rendering propagate them. + +## Repository Search Results + +Repository search for `Guido|Noah` found these occurrence groups: + +- `docs/regression-bugs.md`: tracker entry for BUG-003 mentions both names. +- `tests/gold/responsibility_attribution_negative/transcript.txt`: contains + `Noah` in a synthetic gold scenario. +- `tests/gold/responsibility_attribution_negative/expected.json`: contains + expected `Noah` entries for that synthetic gold scenario. +- `samples/real_live/project_process_meeting/transcript/meeting_speech.json`: + contains `Noah` and `Guido`. +- `samples/real_live/project_process_meeting/transcript/meeting_speech_cleaned.json`: + contains `Noah` and `Guido`. +- `samples/benchmarks/meeting_context_v1/source_reconstruction/chunk_02.txt` + and `chunk_02_normalized.txt`: contain `Noah`. +- `samples/benchmarks/meeting_context_v1/source_reconstruction/chunk_07.txt` + and `chunk_07_normalized.txt`: contain `Guido`. +- `samples/benchmarks/meeting_context_v1/full_context_run/chunk_02_extraction.json` + and `chunk_02_extraction.raw.txt`: contain `Noah`. +- `samples/benchmarks/meeting_context_v1/full_context_run/chunk_07_extraction.json` + and `chunk_07_extraction.raw.txt`: contain `Guido`. +- `samples/benchmarks/meeting_context_v1/canonicalizer/canonicalized_extractions.json`: + contains both names. +- `samples/benchmarks/meeting_context_v1/semantic_consolidator_repair_v1/consolidated_extractions.json`: + contains both names. +- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/raw_model_response.txt` + and `working_protocol.md`: contain both names. +- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/comparison.md`: + mentions `Noah` in the evaluation notes. + +Search results did not find `Guido` or `Noah` in: + +- `prompts/` +- `samples/real_live/project_process_meeting/meeting_context.yaml` +- `src/` + +## Prompt Inspection + +Renderer prompt: + +- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/metadata.json` + records `prompt_file: "prompts/working_protocol.md"` and + `input_source: + "samples/benchmarks/meeting_context_v1/semantic_consolidator_repair_v1/consolidated_extractions.json"`. +- `prompts/working_protocol.md` contains no few-shot examples and no `Guido` or + `Noah` occurrences. + +Extraction prompt assembly: + +- `src/meeting_lab/llm/prompts.py` builds extraction prompts from + `prompts/common.md`, optional Meeting Context, explicitly requested task + prompts and the provided transcript. +- `src/meeting_lab/extraction/extract_chunks.py` calls that shared builder with + `decisions.md` and `todos.md` as task prompts. + +Semantic Consolidator prompt assembly: + +- `src/meeting_lab/consolidation/consolidate_facts.py` builds its prompt from + `prompts/consolidate_facts.md` plus the canonicalized fact payload. + +Gold/test prompt inspection: + +- `tests/gold/responsibility_attribution_negative/` contains synthetic `Noah` + examples. +- `scripts/run_gold_test.py` is a gold-test runner and reads + `scenario_dir / "transcript.txt"` when explicitly invoked. +- The inspected production metadata for the TODO 4 extraction run records + `input_path` values under + `samples/benchmarks/meeting_context_v1/source_reconstruction/`, not + `tests/gold/`. +- The inspected renderer metadata records only the repaired consolidated JSON + and `prompts/working_protocol.md`. + +No evidence was found that few-shot examples, embedded examples, test +transcripts or gold scenarios became part of the production prompt for this +benchmark run. + +## Pipeline Verification + +The persisted metadata verifies the production execution path: + +- `chunk_02_extraction.metadata.json` input: + `samples/benchmarks/meeting_context_v1/source_reconstruction/chunk_02_normalized.txt` +- `chunk_07_extraction.metadata.json` input: + `samples/benchmarks/meeting_context_v1/source_reconstruction/chunk_07_normalized.txt` +- both extraction metadata files record Meeting Context source: + `samples/real_live/project_process_meeting/meeting_context.yaml` +- `canonicalizer/metadata.json` records input directory: + `samples/benchmarks/meeting_context_v1/full_context_run` +- `semantic_consolidator_repair_v1/report.md` records output: + `samples/benchmarks/meeting_context_v1/semantic_consolidator_repair_v1/consolidated_extractions.json` +- `final_protocol_with_context/metadata.json` records renderer input: + `samples/benchmarks/meeting_context_v1/semantic_consolidator_repair_v1/consolidated_extractions.json` + and prompt file `prompts/working_protocol.md` + +No execution metadata points to unrelated examples, templates, tests or gold +scenarios. + +## Verified Root Cause + +The verified root cause for BUG-003 in this pipeline run is upstream transcript +contamination: `Guido` and `Noah` already exist in the Whisper transcript and +the cleaned Whisper transcript used as the benchmark source. + +The later stages did not introduce these names from prompts, Meeting Context, +gold tests or renderer examples. They propagated names already present in the +pipeline input: + +`meeting_speech.json` +-> `meeting_speech_cleaned.json` +-> reconstructed chunks +-> extraction JSON +-> canonicalized JSON +-> repaired consolidated JSON +-> final working protocol + +What cannot be established from repository artifacts alone: + +- whether Whisper hallucinated these names from audio +- whether the source audio actually contains words that sound like these names +- what the correct intended tokens should be + +Those questions require audio-level or human-transcript verification and are +outside the evidence available in this repository trace. + +## Confidence + +High + +The conclusion that the final protocol names originated before extraction is +directly supported by persisted source transcript, chunk, extraction, +canonicalization, consolidation and renderer artifacts. + +The confidence does not extend to identifying the correct replacement words. +That remains unverified. + +## Recommended Fix + +Conceptually, add a transcript/source-quality validation step before extraction +that flags person or organization names not present in Meeting Context or an +approved alias list. The step should preserve the original transcript text, but +mark suspicious entity mentions for review before they become structured +knowledge and final protocol content. + +Do not treat gold scenarios or prompt examples as the root cause for this bug; +the evidence does not support that. diff --git a/docs/bug-005-root-cause.md b/docs/bug-005-root-cause.md new file mode 100644 index 0000000..4a403ee --- /dev/null +++ b/docs/bug-005-root-cause.md @@ -0,0 +1,198 @@ +# BUG-005 Root Cause Analysis + +## Observed Behaviour + +The final working protocol treats Björn as one of the "fehlende Teilnehmer": + +- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/working_protocol.md:53` + +This contradicts Meeting Context, which lists Björn as an actual participant +with `attendance_status: "present"`: + +- `samples/real_live/project_process_meeting/meeting_context.yaml:45` +- `samples/real_live/project_process_meeting/meeting_context.yaml:52` + +## Evidence + +Meeting Context marks Björn present: + +```text +participant_id: "bjoern" +display_name: "Björn" +role: "Leiter Marketing" +attendance_status: "present" +``` + +The extraction prompt path does include Meeting Context: + +- `src/meeting_lab/extraction/extract_chunks.py:161` renders Meeting Context + for the prompt when it is supplied. +- `src/meeting_lab/models/meeting_context.py:108` renders the heading + `MEETING CONTEXT V1 (AUTHORITATIVE METADATA)`. +- `src/meeting_lab/models/meeting_context.py:112` renders the rule + `The participant list is authoritative.` +- `src/meeting_lab/models/meeting_context.py:128` to + `src/meeting_lab/models/meeting_context.py:132` renders actual + participants. + +Execution metadata confirms chunk 08 used the Meeting Context file: + +- `samples/benchmarks/meeting_context_v1/full_context_run/chunk_08_extraction.metadata.json:29` + to `:32` +- `samples/benchmarks/meeting_context_v1/full_context_run/chunk_08_extraction.metadata.json:34` + to `:38` + +The source chunk does not say Björn was absent from the meeting. It says +Giovanna and Björn had not yet understood or reviewed the discussed mail or +process state: + +- `samples/benchmarks/meeting_context_v1/source_reconstruction/chunk_08_normalized.txt:211` + to `:215` +- `samples/benchmarks/meeting_context_v1/source_reconstruction/chunk_08_normalized.txt:229` + to `:245` +- `samples/benchmarks/meeting_context_v1/source_reconstruction/chunk_08_normalized.txt:257` + to `:267` + +The earliest generated artifact that treats Björn as absent is the raw chunk +08 extraction response: + +- `samples/benchmarks/meeting_context_v1/full_context_run/chunk_08_extraction.raw.txt:4` + summarizes "Einbeziehung fehlender Teilnehmer wie Jovana und Björn". +- `samples/benchmarks/meeting_context_v1/full_context_run/chunk_08_extraction.raw.txt:6` + to `:10` lists participants as only Martin, Lars and Malte. +- `samples/benchmarks/meeting_context_v1/full_context_run/chunk_08_extraction.raw.txt:64` + to `:68` emits the open question + `Wie sollen fehlende Teilnehmer (Jovana, Björn) in den Prozess einbezogen werden?` + +The persisted extraction JSON keeps that open question: + +- `samples/benchmarks/meeting_context_v1/full_context_run/chunk_08_extraction.json:14` + to `:16` + +The persisted extraction JSON records only Meeting Context provenance, not the +participant list or attendance status: + +- `samples/benchmarks/meeting_context_v1/full_context_run/chunk_08_extraction.json:23` + to `:27` + +The normalizer drops raw response fields such as `participants` and `topics`: + +- `src/meeting_lab/extraction/extract_chunks.py:318` to `:365` returns only + normalized `facts`, `decisions`, `todos`, `questions`, `positions` and + `technical`. +- `src/meeting_lab/extraction/extract_chunks.py:432` to `:434` adds only + provenance from Meeting Context after normalization. + +Canonicalizer preserves the bad open question: + +- `samples/benchmarks/meeting_context_v1/canonicalizer/canonicalized_extractions.json:1633` + to `:1645` + +Semantic Consolidator output also preserves the bad open question: + +- `samples/benchmarks/meeting_context_v1/semantic_consolidator_repair_v1/consolidated_extractions.json:1669` + to `:1681` + +The renderer receives the repaired consolidated JSON as its only recorded +input: + +- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/metadata.json:2` + to `:4` + +The renderer prompt requires using only the provided input: + +- `prompts/working_protocol.md:7` to `:12` + +Therefore the renderer did not receive Meeting Context attendance metadata +that would let it distinguish a present participant from a mentioned-only or +absent person. + +## Stage-by-Stage Verification + +1. Does extraction receive Björn as a present participant? + +Yes. The extraction code renders Meeting Context into the prompt when supplied, +and the chunk 08 metadata records +`samples/real_live/project_process_meeting/meeting_context.yaml` as the Meeting +Context source. + +2. Does extraction output preserve that information? + +No. The raw chunk 08 response lists participants as Martin, Lars and Malte +only, despite Björn being present in Meeting Context. The official persisted +extraction JSON contains only a minimal `context` provenance object and does +not preserve the authoritative participant list or attendance status. + +3. Does Canonicalizer preserve it? + +No. Canonicalizer input lacks the participant list and attendance status. The +Canonicalizer preserves the already bad open question from chunk 08. + +4. Does Semantic Consolidator preserve it? + +No. The repaired consolidated JSON lacks Meeting Context participant metadata +and preserves the bad open question. + +5. Does the Renderer receive enough information to distinguish present from +mentioned-only participants? + +No. Renderer metadata records only +`semantic_consolidator_repair_v1/consolidated_extractions.json` as input, and +that file contains no Meeting Context participant list or attendance status. + +## Earliest Failing Pipeline Stage + +The earliest failing stage is chunk extraction for +`chunk_08_normalized.txt`. + +More precisely, the raw model response for chunk 08 is the first artifact that +both omits Björn from the participant list and groups Björn with absent or +"fehlende" participants. + +## Verified Root Cause + +The immediate root cause is ignored Meeting Context metadata during extraction +for chunk 08. The extraction model received authoritative Meeting Context, but +interpreted transcript evidence about not having reviewed or understood an +email/process state as meeting absence. + +The downstream cause is missing propagation of participant attendance metadata. +After extraction, only Meeting Context provenance is persisted. Canonicalizer, +Semantic Consolidator and Renderer do not receive the authoritative participant +list or attendance status, so they cannot detect or repair the contradiction. + +This is not primarily a renderer inference bug. The renderer preserved an open +question already present in its input, and its prompt explicitly restricts it +to the supplied consolidated input. + +## Cause Classification + +- Missing propagation: verified. +- Ignored metadata: verified at extraction. +- Prompt wording: not verified as the direct cause. +- Renderer inference: disproved as the earliest cause. +- Lost provenance: partially verified; provenance remains, but the substantive + participant metadata is not propagated. +- Another cause: not identified. + +## Confidence + +High. + +The conclusion is supported by Meeting Context, extraction metadata, raw chunk +08 model output, persisted extraction JSON, Canonicalizer output, Semantic +Consolidator output and renderer metadata. + +## Recommended Fix + +Concept only: + +Propagate authoritative Meeting Context participant metadata, including +attendance status, beyond extraction as structured data. Add validation that +flags contradictions where generated items classify an actual participant as +absent or mentioned-only. The validation should run before rendering, and +ideally immediately after extraction so the contradiction is caught at the +earliest stage. + +Do not rely on the Working Protocol Renderer to correct attendance semantics +from prose-only consolidated items. diff --git a/docs/bug-006-root-cause.md b/docs/bug-006-root-cause.md new file mode 100644 index 0000000..31d44bd --- /dev/null +++ b/docs/bug-006-root-cause.md @@ -0,0 +1,426 @@ +# BUG-006 Root Cause Analysis + +## Observed Behaviour + +In benchmark `meeting_context_v1/e2e_current_20260803_151000`, the Working +Protocol Renderer output is not a faithful view of the matching consolidated +representation. + +Compared files: + +- Consolidated input: + `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/semantic_consolidator/consolidated_extractions.json` +- Rendered output: + `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/working_protocol/working_protocol.md` +- Raw renderer response: + `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/working_protocol/raw_model_response.txt` + +Observed problems: + +- consolidated decisions are omitted from the rendered `Decisions` sections +- consolidated open questions are omitted from rendered `Open Questions` + sections +- consolidated facts are rendered as decisions +- some consolidated action items are omitted +- some action-item and decision meanings are merged into other rendered + sections + +The raw renderer response and `working_protocol.md` are byte-identical +(`cmp` exit code `0`), so the mismatch is already present in the LLM renderer +response. It is not introduced by Markdown file writing. + +## Renderer Input And Prompt Evidence + +Renderer metadata records the exact input and prompt: + +- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/working_protocol/metadata.json:2` + to `:4` + +The renderer prompt explicitly says: + +- use only the provided input: + `prompts/working_protocol.md:9` to `:10` +- preserve all decisions, action items and open questions: + `prompts/working_protocol.md:24` to `:26` +- remove presentation-level redundancy: + `prompts/working_protocol.md:23` +- do not repeat information merely because it appears in multiple categories: + `prompts/working_protocol.md:65` +- condense everything else: + `prompts/working_protocol.md:75` + +The prompt is therefore present and relevant, but the direct observed failure +occurs in the model-generated renderer response. + +## Faithfulness Analysis + +### Rendered Section: Prozessrahmen und Projektideen-Eingang / Background + +Status: modified. + +Rendered lines: + +- `working_protocol.md:7` +- `working_protocol.md:9` + +The section condenses several consolidated facts and decisions into background +paragraphs. Examples: + +- `decision_0001` is restated in background and again as a decision. +- `fact`: `Der Prozess ist ein F&E-Prozess gewesen, einfach historisch.` +- `fact`: `Die Bearbeitung der Projekte muss und kann nur in den + Fachabteilungen passieren.` +- `fact`: `Bevor F&E ein Projekt durchdenkt, müssen im Vorfeld Kriterien wie + Marktexistenz und Zulassungen geprüft werden.` +- `decision_0003` and `decision_0005` are partially restated as background. + +This is a condensation, not a renderer-only invention. + +### Rendered Section: Prozessrahmen und Projektideen-Eingang / Decisions + +Status: partially faithful, partially omitted. + +Faithfully or near-faithfully preserved: + +- `decision_0001` -> `working_protocol.md:13` +- `decision_0002` -> `working_protocol.md:14` +- `decision_0003` -> `working_protocol.md:15` +- `decision_0005` -> `working_protocol.md:16` +- `decision_0006` -> `working_protocol.md:17` +- `decision_0009` -> `working_protocol.md:18` +- `decision_0011` -> `working_protocol.md:19` + +Omitted as decisions: + +- `decision_0004`: `Festlegung eines Reporting-Zyklus für abgelehnte + Projekte.` + - Consolidated evidence: `consolidated_extractions.json:1217` to `:1227` + - Rendered only as background: `working_protocol.md:72` +- `decision_0007`: `Zusammengetragen und beantwortete Fragen führen zur + Sitzung mit fünf Leuten...` + - Consolidated evidence: `consolidated_extractions.json:1329` to `:1339` + - No matching rendered decision. +- `decision_0008`: `Es wird vereinbart, den bestehenden Prozessrahmen zu + nutzen und auf die Bedürfnisse der F&E abzustimmen...` + - Consolidated evidence present in the input. + - No matching rendered decision. +- `decision_0010`: `Es wird vereinbart, dass betriebliche + Verbesserungsvorschläge... ausgeschleust werden.` + - Consolidated evidence present in the input. + - No matching rendered decision. +- `decision_0012`: `Die Diskussion wird beendet und die Änderungen werden + verschickt; Abwarten auf Rückmeldung von Johanna/Björn.` + - Consolidated evidence: `consolidated_extractions.json:1747` to `:1757` + - No matching rendered decision. + +### Rendered Section: Prozessrahmen und Projektideen-Eingang / Action Items + +Status: partially faithful, partially omitted, with upstream semantic problems +preserved. + +Rendered lines: + +- `working_protocol.md:23` to `:32` + +Preserved action items include: + +- `action_item_0001` -> `working_protocol.md:23` +- `action_item_0002` -> `working_protocol.md:24` +- `action_item_0003` -> `working_protocol.md:25` +- `action_item_0004` -> `working_protocol.md:26` +- `action_item_0005` -> `working_protocol.md:27` +- `action_item_0006` -> `working_protocol.md:28` +- `action_item_0007` -> `working_protocol.md:29` +- `action_item_0009` -> `working_protocol.md:30` +- `action_item_0011` -> `working_protocol.md:31` +- `action_item_0012` -> `working_protocol.md:32` + +Omitted or merged action items: + +- `action_item_0008`: `Diskussion über die Relevanz eines speziellen Marktes + (z.B. Turkmenistan) führen.` + - Present in consolidated input. + - No matching rendered action item. +- `action_item_0010`: `Filterkriterien für verschiedene Fälle... definieren.` + - Present in consolidated input. + - Its meaning overlaps with rendered decision `working_protocol.md:18` and + action item `working_protocol.md:30`, but it is not preserved as its own + action item. + +Important boundary: + +Items such as `Feedback zu den Kriterien...`, `Einbinden der Fachbereiche...` +and `Mail an Jovana und Björn...` are questionable todos, but they are already +`action_item` entries in the consolidated input: + +- `consolidated_extractions.json:1007` to `:1019` +- `consolidated_extractions.json:1537` to `:1549` +- `consolidated_extractions.json:1651` to `:1663` + +The renderer preserves those upstream action-item categories. It does not +create those todos from facts or prose. + +### Rendered Section: Prozessrahmen und Projektideen-Eingang / Open Questions + +Status: partially faithful, partially omitted. + +Rendered lines: + +- `working_protocol.md:36` to `:47` + +Preserved open questions: + +- `open_question_0001` -> `working_protocol.md:36` +- `open_question_0002` -> `working_protocol.md:37` +- `open_question_0003` -> `working_protocol.md:38` +- `open_question_0004` -> `working_protocol.md:39` +- `open_question_0005` -> `working_protocol.md:40` +- `open_question_0006` -> `working_protocol.md:41` +- `open_question_0007` -> `working_protocol.md:42` +- `open_question_0009` -> `working_protocol.md:43` +- `open_question_0010` -> `working_protocol.md:44` +- `open_question_0011` -> `working_protocol.md:45` +- `open_question_0012` -> `working_protocol.md:46` +- `open_question_0013` -> `working_protocol.md:47` + +Omitted open questions: + +- `open_question_0008`: `Wer soll die Rolle des Gatekeepers übernehmen und wie + wird der Prozess konkret implementiert?` + - Consolidated evidence: `consolidated_extractions.json:1387` to `:1397` + - No matching rendered open question. +- `open_question_0014`: `Welche Prüfsteine sind relevant für die + Fachabteilung?` + - Consolidated evidence: `consolidated_extractions.json:1765` to `:1775` + - No matching rendered open question. + +### Rendered Section: Projektkategorisierung und Status / Background + +Status: modified. + +Rendered line: + +- `working_protocol.md:53` + +This condenses: + +- fact: `Der Leiter F&E führt die Projektliste...` +- `technical_detail_0010`: project type definition +- fact: `Ein Projekt wandert direkt wieder in Business Development...` + +No renderer-only invention was verified in this section. + +### Rendered Section: Projektkategorisierung und Status / Decisions + +Status: category-changed. + +Rendered lines: + +- `working_protocol.md:57` +- `working_protocol.md:58` + +Both rendered decisions are facts in the consolidated input: + +- fact: `Im Zweifel gehen Projekte durch.` + - Consolidated evidence: `consolidated_extractions.json:775` to `:784` + - Rendered as decision: `working_protocol.md:57` +- fact: `Ein Projekt kann abgelehnt werden, aber es muss sichergestellt sein, + dass relevante Projekte nicht weggeschmissen werden.` + - Consolidated evidence: `consolidated_extractions.json:755` to `:764` + - Rendered as decision: `working_protocol.md:58` + +This is the clearest category change in the renderer output. + +### Rendered Section: Projektkategorisierung und Status / Action Items + +Status: omitted. + +Rendered line: + +- `working_protocol.md:62`: `*Keine.*` + +The consolidated input contains at least one action item that belongs to this +topic area: + +- `action_item_0008`: `Diskussion über die Relevanz eines speziellen Marktes + (z.B. Turkmenistan) führen.` + +The renderer omitted it. + +### Rendered Section: Projektkategorisierung und Status / Open Questions + +Status: faithful to rendered grouping, but unfaithful to complete input. + +Rendered line: + +- `working_protocol.md:66`: `*Keine.*` + +The complete consolidated input still contains project-process open questions, +including omitted `open_question_0008` and `open_question_0014`. They were not +rendered elsewhere. + +### Rendered Section: Technische Details und Infrastruktur / Background + +Status: modified. + +Rendered line: + +- `working_protocol.md:72` + +This condenses facts and technical details: + +- fact: `Das Netzwerk bei Gemeinden ist momentan langsam.` +- `technical_detail_0006`: `Das Netzwerk ist langsam (wie zu ISDN-Zeiten).` +- `technical_detail_0007`: `Der Rechner schmiert ab...` +- fact: `Das Dokument ist das aktuelle Projektdeckblatt...` +- fact: `Es gibt einen Reporting-Zyklus.` +- fact: `Sekretariate sind sehr abweisend...` + +The problem is not invention; the problem is that `decision_0004` +(`Festlegung eines Reporting-Zyklus...`) is downgraded from decision to +background. + +### Rendered Section: Technische Details und Infrastruktur / Decisions + +Status: omitted. + +Rendered line: + +- `working_protocol.md:76`: `*Keine.*` + +The consolidated input contains `decision_0004`, which is about a reporting +cycle and is rendered only as background. + +### Rendered Section: Technische Details und Infrastruktur / Action Items + +Status: faithful to rendered items. + +Rendered lines: + +- `working_protocol.md:80` +- `working_protocol.md:81` + +These correspond to: + +- `action_item_0013` +- `action_item_0014` + +The semantic quality of `action_item_0014` is questionable, but that issue is +upstream because the consolidated input already categorizes it as an action +item. + +### Rendered Section: Technische Details und Infrastruktur / Open Questions + +Status: omitted relative to complete input. + +Rendered line: + +- `working_protocol.md:85`: `*Keine.*` + +No technical open questions are clearly required here, but the overall rendered +protocol still omits consolidated open questions elsewhere. + +## Exact Mismatches + +| Consolidated item | Category before | Rendered text | Category after | Result | +| --- | --- | --- | --- | --- | +| `decision_0004`: `Festlegung eines Reporting-Zyklus...` | decision | `Es gibt einen Reporting-Zyklus...` | background | category weakened | +| `decision_0007`: `...Sitzung mit fünf Leuten...` | decision | none | omitted | omitted | +| `decision_0008`: `...Prozessrahmen nutzen...` | decision | none | omitted | omitted | +| `decision_0010`: `...Verbesserungsvorschläge... ausgeschleust...` | decision | none | omitted | omitted | +| `decision_0012`: `Die Diskussion wird beendet...` | decision | none | omitted | omitted | +| `open_question_0008`: `Wer soll die Rolle des Gatekeepers übernehmen...` | open_question | none | omitted | omitted | +| `open_question_0014`: `Welche Prüfsteine sind relevant...` | open_question | none | omitted | omitted | +| fact: `Im Zweifel gehen Projekte durch.` | fact | `Im Zweifel gehen Projekte durch.` | decision | promoted | +| fact: `Ein Projekt kann abgelehnt werden...` | fact | same meaning | decision | promoted | +| `action_item_0008`: `Diskussion über die Relevanz... Turkmenistan...` | action_item | none | omitted | omitted | +| `action_item_0010`: `Filterkriterien für verschiedene Fälle... definieren.` | action_item | overlapped by decision/action item text | merged / not preserved as action item | modified | + +## Invented Content Check + +No renderer-only invented people were verified in this comparison. `Guido` and +`Noah` appear in the rendered protocol, but they are already present in the +consolidated input. This belongs to BUG-003, not BUG-006. + +No clear renderer-only invented responsibility was verified in this comparison. +Questionable todos in the rendered protocol are already categorized as action +items in the consolidated input. + +## Merged And Split Items + +Verified merges: + +- Multiple facts and decisions are merged into background paragraphs in + `working_protocol.md:7`, `:9`, `:53` and `:72`. +- `action_item_0010` is not preserved as a separate action item and overlaps + with rendered filter-criteria decision/action-item wording. + +Verified splits: + +- No clear split from one consolidated item into multiple materially different + rendered items was identified. + +## Earliest Point Where The Mismatch Occurs + +The earliest point is the Working Protocol Renderer LLM response. + +Evidence: + +- `metadata.json` records renderer input as + `semantic_consolidator/consolidated_extractions.json` and prompt file + `prompts/working_protocol.md`. +- `raw_model_response.txt` already contains the omissions and category + changes. +- `raw_model_response.txt` and `working_protocol.md` are byte-identical. + +Therefore the mismatch is not introduced by renderer post-processing, output +normalization or Markdown generation. + +## Verified Root Cause + +Verified cause: + +The single-pass LLM Working Protocol Renderer does not reliably preserve the +category and coverage constraints of the consolidated input. It generates a +Markdown view that condenses and reorganizes the input, but the generated raw +response omits required items and changes some categories. + +Cause classification: + +- Renderer prompt: not proven as the sole root cause. The prompt includes + explicit preservation rules, but also includes condensation and redundancy + instructions. The evidence proves the model output violates the preservation + rules; it does not prove which prompt sentence caused the violation. +- Renderer post-processing: disproved. Raw response and final Markdown are + identical. +- Renderer output normalization: disproved. No separate normalization changed + the raw response. +- Markdown generation: disproved. The final file is the raw Markdown response. +- Another verified cause: single-pass LLM renderer generation violates + structural faithfulness requirements. + +## Confidence + +High for the location of the failure and for disproving post-processing, +normalization and Markdown generation as causes. + +Medium for root-cause classification beyond that. The evidence identifies the +renderer LLM response as the failing point, but does not isolate one prompt +sentence as the cause. + +## Conceptual Repair Strategy + +Do not attempt to repair this with free-form post-processing. + +Conceptually, renderer faithfulness needs a structural coverage contract: + +- every consolidated decision must be accounted for as a decision +- every consolidated action item must be accounted for as an action item +- every consolidated open question must be accounted for as an open question +- facts may appear in background, but must not be promoted to decisions +- omissions and category changes should be validator-detectable before the + protocol is accepted + +The renderer output should be validated against the consolidated input by item +ID or another stable structured reference, rather than relying on prose-only +Markdown to preserve category semantics implicitly. diff --git a/docs/constraint-repair.md b/docs/constraint-repair.md new file mode 100644 index 0000000..0d9e6bc --- /dev/null +++ b/docs/constraint-repair.md @@ -0,0 +1,348 @@ +# Constraint Repair Engine V1 + +Constraint Repair Engine V1 is a reusable structural repair stage for JSON +outputs produced by local LLM pipeline steps. + +It exists for cases where a model produced semantically usable JSON, but a +strict validator rejected the document because a structural invariant was +violated. Typical examples are repeated identifiers, missing identifiers, empty +containers or unstable item order. + +The repair stage is intentionally not integrated into the production pipeline +yet. It is a standalone module that future stages can opt into explicitly. + +## Architecture + +Package: + +```text +src/meeting_lab/constraint_repair/ +``` + +The package contains: + +- `engine.py`: generic JSON repair logic and the generic repair prompt +- `adapters/`: thin adapters from stage-specific validator results to the + generic validator report format + +The engine receives exactly two logical inputs: + +1. the complete original JSON output +2. a machine-generated validator report + +It never reads the original transcript and never receives Meeting Context or +other source material. The engine is domain-neutral: it operates on JSON +pointers, list keys and validator-reported identifiers. + +## Validator Interface + +The generic validator report has this shape: + +```json +{ + "valid": false, + "violations": [ + { + "type": "duplicate_id", + "collection_pointer": "/groups", + "id_list_key": "source_item_ids", + "id": "item_0001", + "occurrences": [ + {"item_index": 0, "id_index": 1}, + {"item_index": 3, "id_index": 0} + ], + "keep_occurrence": 0 + } + ] +} +``` + +Supported V1 violation types: + +- `duplicate_id`: remove repeated identifier occurrences from a list field +- `missing_id`: restore an identifier by appending a validator-provided item + template or adding it to a validator-specified existing item +- `empty_group`: remove an item whose identifier list is empty +- `reorder_items`: reorder a collection by validator-provided keys + +The report must provide enough structural information for the engine to repair +the document without interpreting content. + +## Meeting Context Constraint Validation Extension + +Meeting Context constraints can be added without making the repair engine +Meeting Context aware. + +The validator may inspect a stage output together with authoritative +`meeting_context.yaml`. It then converts violations into the generic validator +report format. The Constraint Repair Engine still receives only the original +JSON document and the validator report. + +Authoritative metadata is configuration, not meeting content. Examples include: + +- participant attendance +- canonical participant identity +- aliases +- departments +- roles + +Authoritative metadata may be consumed by LLMs and downstream pipeline stages. +It must never be redefined, overwritten or inferred by any pipeline component. + +This includes, but is not limited to: + +- LLM extraction +- Canonicalizer +- Semantic Consolidator +- Constraint Repair +- Renderer + +These components may consume authoritative metadata, but they must treat it as +immutable configuration. + +### Constraint Types + +`participant_attendance_conflict` + +A known participant from Meeting Context is represented in structured output +as absent, mentioned-only, external or otherwise not attending, contradicting +the authoritative attendance status. + +Example: + +```json +{ + "valid": false, + "constraint_source": { + "type": "meeting_context", + "meeting_id": "2026-07-27-projektprozess", + "source_file": "samples/real_live/project_process_meeting/meeting_context.yaml", + "schema_version": "1" + }, + "violations": [ + { + "type": "participant_attendance_conflict", + "severity": "error", + "collection_pointer": "/participants", + "item_index": 3, + "entity_id": "bjoern", + "display_name": "Björn", + "matched_alias": "Björn", + "field": "attendance_status", + "actual_value": "absent", + "expected_value": "present", + "repair": { + "operation": "set_field", + "field": "attendance_status", + "value": "present" + } + } + ] +} +``` + +`participant_unknown` + +A person-like entity appears in a structured participant or mentioned-person +field but is not known in Meeting Context as a participant, mentioned person or +alias. + +Example: + +```json +{ + "type": "participant_unknown", + "severity": "error", + "collection_pointer": "/participants", + "item_index": 4, + "matched_text": "Guido", + "repair": { + "operation": "remove_item" + } +} +``` + +Removal is allowed only when the unknown entity is a standalone structured +participant-like item and removing it does not remove unrelated content. If the +unknown name appears only inside generated prose, the repair must fail closed. + +`participant_alias_conflict` + +A structured entity reference uses an alias that maps to a different +authoritative entity, or uses an ambiguous alias that cannot be resolved to +exactly one Meeting Context entity. + +Example: + +```json +{ + "type": "participant_alias_conflict", + "severity": "error", + "json_pointer": "/participants/2/entity_id", + "matched_alias": "Johanna", + "expected_entity_id": "jovana", + "expected_display_name": "Jovana", + "actual_entity_id": "unknown", + "repair": { + "operation": "set_field", + "field": "entity_id", + "value": "jovana" + } +} +``` + +Alias repair is allowed only for structured identity fields and only when +Meeting Context maps the alias to exactly one entity. + +### Validation Principle + +The validator validates structured data whenever possible. Lexical analysis is +a fallback only when no structured representation exists. The long-term +objective is to reduce lexical validation over time by preserving Meeting +Context metadata as structured data throughout the pipeline. + +Lexical fallback may report a violation, but it must not create repair +instructions that require editing generated prose. + +### Repair Workflow + +```text +Stage output + ↓ +Meeting Context Constraint Validator + ↓ +Generic validator report + ↓ +Constraint Repair Engine + ↓ +Meeting Context Constraint Validator +``` + +If Validator #1 succeeds, repair is skipped. If Validator #1 fails, exactly one +repair pass may run when all violations map to deterministic structural +operations. Validator #2 then checks the repaired document. If Validator #2 +fails, the pipeline stops and reports the remaining violations. + +### Allowed Repairs + +Allowed repairs are deterministic structural operations only: + +- set a structured attendance field to the authoritative value +- set a structured entity ID to the authoritative entity ID +- move a structured participant item between participant collections +- remove a standalone structured unknown participant item +- remove empty structured participant containers +- reorder structured participant collections deterministically + +### Forbidden Repairs + +The repair engine must never perform free-form text editing. + +It must not: + +- remove names from generated prose +- replace text inside prose fields +- rewrite sentences +- infer attendance from transcript content +- invent participants +- invent aliases +- merge people +- split people +- change responsibility attribution +- change fact, decision, action-item or open-question meaning +- use Meeting Context directly + +If a violation exists only inside generated prose and cannot be repaired by a +deterministic structural operation, the repair must fail closed and report the +violation. + +### Integration Strategy + +Minimum useful integration for the current architecture: + +```text +Semantic Consolidator + ↓ +Meeting Context Constraint Validator + ↓ +Constraint Repair + ↓ +Meeting Context Constraint Validator + ↓ +Renderer +``` + +This catches contradictions in the consolidated representation before the +Working Protocol Renderer receives it. It does not require the renderer to +infer attendance semantics from prose. + +Future architecture: + +Meeting Context metadata should remain structured throughout the pipeline so +that the Renderer receives authoritative participant metadata directly instead +of having to infer it from generated prose. In that architecture, Meeting +Context validation can operate primarily on structured fields and use lexical +analysis only as a diagnostic fallback. + +## Allowed Operations + +The engine may: + +- move identifiers +- remove duplicate identifiers +- restore missing identifiers +- remove empty containers +- reorder items + +The engine must not: + +- invent information +- rewrite extracted text +- reinterpret reasons or explanations +- create new semantic relationships +- split semantic relationships + +Missing identifier repair is intentionally conservative. If the validator does +not provide an explicit target item or an item template, the engine refuses the +repair. + +## Generic Repair Prompt + +The module defines a generic prompt contract for future model-backed repair +backends. The prompt describes the task as repairing a structured JSON document +from a validator report and deliberately avoids stage-specific vocabulary. + +V1 unit tests assert that the prompt does not mention the current consolidation +stage, Meeting Context, facts, or merge groups. + +## Semantic Consolidator Adapter + +`constraint_repair.adapters.semantic_consolidator` converts the current +consolidation validator shape into the generic validator report. + +The adapter knows the current consolidation output field names such as +`groups` and `source_item_ids`. The generic repair engine does not. This keeps +stage-specific schema knowledge at the edge and preserves the repair engine as +a reusable JSON utility. + +The adapter currently reports: + +- repeated source IDs +- missing expected source IDs +- empty source-ID containers + +It does not call Ollama and does not change the production consolidation path. + +## Future Reuse + +Future pipeline stages can reuse the same engine by writing a small adapter +that maps their validator failures to the generic report format. The required +contract is that the adapter reports structural locations and does not ask the +engine to infer domain meaning. + +Potential future uses: + +- enforcing exact identifier coverage in LLM-generated grouping output +- removing empty generated containers before strict parsing +- restoring validator-known singleton items +- normalizing deterministic order after otherwise valid generation diff --git a/docs/data-models.md b/docs/data-models.md index bbc6aad..153e87f 100644 --- a/docs/data-models.md +++ b/docs/data-models.md @@ -266,6 +266,90 @@ renderer integration remains planned. --- +# Entity Registry + +Accepted Architecture. Implementation deferred. + +The Entity Registry is the persistent cross-meeting knowledge source for +confirmed entities, aliases and organizational metadata. It is independent from +individual meetings and is the planned long-term source used to prepare Meeting +Context V2. + +Entity types include: + +- people +- organizations +- departments +- products +- projects +- locations +- abbreviations + +Each entity has a stable internal identifier. The displayed name may change +over time, but the internal identifier must remain stable. + +Conceptual shape: + +```json +{ + "entity_id": "person_0001", + "entity_type": "person", + "display_name": "Jovana", + "aliases": [ + "Jovana", + "Giovanna", + "Jovanna", + "Giovana" + ], + "status": "confirmed" +} +``` + +The registry never learns automatically. It may propose matches, but only +confirmed user actions update it. Similarity search may suggest spelling +variants, Whisper transcription variants, umlaut variants or OCR-like mistakes, +but suggestions require explicit confirmation. + +Previously unseen names should be classified by the user as one of: + +- meeting participant +- mentioned person +- external person +- transcription error +- ignore + +The Entity Registry must not infer responsibility, decisions, attendance or +ownership. + +--- + +# Meeting Context V2 + +Accepted Architecture. Implementation deferred. + +Meeting Context V2 is an authoritative meeting-specific YAML Point of Truth +generated or assisted from: + +- Entity Registry +- user confirmations +- meeting metadata + +The YAML remains the extraction pipeline interface and the authoritative +meeting-specific Point of Truth for that meeting run. It is also a reproducible +input artifact: changes to the Entity Registry after a meeting run must not +silently change the historical Meeting Context used for that run. + +The Entity Registry remains the persistent cross-meeting knowledge source. It +must not override explicit meeting-specific confirmations. + +Meeting Context V2 should reduce manual work, improve alias handling, detect +transcription errors earlier and make Meeting Context quality scalable across +many meetings. + +See `docs/adr-meeting-context-v2-entity-registry.md`. + +--- + # Topic Result After extraction, every topic contains the collected information. diff --git a/docs/meeting-context.md b/docs/meeting-context.md index 6f20c52..ce62a77 100644 --- a/docs/meeting-context.md +++ b/docs/meeting-context.md @@ -26,6 +26,48 @@ Supported metadata: - abbreviations - relevant products, projects, systems, locations and technical terms +## Future Direction: V2 + +Accepted Architecture. Implementation deferred. + +Meeting Context V2 should be generated from an interactive entity confirmation +workflow after Whisper transcription: + +```text +Whisper + ↓ +Entity Detection + ↓ +User Confirmation + ↓ +Entity Registry Update + ↓ +Meeting Context Builder + ↓ +meeting_context.yaml + ↓ +Extraction Pipeline +``` + +The YAML remains the extraction interface and the authoritative +meeting-specific Point of Truth for a meeting run. It should become a +meeting-specific snapshot generated or assisted from the Entity Registry, user +confirmations and meeting metadata. + +The Entity Registry is the persistent cross-meeting knowledge source for +confirmed entities, aliases and organizational metadata. It stores stable +internal identifiers and never learns automatically. The Registry must not +override explicit meeting-specific confirmations, and Registry changes after a +meeting run must not silently change the historical Meeting Context used for +that run. + +Unknown names should be explicitly classified by the user as meeting +participant, mentioned person, external person, transcription error or ignore. +Similarity suggestions for spelling variants, Whisper variants, umlaut +handling and OCR-like mistakes require explicit confirmation. + +See `docs/adr-meeting-context-v2-entity-registry.md`. + ## File Locations - Generic template: `samples/templates/meeting_context.template.yaml` diff --git a/docs/pipeline.md b/docs/pipeline.md index 04d8034..944bc03 100644 --- a/docs/pipeline.md +++ b/docs/pipeline.md @@ -30,6 +30,8 @@ Canonical Meeting Knowledge and every Output View renderer. # Pipeline Overview +Current analysis pipeline: + ```text Whisper Transcript │ @@ -73,6 +75,46 @@ extraction, deterministic prompt injection and minimal extraction JSON provenance. Later Canonicalizer, Semantic Consolidator, Canonical Meeting Knowledge and renderer integration remains future work. +Accepted future Meeting Context V2 preparation flow: + +```text +Whisper + │ + ▼ +Entity Detection + │ + ▼ +User Confirmation + │ + ▼ +Entity Registry Update + │ + ▼ +Meeting Context Builder + │ + ▼ +meeting_context.yaml + │ + ▼ +Extraction Pipeline +``` + +This preparation flow is not implemented. It is the accepted long-term +direction for reducing manual Meeting Context work while preserving explicit +user control. The Entity Registry is the persistent cross-meeting knowledge +source for confirmed entities, aliases and organizational metadata. It stores +stable internal IDs and confirmed aliases, and never updates itself +automatically. + +For each meeting run, `meeting_context.yaml` remains the authoritative +meeting-specific Point of Truth and reproducible input artifact consumed by the +pipeline. V2 changes how that artifact is prepared: it may be generated or +assisted from Registry data, user confirmations and meeting metadata. The +Registry must not override explicit meeting-specific confirmations, and +Registry changes after a meeting run must not silently change the historical +Meeting Context used for that run. Similarity suggestions and unknown entity +classifications require explicit user confirmation. + --- # Stage 1 – Normalization diff --git a/docs/quality-readiness.md b/docs/quality-readiness.md new file mode 100644 index 0000000..fe6a8cf --- /dev/null +++ b/docs/quality-readiness.md @@ -0,0 +1,54 @@ +# Quality Readiness + +Meeting Lab quality status should distinguish engineering readiness from +practical usability. + +## Engineering Readiness + +Status: NOT READY + +The current end-to-end pipeline is not ready to replace the previous +extraction pipeline. The latest quality milestone records seven known +regression bugs, two documented root-cause investigations and a renderer +faithfulness failure where consolidated items are omitted or reclassified. + +Engineering readiness requires at least: + +- faithful rendering of consolidated decisions, action items and open questions +- no false responsibility attribution +- no participant-attendance contradictions against Meeting Context +- no unverified person/entity names entering final protocols unchecked +- regression tests or validation coverage for fixed bugs + +## Practical Usability + +Status: READY FOR MANUAL EDIT + +The generated working protocol can still be useful as a draft when reviewed by +a human editor against the known meeting content. This status does not imply +engineering readiness and must not be used as evidence that the pipeline is +faithful or production-ready. + +Practical usability means: + +- the broad meeting structure is recognizable +- many relevant topics and process points are present +- manual correction is still required before the protocol can be trusted + +Known manual correction areas include: + +- false or over-strong responsibility attribution +- incorrect open questions +- synthetic or unverified names from transcript artifacts +- participant-attendance contradictions +- reversed or over-broad process meaning + +## Current Benchmark Reference + +Latest benchmark reviewed for this distinction: + +- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/benchmark_report.md` + +The benchmark report concludes `NOT READY` for engineering replacement. This +document adds the separate practical-usability classification for manual-edit +workflows only. diff --git a/docs/regression-bugs.md b/docs/regression-bugs.md new file mode 100644 index 0000000..9f5ae22 --- /dev/null +++ b/docs/regression-bugs.md @@ -0,0 +1,478 @@ +# Regression Bug Tracker + +Living tracker for real bugs discovered during end-to-end Meeting Lab +evaluation. + +Purpose: + +- prevent forgotten regressions +- document known root causes +- record fixes when they happen +- verify that bugs never silently return + +## BUG-001 + +ID: BUG-001 + +Title: Renderer reverses the intended process direction + +Pipeline stage: Working Protocol Renderer + +Severity: High + +Status: Open + +Date discovered: 2026-08-03 + +Version first observed: `meeting_context_v1/final_protocol_with_context` + +Description: + +The generated working protocol phrases the process direction as if the +existing F&E process is broadly taken over for all project types. The intended +meaning from the evaluation is more constrained: the existing F&E process is a +framework that may be adapted or extended for broader project handling. + +Expected behaviour: + +The protocol should preserve the direction and uncertainty of the source +discussion: F&E process/framework adapted for other project types where +appropriate, without implying unconditional adoption for all project types. + +Actual behaviour: + +The protocol states that the existing F&E process is generally adopted for all +project types. + +Likely root cause: + +Unknown. + +Related files: + +- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/working_protocol.md` +- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/comparison.md` +- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/working_protocol/working_protocol.md` +- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/semantic_consolidator/consolidated_extractions.json` +- `prompts/working_protocol.md` + +Regression test available (yes/no): no + +Current status: + +Open. Documented from end-to-end evaluation output. + +Notes: + +Do not treat this entry as a prompt-change instruction. It records the observed +bug only. + +Current benchmark `meeting_context_v1/e2e_current_20260803_151000` confirms +the issue remains. The new protocol still says the existing F&E process is +basically taken over for all project types. + +## BUG-002 + +ID: BUG-002 + +Title: Jovana todo is strengthened beyond the meeting content + +Pipeline stage: Extraction / Working Protocol Renderer + +Severity: High + +Status: Open + +Date discovered: 2026-08-03 + +Version first observed: `meeting_context_v1/final_protocol_with_context` + +Description: + +The generated output strengthens a proposal or discussion about Jovana +assembling criteria into an assigned todo. The evaluation identified this as +overstating the meeting content. + +Expected behaviour: + +The pipeline should preserve the weaker source meaning unless the transcript +explicitly assigns, accepts or confirms responsibility. + +Actual behaviour: + +The consolidated input and final protocol include an action item assigning +criterion compilation to Jovana. + +Likely root cause: + +Unknown. + +Related files: + +- `samples/benchmarks/meeting_context_v1/semantic_consolidator_repair_v1/consolidated_extractions.json` +- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/working_protocol.md` +- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/comparison.md` +- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/semantic_consolidator/consolidated_extractions.json` +- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/working_protocol/working_protocol.md` +- `prompts/todos.md` +- `prompts/working_protocol.md` + +Regression test available (yes/no): no + +Current status: + +Open. Documented from end-to-end evaluation output. + +Notes: + +This bug is subject to the responsibility attribution invariant in +`AGENTS.md`. + +Current benchmark `meeting_context_v1/e2e_current_20260803_151000` confirms +the issue remains. The consolidated output still contains `Zusammenstellen der +Kriterien durch Jovana` with `responsible: "Jovana"`, and the final protocol +renders it as an action item. + +## BUG-003 + +ID: BUG-003 + +Title: Unexpected synthetic person names appear in generated protocols + +Pipeline stage: Extraction / Working Protocol Renderer + +Severity: Medium + +Status: Root Cause Identified + +Date discovered: 2026-08-03 + +Version first observed: `meeting_context_v1/final_protocol_with_context` + +Description: + +The generated protocol contains unexpected or risky names/aliases such as +`Guido` and `Noah`. These names were flagged during evaluation as synthetic or +inconsistent in the generated protocol context. + +Expected behaviour: + +The protocol should use only names supported by the transcript, consolidated +input or Meeting Context, and should preserve uncertainty when a name or alias +is unclear. + +Actual behaviour: + +The final protocol includes names or name variants that were flagged as +unexpected in evaluation. + +Likely root cause: + +The unexpected names already exist in the Whisper transcript before extraction. +They are propagated through the pipeline and are not introduced by prompts, +gold tests or the renderer. + +Related files: + +- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/working_protocol.md` +- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/comparison.md` +- `samples/benchmarks/meeting_context_v1/semantic_consolidator_repair_v1/consolidated_extractions.json` +- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/working_protocol/working_protocol.md` +- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/semantic_consolidator/consolidated_extractions.json` +- `samples/real_live/project_process_meeting/meeting_context.yaml` + +Regression test available (yes/no): no + +Current status: + +Root Cause Identified. Planned resolution is deferred until Meeting Context V2 +/ Entity Registry implementation. + +Notes: + +Resolution strategy: + +Introduce an interactive entity verification step after Whisper transcription: + +```text +Whisper + ↓ +Entity Detection + ↓ +User Verification + ↓ +Meeting Context Builder + ↓ +Extraction +``` + +Expected effect: + +Unknown or suspicious person names are detected before extraction begins and +can be classified by the user as participant, mentioned person, external +person, transcription error or ignore. + +Implementation status: + +Architecture accepted. Implementation deferred. + +Regression test: + +Not yet possible before Meeting Context V2 exists. + +Current benchmark `meeting_context_v1/e2e_current_20260803_151000` confirms +the issue remains. `Guido` and `Noah` both appear in the generated working +protocol. + +## BUG-004 + +ID: BUG-004 + +Title: Rhetorical question extracted as an open question + +Pipeline stage: Extraction + +Severity: Medium + +Status: Open + +Date discovered: 2026-08-03 + +Version first observed: `meeting_context_v1/final_protocol_with_context` + +Description: + +A rhetorical or discussion-framing question is extracted and later rendered as +an open question. This makes the protocol imply that the meeting left a real +follow-up question unresolved. + +Expected behaviour: + +Only genuine unresolved questions should be extracted as open questions. +Rhetorical questions or conversational framing should not become protocol open +questions. + +Actual behaviour: + +The final protocol includes at least one open question identified during +evaluation as rhetorical rather than genuinely open. + +Likely root cause: + +Unknown. + +Related files: + +- `samples/benchmarks/meeting_context_v1/semantic_consolidator_repair_v1/consolidated_extractions.json` +- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/working_protocol.md` +- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/comparison.md` +- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/semantic_consolidator/consolidated_extractions.json` +- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/working_protocol/working_protocol.md` +- `prompts/questions.md` +- `prompts/working_protocol.md` + +Regression test available (yes/no): no + +Current status: + +Open. Documented from end-to-end evaluation output. + +Notes: + +No specific fix has been proposed. + +Current benchmark `meeting_context_v1/e2e_current_20260803_151000` confirms +the issue remains. The final protocol still contains discussion-framing or +underspecified questions as open protocol questions, including the question +about exactly who is responsible for defining filter criteria. + +## BUG-005 + +ID: BUG-005 + +Title: Renderer treats Björn as absent although Meeting Context marks him present + +Pipeline stage: Working Protocol Renderer + +Severity: Medium + +Status: Investigating + +Date discovered: 2026-08-03 + +Version first observed: `meeting_context_v1/final_protocol_with_context` + +Description: + +The generated protocol asks how missing participants such as Jovana and Björn +should be included in the process, even though Meeting Context marks Björn as +present. This creates a misleading participant-status implication. + +Expected behaviour: + +The protocol should not mark or imply Björn as absent when Meeting Context +records him as present. If the source discussion concerns whether a person had +reviewed materials or been included in a process, that should not be converted +into meeting absence. + +Actual behaviour: + +The final protocol renders an open question about including missing +participants, including Björn. + +Likely root cause: + +The extraction model receives Meeting Context but still classifies Björn with +missing participants in chunk 08. Downstream artifacts preserve only minimal +Meeting Context provenance, not structured participant attendance metadata, so +Canonicalizer, Semantic Consolidator and Renderer cannot correct the conflict. + +Related files: + +- `samples/real_live/project_process_meeting/meeting_context.yaml` +- `samples/benchmarks/meeting_context_v1/semantic_consolidator_repair_v1/consolidated_extractions.json` +- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/working_protocol.md` +- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/comparison.md` +- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/semantic_consolidator/consolidated_extractions.json` +- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/working_protocol/working_protocol.md` +- `docs/bug-005-root-cause.md` +- `prompts/working_protocol.md` + +Regression test available (yes/no): no + +Current status: + +Investigating. Root-cause evidence has been documented in +`docs/bug-005-root-cause.md`, but no fix has been implemented or verified. +Current benchmark `meeting_context_v1/e2e_current_20260803_151000` confirms +the issue remains. + +Notes: + +This bug is distinct from responsibility attribution. It concerns participant +presence/status handling. + +Recommended resolution remains a Meeting Context constraint validator over +structured participant metadata before rendering. + +## BUG-006 + +ID: BUG-006 + +Title: Renderer drops required consolidated items and changes item categories + +Pipeline stage: Working Protocol Renderer + +Severity: High + +Status: Open + +Date discovered: 2026-08-03 + +Version first observed: `meeting_context_v1/e2e_current_20260803_151000` + +Description: + +The Working Protocol Renderer does not preserve every consolidated decision and +open question even though the renderer prompt requires preserving all +decisions, action items and open questions. It also renders some consolidated +facts as decisions. + +Expected behaviour: + +Every consolidated decision and open question should be represented in the +final protocol unless there is an explicit, validated reason to omit it. +Background facts must not be promoted into decisions. + +Actual behaviour: + +The consolidated input contains decisions such as `Festlegung eines +Reporting-Zyklus für abgelehnte Projekte`, `Zusammengetragen und beantwortete +Fragen führen zur Sitzung mit fünf Leuten...`, and `Die Diskussion wird +beendet und die Änderungen werden verschickt...`; these are not preserved as +decisions in the final protocol. The consolidated input also contains open +questions such as `Wer soll die Rolle des Gatekeepers übernehmen...` and +`Welche Prüfsteine sind relevant für die Fachabteilung?`; these are omitted +from the final protocol. Conversely, consolidated facts such as `Im Zweifel +gehen Projekte durch` and `Ein Projekt kann abgelehnt werden...` are rendered +as decisions. + +Likely root cause: + +Unknown. + +Related files: + +- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/semantic_consolidator/consolidated_extractions.json` +- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/working_protocol/working_protocol.md` +- `prompts/working_protocol.md` + +Regression test available (yes/no): no + +Current status: + +Open. Documented from current end-to-end benchmark output. + +Notes: + +This is a renderer faithfulness problem independent from whether the upstream +extracted items are themselves correct. + +## BUG-007 + +ID: BUG-007 + +Title: Non-committed discussion items are extracted and rendered as todos + +Pipeline stage: Extraction / Working Protocol Renderer + +Severity: High + +Status: Open + +Date discovered: 2026-08-03 + +Version first observed: `meeting_context_v1/e2e_current_20260803_151000` + +Description: + +The pipeline extracts and renders several action items that are not clearly +assigned, accepted or confirmed commitments in the meeting evidence. + +Expected behaviour: + +Only explicit assignments, accepted responsibilities or confirmed follow-up +actions should become todos. Discussion, examples, proposals, broad process +needs or unclear alternatives should remain facts/questions or be marked +unclear where supported. + +Actual behaviour: + +The current consolidated output and final protocol contain todos such as +`Feedback zu den Kriterien und Änderungen am Prozess abwarten`, `Einbinden der +Fachbereiche (stellvertretend durch Guido) zur Definition von Prüfsteinen`, +and `Mail an Jovana und Björn erneut senden oder Prozessanpassung vornehmen`. +The evidence for these items is discussion-level or ambiguous rather than a +clear todo commitment. + +Likely root cause: + +Unknown. + +Related files: + +- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/semantic_consolidator/consolidated_extractions.json` +- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/working_protocol/working_protocol.md` +- `prompts/todos.md` +- `prompts/working_protocol.md` + +Regression test available (yes/no): no + +Current status: + +Open. Documented from current end-to-end benchmark output. + +Notes: + +This generalizes BUG-002 beyond the specific Jovana assignment case and should +be evaluated against the responsibility attribution invariant.