Document validation architecture and renderer faithfulness findings
- document Entity Registry and Meeting Context V2 architecture - preserve meeting_context.yaml as the authoritative meeting-specific input - define immutable authoritative metadata across all pipeline stages - restrict Constraint Repair to deterministic structured-data operations - record BUG-003 root cause and deferred entity-verification resolution - document BUG-005 attendance-consistency design - add BUG-006 renderer faithfulness root-cause analysis - distinguish Engineering Readiness from Practical Usability - update the persistent regression bug tracker
This commit is contained in:
@@ -0,0 +1,153 @@
|
||||
# ADR: Meeting Context V2 and Entity Registry
|
||||
|
||||
Status: Accepted Architecture
|
||||
|
||||
Implementation: Deferred
|
||||
|
||||
Date: 2026-08-03
|
||||
|
||||
## Context
|
||||
|
||||
Meeting Context V1 proved that authoritative meeting context can significantly
|
||||
improve extraction quality. It helps the extractor normalize known aliases,
|
||||
identify participants and avoid treating context metadata as evidence for
|
||||
responsibility or decisions.
|
||||
|
||||
End-to-end evaluation also showed that manually writing Meeting Context is not
|
||||
the right long-term primary workflow. The pipeline needs an earlier,
|
||||
interactive entity confirmation step after Whisper transcription. That step
|
||||
should identify candidate entities, ask the user to confirm them and then build
|
||||
the meeting-specific context from confirmed data.
|
||||
|
||||
## Decision
|
||||
|
||||
Meeting Context should evolve toward V2 as an authoritative,
|
||||
meeting-specific Point of Truth generated or assisted from a persistent Entity
|
||||
Registry, user confirmations and meeting metadata.
|
||||
|
||||
Preferred future pipeline:
|
||||
|
||||
```text
|
||||
Whisper
|
||||
↓
|
||||
Entity Detection
|
||||
↓
|
||||
User Confirmation
|
||||
↓
|
||||
Entity Registry Update
|
||||
↓
|
||||
Meeting Context Builder
|
||||
↓
|
||||
meeting_context.yaml
|
||||
↓
|
||||
Extraction Pipeline
|
||||
```
|
||||
|
||||
The Entity Registry is the persistent cross-meeting knowledge source for
|
||||
confirmed entities, aliases and organizational metadata.
|
||||
|
||||
For each individual meeting, `meeting_context.yaml` remains the authoritative
|
||||
meeting-specific Point of Truth and reproducible input artifact consumed by
|
||||
the extraction pipeline. Meeting Context V2 changes how this YAML is prepared,
|
||||
not its authority for a meeting run.
|
||||
|
||||
The generated `meeting_context.yaml` is a meeting-specific snapshot. The
|
||||
Registry must not override explicit meeting-specific confirmations. Changes to
|
||||
the Registry after a meeting run must not silently change the historical
|
||||
Meeting Context used for that run.
|
||||
|
||||
## Entity Registry
|
||||
|
||||
The Entity Registry is the persistent cross-meeting knowledge source and is
|
||||
independent from individual meetings. It stores confirmed entities such as:
|
||||
|
||||
- people
|
||||
- organizations
|
||||
- departments
|
||||
- products
|
||||
- projects
|
||||
- locations
|
||||
- abbreviations
|
||||
|
||||
Each entity receives a stable internal identifier. The displayed name may
|
||||
change over time, but the identifier must remain stable.
|
||||
|
||||
## Learning Principle
|
||||
|
||||
The registry never learns automatically.
|
||||
|
||||
It may propose matches, but only confirmed user actions update the registry.
|
||||
No autonomous learning is allowed.
|
||||
|
||||
## Alias Handling
|
||||
|
||||
Aliases are first-class data.
|
||||
|
||||
Examples:
|
||||
|
||||
```text
|
||||
Jovana
|
||||
Giovanna
|
||||
Jovanna
|
||||
Giovana
|
||||
```
|
||||
|
||||
These variants may all refer to one confirmed entity. Future runs should
|
||||
automatically suggest previously confirmed aliases, but those suggestions still
|
||||
require explicit confirmation when they would update registry data.
|
||||
|
||||
## Unknown Entities
|
||||
|
||||
Previously unseen names are presented to the user for classification.
|
||||
|
||||
Possible classifications:
|
||||
|
||||
- meeting participant
|
||||
- mentioned person
|
||||
- external person
|
||||
- transcription error
|
||||
- ignore
|
||||
|
||||
Nothing is automatically accepted.
|
||||
|
||||
## Similarity Search
|
||||
|
||||
Similarity search is a future extension. It can propose likely matches for:
|
||||
|
||||
- spelling variants
|
||||
- Whisper transcription variants
|
||||
- umlaut handling
|
||||
- OCR-like mistakes
|
||||
|
||||
Similarity suggestions require explicit confirmation.
|
||||
|
||||
## Rationale
|
||||
|
||||
Expected advantages:
|
||||
|
||||
- significantly less manual work
|
||||
- earlier detection of transcription errors
|
||||
- robust alias handling
|
||||
- reusable organizational knowledge
|
||||
- improved Meeting Context quality
|
||||
- easier GUI workflow
|
||||
- better scalability across many meetings
|
||||
|
||||
## Consequences
|
||||
|
||||
Meeting Context V1 remains the current implemented interface.
|
||||
|
||||
Meeting Context V2 should preserve the YAML interface for extraction.
|
||||
`meeting_context.yaml` remains the authoritative meeting-specific Point of
|
||||
Truth and reproducible input artifact for a meeting run. The Entity Registry is
|
||||
the persistent cross-meeting knowledge source used to prepare that artifact.
|
||||
|
||||
The Registry must not override explicit meeting-specific confirmations. Changes
|
||||
to the Registry after a meeting run must not silently change the historical
|
||||
Meeting Context used for that run.
|
||||
|
||||
The registry must not infer responsibility, decisions, attendance or ownership.
|
||||
Those still require meeting evidence and remain governed by the existing
|
||||
responsibility attribution invariant.
|
||||
|
||||
No implementation is part of this ADR.
|
||||
Reference in New Issue
Block a user