Files
meeting-lab/docs/constraint-repair.md
T
admin 9446c6e0be Document validation architecture and renderer faithfulness findings
- document Entity Registry and Meeting Context V2 architecture
- preserve meeting_context.yaml as the authoritative meeting-specific input
- define immutable authoritative metadata across all pipeline stages
- restrict Constraint Repair to deterministic structured-data operations
- record BUG-003 root cause and deferred entity-verification resolution
- document BUG-005 attendance-consistency design
- add BUG-006 renderer faithfulness root-cause analysis
- distinguish Engineering Readiness from Practical Usability
- update the persistent regression bug tracker
2026-08-03 16:05:47 +02:00

349 lines
10 KiB
Markdown

# Constraint Repair Engine V1
Constraint Repair Engine V1 is a reusable structural repair stage for JSON
outputs produced by local LLM pipeline steps.
It exists for cases where a model produced semantically usable JSON, but a
strict validator rejected the document because a structural invariant was
violated. Typical examples are repeated identifiers, missing identifiers, empty
containers or unstable item order.
The repair stage is intentionally not integrated into the production pipeline
yet. It is a standalone module that future stages can opt into explicitly.
## Architecture
Package:
```text
src/meeting_lab/constraint_repair/
```
The package contains:
- `engine.py`: generic JSON repair logic and the generic repair prompt
- `adapters/`: thin adapters from stage-specific validator results to the
generic validator report format
The engine receives exactly two logical inputs:
1. the complete original JSON output
2. a machine-generated validator report
It never reads the original transcript and never receives Meeting Context or
other source material. The engine is domain-neutral: it operates on JSON
pointers, list keys and validator-reported identifiers.
## Validator Interface
The generic validator report has this shape:
```json
{
"valid": false,
"violations": [
{
"type": "duplicate_id",
"collection_pointer": "/groups",
"id_list_key": "source_item_ids",
"id": "item_0001",
"occurrences": [
{"item_index": 0, "id_index": 1},
{"item_index": 3, "id_index": 0}
],
"keep_occurrence": 0
}
]
}
```
Supported V1 violation types:
- `duplicate_id`: remove repeated identifier occurrences from a list field
- `missing_id`: restore an identifier by appending a validator-provided item
template or adding it to a validator-specified existing item
- `empty_group`: remove an item whose identifier list is empty
- `reorder_items`: reorder a collection by validator-provided keys
The report must provide enough structural information for the engine to repair
the document without interpreting content.
## Meeting Context Constraint Validation Extension
Meeting Context constraints can be added without making the repair engine
Meeting Context aware.
The validator may inspect a stage output together with authoritative
`meeting_context.yaml`. It then converts violations into the generic validator
report format. The Constraint Repair Engine still receives only the original
JSON document and the validator report.
Authoritative metadata is configuration, not meeting content. Examples include:
- participant attendance
- canonical participant identity
- aliases
- departments
- roles
Authoritative metadata may be consumed by LLMs and downstream pipeline stages.
It must never be redefined, overwritten or inferred by any pipeline component.
This includes, but is not limited to:
- LLM extraction
- Canonicalizer
- Semantic Consolidator
- Constraint Repair
- Renderer
These components may consume authoritative metadata, but they must treat it as
immutable configuration.
### Constraint Types
`participant_attendance_conflict`
A known participant from Meeting Context is represented in structured output
as absent, mentioned-only, external or otherwise not attending, contradicting
the authoritative attendance status.
Example:
```json
{
"valid": false,
"constraint_source": {
"type": "meeting_context",
"meeting_id": "2026-07-27-projektprozess",
"source_file": "samples/real_live/project_process_meeting/meeting_context.yaml",
"schema_version": "1"
},
"violations": [
{
"type": "participant_attendance_conflict",
"severity": "error",
"collection_pointer": "/participants",
"item_index": 3,
"entity_id": "bjoern",
"display_name": "Björn",
"matched_alias": "Björn",
"field": "attendance_status",
"actual_value": "absent",
"expected_value": "present",
"repair": {
"operation": "set_field",
"field": "attendance_status",
"value": "present"
}
}
]
}
```
`participant_unknown`
A person-like entity appears in a structured participant or mentioned-person
field but is not known in Meeting Context as a participant, mentioned person or
alias.
Example:
```json
{
"type": "participant_unknown",
"severity": "error",
"collection_pointer": "/participants",
"item_index": 4,
"matched_text": "Guido",
"repair": {
"operation": "remove_item"
}
}
```
Removal is allowed only when the unknown entity is a standalone structured
participant-like item and removing it does not remove unrelated content. If the
unknown name appears only inside generated prose, the repair must fail closed.
`participant_alias_conflict`
A structured entity reference uses an alias that maps to a different
authoritative entity, or uses an ambiguous alias that cannot be resolved to
exactly one Meeting Context entity.
Example:
```json
{
"type": "participant_alias_conflict",
"severity": "error",
"json_pointer": "/participants/2/entity_id",
"matched_alias": "Johanna",
"expected_entity_id": "jovana",
"expected_display_name": "Jovana",
"actual_entity_id": "unknown",
"repair": {
"operation": "set_field",
"field": "entity_id",
"value": "jovana"
}
}
```
Alias repair is allowed only for structured identity fields and only when
Meeting Context maps the alias to exactly one entity.
### Validation Principle
The validator validates structured data whenever possible. Lexical analysis is
a fallback only when no structured representation exists. The long-term
objective is to reduce lexical validation over time by preserving Meeting
Context metadata as structured data throughout the pipeline.
Lexical fallback may report a violation, but it must not create repair
instructions that require editing generated prose.
### Repair Workflow
```text
Stage output
↓
Meeting Context Constraint Validator
↓
Generic validator report
↓
Constraint Repair Engine
↓
Meeting Context Constraint Validator
```
If Validator #1 succeeds, repair is skipped. If Validator #1 fails, exactly one
repair pass may run when all violations map to deterministic structural
operations. Validator #2 then checks the repaired document. If Validator #2
fails, the pipeline stops and reports the remaining violations.
### Allowed Repairs
Allowed repairs are deterministic structural operations only:
- set a structured attendance field to the authoritative value
- set a structured entity ID to the authoritative entity ID
- move a structured participant item between participant collections
- remove a standalone structured unknown participant item
- remove empty structured participant containers
- reorder structured participant collections deterministically
### Forbidden Repairs
The repair engine must never perform free-form text editing.
It must not:
- remove names from generated prose
- replace text inside prose fields
- rewrite sentences
- infer attendance from transcript content
- invent participants
- invent aliases
- merge people
- split people
- change responsibility attribution
- change fact, decision, action-item or open-question meaning
- use Meeting Context directly
If a violation exists only inside generated prose and cannot be repaired by a
deterministic structural operation, the repair must fail closed and report the
violation.
### Integration Strategy
Minimum useful integration for the current architecture:
```text
Semantic Consolidator
↓
Meeting Context Constraint Validator
↓
Constraint Repair
↓
Meeting Context Constraint Validator
↓
Renderer
```
This catches contradictions in the consolidated representation before the
Working Protocol Renderer receives it. It does not require the renderer to
infer attendance semantics from prose.
Future architecture:
Meeting Context metadata should remain structured throughout the pipeline so
that the Renderer receives authoritative participant metadata directly instead
of having to infer it from generated prose. In that architecture, Meeting
Context validation can operate primarily on structured fields and use lexical
analysis only as a diagnostic fallback.
## Allowed Operations
The engine may:
- move identifiers
- remove duplicate identifiers
- restore missing identifiers
- remove empty containers
- reorder items
The engine must not:
- invent information
- rewrite extracted text
- reinterpret reasons or explanations
- create new semantic relationships
- split semantic relationships
Missing identifier repair is intentionally conservative. If the validator does
not provide an explicit target item or an item template, the engine refuses the
repair.
## Generic Repair Prompt
The module defines a generic prompt contract for future model-backed repair
backends. The prompt describes the task as repairing a structured JSON document
from a validator report and deliberately avoids stage-specific vocabulary.
V1 unit tests assert that the prompt does not mention the current consolidation
stage, Meeting Context, facts, or merge groups.
## Semantic Consolidator Adapter
`constraint_repair.adapters.semantic_consolidator` converts the current
consolidation validator shape into the generic validator report.
The adapter knows the current consolidation output field names such as
`groups` and `source_item_ids`. The generic repair engine does not. This keeps
stage-specific schema knowledge at the edge and preserves the repair engine as
a reusable JSON utility.
The adapter currently reports:
- repeated source IDs
- missing expected source IDs
- empty source-ID containers
It does not call Ollama and does not change the production consolidation path.
## Future Reuse
Future pipeline stages can reuse the same engine by writing a small adapter
that maps their validator failures to the generic report format. The required
contract is that the adapter reports structural locations and does not ask the
engine to infer domain meaning.
Potential future uses:
- enforcing exact identifier coverage in LLM-generated grouping output
- removing empty generated containers before strict parsing
- restoring validator-known singleton items
- normalizing deterministic order after otherwise valid generation