Implement Semantic Consolidator V0

- add deterministic canonicalization support for extraction items
- add facts-only semantic consolidation using local Ollama
- preserve source evidence and validate complete fact coverage
- add conservative merge rules and non-LLM tests
- record the first validated real-life consolidation benchmark
- document current scope, limitations and next evaluation step
This commit is contained in:
2026-07-31 11:25:46 +02:00
parent 6e34334506
commit 90aa34d5d0
20 changed files with 5517 additions and 93 deletions
+34 -15
View File
@@ -135,7 +135,15 @@ Deterministic
## Current Status
Planned
Implemented as Canonicalizer V1.
CLI:
```text
PYTHONPATH=src .venv/bin/python -m meeting_lab.consolidation.canonicalize \
samples/whisper/meeting_speech_cleaned_chunks \
-o /tmp/canonicalized_extractions.json
```
---
@@ -231,7 +239,7 @@ LLM
## Current Status
Planned
Implemented as Canonicalizer V1.
---
@@ -325,16 +333,23 @@ Canonicalized extraction objects.
## Output
Canonical Meeting Knowledge.
Semantic Consolidator V0 output is `consolidated_extractions.json` with fact
groups and unchanged non-fact items.
Future broader semantic consolidation should produce Canonical Meeting
Knowledge.
## Responsibilities
- Merge semantically equivalent statements
- Group content by topic
- V0: merge semantically equivalent fact items only
- V0: preserve all non-fact categories unchanged
- V0: validate that every source fact ID appears exactly once
- Future: merge semantically equivalent statements across categories
- Future: group content by topic
- Preserve evidence from all contributing chunks
- Mark contradictions and uncertainty
- Separate durable information from transient discussion
- Reconcile category shifts where supported by evidence
- Future: mark contradictions and uncertainty
- Future: separate durable information from transient discussion
- Future: reconcile category shifts where supported by evidence
## Must Not
@@ -348,7 +363,9 @@ Local LLM, with deterministic pre/post-processing where useful.
## Current Status
Planned
Semantic Consolidator V0 is implemented and experimentally validated for
facts-only conservative duplicate detection. Broader semantic consolidation and
Canonical Meeting Knowledge generation remain planned.
---
@@ -507,16 +524,18 @@ A processing stage may be replaced by another implementation as long as it prese
⬜ Specialized Extraction
⬜ Deterministic Canonicalization
✔ Deterministic Canonicalization
⬜ Semantic Consolidation
✅ Semantic Consolidation V0 - facts-only duplicate detection
⬜ Canonical Meeting Knowledge
⬜ Output View Rendering
```
The immediate architecture focus is the Deterministic Canonicalizer followed by
the Semantic Consolidator. These stages preserve source evidence, recover
global context from independent chunk extractions and prepare Canonical Meeting
Knowledge for parallel Output View rendering.
The immediate evaluation focus is using the Semantic Consolidator V0 output as
input for the unchanged Working Protocol renderer. Canonicalizer V1 now
preserves source evidence, and Semantic Consolidator V0 conservatively merges
semantically equivalent fact items. Broader semantic consolidation should later
recover global context from independent chunk extractions and prepare Canonical
Meeting Knowledge for parallel Output View rendering.