Implement Semantic Consolidator V0

- add deterministic canonicalization support for extraction items
- add facts-only semantic consolidation using local Ollama
- preserve source evidence and validate complete fact coverage
- add conservative merge rules and non-LLM tests
- record the first validated real-life consolidation benchmark
- document current scope, limitations and next evaluation step
This commit is contained in:
2026-07-31 11:25:46 +02:00
parent 6e34334506
commit 90aa34d5d0
20 changed files with 5517 additions and 93 deletions
+20 -15
View File
@@ -30,10 +30,11 @@ Audio
Status:
- Implemented: Whisper JSON cleanup script, normalization, technical chunking,
local chunk extraction, interim Markdown protocol builder.
local chunk extraction, Canonicalizer V1, Semantic Consolidator V0
facts-only duplicate detection, interim Markdown protocol builder.
- Experimental/prototype: topic segmentation and review tooling.
- Planned: Deterministic Canonicalizer, Semantic Consolidator, Canonical
Meeting Knowledge implementation, final Output Views.
- Planned: broader semantic consolidation, Canonical Meeting Knowledge
implementation, final Output Views.
## Architectural Principles
@@ -46,16 +47,20 @@ Status:
- Extraction, consolidation, synthesis and rendering are separate concerns.
- Deterministic canonicalization and semantic consolidation are separate
concerns.
- The Deterministic Canonicalizer is planned Python code. It validates and
normalizes extraction objects, assigns stable source references and IDs,
normalizes category names and basic field structure, performs only safe
deterministic cleanup, may group exact duplicates, and must preserve all
source evidence. It must not perform uncertain semantic merging.
- The Semantic Consolidator is planned local-LLM work. It merges semantically
equivalent statements, groups content by topic, preserves evidence from all
contributing chunks, marks contradictions and uncertainty, separates durable
information from transient discussion, and produces Canonical Meeting
Knowledge. It does not directly write a protocol.
- Canonicalizer V1 is implemented Python code. It validates and normalizes
extraction objects, assigns stable source references and IDs, normalizes
category names and basic field structure, performs only safe deterministic
cleanup, may group exact duplicates, and must preserve all source evidence.
It must not perform uncertain semantic merging.
- Semantic Consolidator V0 is implemented as local-LLM facts-only duplicate
detection. It merges semantically equivalent fact items conservatively,
preserves source references and evidence, and validates complete source fact
coverage. It is not a summarizer, topic grouper, protocol renderer or
complete Canonical Meeting Knowledge stage.
- Broader semantic consolidation remains planned. It should group content by
topic, mark contradictions and uncertainty, separate durable information from
transient discussion, and prepare Canonical Meeting Knowledge. It does not
directly write a protocol.
- Prefer small, testable processing stages over one monolithic LLM prompt.
- Current extraction strategy is one normalized chunk per LLM call.
- Do not expand context windows or redesign the extraction strategy without an
@@ -132,8 +137,8 @@ Current documented Prompt Version 2 decision baseline:
- `src/meeting_lab/segmentation/`: experimental topic segmentation tooling.
- `src/meeting_lab/extraction/`: local LLM extraction flow and category
extractor modules.
- `src/meeting_lab/consolidation/`: planned deterministic canonicalization and
semantic consolidation area.
- `src/meeting_lab/consolidation/`: Canonicalizer V1 and Semantic
Consolidator V0.
- `src/meeting_lab/protocol/`: interim protocol rendering.
- `src/meeting_lab/models/`: current lightweight data models.
- `samples/`: sample inputs and generated/experimental artifacts; do not treat