Implement Semantic Consolidator V0

- add deterministic canonicalization support for extraction items
- add facts-only semantic consolidation using local Ollama
- preserve source evidence and validate complete fact coverage
- add conservative merge rules and non-LLM tests
- record the first validated real-life consolidation benchmark
- document current scope, limitations and next evaluation step
This commit is contained in:
2026-07-31 11:25:46 +02:00
parent 6e34334506
commit 90aa34d5d0
20 changed files with 5517 additions and 93 deletions
+42 -15
View File
@@ -17,6 +17,11 @@ Implemented:
- Transcript normalization in `src/meeting_lab/normalization/`.
- Technical chunking in `src/meeting_lab/chunking/`.
- Local per-chunk extraction in `src/meeting_lab/extraction/extract_chunks.py`.
- Deterministic Canonicalizer V1 in
`src/meeting_lab/consolidation/canonicalize.py`.
- Semantic Consolidator V0 in
`src/meeting_lab/consolidation/consolidate_facts.py` for facts-only
semantic duplicate detection.
- Prompt loading from `src/meeting_lab/llm/prompts.py`.
- Interim Markdown protocol generation in `src/meeting_lab/protocol/`.
- Non-LLM unit tests for chunking, extraction helpers, protocol rendering and
@@ -30,8 +35,8 @@ Experimental/prototype:
Planned:
- Deterministic Canonicalizer for extraction results.
- Semantic Consolidator for evidence-preserving semantic merging.
- Broader semantic consolidation for topic grouping, contradiction handling,
uncertainty marking and durable/transient separation.
- Canonical Meeting Knowledge implementation as the semantic source of truth.
- Final Working Protocol / Arbeitsprotokoll, Distribution Protocol /
Verteilerprotokoll and Knowledge Objects / Wissensdatenbankeintrag renderers.
@@ -123,17 +128,37 @@ Chunk Extractions
-> Output View Renderers
```
The Deterministic Canonicalizer is planned Python code with no LLM. It should
validate and normalize extraction objects, assign stable source references and
IDs, normalize category names and basic field structure, perform only safe
deterministic cleanup, optionally group exact duplicates, and preserve all
Canonicalizer V1 is implemented Python code with no LLM. It validates and
normalizes extraction objects, assigns stable source references and IDs,
normalizes category names and basic field structure, performs only safe
deterministic cleanup, optionally groups exact duplicates, and preserves all
source evidence. It must not perform uncertain semantic merging.
The Semantic Consolidator is planned local-LLM work. It should merge
semantically equivalent statements, group content by topic, preserve evidence
from all contributing chunks, mark contradictions and uncertainty, separate
durable information from transient discussion, and produce Canonical Meeting
Knowledge. It does not directly write a protocol.
CLI:
```text
PYTHONPATH=src .venv/bin/python -m meeting_lab.consolidation.canonicalize \
samples/whisper/meeting_speech_cleaned_chunks \
-o /tmp/canonicalized_extractions.json
```
Semantic Consolidator V0 is implemented as narrow local-LLM work for fact
items only. It merges semantically equivalent fact statements, preserves source
references and evidence, and prefers false negatives over false-positive
merges. It is not a summarizer, topic grouper, protocol renderer or complete
Canonical Meeting Knowledge stage.
The first accepted V0 benchmark used `qwen3.5:9B` in one Ollama call over 33
fact items. Runtime on the current machine was 390.119 seconds. One correct
merge was accepted, involving `fact_0025` and `fact_0031`; 31 facts remained
singletons, validation passed, no source fact was lost or duplicated, and
non-fact categories remained unchanged. This is a local benchmark, not a
general hardware claim.
Broader semantic consolidation remains planned. It should group content by
topic, mark contradictions and uncertainty, separate durable information from
transient discussion, and produce Canonical Meeting Knowledge. It does not
directly write a protocol.
Canonical Meeting Knowledge is the planned semantic intermediate model and
future single source of truth. It should be structured, preferably JSON, and
@@ -164,8 +189,8 @@ language is requested.
- Topic segmentation exists as prototype tooling, not a stable pipeline stage.
- Extraction is still a combined current flow, even though separate extractors
are the intended architecture.
- Deterministic Canonicalizer is documented but not implemented.
- Semantic Consolidator is documented but not implemented.
- Semantic Consolidator V0 is implemented only for facts-only duplicate
detection.
- Canonical Meeting Knowledge is documented but not implemented.
- Final output views are documented but not implemented.
- Most prompt files are placeholders except the common and decision prompts.
@@ -176,5 +201,7 @@ language is requested.
Stabilize repeatable local extraction evaluation before broadening the pipeline:
expand Gold Standard coverage by category, keep one-chunk extraction as the
baseline, and use small prompt experiments with immediate non-regression checks.
After extraction behavior is stable enough, implement the Deterministic
Canonicalizer first, then the Semantic Consolidator with evidence retention.
The next recommended evaluation step is to use the consolidated V0 result as
input for the unchanged Working Protocol renderer and compare that output
against the Working Protocol Synthesizer V0 baseline and the human reference
protocol.