Implement Semantic Consolidator V0

- add deterministic canonicalization support for extraction items
- add facts-only semantic consolidation using local Ollama
- preserve source evidence and validate complete fact coverage
- add conservative merge rules and non-LLM tests
- record the first validated real-life consolidation benchmark
- document current scope, limitations and next evaluation step
This commit is contained in:
2026-07-31 11:25:46 +02:00
parent 6e34334506
commit 90aa34d5d0
20 changed files with 5517 additions and 93 deletions
+135 -3
View File
@@ -954,9 +954,10 @@ marks contradictions or uncertainty.
Decision:
Deterministic canonicalization is the next implementation step after stable
local extraction, followed by semantic consolidation. Canonical Meeting
Knowledge and final output views are planned, not implemented.
Deterministic canonicalization is implemented as Canonicalizer V1. The first
semantic consolidation milestone is implemented as Semantic Consolidator V0 for
facts-only duplicate detection. Broader semantic consolidation, Canonical
Meeting Knowledge and final output views remain planned.
Lessons learned:
@@ -1037,3 +1038,134 @@ Evidence:
- `PROJECT_KNOWLEDGE.md`
- `docs/output-views.md`
- See EXP-0017 and EXP-0019.
## EXP-0021 - Canonicalizer V1
Status: Accepted
Date or period: 2026-07-31
Hypothesis:
Independent chunk extraction JSON can be converted into a stable deterministic
intermediate format before any semantic LLM consolidation is attempted.
Setup:
Canonicalizer V1 discovers `chunk_*_extraction.json` files in stable chunk
order, validates required categories, normalizes category names and basic field
structure, parses existing legacy string formats where safe, trims redundant
whitespace, assigns deterministic IDs, preserves original values and source
references, and merges only exact duplicates when all semantic fields are
identical.
Inputs:
- Synthetic unit-test fixtures.
- Existing nine extraction JSON files under
`samples/whisper/meeting_speech_cleaned_chunks/`.
Model / configuration:
- No LLM.
- CLI module: `meeting_lab.consolidation.canonicalize`.
Result:
Canonicalizer V1 produces `schema_version`, `source_files`, `stats` and
`items`. It is deterministic preparation for the future Semantic Consolidator
and is not Canonical Meeting Knowledge.
Decision:
Canonicalizer V1 is the current implemented deterministic canonicalization
stage. Semantic Consolidator V0 now uses this representation for facts-only
semantic duplicate detection; broader semantic consolidation and Canonical
Meeting Knowledge remain planned.
Lessons learned:
Exact duplicate handling, source-reference preservation and legacy string
parsing can be tested without model calls. Any uncertain semantic merge remains
out of scope for this stage.
Evidence:
- `src/meeting_lab/consolidation/canonicalize.py`
- `tests/test_canonicalize.py`
- `docs/data-models.md`
- `docs/pipeline.md`
## EXP-0022 - Semantic Consolidator V0 facts-only merge
Status: Accepted
Date or period: 2026-07-31
Hypothesis:
The Canonicalizer V1 output contains enough stable structure for a local LLM to
identify semantically equivalent fact items without losing source coverage or
changing non-fact categories.
Setup:
Semantic Consolidator V0 consumed the existing Canonicalizer V1 benchmark JSON,
selected only items with `category: "fact"`, and sent one bounded consolidation
request to local Ollama. The merge rules required semantic equivalence, not
topic similarity, and validation required every source fact ID to appear
exactly once.
Inputs:
- `samples/benchmarks/canonicalizer_v1/canonicalized_extractions.json`
- 33 fact items.
Model / configuration:
- `qwen3.5:9B`
- Ollama endpoint: `http://127.0.0.1:11434/api/generate`
- Thinking disabled.
- One LLM call.
- `num_ctx=32768`
- `num_predict=4096`
Result:
- Runtime: 390.119 seconds on the current machine.
- Merged fact groups: 1.
- Source facts involved in merges: 2.
- Singleton fact groups: 31.
- Validation: passed.
- No source fact was lost or duplicated.
- Non-fact categories remained unchanged.
Accepted merge:
- `fact_0025` + `fact_0031`
- Canonical statement: "Der Leiter F&E führt die Projektliste auf dem
zweiwöchentlichen Schnittstellen-Stand-Up."
Decision:
Semantic Consolidator V0 is complete for its current narrow scope:
conservative facts-only semantic duplicate detection with source evidence
preserved. The selected `report.md` and `consolidated_extractions.json`
benchmark artifacts should be versioned for later comparison. The raw model
response remains a local diagnostic artifact and is not versioned.
Lessons learned:
Semantic duplicate consolidation is technically viable and conservative enough
for continued evaluation, but broader semantic synthesis remains a separate
future stage. The measured runtime is useful for this machine and run, but
should not be generalized into a universal benchmark.
Evidence:
- `src/meeting_lab/consolidation/consolidate_facts.py`
- `prompts/consolidate_facts.md`
- `tests/test_consolidate_facts.py`
- `samples/benchmarks/semantic_consolidator_v0/report.md`
- `samples/benchmarks/semantic_consolidator_v0/consolidated_extractions.json`
- Local diagnostic only: `samples/benchmarks/semantic_consolidator_v0/raw_model_response.txt`