Implement Semantic Consolidator V0
- add deterministic canonicalization support for extraction items - add facts-only semantic consolidation using local Ollama - preserve source evidence and validate complete fact coverage - add conservative merge rules and non-LLM tests - record the first validated real-life consolidation benchmark - document current scope, limitations and next evaluation step
This commit is contained in:
+42
-15
@@ -17,6 +17,11 @@ Implemented:
|
||||
- Transcript normalization in `src/meeting_lab/normalization/`.
|
||||
- Technical chunking in `src/meeting_lab/chunking/`.
|
||||
- Local per-chunk extraction in `src/meeting_lab/extraction/extract_chunks.py`.
|
||||
- Deterministic Canonicalizer V1 in
|
||||
`src/meeting_lab/consolidation/canonicalize.py`.
|
||||
- Semantic Consolidator V0 in
|
||||
`src/meeting_lab/consolidation/consolidate_facts.py` for facts-only
|
||||
semantic duplicate detection.
|
||||
- Prompt loading from `src/meeting_lab/llm/prompts.py`.
|
||||
- Interim Markdown protocol generation in `src/meeting_lab/protocol/`.
|
||||
- Non-LLM unit tests for chunking, extraction helpers, protocol rendering and
|
||||
@@ -30,8 +35,8 @@ Experimental/prototype:
|
||||
|
||||
Planned:
|
||||
|
||||
- Deterministic Canonicalizer for extraction results.
|
||||
- Semantic Consolidator for evidence-preserving semantic merging.
|
||||
- Broader semantic consolidation for topic grouping, contradiction handling,
|
||||
uncertainty marking and durable/transient separation.
|
||||
- Canonical Meeting Knowledge implementation as the semantic source of truth.
|
||||
- Final Working Protocol / Arbeitsprotokoll, Distribution Protocol /
|
||||
Verteilerprotokoll and Knowledge Objects / Wissensdatenbankeintrag renderers.
|
||||
@@ -123,17 +128,37 @@ Chunk Extractions
|
||||
-> Output View Renderers
|
||||
```
|
||||
|
||||
The Deterministic Canonicalizer is planned Python code with no LLM. It should
|
||||
validate and normalize extraction objects, assign stable source references and
|
||||
IDs, normalize category names and basic field structure, perform only safe
|
||||
deterministic cleanup, optionally group exact duplicates, and preserve all
|
||||
Canonicalizer V1 is implemented Python code with no LLM. It validates and
|
||||
normalizes extraction objects, assigns stable source references and IDs,
|
||||
normalizes category names and basic field structure, performs only safe
|
||||
deterministic cleanup, optionally groups exact duplicates, and preserves all
|
||||
source evidence. It must not perform uncertain semantic merging.
|
||||
|
||||
The Semantic Consolidator is planned local-LLM work. It should merge
|
||||
semantically equivalent statements, group content by topic, preserve evidence
|
||||
from all contributing chunks, mark contradictions and uncertainty, separate
|
||||
durable information from transient discussion, and produce Canonical Meeting
|
||||
Knowledge. It does not directly write a protocol.
|
||||
CLI:
|
||||
|
||||
```text
|
||||
PYTHONPATH=src .venv/bin/python -m meeting_lab.consolidation.canonicalize \
|
||||
samples/whisper/meeting_speech_cleaned_chunks \
|
||||
-o /tmp/canonicalized_extractions.json
|
||||
```
|
||||
|
||||
Semantic Consolidator V0 is implemented as narrow local-LLM work for fact
|
||||
items only. It merges semantically equivalent fact statements, preserves source
|
||||
references and evidence, and prefers false negatives over false-positive
|
||||
merges. It is not a summarizer, topic grouper, protocol renderer or complete
|
||||
Canonical Meeting Knowledge stage.
|
||||
|
||||
The first accepted V0 benchmark used `qwen3.5:9B` in one Ollama call over 33
|
||||
fact items. Runtime on the current machine was 390.119 seconds. One correct
|
||||
merge was accepted, involving `fact_0025` and `fact_0031`; 31 facts remained
|
||||
singletons, validation passed, no source fact was lost or duplicated, and
|
||||
non-fact categories remained unchanged. This is a local benchmark, not a
|
||||
general hardware claim.
|
||||
|
||||
Broader semantic consolidation remains planned. It should group content by
|
||||
topic, mark contradictions and uncertainty, separate durable information from
|
||||
transient discussion, and produce Canonical Meeting Knowledge. It does not
|
||||
directly write a protocol.
|
||||
|
||||
Canonical Meeting Knowledge is the planned semantic intermediate model and
|
||||
future single source of truth. It should be structured, preferably JSON, and
|
||||
@@ -164,8 +189,8 @@ language is requested.
|
||||
- Topic segmentation exists as prototype tooling, not a stable pipeline stage.
|
||||
- Extraction is still a combined current flow, even though separate extractors
|
||||
are the intended architecture.
|
||||
- Deterministic Canonicalizer is documented but not implemented.
|
||||
- Semantic Consolidator is documented but not implemented.
|
||||
- Semantic Consolidator V0 is implemented only for facts-only duplicate
|
||||
detection.
|
||||
- Canonical Meeting Knowledge is documented but not implemented.
|
||||
- Final output views are documented but not implemented.
|
||||
- Most prompt files are placeholders except the common and decision prompts.
|
||||
@@ -176,5 +201,7 @@ language is requested.
|
||||
Stabilize repeatable local extraction evaluation before broadening the pipeline:
|
||||
expand Gold Standard coverage by category, keep one-chunk extraction as the
|
||||
baseline, and use small prompt experiments with immediate non-regression checks.
|
||||
After extraction behavior is stable enough, implement the Deterministic
|
||||
Canonicalizer first, then the Semantic Consolidator with evidence retention.
|
||||
The next recommended evaluation step is to use the consolidated V0 result as
|
||||
input for the unchanged Working Protocol renderer and compare that output
|
||||
against the Working Protocol Synthesizer V0 baseline and the human reference
|
||||
protocol.
|
||||
|
||||
Reference in New Issue
Block a user