Implement Semantic Consolidator V0

- add deterministic canonicalization support for extraction items
- add facts-only semantic consolidation using local Ollama
- preserve source evidence and validate complete fact coverage
- add conservative merge rules and non-LLM tests
- record the first validated real-life consolidation benchmark
- document current scope, limitations and next evaluation step
This commit is contained in:
2026-07-31 11:25:46 +02:00
parent 6e34334506
commit 90aa34d5d0
20 changed files with 5517 additions and 93 deletions
+1
View File
@@ -39,6 +39,7 @@ samples/whisper/
**/meeting_protocol.md **/meeting_protocol.md
**/*.raw.txt **/*.raw.txt
tests/gold/**/actual.json tests/gold/**/actual.json
samples/benchmarks/semantic_consolidator_v0/raw_model_response.txt
# Lokale Meetings (niemals versionieren) # Lokale Meetings (niemals versionieren)
meeting_data/ meeting_data/
+20 -15
View File
@@ -30,10 +30,11 @@ Audio
Status: Status:
- Implemented: Whisper JSON cleanup script, normalization, technical chunking, - Implemented: Whisper JSON cleanup script, normalization, technical chunking,
local chunk extraction, interim Markdown protocol builder. local chunk extraction, Canonicalizer V1, Semantic Consolidator V0
facts-only duplicate detection, interim Markdown protocol builder.
- Experimental/prototype: topic segmentation and review tooling. - Experimental/prototype: topic segmentation and review tooling.
- Planned: Deterministic Canonicalizer, Semantic Consolidator, Canonical - Planned: broader semantic consolidation, Canonical Meeting Knowledge
Meeting Knowledge implementation, final Output Views. implementation, final Output Views.
## Architectural Principles ## Architectural Principles
@@ -46,16 +47,20 @@ Status:
- Extraction, consolidation, synthesis and rendering are separate concerns. - Extraction, consolidation, synthesis and rendering are separate concerns.
- Deterministic canonicalization and semantic consolidation are separate - Deterministic canonicalization and semantic consolidation are separate
concerns. concerns.
- The Deterministic Canonicalizer is planned Python code. It validates and - Canonicalizer V1 is implemented Python code. It validates and normalizes
normalizes extraction objects, assigns stable source references and IDs, extraction objects, assigns stable source references and IDs, normalizes
normalizes category names and basic field structure, performs only safe category names and basic field structure, performs only safe deterministic
deterministic cleanup, may group exact duplicates, and must preserve all cleanup, may group exact duplicates, and must preserve all source evidence.
source evidence. It must not perform uncertain semantic merging. It must not perform uncertain semantic merging.
- The Semantic Consolidator is planned local-LLM work. It merges semantically - Semantic Consolidator V0 is implemented as local-LLM facts-only duplicate
equivalent statements, groups content by topic, preserves evidence from all detection. It merges semantically equivalent fact items conservatively,
contributing chunks, marks contradictions and uncertainty, separates durable preserves source references and evidence, and validates complete source fact
information from transient discussion, and produces Canonical Meeting coverage. It is not a summarizer, topic grouper, protocol renderer or
Knowledge. It does not directly write a protocol. complete Canonical Meeting Knowledge stage.
- Broader semantic consolidation remains planned. It should group content by
topic, mark contradictions and uncertainty, separate durable information from
transient discussion, and prepare Canonical Meeting Knowledge. It does not
directly write a protocol.
- Prefer small, testable processing stages over one monolithic LLM prompt. - Prefer small, testable processing stages over one monolithic LLM prompt.
- Current extraction strategy is one normalized chunk per LLM call. - Current extraction strategy is one normalized chunk per LLM call.
- Do not expand context windows or redesign the extraction strategy without an - Do not expand context windows or redesign the extraction strategy without an
@@ -132,8 +137,8 @@ Current documented Prompt Version 2 decision baseline:
- `src/meeting_lab/segmentation/`: experimental topic segmentation tooling. - `src/meeting_lab/segmentation/`: experimental topic segmentation tooling.
- `src/meeting_lab/extraction/`: local LLM extraction flow and category - `src/meeting_lab/extraction/`: local LLM extraction flow and category
extractor modules. extractor modules.
- `src/meeting_lab/consolidation/`: planned deterministic canonicalization and - `src/meeting_lab/consolidation/`: Canonicalizer V1 and Semantic
semantic consolidation area. Consolidator V0.
- `src/meeting_lab/protocol/`: interim protocol rendering. - `src/meeting_lab/protocol/`: interim protocol rendering.
- `src/meeting_lab/models/`: current lightweight data models. - `src/meeting_lab/models/`: current lightweight data models.
- `samples/`: sample inputs and generated/experimental artifacts; do not treat - `samples/`: sample inputs and generated/experimental artifacts; do not treat
+11
View File
@@ -19,6 +19,14 @@
Verteilerprotokoll and Knowledge Objects / Wissensdatenbankeintrag. Verteilerprotokoll and Knowledge Objects / Wissensdatenbankeintrag.
- Working Protocol Synthesizer V0 benchmark artifact for future - Working Protocol Synthesizer V0 benchmark artifact for future
canonicalizer/consolidator comparisons. canonicalizer/consolidator comparisons.
- Canonicalizer V1 deterministic CLI for canonicalizing chunk extraction JSON.
- Non-LLM Canonicalizer V1 tests covering ordering, validation, IDs, parsing,
source references, exact duplicates, invalid JSON and empty categories.
- Semantic Consolidator V0 CLI for conservative facts-only semantic duplicate
detection using local Ollama.
- Consolidation prompt and non-LLM tests for payload construction, grouping
validation and source fact coverage.
- Semantic Consolidator V0 benchmark report and consolidated extraction JSON.
### Changed ### Changed
@@ -28,6 +36,9 @@
are parallel renderings, not derived from one another. are parallel renderings, not derived from one another.
- Documented the planned Deterministic Canonicalizer and Semantic Consolidator - Documented the planned Deterministic Canonicalizer and Semantic Consolidator
stages before Canonical Meeting Knowledge. stages before Canonical Meeting Knowledge.
- Clarified that Semantic Consolidator V0 is implemented only for fact
duplicate detection; broader semantic synthesis and Canonical Meeting
Knowledge remain planned.
- Documented that rendered protocol language should normally match the source - Documented that rendered protocol language should normally match the source
transcript or consolidated meeting knowledge unless explicitly requested transcript or consolidated meeting knowledge unless explicitly requested
otherwise. otherwise.
+42 -15
View File
@@ -17,6 +17,11 @@ Implemented:
- Transcript normalization in `src/meeting_lab/normalization/`. - Transcript normalization in `src/meeting_lab/normalization/`.
- Technical chunking in `src/meeting_lab/chunking/`. - Technical chunking in `src/meeting_lab/chunking/`.
- Local per-chunk extraction in `src/meeting_lab/extraction/extract_chunks.py`. - Local per-chunk extraction in `src/meeting_lab/extraction/extract_chunks.py`.
- Deterministic Canonicalizer V1 in
`src/meeting_lab/consolidation/canonicalize.py`.
- Semantic Consolidator V0 in
`src/meeting_lab/consolidation/consolidate_facts.py` for facts-only
semantic duplicate detection.
- Prompt loading from `src/meeting_lab/llm/prompts.py`. - Prompt loading from `src/meeting_lab/llm/prompts.py`.
- Interim Markdown protocol generation in `src/meeting_lab/protocol/`. - Interim Markdown protocol generation in `src/meeting_lab/protocol/`.
- Non-LLM unit tests for chunking, extraction helpers, protocol rendering and - Non-LLM unit tests for chunking, extraction helpers, protocol rendering and
@@ -30,8 +35,8 @@ Experimental/prototype:
Planned: Planned:
- Deterministic Canonicalizer for extraction results. - Broader semantic consolidation for topic grouping, contradiction handling,
- Semantic Consolidator for evidence-preserving semantic merging. uncertainty marking and durable/transient separation.
- Canonical Meeting Knowledge implementation as the semantic source of truth. - Canonical Meeting Knowledge implementation as the semantic source of truth.
- Final Working Protocol / Arbeitsprotokoll, Distribution Protocol / - Final Working Protocol / Arbeitsprotokoll, Distribution Protocol /
Verteilerprotokoll and Knowledge Objects / Wissensdatenbankeintrag renderers. Verteilerprotokoll and Knowledge Objects / Wissensdatenbankeintrag renderers.
@@ -123,17 +128,37 @@ Chunk Extractions
-> Output View Renderers -> Output View Renderers
``` ```
The Deterministic Canonicalizer is planned Python code with no LLM. It should Canonicalizer V1 is implemented Python code with no LLM. It validates and
validate and normalize extraction objects, assign stable source references and normalizes extraction objects, assigns stable source references and IDs,
IDs, normalize category names and basic field structure, perform only safe normalizes category names and basic field structure, performs only safe
deterministic cleanup, optionally group exact duplicates, and preserve all deterministic cleanup, optionally groups exact duplicates, and preserves all
source evidence. It must not perform uncertain semantic merging. source evidence. It must not perform uncertain semantic merging.
The Semantic Consolidator is planned local-LLM work. It should merge CLI:
semantically equivalent statements, group content by topic, preserve evidence
from all contributing chunks, mark contradictions and uncertainty, separate ```text
durable information from transient discussion, and produce Canonical Meeting PYTHONPATH=src .venv/bin/python -m meeting_lab.consolidation.canonicalize \
Knowledge. It does not directly write a protocol. samples/whisper/meeting_speech_cleaned_chunks \
-o /tmp/canonicalized_extractions.json
```
Semantic Consolidator V0 is implemented as narrow local-LLM work for fact
items only. It merges semantically equivalent fact statements, preserves source
references and evidence, and prefers false negatives over false-positive
merges. It is not a summarizer, topic grouper, protocol renderer or complete
Canonical Meeting Knowledge stage.
The first accepted V0 benchmark used `qwen3.5:9B` in one Ollama call over 33
fact items. Runtime on the current machine was 390.119 seconds. One correct
merge was accepted, involving `fact_0025` and `fact_0031`; 31 facts remained
singletons, validation passed, no source fact was lost or duplicated, and
non-fact categories remained unchanged. This is a local benchmark, not a
general hardware claim.
Broader semantic consolidation remains planned. It should group content by
topic, mark contradictions and uncertainty, separate durable information from
transient discussion, and produce Canonical Meeting Knowledge. It does not
directly write a protocol.
Canonical Meeting Knowledge is the planned semantic intermediate model and Canonical Meeting Knowledge is the planned semantic intermediate model and
future single source of truth. It should be structured, preferably JSON, and future single source of truth. It should be structured, preferably JSON, and
@@ -164,8 +189,8 @@ language is requested.
- Topic segmentation exists as prototype tooling, not a stable pipeline stage. - Topic segmentation exists as prototype tooling, not a stable pipeline stage.
- Extraction is still a combined current flow, even though separate extractors - Extraction is still a combined current flow, even though separate extractors
are the intended architecture. are the intended architecture.
- Deterministic Canonicalizer is documented but not implemented. - Semantic Consolidator V0 is implemented only for facts-only duplicate
- Semantic Consolidator is documented but not implemented. detection.
- Canonical Meeting Knowledge is documented but not implemented. - Canonical Meeting Knowledge is documented but not implemented.
- Final output views are documented but not implemented. - Final output views are documented but not implemented.
- Most prompt files are placeholders except the common and decision prompts. - Most prompt files are placeholders except the common and decision prompts.
@@ -176,5 +201,7 @@ language is requested.
Stabilize repeatable local extraction evaluation before broadening the pipeline: Stabilize repeatable local extraction evaluation before broadening the pipeline:
expand Gold Standard coverage by category, keep one-chunk extraction as the expand Gold Standard coverage by category, keep one-chunk extraction as the
baseline, and use small prompt experiments with immediate non-regression checks. baseline, and use small prompt experiments with immediate non-regression checks.
After extraction behavior is stable enough, implement the Deterministic The next recommended evaluation step is to use the consolidated V0 result as
Canonicalizer first, then the Semantic Consolidator with evidence retention. input for the unchanged Working Protocol renderer and compare that output
against the Working Protocol Synthesizer V0 baseline and the human reference
protocol.
+29 -10
View File
@@ -102,12 +102,11 @@ Canonical Meeting Knowledge
Output-Ansichten Output-Ansichten
``` ```
Der nächste Architekturmeilenstein ist die Trennung zwischen deterministischer Der aktuelle Architekturmeilenstein trennt deterministische Kanonisierung der
Kanonisierung der Chunk-Extraktionen und semantischer Konsolidierung. Chunk-Extraktionen von semantischer Konsolidierung. Die Kanonisierung validiert
Die Kanonisierung validiert und normalisiert Extraktionsobjekte ohne LLM. Die und normalisiert Extraktionsobjekte ohne LLM. Semantic Consolidator V0 nutzt
semantische Konsolidierung nutzt das lokale LLM, um gleichbedeutende Aussagen das lokale LLM nur fuer konservative facts-only Duplikaterkennung, erhaelt
zusammenzuführen, Evidenz zu erhalten und die Canonical Meeting Knowledge zu Evidenz und erzeugt noch keine Canonical Meeting Knowledge.
erzeugen.
Das Meeting Lab behandelt "das Protokoll" nicht mehr als ein einzelnes Das Meeting Lab behandelt "das Protokoll" nicht mehr als ein einzelnes
Endprodukt. Das konsolidierte Meeting-Wissen ist die **Canonical Meeting Endprodukt. Das konsolidierte Meeting-Wissen ist die **Canonical Meeting
@@ -138,10 +137,30 @@ werden, sofern keine explizite Ausgabesprache angefordert wurde.
## Projektstatus ## Projektstatus
Aktuell liegt der Schwerpunkt auf der Entwicklung eines modularen Aktuell liegt der Schwerpunkt auf der Entwicklung eines modularen
Diskussionsanalyzers. Implementiert sind Vorverarbeitung, technische Chunking- Diskussionsanalyzers. Implementiert sind Vorverarbeitung, technisches Chunking,
und lokale Chunk-Extraktionsschritte. Deterministic Canonicalizer, Semantic lokale Chunk-Extraktion, Canonicalizer V1 als deterministische Vorbereitung der
Consolidator, Canonical Meeting Knowledge und finale Output-View-Renderer sind Extraktionsergebnisse und Semantic Consolidator V0 fuer facts-only
geplante nächste Schritte. Duplikaterkennung. Canonical Meeting Knowledge, breitere semantische Synthese
und finale Output-View-Renderer sind geplante nächste Schritte.
Canonicalizer V1 kann aus dem Repository heraus so ausgeführt werden:
```text
PYTHONPATH=src .venv/bin/python -m meeting_lab.consolidation.canonicalize \
samples/whisper/meeting_speech_cleaned_chunks \
-o /tmp/canonicalized_extractions.json
```
Semantic Consolidator V0 kann auf dem Canonicalizer-Output ausgefuehrt werden:
```text
PYTHONPATH=src .venv/bin/python -m meeting_lab.consolidation.consolidate_facts \
samples/benchmarks/canonicalizer_v1/canonicalized_extractions.json \
-o samples/benchmarks/semantic_consolidator_v0 \
--model qwen3.5:9B \
--endpoint http://127.0.0.1:11434/api/generate \
--no-think
```
Die eigentliche Ausgabeerzeugung ist bewusst der letzte Verarbeitungsschritt. Die eigentliche Ausgabeerzeugung ist bewusst der letzte Verarbeitungsschritt.
+24 -8
View File
@@ -36,7 +36,7 @@ Goal:
Deliverables: Deliverables:
- Deterministic Canonicalizer implemented in Python. - Canonicalizer V1 implemented in Python.
- Stable source references and IDs. - Stable source references and IDs.
- Normalized category names and basic field structure. - Normalized category names and basic field structure.
- Safe deterministic cleanup. - Safe deterministic cleanup.
@@ -48,6 +48,10 @@ Prerequisites:
- Stable local extraction baseline. - Stable local extraction baseline.
- Agreement on the extraction object shape that should be canonicalized. - Agreement on the extraction object shape that should be canonicalized.
Current status:
- Implemented as `meeting_lab.consolidation.canonicalize`.
Out of scope: Out of scope:
- Uncertain semantic merging. - Uncertain semantic merging.
@@ -64,22 +68,34 @@ Goal:
Deliverables: Deliverables:
- Semantic Consolidator using the local LLM. - Semantic Consolidator V0 using the local LLM for facts-only duplicate
- Semantically equivalent statement merging. detection.
- Topic grouping. - Semantically equivalent fact statement merging.
- Evidence preserved from all contributing chunks. - Evidence preserved from all contributing chunks.
- Contradiction and uncertainty markers. - Complete source fact coverage validation.
- Separation of durable information from transient discussion. - Later broader semantic consolidation with topic grouping, contradiction and
- Canonical Meeting Knowledge output. uncertainty markers, durable/transient separation and Canonical Meeting
Knowledge preparation.
Prerequisites: Prerequisites:
- Deterministic Canonicalizer output with stable IDs and source references. - Canonicalizer V1 output with stable IDs and source references.
- Gold or benchmark cases that expose duplication and category shifts. - Gold or benchmark cases that expose duplication and category shifts.
Current status:
- Semantic Consolidator V0 is implemented and experimentally validated for
facts-only conservative merging.
- The first accepted benchmark merged one correct pair among 33 facts and left
31 singleton groups.
Out of scope: Out of scope:
- Direct protocol writing. - Direct protocol writing.
- Topic synthesis in V0.
- Processing decisions, action items, questions, positions or technical details
in V0.
- Canonical Meeting Knowledge generation in V0.
- Deriving output views from one another. - Deriving output views from one another.
- Retrieval or RAG integration. - Retrieval or RAG integration.
+9 -6
View File
@@ -195,12 +195,14 @@ Deterministic Canonicalizer:
Semantic Consolidator: Semantic Consolidator:
- uses the local LLM - uses the local LLM
- merges semantically equivalent statements - V0 is implemented for facts-only semantic duplicate detection
- groups content by topic - V0 merges semantically equivalent fact items conservatively
- preserves evidence from all contributing chunks - V0 preserves source references and evidence
- marks contradictions and uncertainty - V0 validates that every source fact appears exactly once
- separates durable information from transient discussion - V0 does not process non-fact categories semantically
- produces Canonical Meeting Knowledge - later versions should group content by topic, mark contradictions and
uncertainty, separate durable information from transient discussion and
prepare Canonical Meeting Knowledge
- does not directly write a protocol - does not directly write a protocol
--- ---
@@ -267,6 +269,7 @@ Implemented:
- Transcript normalization - Transcript normalization
- Technical chunk generation - Technical chunk generation
- Experimental LLM-based information extraction - Experimental LLM-based information extraction
- Canonicalizer V1 deterministic extraction canonicalization
The current extraction step still performs multiple tasks simultaneously. The current extraction step still performs multiple tasks simultaneously.
+63 -12
View File
@@ -242,10 +242,35 @@ This object feeds the Canonical Meeting Knowledge representation.
--- ---
# Canonical Extraction Object # Canonicalized Extractions
Planned deterministic intermediate object created from raw chunk extraction Implemented deterministic intermediate file created from raw chunk extraction
JSON. JSON by Canonicalizer V1. This is not Canonical Meeting Knowledge.
Top-level structure:
```json
{
"schema_version": "1",
"source_files": [],
"stats": {},
"items": []
}
```
Each item contains at least:
- item_id
- category
- text
- evidence
- source_file
- source_index
- original_value
- source_references
Action items also preserve deterministic fields such as `responsible` and
`deadline` when present.
Example: Example:
@@ -264,16 +289,42 @@ Example:
} }
``` ```
The Deterministic Canonicalizer should create this kind of object without an Canonicalizer V1 creates this kind of object without an LLM. It validates and
LLM. It validates and normalizes raw extraction objects, assigns stable IDs and normalizes raw extraction objects, assigns stable IDs and source references,
source references, normalizes category names and basic field structure, normalizes category names and basic field structure, performs only safe
performs only safe deterministic cleanup, may group exact duplicates and must deterministic cleanup, may group exact duplicates and must preserve all source
preserve all source evidence. evidence.
It must not perform uncertain semantic merging. It must not perform uncertain semantic merging.
--- ---
# Semantic Fact Group
Implemented by Semantic Consolidator V0.
Example:
```json
{
"consolidated_id": "fact_group_0001",
"category": "fact",
"canonical_text": "...",
"source_item_ids": ["fact_0001"],
"source_references": [],
"evidence": [],
"merge_reason": "Singleton; no semantically equivalent fact found."
}
```
Semantic Consolidator V0 only processes fact items. It merges semantically
equivalent facts conservatively, preserves source references and evidence, and
validates that every source fact appears exactly once. Non-fact categories are
copied unchanged. It is not a summarizer, topic grouper, protocol renderer or
Canonical Meeting Knowledge generator.
---
# Consolidated Topic # Consolidated Topic
Planned semantic object produced by the Semantic Consolidator. Planned semantic object produced by the Semantic Consolidator.
@@ -294,10 +345,10 @@ Example:
} }
``` ```
The Semantic Consolidator may use the local LLM to merge semantically Future Semantic Consolidator versions may use the local LLM to merge
equivalent statements, group content by topic, preserve evidence from all semantically equivalent statements beyond facts, group content by topic,
contributing chunks, mark contradictions and uncertainty and separate durable preserve evidence from all contributing chunks, mark contradictions and
information from transient discussion. uncertainty and separate durable information from transient discussion.
It produces Canonical Meeting Knowledge. It does not directly write a protocol. It produces Canonical Meeting Knowledge. It does not directly write a protocol.
+135 -3
View File
@@ -954,9 +954,10 @@ marks contradictions or uncertainty.
Decision: Decision:
Deterministic canonicalization is the next implementation step after stable Deterministic canonicalization is implemented as Canonicalizer V1. The first
local extraction, followed by semantic consolidation. Canonical Meeting semantic consolidation milestone is implemented as Semantic Consolidator V0 for
Knowledge and final output views are planned, not implemented. facts-only duplicate detection. Broader semantic consolidation, Canonical
Meeting Knowledge and final output views remain planned.
Lessons learned: Lessons learned:
@@ -1037,3 +1038,134 @@ Evidence:
- `PROJECT_KNOWLEDGE.md` - `PROJECT_KNOWLEDGE.md`
- `docs/output-views.md` - `docs/output-views.md`
- See EXP-0017 and EXP-0019. - See EXP-0017 and EXP-0019.
## EXP-0021 - Canonicalizer V1
Status: Accepted
Date or period: 2026-07-31
Hypothesis:
Independent chunk extraction JSON can be converted into a stable deterministic
intermediate format before any semantic LLM consolidation is attempted.
Setup:
Canonicalizer V1 discovers `chunk_*_extraction.json` files in stable chunk
order, validates required categories, normalizes category names and basic field
structure, parses existing legacy string formats where safe, trims redundant
whitespace, assigns deterministic IDs, preserves original values and source
references, and merges only exact duplicates when all semantic fields are
identical.
Inputs:
- Synthetic unit-test fixtures.
- Existing nine extraction JSON files under
`samples/whisper/meeting_speech_cleaned_chunks/`.
Model / configuration:
- No LLM.
- CLI module: `meeting_lab.consolidation.canonicalize`.
Result:
Canonicalizer V1 produces `schema_version`, `source_files`, `stats` and
`items`. It is deterministic preparation for the future Semantic Consolidator
and is not Canonical Meeting Knowledge.
Decision:
Canonicalizer V1 is the current implemented deterministic canonicalization
stage. Semantic Consolidator V0 now uses this representation for facts-only
semantic duplicate detection; broader semantic consolidation and Canonical
Meeting Knowledge remain planned.
Lessons learned:
Exact duplicate handling, source-reference preservation and legacy string
parsing can be tested without model calls. Any uncertain semantic merge remains
out of scope for this stage.
Evidence:
- `src/meeting_lab/consolidation/canonicalize.py`
- `tests/test_canonicalize.py`
- `docs/data-models.md`
- `docs/pipeline.md`
## EXP-0022 - Semantic Consolidator V0 facts-only merge
Status: Accepted
Date or period: 2026-07-31
Hypothesis:
The Canonicalizer V1 output contains enough stable structure for a local LLM to
identify semantically equivalent fact items without losing source coverage or
changing non-fact categories.
Setup:
Semantic Consolidator V0 consumed the existing Canonicalizer V1 benchmark JSON,
selected only items with `category: "fact"`, and sent one bounded consolidation
request to local Ollama. The merge rules required semantic equivalence, not
topic similarity, and validation required every source fact ID to appear
exactly once.
Inputs:
- `samples/benchmarks/canonicalizer_v1/canonicalized_extractions.json`
- 33 fact items.
Model / configuration:
- `qwen3.5:9B`
- Ollama endpoint: `http://127.0.0.1:11434/api/generate`
- Thinking disabled.
- One LLM call.
- `num_ctx=32768`
- `num_predict=4096`
Result:
- Runtime: 390.119 seconds on the current machine.
- Merged fact groups: 1.
- Source facts involved in merges: 2.
- Singleton fact groups: 31.
- Validation: passed.
- No source fact was lost or duplicated.
- Non-fact categories remained unchanged.
Accepted merge:
- `fact_0025` + `fact_0031`
- Canonical statement: "Der Leiter F&E führt die Projektliste auf dem
zweiwöchentlichen Schnittstellen-Stand-Up."
Decision:
Semantic Consolidator V0 is complete for its current narrow scope:
conservative facts-only semantic duplicate detection with source evidence
preserved. The selected `report.md` and `consolidated_extractions.json`
benchmark artifacts should be versioned for later comparison. The raw model
response remains a local diagnostic artifact and is not versioned.
Lessons learned:
Semantic duplicate consolidation is technically viable and conservative enough
for continued evaluation, but broader semantic synthesis remains a separate
future stage. The measured runtime is useful for this machine and run, but
should not be generalized into a universal benchmark.
Evidence:
- `src/meeting_lab/consolidation/consolidate_facts.py`
- `prompts/consolidate_facts.md`
- `tests/test_consolidate_facts.py`
- `samples/benchmarks/semantic_consolidator_v0/report.md`
- `samples/benchmarks/semantic_consolidator_v0/consolidated_extractions.json`
- Local diagnostic only: `samples/benchmarks/semantic_consolidator_v0/raw_model_response.txt`
+10 -9
View File
@@ -51,15 +51,16 @@ technical validation.
The next planned architecture stage before this representation is explicit: The next planned architecture stage before this representation is explicit:
- The Deterministic Canonicalizer validates and normalizes extraction objects, - Canonicalizer V1 validates and normalizes extraction objects, assigns stable
assigns stable source references and IDs, performs only safe deterministic source references and IDs, performs only safe deterministic cleanup and
cleanup and preserves all source evidence. It uses no LLM and must not make preserves all source evidence. It uses no LLM and must not make uncertain
uncertain semantic merges. semantic merges.
- The Semantic Consolidator uses the local LLM to merge semantically equivalent - Semantic Consolidator V0 uses the local LLM only for facts-only semantic
statements, group content by topic, preserve evidence from all contributing duplicate detection. It preserves source evidence and does not directly write
chunks, mark contradictions and uncertainty, separate durable information from a protocol or produce Canonical Meeting Knowledge.
transient discussion and produce Canonical Meeting Knowledge. It does not - Future Semantic Consolidator versions should group content by topic, mark
directly write a protocol. contradictions and uncertainty, separate durable information from transient
discussion and prepare Canonical Meeting Knowledge.
## Renderers ## Renderers
+34 -15
View File
@@ -135,7 +135,15 @@ Deterministic
## Current Status ## Current Status
Planned Implemented as Canonicalizer V1.
CLI:
```text
PYTHONPATH=src .venv/bin/python -m meeting_lab.consolidation.canonicalize \
samples/whisper/meeting_speech_cleaned_chunks \
-o /tmp/canonicalized_extractions.json
```
--- ---
@@ -231,7 +239,7 @@ LLM
## Current Status ## Current Status
Planned Implemented as Canonicalizer V1.
--- ---
@@ -325,16 +333,23 @@ Canonicalized extraction objects.
## Output ## Output
Canonical Meeting Knowledge. Semantic Consolidator V0 output is `consolidated_extractions.json` with fact
groups and unchanged non-fact items.
Future broader semantic consolidation should produce Canonical Meeting
Knowledge.
## Responsibilities ## Responsibilities
- Merge semantically equivalent statements - V0: merge semantically equivalent fact items only
- Group content by topic - V0: preserve all non-fact categories unchanged
- V0: validate that every source fact ID appears exactly once
- Future: merge semantically equivalent statements across categories
- Future: group content by topic
- Preserve evidence from all contributing chunks - Preserve evidence from all contributing chunks
- Mark contradictions and uncertainty - Future: mark contradictions and uncertainty
- Separate durable information from transient discussion - Future: separate durable information from transient discussion
- Reconcile category shifts where supported by evidence - Future: reconcile category shifts where supported by evidence
## Must Not ## Must Not
@@ -348,7 +363,9 @@ Local LLM, with deterministic pre/post-processing where useful.
## Current Status ## Current Status
Planned Semantic Consolidator V0 is implemented and experimentally validated for
facts-only conservative duplicate detection. Broader semantic consolidation and
Canonical Meeting Knowledge generation remain planned.
--- ---
@@ -507,16 +524,18 @@ A processing stage may be replaced by another implementation as long as it prese
⬜ Specialized Extraction ⬜ Specialized Extraction
⬜ Deterministic Canonicalization ✔ Deterministic Canonicalization
⬜ Semantic Consolidation ✅ Semantic Consolidation V0 - facts-only duplicate detection
⬜ Canonical Meeting Knowledge ⬜ Canonical Meeting Knowledge
⬜ Output View Rendering ⬜ Output View Rendering
``` ```
The immediate architecture focus is the Deterministic Canonicalizer followed by The immediate evaluation focus is using the Semantic Consolidator V0 output as
the Semantic Consolidator. These stages preserve source evidence, recover input for the unchanged Working Protocol renderer. Canonicalizer V1 now
global context from independent chunk extractions and prepare Canonical Meeting preserves source evidence, and Semantic Consolidator V0 conservatively merges
Knowledge for parallel Output View rendering. semantically equivalent fact items. Broader semantic consolidation should later
recover global context from independent chunk extractions and prepare Canonical
Meeting Knowledge for parallel Output View rendering.
+52
View File
@@ -0,0 +1,52 @@
You consolidate fact items from Meeting Lab.
Return only valid JSON.
Semantic equivalence is stricter than topical similarity.
Merge fact items only when they express the same core factual proposition.
Related facts are not enough. Facts from the same topic are not enough.
Preserve distinctions in:
- actor
- scope
- timing
- condition
- certainty
- responsibility
- current state versus future intention
Do not merge:
- broader and narrower statements
- cause and effect
- process and responsibility
- fact and interpretation
- related statements from the same topic
- statements that differ in actor, scope, timing, condition, or certainty
Use only the provided fact items.
Do not invent information.
Do not rewrite evidence.
Do not add source item IDs that were not provided.
Every provided fact item ID must appear exactly once.
When uncertain, keep items separate.
False negatives are preferable to false-positive merges.
Return JSON with this exact top-level shape:
{
"groups": [
{
"canonical_text": "One factual statement preserving the shared meaning",
"source_item_ids": ["fact_0001"],
"merge_reason": "Why these items are equivalent, or why this singleton remains separate"
}
]
}
For singleton groups, use the original fact text as canonical_text.
For merged groups, canonical_text must contain only information shared by all
source facts in the group.
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,168 @@
# Canonicalizer V1 Evaluation Run
## Scope
This run evaluates Canonicalizer V1 as deterministic input preparation for a
future semantic consolidator. It does not evaluate protocol prose and does not
run the semantic consolidator.
Input files:
- `samples/whisper/meeting_speech_cleaned_chunks/chunk_01_extraction.json`
- `samples/whisper/meeting_speech_cleaned_chunks/chunk_02_extraction.json`
- `samples/whisper/meeting_speech_cleaned_chunks/chunk_03_extraction.json`
- `samples/whisper/meeting_speech_cleaned_chunks/chunk_04_extraction.json`
- `samples/whisper/meeting_speech_cleaned_chunks/chunk_05_extraction.json`
- `samples/whisper/meeting_speech_cleaned_chunks/chunk_06_extraction.json`
- `samples/whisper/meeting_speech_cleaned_chunks/chunk_07_extraction.json`
- `samples/whisper/meeting_speech_cleaned_chunks/chunk_08_extraction.json`
- `samples/whisper/meeting_speech_cleaned_chunks/chunk_09_extraction.json`
Output:
- `samples/benchmarks/canonicalizer_v1/canonicalized_extractions.json`
## Run Metrics
- Runtime: 0.04 seconds
- Input file count: 9
- Input byte size: 22,695 bytes
- Output byte size: 115,924 bytes
- Exact duplicate count: 0
- Merged exact duplicate count: 0
- Malformed or unparsed items: 0
- Source-reference mismatches found during comparison: 0
Input item count by category:
- fact: 33
- decision: 14
- action_item: 16
- open_question: 10
- position: 0
- technical_detail: 14
Output item count by category:
- fact: 33
- decision: 14
- action_item: 16
- open_question: 10
- position: 0
- technical_detail: 14
## Representative Inspection
### Fact
- Item: `fact_0001`
- Source: `chunk_01_extraction.json`, index 0
- Text: `Das ist das aktuelle Projektdeckblatt, was wir in der F&E benutzen.`
- Speaker: `Martin`
- Status: `clear`
- Evidence preserved from the original extraction item.
Assessment: original information and evidence were preserved, and the source
reference points back to the correct original item.
### Decision
- Item: `decision_0001`
- Source: `chunk_02_extraction.json`, index 0
- Text: `Der bestehende Prozess wird grundsätzlich auch für Business Development (BD)-Projekte genutzt, wobei die spezifischen Dokumente je nach Projekttyp angepasst werden.`
- Evidence preserved from the original extraction item.
Assessment: decision text and evidence were preserved without semantic
rewriting.
### Action Item
- Item: `action_item_0001`
- Source: `chunk_01_extraction.json`, index 0
- Text: `Giovanna stellt die Kriterien zusammen und erstellt einen ersten Entwurf für den Auswahlkatalog, der EDD-, Marketing- und PM-Kriterien integriert.`
- Responsible: `Giovanna`
- Deadline: `null`
- Evidence preserved from the original extraction item.
Assessment: responsibility was parsed from the legacy string and no deadline
was invented.
### Open Question
- Item: `open_question_0001`
- Source: `chunk_01_extraction.json`, index 0
- Text: `Wie werden spezifische Auswahlkriterien für digitale Produkte (z.B. Portal) versus Realprodukte definiert und integriert?`
- Evidence preserved from the original extraction item.
Assessment: question text and evidence were preserved.
### Position
The input extraction files contain zero `positions` items. The canonicalized
output correctly reports zero `position` items.
### Technical Detail
- Item: `technical_detail_0001`
- Source: `chunk_01_extraction.json`, index 0
- Subject: `Projektdeckblatt Struktur`
- Text: `Das ist eine kurze Projektidee, Ziel, Gegenüber, Entwicklungshemmnisse (Patente), Budgetabfrage.`
- Status: `clear`
- Evidence preserved from the original extraction item.
Assessment: subject, statement, status and evidence were parsed without adding
technical interpretation.
## Comparison Findings
- Original information is preserved. All output items retain `original_value`.
- Source references are correct. A comparison against the nine original JSON
files found zero source-reference mismatches.
- No semantic merges occurred. Output item counts match input item counts in
every category.
- Names and responsibilities were not invented. Missing optional values remain
`null`; for example, `action_item_0001` has `deadline: null`.
- Similar but non-identical items remain separate. For example, `fact_0025`
and `fact_0031` both discuss the F&E lead maintaining the project list, but
they have different wording and remain separate items.
- The output is suitable as deterministic input for a later LLM consolidator:
it has stable IDs, normalized categories, source references, evidence and
original values.
## Conclusion
1. Is the deterministic representation lossless enough?
Yes for the current extraction format. The canonicalized output preserves the
original value, parsed text fields, evidence and source references for each
item. The output is larger than the input because it adds deterministic
metadata and source-reference structure.
2. Are exact duplicates handled correctly?
Yes for this run. No exact duplicates were present, so no items were merged.
The input and output category counts are identical.
3. Are legacy extraction strings parsed reliably?
Yes for the inspected current files. Legacy pipe-delimited strings were parsed
into category-specific fields such as `speaker`, `status`, `responsible`,
`deadline` and `subject` where safely available. No malformed or unparsed items
were detected.
4. Does the result reduce avoidable work for the semantic consolidator?
Yes. The future consolidator can consume normalized categories, stable item
IDs, source references, evidence and parsed optional fields instead of
re-reading heterogeneous legacy strings directly.
5. What unresolved transformations must remain an LLM task?
- Merging semantically equivalent but differently worded statements.
- Grouping items into coherent topics.
- Reconciling facts, decisions, action items and questions across chunks.
- Detecting contradictions and uncertainty.
- Separating durable organizational knowledge from transient discussion.
- Deciding whether similar items such as `fact_0025` and `fact_0031` should be
merged, related or kept separate.
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,24 @@
# Semantic Consolidator V0 Report
- Scope: facts only
- Model: `qwen3.5:9B`
- LLM call count: 1
- Runtime: 390.119 seconds
- Fact item count: 33
- Prompt characters: 15393
- Estimated prompt tokens: 3849
- Prompt eval count: 4295
- Eval count: 2344
- Merged fact groups: 1
- Source facts involved in merges: 2
- Singleton fact groups: 31
- Validation: passed
- Output path: `samples/benchmarks/semantic_consolidator_v0/consolidated_extractions.json`
## Actual Merges
### Der Leiter F&E führt die Projektliste auf dem zweiwöchentlichen Schnittstellen-Stand-Up.
- Source fact IDs: fact_0025, fact_0031
- Merge reason: Semantic equivalence: Both items state that the Head of R&D leads the project list at the bi-weekly interface stand-up meeting. The difference in wording ('auf den' vs 'auf dem') and minor elaboration on defining project types does not change the core factual proposition regarding who performs which action when.
- Review decision: accepted as correct.
@@ -0,0 +1,456 @@
#!/usr/bin/env python3
"""Canonicalize chunk extraction JSON without semantic merging."""
from __future__ import annotations
import argparse
import json
import re
import sys
from collections import Counter
from pathlib import Path
from typing import Any
SCHEMA_VERSION = "1"
REQUIRED_CATEGORIES = (
"facts",
"decisions",
"todos",
"questions",
"positions",
"technical",
)
CATEGORY_NAMES = {
"facts": "fact",
"decisions": "decision",
"todos": "action_item",
"questions": "open_question",
"positions": "position",
"technical": "technical_detail",
}
CANONICAL_CATEGORIES = tuple(CATEGORY_NAMES.values())
CHUNK_EXTRACTION_RE = re.compile(r"^chunk_(\d+)_extraction\.json$")
WHITESPACE_RE = re.compile(r"\s+")
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser(
description="Canonicalize chunk extraction JSON files deterministically."
)
parser.add_argument(
"input_dir",
type=Path,
help="Directory containing chunk_XX_extraction.json files.",
)
parser.add_argument(
"-o",
"--output",
type=Path,
default=Path("canonicalized_extractions.json"),
help="Output JSON path (default: canonicalized_extractions.json).",
)
parser.add_argument(
"--no-merge-exact-duplicates",
action="store_true",
help="Preserve exact duplicate items instead of merging them.",
)
return parser.parse_args()
def chunk_sort_key(path: Path) -> tuple[int, str]:
match = CHUNK_EXTRACTION_RE.match(path.name)
if not match:
return (sys.maxsize, path.name)
return (int(match.group(1)), path.name)
def find_extraction_files(input_dir: Path) -> list[Path]:
if not input_dir.is_dir():
raise FileNotFoundError(f"Input directory not found: {input_dir}")
return sorted(input_dir.glob("chunk_*_extraction.json"), key=chunk_sort_key)
def load_json_object(path: Path) -> dict[str, Any]:
try:
data = json.loads(path.read_text(encoding="utf-8-sig"))
except json.JSONDecodeError as exc:
raise ValueError(f"Invalid JSON in {path}: {exc}") from exc
if not isinstance(data, dict):
raise ValueError(f"Extraction file must contain a JSON object: {path}")
return data
def validate_required_categories(data: dict[str, Any], path: Path) -> None:
missing = [category for category in REQUIRED_CATEGORIES if category not in data]
if missing:
raise ValueError(
f"Missing required categories in {path}: {', '.join(missing)}"
)
invalid = [
category
for category in REQUIRED_CATEGORIES
if not isinstance(data.get(category), list)
]
if invalid:
raise ValueError(
f"Required categories must be lists in {path}: {', '.join(invalid)}"
)
def clean_text(value: Any) -> str | None:
if value is None:
return None
text = WHITESPACE_RE.sub(" ", str(value)).strip()
return text if text else None
def split_legacy_string(value: str) -> list[str]:
return [part.strip() for part in value.split("|")]
def first_present(data: dict[str, Any], keys: tuple[str, ...]) -> str | None:
for key in keys:
text = clean_text(data.get(key))
if text:
return text
return None
def parse_fact(value: Any) -> dict[str, Any]:
if isinstance(value, dict):
return {
"speaker": clean_text(value.get("speaker")),
"text": first_present(value, ("statement", "fact", "text")),
"status": clean_text(value.get("status")),
"evidence": clean_text(value.get("evidence")),
}
text = clean_text(value)
parts = split_legacy_string(text or "")
if len(parts) >= 4:
return {
"speaker": clean_text(parts[0]),
"text": clean_text(parts[1]),
"status": clean_text(parts[2]),
"evidence": clean_text(" | ".join(parts[3:])),
}
if len(parts) >= 3:
return {
"speaker": None,
"text": clean_text(parts[0]),
"status": clean_text(parts[1]),
"evidence": clean_text(" | ".join(parts[2:])),
}
return {"speaker": None, "text": text, "status": None, "evidence": None}
def parse_decision(value: Any) -> dict[str, Any]:
if isinstance(value, dict):
return {
"text": first_present(value, ("decision", "text")),
"evidence": clean_text(value.get("evidence")),
}
text = clean_text(value)
parts = split_legacy_string(text or "")
if len(parts) >= 2:
return {
"text": clean_text(parts[0]),
"evidence": clean_text(" | ".join(parts[1:])),
}
return {"text": text, "evidence": None}
def parse_action_item(value: Any) -> dict[str, Any]:
if isinstance(value, dict):
return {
"text": first_present(value, ("task", "todo", "text")),
"responsible": first_present(value, ("responsible", "owner")),
"deadline": clean_text(value.get("deadline")),
"evidence": clean_text(value.get("evidence")),
}
text = clean_text(value)
parts = split_legacy_string(text or "")
if len(parts) >= 4:
return {
"text": clean_text(parts[0]),
"responsible": clean_text(parts[1]),
"deadline": clean_text(parts[2]),
"evidence": clean_text(" | ".join(parts[3:])),
}
if len(parts) == 3:
return {
"text": clean_text(parts[0]),
"responsible": clean_text(parts[1]),
"deadline": None,
"evidence": clean_text(parts[2]),
}
if len(parts) == 2:
return {
"text": clean_text(parts[0]),
"responsible": None,
"deadline": None,
"evidence": clean_text(parts[1]),
}
return {
"text": text,
"responsible": None,
"deadline": None,
"evidence": None,
}
def parse_question(value: Any) -> dict[str, Any]:
if isinstance(value, dict):
return {
"text": first_present(value, ("question", "text")),
"evidence": clean_text(value.get("evidence")),
}
text = clean_text(value)
parts = split_legacy_string(text or "")
if len(parts) >= 2:
return {
"text": clean_text(parts[0]),
"evidence": clean_text(" | ".join(parts[1:])),
}
return {"text": text, "evidence": None}
def parse_position(value: Any) -> dict[str, Any]:
if isinstance(value, dict):
return {
"speaker": clean_text(value.get("speaker")),
"text": first_present(value, ("position", "statement", "text")),
"evidence": clean_text(value.get("evidence")),
}
text = clean_text(value)
parts = split_legacy_string(text or "")
if len(parts) >= 3:
return {
"speaker": clean_text(parts[0]),
"text": clean_text(parts[1]),
"evidence": clean_text(" | ".join(parts[2:])),
}
if len(parts) == 2:
return {
"speaker": None,
"text": clean_text(parts[0]),
"evidence": clean_text(parts[1]),
}
return {"speaker": None, "text": text, "evidence": None}
def parse_technical_detail(value: Any) -> dict[str, Any]:
if isinstance(value, dict):
return {
"subject": clean_text(value.get("subject")),
"text": first_present(value, ("statement", "technical", "text")),
"status": clean_text(value.get("status")),
"evidence": clean_text(value.get("evidence")),
}
text = clean_text(value)
parts = split_legacy_string(text or "")
if len(parts) >= 4:
return {
"subject": clean_text(parts[0]),
"text": clean_text(parts[1]),
"status": clean_text(parts[2]),
"evidence": clean_text(" | ".join(parts[3:])),
}
if len(parts) >= 3:
return {
"subject": None,
"text": clean_text(parts[0]),
"status": clean_text(parts[1]),
"evidence": clean_text(" | ".join(parts[2:])),
}
return {"subject": None, "text": text, "status": None, "evidence": None}
PARSERS = {
"fact": parse_fact,
"decision": parse_decision,
"action_item": parse_action_item,
"open_question": parse_question,
"position": parse_position,
"technical_detail": parse_technical_detail,
}
def source_reference(
source_file: str,
source_index: int,
original_value: Any,
evidence: str | None,
) -> dict[str, Any]:
return {
"source_file": source_file,
"source_index": source_index,
"evidence": evidence,
"original_value": original_value,
}
def semantic_key(item: dict[str, Any]) -> str:
ignored = {
"item_id",
"source_file",
"source_index",
"original_value",
"source_references",
"duplicate_count",
}
comparable = {key: value for key, value in item.items() if key not in ignored}
return json.dumps(comparable, ensure_ascii=False, sort_keys=True)
def item_id_for(category: str, counts: Counter[str]) -> str:
counts[category] += 1
return f"{category}_{counts[category]:04d}"
def canonicalize_value(
category: str,
value: Any,
source_file: str,
source_index: int,
counts: Counter[str],
) -> dict[str, Any]:
parsed = PARSERS[category](value)
item: dict[str, Any] = {
"item_id": item_id_for(category, counts),
"category": category,
"text": parsed.pop("text", None),
"evidence": parsed.pop("evidence", None),
"source_file": source_file,
"source_index": source_index,
"original_value": value,
}
for key, parsed_value in parsed.items():
item[key] = parsed_value
item["source_references"] = [
source_reference(source_file, source_index, value, item["evidence"])
]
return item
def merge_exact_duplicates(items: list[dict[str, Any]]) -> tuple[list[dict[str, Any]], int]:
merged: list[dict[str, Any]] = []
seen: dict[str, dict[str, Any]] = {}
duplicates = 0
for item in items:
key = semantic_key(item)
existing = seen.get(key)
if existing is None:
item["duplicate_count"] = 1
seen[key] = item
merged.append(item)
continue
duplicates += 1
existing["duplicate_count"] += 1
existing["source_references"].extend(item["source_references"])
return merged, duplicates
def canonicalize_extractions(
input_dir: Path,
merge_duplicates: bool = True,
) -> dict[str, Any]:
files = find_extraction_files(input_dir)
if not files:
raise ValueError(f"No chunk extraction JSON files found in {input_dir}")
counts: Counter[str] = Counter()
input_counts: Counter[str] = Counter()
items: list[dict[str, Any]] = []
for path in files:
data = load_json_object(path)
validate_required_categories(data, path)
for raw_category in REQUIRED_CATEGORIES:
category = CATEGORY_NAMES[raw_category]
values = data[raw_category]
input_counts[category] += len(values)
for source_index, value in enumerate(values):
items.append(
canonicalize_value(
category=category,
value=value,
source_file=path.name,
source_index=source_index,
counts=counts,
)
)
exact_duplicates = 0
if merge_duplicates:
items, exact_duplicates = merge_exact_duplicates(items)
output_counts = Counter(item["category"] for item in items)
input_counts_by_category = {
category: input_counts[category] for category in CANONICAL_CATEGORIES
}
output_counts_by_category = {
category: output_counts[category] for category in CANONICAL_CATEGORIES
}
return {
"schema_version": SCHEMA_VERSION,
"source_files": [path.name for path in files],
"stats": {
"input_item_count": sum(input_counts.values()),
"output_item_count": len(items),
"input_item_count_by_category": input_counts_by_category,
"output_item_count_by_category": output_counts_by_category,
"exact_duplicates_merged": exact_duplicates,
},
"items": items,
}
def write_canonicalized(output: dict[str, Any], output_path: Path) -> Path:
output_path.parent.mkdir(parents=True, exist_ok=True)
output_path.write_text(
json.dumps(output, ensure_ascii=False, indent=2) + "\n",
encoding="utf-8",
)
return output_path
def main() -> int:
args = parse_args()
try:
output = canonicalize_extractions(
args.input_dir,
merge_duplicates=not args.no_merge_exact_duplicates,
)
output_path = write_canonicalized(output, args.output)
except (OSError, UnicodeError, ValueError) as exc:
print(f"Error: {exc}", file=sys.stderr)
return 1
stats = output["stats"]
print(f"Input directory: {args.input_dir}")
print(f"Input files processed: {len(output['source_files'])}")
print(f"Input item count by category: {stats['input_item_count_by_category']}")
print(f"Output item count by category: {stats['output_item_count_by_category']}")
print(f"Exact duplicates merged: {stats['exact_duplicates_merged']}")
print(f"Output: {output_path}")
print(f"Output JSON size: {output_path.stat().st_size} bytes")
return 0
if __name__ == "__main__":
raise SystemExit(main())
@@ -0,0 +1,603 @@
#!/usr/bin/env python3
"""LLM-backed V0 semantic consolidation for fact items only."""
from __future__ import annotations
import argparse
import json
import sys
import time
from collections import Counter
from datetime import datetime
from pathlib import Path
from queue import Empty, Queue
from threading import Thread
from typing import Any
import requests
try:
from meeting_lab.llm.prompts import load_prompt
except ModuleNotFoundError: # pragma: no cover - used by repository-root tests.
from src.meeting_lab.llm.prompts import load_prompt
DEFAULT_MODEL = "qwen3.5:9b"
DEFAULT_ENDPOINT = "http://127.0.0.1:11434/api/generate"
DEFAULT_NUM_CTX = 32768
DEFAULT_NUM_PREDICT = 4096
DEFAULT_PROGRESS_INTERVAL = 30
PROMPT_NAME = "consolidate_facts.md"
class ConsolidationValidationError(ValueError):
"""Raised when model consolidation output violates strict invariants."""
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser(
description="Consolidate semantically equivalent fact items only."
)
parser.add_argument(
"canonicalized_input",
type=Path,
help="Canonicalizer V1 JSON file.",
)
parser.add_argument(
"-o",
"--output-dir",
type=Path,
required=True,
help="Directory for consolidated_extractions.json, report.md and raw response.",
)
parser.add_argument(
"--model",
default=DEFAULT_MODEL,
help=f"Ollama model name (default: {DEFAULT_MODEL}).",
)
parser.add_argument(
"--endpoint",
default=DEFAULT_ENDPOINT,
help=f"Ollama generate endpoint (default: {DEFAULT_ENDPOINT}).",
)
parser.add_argument(
"--timeout",
type=int,
default=1800,
help="HTTP timeout in seconds (default: 1800).",
)
parser.add_argument(
"--num-ctx",
type=int,
default=DEFAULT_NUM_CTX,
help=f"Context window tokens (default: {DEFAULT_NUM_CTX}).",
)
parser.add_argument(
"--num-predict",
type=int,
default=DEFAULT_NUM_PREDICT,
help=(
"Maximum generated tokens. The default is bounded for the expected "
f"fact-group JSON while leaving truncation headroom (default: {DEFAULT_NUM_PREDICT})."
),
)
thinking = parser.add_mutually_exclusive_group()
thinking.add_argument(
"--think",
dest="think",
action="store_true",
help="Enable Ollama thinking output when the selected model supports it.",
)
thinking.add_argument(
"--no-think",
dest="think",
action="store_false",
help="Disable Ollama thinking output for structured JSON consolidation.",
)
parser.set_defaults(think=False)
parser.add_argument(
"--progress-interval",
type=int,
default=DEFAULT_PROGRESS_INTERVAL,
help=(
"Seconds between waiting-status messages while the non-streaming "
f"Ollama request is in flight (default: {DEFAULT_PROGRESS_INTERVAL})."
),
)
return parser.parse_args()
def load_json_object(path: Path) -> dict[str, Any]:
try:
data = json.loads(path.read_text(encoding="utf-8-sig"))
except json.JSONDecodeError as exc:
raise ValueError(f"Invalid JSON in {path}: {exc}") from exc
if not isinstance(data, dict):
raise ValueError(f"JSON file must contain an object: {path}")
return data
def fact_items(canonicalized: dict[str, Any]) -> list[dict[str, Any]]:
items = canonicalized.get("items")
if not isinstance(items, list):
raise ValueError("Canonicalized input must contain an items list.")
return [item for item in items if item.get("category") == "fact"]
def non_fact_items(canonicalized: dict[str, Any]) -> list[dict[str, Any]]:
items = canonicalized.get("items")
if not isinstance(items, list):
raise ValueError("Canonicalized input must contain an items list.")
return [item for item in items if item.get("category") != "fact"]
def model_fact_payload(facts: list[dict[str, Any]]) -> list[dict[str, Any]]:
payload: list[dict[str, Any]] = []
for item in facts:
payload.append(
{
"item_id": item.get("item_id"),
"text": item.get("text"),
"evidence": item.get("evidence"),
"speaker": item.get("speaker"),
"status": item.get("status"),
"source_file": item.get("source_file"),
"source_index": item.get("source_index"),
}
)
return payload
def build_consolidation_prompt(facts: list[dict[str, Any]]) -> str:
task_prompt = load_prompt(PROMPT_NAME)
payload = json.dumps(
{"fact_items": model_fact_payload(facts)},
ensure_ascii=False,
indent=2,
)
return f"{task_prompt}\n\nFACT ITEMS:\n{payload}\n"
def response_text_from_ollama_data(data: dict[str, Any]) -> str | None:
text = data.get("response")
if isinstance(text, str) and text.strip():
return text
message = data.get("message")
if isinstance(message, dict):
content = message.get("content")
if isinstance(content, str) and content.strip():
return content
return text if isinstance(text, str) else None
def build_ollama_payload(
model: str,
prompt: str,
num_ctx: int,
num_predict: int,
think: bool,
) -> dict[str, Any]:
return {
"model": model,
"prompt": prompt,
"think": think,
"stream": False,
"format": "json",
"options": {
"temperature": 0.0,
"num_ctx": num_ctx,
"num_predict": num_predict,
},
}
def print_response_metadata(data: dict[str, Any]) -> None:
fields = [
"total_duration",
"load_duration",
"prompt_eval_count",
"prompt_eval_duration",
"eval_count",
"eval_duration",
]
present = [(field, data.get(field)) for field in fields if field in data]
if not present:
print("Ollama response metadata: unavailable")
return
print("Ollama response metadata:")
for field, value in present:
print(f" {field}: {value}")
def post_with_progress(
endpoint: str,
payload: dict[str, Any],
timeout: int,
progress_interval: int,
) -> tuple[requests.Response, float]:
results: Queue[tuple[str, requests.Response | BaseException, float]] = Queue()
def worker() -> None:
started = time.perf_counter()
try:
response = requests.post(endpoint, json=payload, timeout=timeout)
except BaseException as exc: # noqa: BLE001 - forwarded to main thread.
results.put(("error", exc, time.perf_counter() - started))
else:
results.put(("response", response, time.perf_counter() - started))
started_wait = time.perf_counter()
thread = Thread(target=worker, daemon=True)
thread.start()
while True:
try:
status, result, elapsed = results.get(timeout=max(progress_interval, 1))
except Empty:
print(
"Waiting for Ollama response: "
f"{time.perf_counter() - started_wait:.1f} seconds elapsed",
flush=True,
)
continue
if status == "error":
raise result
return result, elapsed
def call_ollama(
endpoint: str,
model: str,
prompt: str,
timeout: int,
num_ctx: int,
num_predict: int,
think: bool,
progress_interval: int,
) -> tuple[str, dict[str, Any], float]:
payload = build_ollama_payload(
model=model,
prompt=prompt,
num_ctx=num_ctx,
num_predict=num_predict,
think=think,
)
print(f"Ollama request start: {datetime.now().isoformat(timespec='seconds')}")
print(f"Ollama endpoint: {endpoint}")
print(f"Ollama model: {model}")
print(f"Ollama stream: {payload['stream']}")
print(f"Ollama think: {payload['think']}")
print(f"Ollama timeout seconds: {timeout}")
print(f"Ollama options: {json.dumps(payload['options'], sort_keys=True)}")
response, elapsed = post_with_progress(
endpoint=endpoint,
payload=payload,
timeout=timeout,
progress_interval=progress_interval,
)
response.raise_for_status()
data = response.json()
print_response_metadata(data)
text = response_text_from_ollama_data(data)
if not isinstance(text, str) or not text.strip():
raise ValueError("Ollama returned no usable response text.")
return text.strip(), data, elapsed
def parse_model_json(text: str) -> dict[str, Any]:
try:
parsed = json.loads(text)
except json.JSONDecodeError as exc:
raise ConsolidationValidationError(f"Invalid model JSON: {exc}") from exc
if not isinstance(parsed, dict):
raise ConsolidationValidationError("Model JSON must be an object.")
return parsed
def validate_model_groups(
model_output: dict[str, Any],
expected_fact_ids: set[str],
) -> list[dict[str, Any]]:
groups = model_output.get("groups")
if not isinstance(groups, list):
raise ConsolidationValidationError("Model output must contain a groups list.")
seen: list[str] = []
validated: list[dict[str, Any]] = []
for index, group in enumerate(groups):
if not isinstance(group, dict):
raise ConsolidationValidationError(f"Group {index} must be an object.")
source_ids = group.get("source_item_ids")
if not isinstance(source_ids, list) or not source_ids:
raise ConsolidationValidationError(
f"Group {index} must contain source_item_ids."
)
if not all(isinstance(item_id, str) for item_id in source_ids):
raise ConsolidationValidationError(
f"Group {index} source_item_ids must be strings."
)
unknown = sorted(set(source_ids) - expected_fact_ids)
if unknown:
raise ConsolidationValidationError(
f"Group {index} contains unknown source item IDs: {unknown}"
)
duplicates_in_group = [
item_id for item_id, count in Counter(source_ids).items() if count > 1
]
if duplicates_in_group:
raise ConsolidationValidationError(
f"Group {index} repeats source item IDs: {duplicates_in_group}"
)
canonical_text = group.get("canonical_text")
if not isinstance(canonical_text, str) or not canonical_text.strip():
raise ConsolidationValidationError(
f"Group {index} must contain canonical_text."
)
merge_reason = group.get("merge_reason")
if not isinstance(merge_reason, str) or not merge_reason.strip():
raise ConsolidationValidationError(
f"Group {index} must contain merge_reason."
)
seen.extend(source_ids)
validated.append(
{
"canonical_text": canonical_text.strip(),
"source_item_ids": source_ids,
"merge_reason": merge_reason.strip(),
}
)
seen_counts = Counter(seen)
duplicated = sorted(item_id for item_id, count in seen_counts.items() if count > 1)
if duplicated:
raise ConsolidationValidationError(
f"Source item IDs appear in multiple groups: {duplicated}"
)
missing = sorted(expected_fact_ids - set(seen))
if missing:
raise ConsolidationValidationError(f"Missing source item IDs: {missing}")
return validated
def validate_group_shapes(groups: list[dict[str, Any]]) -> None:
for group in groups:
source_count = len(group["source_item_ids"])
if source_count < 1:
raise ConsolidationValidationError("Groups must not be empty.")
if source_count == 1:
continue
if source_count < 2:
raise ConsolidationValidationError("Merged groups need at least two IDs.")
def build_consolidated_fact_item(
group: dict[str, Any],
fact_by_id: dict[str, dict[str, Any]],
sequence: int,
) -> dict[str, Any]:
source_ids = group["source_item_ids"]
source_items = [fact_by_id[item_id] for item_id in source_ids]
source_references: list[dict[str, Any]] = []
evidence: list[str] = []
for item in source_items:
source_references.extend(item.get("source_references", []))
item_evidence = item.get("evidence")
if isinstance(item_evidence, str):
evidence.append(item_evidence)
return {
"consolidated_id": f"fact_group_{sequence:04d}",
"category": "fact",
"canonical_text": group["canonical_text"],
"source_item_ids": source_ids,
"source_references": source_references,
"evidence": evidence,
"merge_reason": group["merge_reason"],
}
def build_consolidated_output(
canonicalized: dict[str, Any],
groups: list[dict[str, Any]],
) -> dict[str, Any]:
facts = fact_items(canonicalized)
fact_by_id = {item["item_id"]: item for item in facts}
consolidated_facts = [
build_consolidated_fact_item(group, fact_by_id, index)
for index, group in enumerate(groups, start=1)
]
output_items = consolidated_facts + non_fact_items(canonicalized)
merged_groups = [item for item in consolidated_facts if len(item["source_item_ids"]) > 1]
singleton_groups = [
item for item in consolidated_facts if len(item["source_item_ids"]) == 1
]
output = dict(canonicalized)
output["schema_version"] = "semantic_consolidator_v0"
output["items"] = output_items
output["semantic_consolidation"] = {
"scope": "facts_only",
"merged_fact_group_count": len(merged_groups),
"source_facts_in_merged_groups": sum(
len(item["source_item_ids"]) for item in merged_groups
),
"singleton_fact_group_count": len(singleton_groups),
}
return output
def validate_consolidated_output(
canonicalized: dict[str, Any],
output: dict[str, Any],
) -> None:
original_facts = fact_items(canonicalized)
expected_fact_ids = {item["item_id"] for item in original_facts}
output_items = output.get("items")
if not isinstance(output_items, list):
raise ConsolidationValidationError("Output items must be a list.")
output_fact_groups = [item for item in output_items if item.get("category") == "fact"]
seen: list[str] = []
for item in output_fact_groups:
source_ids = item.get("source_item_ids")
if not isinstance(source_ids, list):
raise ConsolidationValidationError("Fact groups need source_item_ids.")
if len(source_ids) == 0:
raise ConsolidationValidationError("Fact groups must not be empty.")
if len(source_ids) > 1 and not item.get("merge_reason"):
raise ConsolidationValidationError("Merged fact groups need merge_reason.")
seen.extend(source_ids)
source_refs = item.get("source_references")
if not isinstance(source_refs, list) or not source_refs:
raise ConsolidationValidationError("Fact groups need source references.")
seen_counts = Counter(seen)
duplicated = sorted(item_id for item_id, count in seen_counts.items() if count > 1)
if duplicated:
raise ConsolidationValidationError(
f"Output duplicates source fact IDs: {duplicated}"
)
missing = sorted(expected_fact_ids - set(seen))
if missing:
raise ConsolidationValidationError(f"Output misses source fact IDs: {missing}")
unknown = sorted(set(seen) - expected_fact_ids)
if unknown:
raise ConsolidationValidationError(f"Output has unknown source IDs: {unknown}")
original_non_facts = non_fact_items(canonicalized)
output_non_facts = [item for item in output_items if item.get("category") != "fact"]
if output_non_facts != original_non_facts:
raise ConsolidationValidationError("Non-fact categories changed.")
def write_json(path: Path, data: Any) -> None:
path.parent.mkdir(parents=True, exist_ok=True)
path.write_text(
json.dumps(data, ensure_ascii=False, indent=2) + "\n",
encoding="utf-8",
)
def write_report(
path: Path,
model: str,
runtime: float,
prompt_chars: int,
prompt_token_estimate: int,
fact_count: int,
groups: list[dict[str, Any]],
output_path: Path,
) -> None:
merged = [group for group in groups if len(group["source_item_ids"]) > 1]
singletons = [group for group in groups if len(group["source_item_ids"]) == 1]
lines = [
"# Semantic Consolidator V0 Report",
"",
"- Scope: facts only",
f"- Model: `{model}`",
"- LLM call count: 1",
f"- Runtime: {runtime:.3f} seconds",
f"- Fact item count: {fact_count}",
f"- Prompt characters: {prompt_chars}",
f"- Estimated prompt tokens: {prompt_token_estimate}",
f"- Merged fact groups: {len(merged)}",
f"- Source facts involved in merges: {sum(len(group['source_item_ids']) for group in merged)}",
f"- Singleton fact groups: {len(singletons)}",
f"- Output path: `{output_path}`",
"",
"## Actual Merges",
"",
]
if not merged:
lines.append("- None.")
else:
for group in merged:
lines.extend(
[
f"### {group['canonical_text']}",
"",
f"- Source fact IDs: {', '.join(group['source_item_ids'])}",
f"- Merge reason: {group['merge_reason']}",
"",
]
)
path.write_text("\n".join(lines).rstrip() + "\n", encoding="utf-8")
def main() -> int:
args = parse_args()
args.output_dir.mkdir(parents=True, exist_ok=True)
raw_response_path = args.output_dir / "raw_model_response.txt"
output_path = args.output_dir / "consolidated_extractions.json"
report_path = args.output_dir / "report.md"
try:
canonicalized = load_json_object(args.canonicalized_input)
facts = fact_items(canonicalized)
prompt = build_consolidation_prompt(facts)
prompt_chars = len(prompt)
prompt_token_estimate = (prompt_chars + 3) // 4
print(f"Fact item count: {len(facts)}")
print(f"Estimated prompt size chars: {prompt_chars}")
print(f"Estimated prompt tokens: {prompt_token_estimate}")
print("Expected LLM call count: 1")
print("Expected runtime: 5-10 minutes on current local benchmark basis")
raw_text, _response_data, runtime = call_ollama(
endpoint=args.endpoint,
model=args.model,
prompt=prompt,
timeout=args.timeout,
num_ctx=args.num_ctx,
num_predict=args.num_predict,
think=args.think,
progress_interval=args.progress_interval,
)
raw_response_path.write_text(raw_text + "\n", encoding="utf-8")
model_output = parse_model_json(raw_text)
expected_fact_ids = {item["item_id"] for item in facts}
groups = validate_model_groups(model_output, expected_fact_ids)
validate_group_shapes(groups)
output = build_consolidated_output(canonicalized, groups)
validate_consolidated_output(canonicalized, output)
write_json(output_path, output)
write_report(
path=report_path,
model=args.model,
runtime=runtime,
prompt_chars=prompt_chars,
prompt_token_estimate=prompt_token_estimate,
fact_count=len(facts),
groups=groups,
output_path=output_path,
)
except requests.ConnectionError as exc:
print(f"Error: Ollama is not reachable at {args.endpoint}: {exc}", file=sys.stderr)
return 1
except requests.Timeout as exc:
print(
f"Error: Ollama request timed out after {args.timeout} seconds: {exc}",
file=sys.stderr,
)
return 1
except requests.HTTPError as exc:
print(f"Error: Ollama returned an HTTP error: {exc}", file=sys.stderr)
return 1
except (OSError, UnicodeError, ValueError) as exc:
print(f"Error: {exc}", file=sys.stderr)
return 1
merged = [group for group in groups if len(group["source_item_ids"]) > 1]
print(f"Runtime seconds: {runtime:.3f}")
print(f"Merged fact groups: {len(merged)}")
print(
"Source facts involved in merges: "
f"{sum(len(group['source_item_ids']) for group in merged)}"
)
print(f"Singleton fact groups: {len(groups) - len(merged)}")
print("Validation result: passed")
print(f"Output: {output_path}")
print(f"Report: {report_path}")
print(f"Raw model response: {raw_response_path}")
return 0
if __name__ == "__main__":
raise SystemExit(main())
+169
View File
@@ -0,0 +1,169 @@
import json
import tempfile
import unittest
from pathlib import Path
from src.meeting_lab.consolidation.canonicalize import (
canonicalize_extractions,
find_extraction_files,
load_json_object,
)
EMPTY_EXTRACTION = {
"facts": [],
"decisions": [],
"todos": [],
"questions": [],
"positions": [],
"technical": [],
}
def write_extraction(directory: Path, name: str, data: dict) -> Path:
path = directory / name
path.write_text(json.dumps(data), encoding="utf-8")
return path
class CanonicalizeTests(unittest.TestCase):
def test_stable_file_ordering(self) -> None:
with tempfile.TemporaryDirectory() as directory:
root = Path(directory)
write_extraction(root, "chunk_10_extraction.json", EMPTY_EXTRACTION)
write_extraction(root, "chunk_02_extraction.json", EMPTY_EXTRACTION)
write_extraction(root, "chunk_01_extraction.json", EMPTY_EXTRACTION)
self.assertEqual(
[path.name for path in find_extraction_files(root)],
[
"chunk_01_extraction.json",
"chunk_02_extraction.json",
"chunk_10_extraction.json",
],
)
def test_required_category_validation(self) -> None:
with tempfile.TemporaryDirectory() as directory:
root = Path(directory)
data = dict(EMPTY_EXTRACTION)
data.pop("technical")
write_extraction(root, "chunk_01_extraction.json", data)
with self.assertRaisesRegex(ValueError, "technical"):
canonicalize_extractions(root)
def test_stable_item_ids_and_string_parsing(self) -> None:
with tempfile.TemporaryDirectory() as directory:
root = Path(directory)
data = dict(EMPTY_EXTRACTION)
data["facts"] = ["Ada | Revenue increased | clear | Revenue increased"]
data["decisions"] = ["Ship the patch | Agreed to ship"]
data["todos"] = ["Update docs | Mira | Friday | I will update docs"]
data["questions"] = ["Which plan? | Which plan should we use?"]
data["positions"] = ["Kai | Prefer option B | I prefer option B"]
data["technical"] = ["API | Uses v2 mapping | clear | API uses v2"]
write_extraction(root, "chunk_01_extraction.json", data)
output = canonicalize_extractions(root)
items = output["items"]
self.assertEqual(
[item["item_id"] for item in items],
[
"fact_0001",
"decision_0001",
"action_item_0001",
"open_question_0001",
"position_0001",
"technical_detail_0001",
],
)
fact = items[0]
self.assertEqual(fact["speaker"], "Ada")
self.assertEqual(fact["text"], "Revenue increased")
self.assertEqual(fact["status"], "clear")
todo = items[2]
self.assertEqual(todo["responsible"], "Mira")
self.assertEqual(todo["deadline"], "Friday")
self.assertEqual(todo["evidence"], "I will update docs")
def test_source_reference_preservation(self) -> None:
with tempfile.TemporaryDirectory() as directory:
root = Path(directory)
data = dict(EMPTY_EXTRACTION)
data["decisions"] = ["Decision one | Evidence one"]
write_extraction(root, "chunk_01_extraction.json", data)
item = canonicalize_extractions(root)["items"][0]
self.assertEqual(item["source_file"], "chunk_01_extraction.json")
self.assertEqual(item["source_index"], 0)
self.assertEqual(item["original_value"], "Decision one | Evidence one")
self.assertEqual(
item["source_references"],
[
{
"source_file": "chunk_01_extraction.json",
"source_index": 0,
"evidence": "Evidence one",
"original_value": "Decision one | Evidence one",
}
],
)
def test_exact_duplicate_handling(self) -> None:
with tempfile.TemporaryDirectory() as directory:
root = Path(directory)
first = dict(EMPTY_EXTRACTION)
second = dict(EMPTY_EXTRACTION)
first["decisions"] = ["Same decision | Same evidence"]
second["decisions"] = ["Same decision | Same evidence"]
write_extraction(root, "chunk_01_extraction.json", first)
write_extraction(root, "chunk_02_extraction.json", second)
output = canonicalize_extractions(root)
self.assertEqual(len(output["items"]), 1)
self.assertEqual(output["stats"]["exact_duplicates_merged"], 1)
self.assertEqual(output["items"][0]["duplicate_count"], 2)
self.assertEqual(len(output["items"][0]["source_references"]), 2)
def test_no_merging_of_merely_similar_statements(self) -> None:
with tempfile.TemporaryDirectory() as directory:
root = Path(directory)
data = dict(EMPTY_EXTRACTION)
data["decisions"] = [
"Ship the patch Friday | Agreed to ship Friday",
"Ship the patch next week | Agreed to ship next week",
]
write_extraction(root, "chunk_01_extraction.json", data)
output = canonicalize_extractions(root)
self.assertEqual(len(output["items"]), 2)
self.assertEqual(output["stats"]["exact_duplicates_merged"], 0)
def test_invalid_json_handling(self) -> None:
with tempfile.TemporaryDirectory() as directory:
root = Path(directory)
(root / "chunk_01_extraction.json").write_text("{invalid", encoding="utf-8")
with self.assertRaisesRegex(ValueError, "Invalid JSON"):
load_json_object(root / "chunk_01_extraction.json")
def test_empty_categories(self) -> None:
with tempfile.TemporaryDirectory() as directory:
root = Path(directory)
write_extraction(root, "chunk_01_extraction.json", EMPTY_EXTRACTION)
output = canonicalize_extractions(root)
self.assertEqual(output["items"], [])
self.assertEqual(output["stats"]["input_item_count"], 0)
self.assertEqual(output["stats"]["output_item_count"], 0)
if __name__ == "__main__":
unittest.main()
+215
View File
@@ -0,0 +1,215 @@
import unittest
from src.meeting_lab.consolidation.consolidate_facts import (
ConsolidationValidationError,
DEFAULT_NUM_PREDICT,
build_consolidated_output,
build_ollama_payload,
parse_model_json,
validate_consolidated_output,
validate_model_groups,
)
def canonicalized_fixture():
fact_one = {
"item_id": "fact_0001",
"category": "fact",
"text": "The lead maintains the list.",
"evidence": "lead maintains the list",
"source_file": "chunk_01_extraction.json",
"source_index": 0,
"original_value": "The lead maintains the list. | evidence",
"source_references": [
{
"source_file": "chunk_01_extraction.json",
"source_index": 0,
"evidence": "lead maintains the list",
"original_value": "The lead maintains the list. | evidence",
}
],
"duplicate_count": 1,
}
fact_two = {
"item_id": "fact_0002",
"category": "fact",
"text": "The head maintains the project list.",
"evidence": "head maintains the project list",
"source_file": "chunk_02_extraction.json",
"source_index": 0,
"original_value": "The head maintains the project list. | evidence",
"source_references": [
{
"source_file": "chunk_02_extraction.json",
"source_index": 0,
"evidence": "head maintains the project list",
"original_value": "The head maintains the project list. | evidence",
}
],
"duplicate_count": 1,
}
decision = {
"item_id": "decision_0001",
"category": "decision",
"text": "Ship it.",
"evidence": "Agreed.",
"source_file": "chunk_01_extraction.json",
"source_index": 0,
"original_value": "Ship it. | Agreed.",
"source_references": [],
"duplicate_count": 1,
}
return {
"schema_version": "1",
"source_files": ["chunk_01_extraction.json", "chunk_02_extraction.json"],
"items": [fact_one, fact_two, decision],
}
class ConsolidateFactsTests(unittest.TestCase):
def test_payload_construction_disables_streaming_and_thinking_by_default(self):
payload = build_ollama_payload(
model="qwen3.5:9B",
prompt="prompt",
num_ctx=32768,
num_predict=DEFAULT_NUM_PREDICT,
think=False,
)
self.assertEqual(payload["model"], "qwen3.5:9B")
self.assertEqual(payload["prompt"], "prompt")
self.assertIs(payload["stream"], False)
self.assertIs(payload["think"], False)
self.assertEqual(payload["format"], "json")
self.assertEqual(payload["options"]["temperature"], 0.0)
self.assertEqual(payload["options"]["num_ctx"], 32768)
self.assertEqual(payload["options"]["num_predict"], DEFAULT_NUM_PREDICT)
def test_payload_construction_can_enable_thinking_explicitly(self):
payload = build_ollama_payload(
model="qwen3.5:9B",
prompt="prompt",
num_ctx=16384,
num_predict=1024,
think=True,
)
self.assertIs(payload["think"], True)
self.assertEqual(payload["options"]["num_ctx"], 16384)
self.assertEqual(payload["options"]["num_predict"], 1024)
def test_grouping_validation_accepts_complete_singletons(self):
groups = validate_model_groups(
{
"groups": [
{
"canonical_text": "The lead maintains the list.",
"source_item_ids": ["fact_0001"],
"merge_reason": "Singleton.",
},
{
"canonical_text": "The head maintains the project list.",
"source_item_ids": ["fact_0002"],
"merge_reason": "Singleton.",
},
]
},
{"fact_0001", "fact_0002"},
)
self.assertEqual(len(groups), 2)
def test_no_missing_source_ids(self):
with self.assertRaisesRegex(ConsolidationValidationError, "Missing"):
validate_model_groups(
{
"groups": [
{
"canonical_text": "Only one.",
"source_item_ids": ["fact_0001"],
"merge_reason": "Singleton.",
}
]
},
{"fact_0001", "fact_0002"},
)
def test_no_duplicate_source_ids(self):
with self.assertRaisesRegex(ConsolidationValidationError, "multiple"):
validate_model_groups(
{
"groups": [
{
"canonical_text": "One.",
"source_item_ids": ["fact_0001"],
"merge_reason": "Singleton.",
},
{
"canonical_text": "Again.",
"source_item_ids": ["fact_0001"],
"merge_reason": "Singleton.",
},
]
},
{"fact_0001"},
)
def test_merged_group_validation(self):
groups = validate_model_groups(
{
"groups": [
{
"canonical_text": "The lead maintains the project list.",
"source_item_ids": ["fact_0001", "fact_0002"],
"merge_reason": "Same proposition.",
}
]
},
{"fact_0001", "fact_0002"},
)
output = build_consolidated_output(canonicalized_fixture(), groups)
validate_consolidated_output(canonicalized_fixture(), output)
self.assertEqual(output["items"][0]["source_item_ids"], ["fact_0001", "fact_0002"])
def test_preservation_of_non_fact_categories(self):
fixture = canonicalized_fixture()
groups = validate_model_groups(
{
"groups": [
{
"canonical_text": "The lead maintains the project list.",
"source_item_ids": ["fact_0001", "fact_0002"],
"merge_reason": "Same proposition.",
}
]
},
{"fact_0001", "fact_0002"},
)
output = build_consolidated_output(fixture, groups)
self.assertEqual(output["items"][1:], fixture["items"][2:])
def test_invalid_model_json(self):
with self.assertRaisesRegex(ConsolidationValidationError, "Invalid model JSON"):
parse_model_json("{invalid")
def test_unknown_source_item_ids(self):
with self.assertRaisesRegex(ConsolidationValidationError, "unknown"):
validate_model_groups(
{
"groups": [
{
"canonical_text": "Unknown.",
"source_item_ids": ["fact_9999"],
"merge_reason": "Bad ID.",
}
]
},
{"fact_0001"},
)
if __name__ == "__main__":
unittest.main()