Implement Semantic Consolidator V0
- add deterministic canonicalization support for extraction items - add facts-only semantic consolidation using local Ollama - preserve source evidence and validate complete fact coverage - add conservative merge rules and non-LLM tests - record the first validated real-life consolidation benchmark - document current scope, limitations and next evaluation step
This commit is contained in:
@@ -39,6 +39,7 @@ samples/whisper/
|
|||||||
**/meeting_protocol.md
|
**/meeting_protocol.md
|
||||||
**/*.raw.txt
|
**/*.raw.txt
|
||||||
tests/gold/**/actual.json
|
tests/gold/**/actual.json
|
||||||
|
samples/benchmarks/semantic_consolidator_v0/raw_model_response.txt
|
||||||
|
|
||||||
# Lokale Meetings (niemals versionieren)
|
# Lokale Meetings (niemals versionieren)
|
||||||
meeting_data/
|
meeting_data/
|
||||||
|
|||||||
@@ -30,10 +30,11 @@ Audio
|
|||||||
Status:
|
Status:
|
||||||
|
|
||||||
- Implemented: Whisper JSON cleanup script, normalization, technical chunking,
|
- Implemented: Whisper JSON cleanup script, normalization, technical chunking,
|
||||||
local chunk extraction, interim Markdown protocol builder.
|
local chunk extraction, Canonicalizer V1, Semantic Consolidator V0
|
||||||
|
facts-only duplicate detection, interim Markdown protocol builder.
|
||||||
- Experimental/prototype: topic segmentation and review tooling.
|
- Experimental/prototype: topic segmentation and review tooling.
|
||||||
- Planned: Deterministic Canonicalizer, Semantic Consolidator, Canonical
|
- Planned: broader semantic consolidation, Canonical Meeting Knowledge
|
||||||
Meeting Knowledge implementation, final Output Views.
|
implementation, final Output Views.
|
||||||
|
|
||||||
## Architectural Principles
|
## Architectural Principles
|
||||||
|
|
||||||
@@ -46,16 +47,20 @@ Status:
|
|||||||
- Extraction, consolidation, synthesis and rendering are separate concerns.
|
- Extraction, consolidation, synthesis and rendering are separate concerns.
|
||||||
- Deterministic canonicalization and semantic consolidation are separate
|
- Deterministic canonicalization and semantic consolidation are separate
|
||||||
concerns.
|
concerns.
|
||||||
- The Deterministic Canonicalizer is planned Python code. It validates and
|
- Canonicalizer V1 is implemented Python code. It validates and normalizes
|
||||||
normalizes extraction objects, assigns stable source references and IDs,
|
extraction objects, assigns stable source references and IDs, normalizes
|
||||||
normalizes category names and basic field structure, performs only safe
|
category names and basic field structure, performs only safe deterministic
|
||||||
deterministic cleanup, may group exact duplicates, and must preserve all
|
cleanup, may group exact duplicates, and must preserve all source evidence.
|
||||||
source evidence. It must not perform uncertain semantic merging.
|
It must not perform uncertain semantic merging.
|
||||||
- The Semantic Consolidator is planned local-LLM work. It merges semantically
|
- Semantic Consolidator V0 is implemented as local-LLM facts-only duplicate
|
||||||
equivalent statements, groups content by topic, preserves evidence from all
|
detection. It merges semantically equivalent fact items conservatively,
|
||||||
contributing chunks, marks contradictions and uncertainty, separates durable
|
preserves source references and evidence, and validates complete source fact
|
||||||
information from transient discussion, and produces Canonical Meeting
|
coverage. It is not a summarizer, topic grouper, protocol renderer or
|
||||||
Knowledge. It does not directly write a protocol.
|
complete Canonical Meeting Knowledge stage.
|
||||||
|
- Broader semantic consolidation remains planned. It should group content by
|
||||||
|
topic, mark contradictions and uncertainty, separate durable information from
|
||||||
|
transient discussion, and prepare Canonical Meeting Knowledge. It does not
|
||||||
|
directly write a protocol.
|
||||||
- Prefer small, testable processing stages over one monolithic LLM prompt.
|
- Prefer small, testable processing stages over one monolithic LLM prompt.
|
||||||
- Current extraction strategy is one normalized chunk per LLM call.
|
- Current extraction strategy is one normalized chunk per LLM call.
|
||||||
- Do not expand context windows or redesign the extraction strategy without an
|
- Do not expand context windows or redesign the extraction strategy without an
|
||||||
@@ -132,8 +137,8 @@ Current documented Prompt Version 2 decision baseline:
|
|||||||
- `src/meeting_lab/segmentation/`: experimental topic segmentation tooling.
|
- `src/meeting_lab/segmentation/`: experimental topic segmentation tooling.
|
||||||
- `src/meeting_lab/extraction/`: local LLM extraction flow and category
|
- `src/meeting_lab/extraction/`: local LLM extraction flow and category
|
||||||
extractor modules.
|
extractor modules.
|
||||||
- `src/meeting_lab/consolidation/`: planned deterministic canonicalization and
|
- `src/meeting_lab/consolidation/`: Canonicalizer V1 and Semantic
|
||||||
semantic consolidation area.
|
Consolidator V0.
|
||||||
- `src/meeting_lab/protocol/`: interim protocol rendering.
|
- `src/meeting_lab/protocol/`: interim protocol rendering.
|
||||||
- `src/meeting_lab/models/`: current lightweight data models.
|
- `src/meeting_lab/models/`: current lightweight data models.
|
||||||
- `samples/`: sample inputs and generated/experimental artifacts; do not treat
|
- `samples/`: sample inputs and generated/experimental artifacts; do not treat
|
||||||
|
|||||||
@@ -19,6 +19,14 @@
|
|||||||
Verteilerprotokoll and Knowledge Objects / Wissensdatenbankeintrag.
|
Verteilerprotokoll and Knowledge Objects / Wissensdatenbankeintrag.
|
||||||
- Working Protocol Synthesizer V0 benchmark artifact for future
|
- Working Protocol Synthesizer V0 benchmark artifact for future
|
||||||
canonicalizer/consolidator comparisons.
|
canonicalizer/consolidator comparisons.
|
||||||
|
- Canonicalizer V1 deterministic CLI for canonicalizing chunk extraction JSON.
|
||||||
|
- Non-LLM Canonicalizer V1 tests covering ordering, validation, IDs, parsing,
|
||||||
|
source references, exact duplicates, invalid JSON and empty categories.
|
||||||
|
- Semantic Consolidator V0 CLI for conservative facts-only semantic duplicate
|
||||||
|
detection using local Ollama.
|
||||||
|
- Consolidation prompt and non-LLM tests for payload construction, grouping
|
||||||
|
validation and source fact coverage.
|
||||||
|
- Semantic Consolidator V0 benchmark report and consolidated extraction JSON.
|
||||||
|
|
||||||
### Changed
|
### Changed
|
||||||
|
|
||||||
@@ -28,6 +36,9 @@
|
|||||||
are parallel renderings, not derived from one another.
|
are parallel renderings, not derived from one another.
|
||||||
- Documented the planned Deterministic Canonicalizer and Semantic Consolidator
|
- Documented the planned Deterministic Canonicalizer and Semantic Consolidator
|
||||||
stages before Canonical Meeting Knowledge.
|
stages before Canonical Meeting Knowledge.
|
||||||
|
- Clarified that Semantic Consolidator V0 is implemented only for fact
|
||||||
|
duplicate detection; broader semantic synthesis and Canonical Meeting
|
||||||
|
Knowledge remain planned.
|
||||||
- Documented that rendered protocol language should normally match the source
|
- Documented that rendered protocol language should normally match the source
|
||||||
transcript or consolidated meeting knowledge unless explicitly requested
|
transcript or consolidated meeting knowledge unless explicitly requested
|
||||||
otherwise.
|
otherwise.
|
||||||
|
|||||||
+42
-15
@@ -17,6 +17,11 @@ Implemented:
|
|||||||
- Transcript normalization in `src/meeting_lab/normalization/`.
|
- Transcript normalization in `src/meeting_lab/normalization/`.
|
||||||
- Technical chunking in `src/meeting_lab/chunking/`.
|
- Technical chunking in `src/meeting_lab/chunking/`.
|
||||||
- Local per-chunk extraction in `src/meeting_lab/extraction/extract_chunks.py`.
|
- Local per-chunk extraction in `src/meeting_lab/extraction/extract_chunks.py`.
|
||||||
|
- Deterministic Canonicalizer V1 in
|
||||||
|
`src/meeting_lab/consolidation/canonicalize.py`.
|
||||||
|
- Semantic Consolidator V0 in
|
||||||
|
`src/meeting_lab/consolidation/consolidate_facts.py` for facts-only
|
||||||
|
semantic duplicate detection.
|
||||||
- Prompt loading from `src/meeting_lab/llm/prompts.py`.
|
- Prompt loading from `src/meeting_lab/llm/prompts.py`.
|
||||||
- Interim Markdown protocol generation in `src/meeting_lab/protocol/`.
|
- Interim Markdown protocol generation in `src/meeting_lab/protocol/`.
|
||||||
- Non-LLM unit tests for chunking, extraction helpers, protocol rendering and
|
- Non-LLM unit tests for chunking, extraction helpers, protocol rendering and
|
||||||
@@ -30,8 +35,8 @@ Experimental/prototype:
|
|||||||
|
|
||||||
Planned:
|
Planned:
|
||||||
|
|
||||||
- Deterministic Canonicalizer for extraction results.
|
- Broader semantic consolidation for topic grouping, contradiction handling,
|
||||||
- Semantic Consolidator for evidence-preserving semantic merging.
|
uncertainty marking and durable/transient separation.
|
||||||
- Canonical Meeting Knowledge implementation as the semantic source of truth.
|
- Canonical Meeting Knowledge implementation as the semantic source of truth.
|
||||||
- Final Working Protocol / Arbeitsprotokoll, Distribution Protocol /
|
- Final Working Protocol / Arbeitsprotokoll, Distribution Protocol /
|
||||||
Verteilerprotokoll and Knowledge Objects / Wissensdatenbankeintrag renderers.
|
Verteilerprotokoll and Knowledge Objects / Wissensdatenbankeintrag renderers.
|
||||||
@@ -123,17 +128,37 @@ Chunk Extractions
|
|||||||
-> Output View Renderers
|
-> Output View Renderers
|
||||||
```
|
```
|
||||||
|
|
||||||
The Deterministic Canonicalizer is planned Python code with no LLM. It should
|
Canonicalizer V1 is implemented Python code with no LLM. It validates and
|
||||||
validate and normalize extraction objects, assign stable source references and
|
normalizes extraction objects, assigns stable source references and IDs,
|
||||||
IDs, normalize category names and basic field structure, perform only safe
|
normalizes category names and basic field structure, performs only safe
|
||||||
deterministic cleanup, optionally group exact duplicates, and preserve all
|
deterministic cleanup, optionally groups exact duplicates, and preserves all
|
||||||
source evidence. It must not perform uncertain semantic merging.
|
source evidence. It must not perform uncertain semantic merging.
|
||||||
|
|
||||||
The Semantic Consolidator is planned local-LLM work. It should merge
|
CLI:
|
||||||
semantically equivalent statements, group content by topic, preserve evidence
|
|
||||||
from all contributing chunks, mark contradictions and uncertainty, separate
|
```text
|
||||||
durable information from transient discussion, and produce Canonical Meeting
|
PYTHONPATH=src .venv/bin/python -m meeting_lab.consolidation.canonicalize \
|
||||||
Knowledge. It does not directly write a protocol.
|
samples/whisper/meeting_speech_cleaned_chunks \
|
||||||
|
-o /tmp/canonicalized_extractions.json
|
||||||
|
```
|
||||||
|
|
||||||
|
Semantic Consolidator V0 is implemented as narrow local-LLM work for fact
|
||||||
|
items only. It merges semantically equivalent fact statements, preserves source
|
||||||
|
references and evidence, and prefers false negatives over false-positive
|
||||||
|
merges. It is not a summarizer, topic grouper, protocol renderer or complete
|
||||||
|
Canonical Meeting Knowledge stage.
|
||||||
|
|
||||||
|
The first accepted V0 benchmark used `qwen3.5:9B` in one Ollama call over 33
|
||||||
|
fact items. Runtime on the current machine was 390.119 seconds. One correct
|
||||||
|
merge was accepted, involving `fact_0025` and `fact_0031`; 31 facts remained
|
||||||
|
singletons, validation passed, no source fact was lost or duplicated, and
|
||||||
|
non-fact categories remained unchanged. This is a local benchmark, not a
|
||||||
|
general hardware claim.
|
||||||
|
|
||||||
|
Broader semantic consolidation remains planned. It should group content by
|
||||||
|
topic, mark contradictions and uncertainty, separate durable information from
|
||||||
|
transient discussion, and produce Canonical Meeting Knowledge. It does not
|
||||||
|
directly write a protocol.
|
||||||
|
|
||||||
Canonical Meeting Knowledge is the planned semantic intermediate model and
|
Canonical Meeting Knowledge is the planned semantic intermediate model and
|
||||||
future single source of truth. It should be structured, preferably JSON, and
|
future single source of truth. It should be structured, preferably JSON, and
|
||||||
@@ -164,8 +189,8 @@ language is requested.
|
|||||||
- Topic segmentation exists as prototype tooling, not a stable pipeline stage.
|
- Topic segmentation exists as prototype tooling, not a stable pipeline stage.
|
||||||
- Extraction is still a combined current flow, even though separate extractors
|
- Extraction is still a combined current flow, even though separate extractors
|
||||||
are the intended architecture.
|
are the intended architecture.
|
||||||
- Deterministic Canonicalizer is documented but not implemented.
|
- Semantic Consolidator V0 is implemented only for facts-only duplicate
|
||||||
- Semantic Consolidator is documented but not implemented.
|
detection.
|
||||||
- Canonical Meeting Knowledge is documented but not implemented.
|
- Canonical Meeting Knowledge is documented but not implemented.
|
||||||
- Final output views are documented but not implemented.
|
- Final output views are documented but not implemented.
|
||||||
- Most prompt files are placeholders except the common and decision prompts.
|
- Most prompt files are placeholders except the common and decision prompts.
|
||||||
@@ -176,5 +201,7 @@ language is requested.
|
|||||||
Stabilize repeatable local extraction evaluation before broadening the pipeline:
|
Stabilize repeatable local extraction evaluation before broadening the pipeline:
|
||||||
expand Gold Standard coverage by category, keep one-chunk extraction as the
|
expand Gold Standard coverage by category, keep one-chunk extraction as the
|
||||||
baseline, and use small prompt experiments with immediate non-regression checks.
|
baseline, and use small prompt experiments with immediate non-regression checks.
|
||||||
After extraction behavior is stable enough, implement the Deterministic
|
The next recommended evaluation step is to use the consolidated V0 result as
|
||||||
Canonicalizer first, then the Semantic Consolidator with evidence retention.
|
input for the unchanged Working Protocol renderer and compare that output
|
||||||
|
against the Working Protocol Synthesizer V0 baseline and the human reference
|
||||||
|
protocol.
|
||||||
|
|||||||
@@ -102,12 +102,11 @@ Canonical Meeting Knowledge
|
|||||||
Output-Ansichten
|
Output-Ansichten
|
||||||
```
|
```
|
||||||
|
|
||||||
Der nächste Architekturmeilenstein ist die Trennung zwischen deterministischer
|
Der aktuelle Architekturmeilenstein trennt deterministische Kanonisierung der
|
||||||
Kanonisierung der Chunk-Extraktionen und semantischer Konsolidierung.
|
Chunk-Extraktionen von semantischer Konsolidierung. Die Kanonisierung validiert
|
||||||
Die Kanonisierung validiert und normalisiert Extraktionsobjekte ohne LLM. Die
|
und normalisiert Extraktionsobjekte ohne LLM. Semantic Consolidator V0 nutzt
|
||||||
semantische Konsolidierung nutzt das lokale LLM, um gleichbedeutende Aussagen
|
das lokale LLM nur fuer konservative facts-only Duplikaterkennung, erhaelt
|
||||||
zusammenzuführen, Evidenz zu erhalten und die Canonical Meeting Knowledge zu
|
Evidenz und erzeugt noch keine Canonical Meeting Knowledge.
|
||||||
erzeugen.
|
|
||||||
|
|
||||||
Das Meeting Lab behandelt "das Protokoll" nicht mehr als ein einzelnes
|
Das Meeting Lab behandelt "das Protokoll" nicht mehr als ein einzelnes
|
||||||
Endprodukt. Das konsolidierte Meeting-Wissen ist die **Canonical Meeting
|
Endprodukt. Das konsolidierte Meeting-Wissen ist die **Canonical Meeting
|
||||||
@@ -138,10 +137,30 @@ werden, sofern keine explizite Ausgabesprache angefordert wurde.
|
|||||||
## Projektstatus
|
## Projektstatus
|
||||||
|
|
||||||
Aktuell liegt der Schwerpunkt auf der Entwicklung eines modularen
|
Aktuell liegt der Schwerpunkt auf der Entwicklung eines modularen
|
||||||
Diskussionsanalyzers. Implementiert sind Vorverarbeitung, technische Chunking-
|
Diskussionsanalyzers. Implementiert sind Vorverarbeitung, technisches Chunking,
|
||||||
und lokale Chunk-Extraktionsschritte. Deterministic Canonicalizer, Semantic
|
lokale Chunk-Extraktion, Canonicalizer V1 als deterministische Vorbereitung der
|
||||||
Consolidator, Canonical Meeting Knowledge und finale Output-View-Renderer sind
|
Extraktionsergebnisse und Semantic Consolidator V0 fuer facts-only
|
||||||
geplante nächste Schritte.
|
Duplikaterkennung. Canonical Meeting Knowledge, breitere semantische Synthese
|
||||||
|
und finale Output-View-Renderer sind geplante nächste Schritte.
|
||||||
|
|
||||||
|
Canonicalizer V1 kann aus dem Repository heraus so ausgeführt werden:
|
||||||
|
|
||||||
|
```text
|
||||||
|
PYTHONPATH=src .venv/bin/python -m meeting_lab.consolidation.canonicalize \
|
||||||
|
samples/whisper/meeting_speech_cleaned_chunks \
|
||||||
|
-o /tmp/canonicalized_extractions.json
|
||||||
|
```
|
||||||
|
|
||||||
|
Semantic Consolidator V0 kann auf dem Canonicalizer-Output ausgefuehrt werden:
|
||||||
|
|
||||||
|
```text
|
||||||
|
PYTHONPATH=src .venv/bin/python -m meeting_lab.consolidation.consolidate_facts \
|
||||||
|
samples/benchmarks/canonicalizer_v1/canonicalized_extractions.json \
|
||||||
|
-o samples/benchmarks/semantic_consolidator_v0 \
|
||||||
|
--model qwen3.5:9B \
|
||||||
|
--endpoint http://127.0.0.1:11434/api/generate \
|
||||||
|
--no-think
|
||||||
|
```
|
||||||
|
|
||||||
Die eigentliche Ausgabeerzeugung ist bewusst der letzte Verarbeitungsschritt.
|
Die eigentliche Ausgabeerzeugung ist bewusst der letzte Verarbeitungsschritt.
|
||||||
|
|
||||||
|
|||||||
+24
-8
@@ -36,7 +36,7 @@ Goal:
|
|||||||
|
|
||||||
Deliverables:
|
Deliverables:
|
||||||
|
|
||||||
- Deterministic Canonicalizer implemented in Python.
|
- Canonicalizer V1 implemented in Python.
|
||||||
- Stable source references and IDs.
|
- Stable source references and IDs.
|
||||||
- Normalized category names and basic field structure.
|
- Normalized category names and basic field structure.
|
||||||
- Safe deterministic cleanup.
|
- Safe deterministic cleanup.
|
||||||
@@ -48,6 +48,10 @@ Prerequisites:
|
|||||||
- Stable local extraction baseline.
|
- Stable local extraction baseline.
|
||||||
- Agreement on the extraction object shape that should be canonicalized.
|
- Agreement on the extraction object shape that should be canonicalized.
|
||||||
|
|
||||||
|
Current status:
|
||||||
|
|
||||||
|
- Implemented as `meeting_lab.consolidation.canonicalize`.
|
||||||
|
|
||||||
Out of scope:
|
Out of scope:
|
||||||
|
|
||||||
- Uncertain semantic merging.
|
- Uncertain semantic merging.
|
||||||
@@ -64,22 +68,34 @@ Goal:
|
|||||||
|
|
||||||
Deliverables:
|
Deliverables:
|
||||||
|
|
||||||
- Semantic Consolidator using the local LLM.
|
- Semantic Consolidator V0 using the local LLM for facts-only duplicate
|
||||||
- Semantically equivalent statement merging.
|
detection.
|
||||||
- Topic grouping.
|
- Semantically equivalent fact statement merging.
|
||||||
- Evidence preserved from all contributing chunks.
|
- Evidence preserved from all contributing chunks.
|
||||||
- Contradiction and uncertainty markers.
|
- Complete source fact coverage validation.
|
||||||
- Separation of durable information from transient discussion.
|
- Later broader semantic consolidation with topic grouping, contradiction and
|
||||||
- Canonical Meeting Knowledge output.
|
uncertainty markers, durable/transient separation and Canonical Meeting
|
||||||
|
Knowledge preparation.
|
||||||
|
|
||||||
Prerequisites:
|
Prerequisites:
|
||||||
|
|
||||||
- Deterministic Canonicalizer output with stable IDs and source references.
|
- Canonicalizer V1 output with stable IDs and source references.
|
||||||
- Gold or benchmark cases that expose duplication and category shifts.
|
- Gold or benchmark cases that expose duplication and category shifts.
|
||||||
|
|
||||||
|
Current status:
|
||||||
|
|
||||||
|
- Semantic Consolidator V0 is implemented and experimentally validated for
|
||||||
|
facts-only conservative merging.
|
||||||
|
- The first accepted benchmark merged one correct pair among 33 facts and left
|
||||||
|
31 singleton groups.
|
||||||
|
|
||||||
Out of scope:
|
Out of scope:
|
||||||
|
|
||||||
- Direct protocol writing.
|
- Direct protocol writing.
|
||||||
|
- Topic synthesis in V0.
|
||||||
|
- Processing decisions, action items, questions, positions or technical details
|
||||||
|
in V0.
|
||||||
|
- Canonical Meeting Knowledge generation in V0.
|
||||||
- Deriving output views from one another.
|
- Deriving output views from one another.
|
||||||
- Retrieval or RAG integration.
|
- Retrieval or RAG integration.
|
||||||
|
|
||||||
|
|||||||
@@ -195,12 +195,14 @@ Deterministic Canonicalizer:
|
|||||||
Semantic Consolidator:
|
Semantic Consolidator:
|
||||||
|
|
||||||
- uses the local LLM
|
- uses the local LLM
|
||||||
- merges semantically equivalent statements
|
- V0 is implemented for facts-only semantic duplicate detection
|
||||||
- groups content by topic
|
- V0 merges semantically equivalent fact items conservatively
|
||||||
- preserves evidence from all contributing chunks
|
- V0 preserves source references and evidence
|
||||||
- marks contradictions and uncertainty
|
- V0 validates that every source fact appears exactly once
|
||||||
- separates durable information from transient discussion
|
- V0 does not process non-fact categories semantically
|
||||||
- produces Canonical Meeting Knowledge
|
- later versions should group content by topic, mark contradictions and
|
||||||
|
uncertainty, separate durable information from transient discussion and
|
||||||
|
prepare Canonical Meeting Knowledge
|
||||||
- does not directly write a protocol
|
- does not directly write a protocol
|
||||||
|
|
||||||
---
|
---
|
||||||
@@ -267,6 +269,7 @@ Implemented:
|
|||||||
- Transcript normalization
|
- Transcript normalization
|
||||||
- Technical chunk generation
|
- Technical chunk generation
|
||||||
- Experimental LLM-based information extraction
|
- Experimental LLM-based information extraction
|
||||||
|
- Canonicalizer V1 deterministic extraction canonicalization
|
||||||
|
|
||||||
The current extraction step still performs multiple tasks simultaneously.
|
The current extraction step still performs multiple tasks simultaneously.
|
||||||
|
|
||||||
|
|||||||
+63
-12
@@ -242,10 +242,35 @@ This object feeds the Canonical Meeting Knowledge representation.
|
|||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
# Canonical Extraction Object
|
# Canonicalized Extractions
|
||||||
|
|
||||||
Planned deterministic intermediate object created from raw chunk extraction
|
Implemented deterministic intermediate file created from raw chunk extraction
|
||||||
JSON.
|
JSON by Canonicalizer V1. This is not Canonical Meeting Knowledge.
|
||||||
|
|
||||||
|
Top-level structure:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"schema_version": "1",
|
||||||
|
"source_files": [],
|
||||||
|
"stats": {},
|
||||||
|
"items": []
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
Each item contains at least:
|
||||||
|
|
||||||
|
- item_id
|
||||||
|
- category
|
||||||
|
- text
|
||||||
|
- evidence
|
||||||
|
- source_file
|
||||||
|
- source_index
|
||||||
|
- original_value
|
||||||
|
- source_references
|
||||||
|
|
||||||
|
Action items also preserve deterministic fields such as `responsible` and
|
||||||
|
`deadline` when present.
|
||||||
|
|
||||||
Example:
|
Example:
|
||||||
|
|
||||||
@@ -264,16 +289,42 @@ Example:
|
|||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
The Deterministic Canonicalizer should create this kind of object without an
|
Canonicalizer V1 creates this kind of object without an LLM. It validates and
|
||||||
LLM. It validates and normalizes raw extraction objects, assigns stable IDs and
|
normalizes raw extraction objects, assigns stable IDs and source references,
|
||||||
source references, normalizes category names and basic field structure,
|
normalizes category names and basic field structure, performs only safe
|
||||||
performs only safe deterministic cleanup, may group exact duplicates and must
|
deterministic cleanup, may group exact duplicates and must preserve all source
|
||||||
preserve all source evidence.
|
evidence.
|
||||||
|
|
||||||
It must not perform uncertain semantic merging.
|
It must not perform uncertain semantic merging.
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
|
# Semantic Fact Group
|
||||||
|
|
||||||
|
Implemented by Semantic Consolidator V0.
|
||||||
|
|
||||||
|
Example:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"consolidated_id": "fact_group_0001",
|
||||||
|
"category": "fact",
|
||||||
|
"canonical_text": "...",
|
||||||
|
"source_item_ids": ["fact_0001"],
|
||||||
|
"source_references": [],
|
||||||
|
"evidence": [],
|
||||||
|
"merge_reason": "Singleton; no semantically equivalent fact found."
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
Semantic Consolidator V0 only processes fact items. It merges semantically
|
||||||
|
equivalent facts conservatively, preserves source references and evidence, and
|
||||||
|
validates that every source fact appears exactly once. Non-fact categories are
|
||||||
|
copied unchanged. It is not a summarizer, topic grouper, protocol renderer or
|
||||||
|
Canonical Meeting Knowledge generator.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
# Consolidated Topic
|
# Consolidated Topic
|
||||||
|
|
||||||
Planned semantic object produced by the Semantic Consolidator.
|
Planned semantic object produced by the Semantic Consolidator.
|
||||||
@@ -294,10 +345,10 @@ Example:
|
|||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
The Semantic Consolidator may use the local LLM to merge semantically
|
Future Semantic Consolidator versions may use the local LLM to merge
|
||||||
equivalent statements, group content by topic, preserve evidence from all
|
semantically equivalent statements beyond facts, group content by topic,
|
||||||
contributing chunks, mark contradictions and uncertainty and separate durable
|
preserve evidence from all contributing chunks, mark contradictions and
|
||||||
information from transient discussion.
|
uncertainty and separate durable information from transient discussion.
|
||||||
|
|
||||||
It produces Canonical Meeting Knowledge. It does not directly write a protocol.
|
It produces Canonical Meeting Knowledge. It does not directly write a protocol.
|
||||||
|
|
||||||
|
|||||||
+135
-3
@@ -954,9 +954,10 @@ marks contradictions or uncertainty.
|
|||||||
|
|
||||||
Decision:
|
Decision:
|
||||||
|
|
||||||
Deterministic canonicalization is the next implementation step after stable
|
Deterministic canonicalization is implemented as Canonicalizer V1. The first
|
||||||
local extraction, followed by semantic consolidation. Canonical Meeting
|
semantic consolidation milestone is implemented as Semantic Consolidator V0 for
|
||||||
Knowledge and final output views are planned, not implemented.
|
facts-only duplicate detection. Broader semantic consolidation, Canonical
|
||||||
|
Meeting Knowledge and final output views remain planned.
|
||||||
|
|
||||||
Lessons learned:
|
Lessons learned:
|
||||||
|
|
||||||
@@ -1037,3 +1038,134 @@ Evidence:
|
|||||||
- `PROJECT_KNOWLEDGE.md`
|
- `PROJECT_KNOWLEDGE.md`
|
||||||
- `docs/output-views.md`
|
- `docs/output-views.md`
|
||||||
- See EXP-0017 and EXP-0019.
|
- See EXP-0017 and EXP-0019.
|
||||||
|
|
||||||
|
## EXP-0021 - Canonicalizer V1
|
||||||
|
|
||||||
|
Status: Accepted
|
||||||
|
|
||||||
|
Date or period: 2026-07-31
|
||||||
|
|
||||||
|
Hypothesis:
|
||||||
|
|
||||||
|
Independent chunk extraction JSON can be converted into a stable deterministic
|
||||||
|
intermediate format before any semantic LLM consolidation is attempted.
|
||||||
|
|
||||||
|
Setup:
|
||||||
|
|
||||||
|
Canonicalizer V1 discovers `chunk_*_extraction.json` files in stable chunk
|
||||||
|
order, validates required categories, normalizes category names and basic field
|
||||||
|
structure, parses existing legacy string formats where safe, trims redundant
|
||||||
|
whitespace, assigns deterministic IDs, preserves original values and source
|
||||||
|
references, and merges only exact duplicates when all semantic fields are
|
||||||
|
identical.
|
||||||
|
|
||||||
|
Inputs:
|
||||||
|
|
||||||
|
- Synthetic unit-test fixtures.
|
||||||
|
- Existing nine extraction JSON files under
|
||||||
|
`samples/whisper/meeting_speech_cleaned_chunks/`.
|
||||||
|
|
||||||
|
Model / configuration:
|
||||||
|
|
||||||
|
- No LLM.
|
||||||
|
- CLI module: `meeting_lab.consolidation.canonicalize`.
|
||||||
|
|
||||||
|
Result:
|
||||||
|
|
||||||
|
Canonicalizer V1 produces `schema_version`, `source_files`, `stats` and
|
||||||
|
`items`. It is deterministic preparation for the future Semantic Consolidator
|
||||||
|
and is not Canonical Meeting Knowledge.
|
||||||
|
|
||||||
|
Decision:
|
||||||
|
|
||||||
|
Canonicalizer V1 is the current implemented deterministic canonicalization
|
||||||
|
stage. Semantic Consolidator V0 now uses this representation for facts-only
|
||||||
|
semantic duplicate detection; broader semantic consolidation and Canonical
|
||||||
|
Meeting Knowledge remain planned.
|
||||||
|
|
||||||
|
Lessons learned:
|
||||||
|
|
||||||
|
Exact duplicate handling, source-reference preservation and legacy string
|
||||||
|
parsing can be tested without model calls. Any uncertain semantic merge remains
|
||||||
|
out of scope for this stage.
|
||||||
|
|
||||||
|
Evidence:
|
||||||
|
|
||||||
|
- `src/meeting_lab/consolidation/canonicalize.py`
|
||||||
|
- `tests/test_canonicalize.py`
|
||||||
|
- `docs/data-models.md`
|
||||||
|
- `docs/pipeline.md`
|
||||||
|
|
||||||
|
## EXP-0022 - Semantic Consolidator V0 facts-only merge
|
||||||
|
|
||||||
|
Status: Accepted
|
||||||
|
|
||||||
|
Date or period: 2026-07-31
|
||||||
|
|
||||||
|
Hypothesis:
|
||||||
|
|
||||||
|
The Canonicalizer V1 output contains enough stable structure for a local LLM to
|
||||||
|
identify semantically equivalent fact items without losing source coverage or
|
||||||
|
changing non-fact categories.
|
||||||
|
|
||||||
|
Setup:
|
||||||
|
|
||||||
|
Semantic Consolidator V0 consumed the existing Canonicalizer V1 benchmark JSON,
|
||||||
|
selected only items with `category: "fact"`, and sent one bounded consolidation
|
||||||
|
request to local Ollama. The merge rules required semantic equivalence, not
|
||||||
|
topic similarity, and validation required every source fact ID to appear
|
||||||
|
exactly once.
|
||||||
|
|
||||||
|
Inputs:
|
||||||
|
|
||||||
|
- `samples/benchmarks/canonicalizer_v1/canonicalized_extractions.json`
|
||||||
|
- 33 fact items.
|
||||||
|
|
||||||
|
Model / configuration:
|
||||||
|
|
||||||
|
- `qwen3.5:9B`
|
||||||
|
- Ollama endpoint: `http://127.0.0.1:11434/api/generate`
|
||||||
|
- Thinking disabled.
|
||||||
|
- One LLM call.
|
||||||
|
- `num_ctx=32768`
|
||||||
|
- `num_predict=4096`
|
||||||
|
|
||||||
|
Result:
|
||||||
|
|
||||||
|
- Runtime: 390.119 seconds on the current machine.
|
||||||
|
- Merged fact groups: 1.
|
||||||
|
- Source facts involved in merges: 2.
|
||||||
|
- Singleton fact groups: 31.
|
||||||
|
- Validation: passed.
|
||||||
|
- No source fact was lost or duplicated.
|
||||||
|
- Non-fact categories remained unchanged.
|
||||||
|
|
||||||
|
Accepted merge:
|
||||||
|
|
||||||
|
- `fact_0025` + `fact_0031`
|
||||||
|
- Canonical statement: "Der Leiter F&E führt die Projektliste auf dem
|
||||||
|
zweiwöchentlichen Schnittstellen-Stand-Up."
|
||||||
|
|
||||||
|
Decision:
|
||||||
|
|
||||||
|
Semantic Consolidator V0 is complete for its current narrow scope:
|
||||||
|
conservative facts-only semantic duplicate detection with source evidence
|
||||||
|
preserved. The selected `report.md` and `consolidated_extractions.json`
|
||||||
|
benchmark artifacts should be versioned for later comparison. The raw model
|
||||||
|
response remains a local diagnostic artifact and is not versioned.
|
||||||
|
|
||||||
|
Lessons learned:
|
||||||
|
|
||||||
|
Semantic duplicate consolidation is technically viable and conservative enough
|
||||||
|
for continued evaluation, but broader semantic synthesis remains a separate
|
||||||
|
future stage. The measured runtime is useful for this machine and run, but
|
||||||
|
should not be generalized into a universal benchmark.
|
||||||
|
|
||||||
|
Evidence:
|
||||||
|
|
||||||
|
- `src/meeting_lab/consolidation/consolidate_facts.py`
|
||||||
|
- `prompts/consolidate_facts.md`
|
||||||
|
- `tests/test_consolidate_facts.py`
|
||||||
|
- `samples/benchmarks/semantic_consolidator_v0/report.md`
|
||||||
|
- `samples/benchmarks/semantic_consolidator_v0/consolidated_extractions.json`
|
||||||
|
- Local diagnostic only: `samples/benchmarks/semantic_consolidator_v0/raw_model_response.txt`
|
||||||
|
|||||||
+10
-9
@@ -51,15 +51,16 @@ technical validation.
|
|||||||
|
|
||||||
The next planned architecture stage before this representation is explicit:
|
The next planned architecture stage before this representation is explicit:
|
||||||
|
|
||||||
- The Deterministic Canonicalizer validates and normalizes extraction objects,
|
- Canonicalizer V1 validates and normalizes extraction objects, assigns stable
|
||||||
assigns stable source references and IDs, performs only safe deterministic
|
source references and IDs, performs only safe deterministic cleanup and
|
||||||
cleanup and preserves all source evidence. It uses no LLM and must not make
|
preserves all source evidence. It uses no LLM and must not make uncertain
|
||||||
uncertain semantic merges.
|
semantic merges.
|
||||||
- The Semantic Consolidator uses the local LLM to merge semantically equivalent
|
- Semantic Consolidator V0 uses the local LLM only for facts-only semantic
|
||||||
statements, group content by topic, preserve evidence from all contributing
|
duplicate detection. It preserves source evidence and does not directly write
|
||||||
chunks, mark contradictions and uncertainty, separate durable information from
|
a protocol or produce Canonical Meeting Knowledge.
|
||||||
transient discussion and produce Canonical Meeting Knowledge. It does not
|
- Future Semantic Consolidator versions should group content by topic, mark
|
||||||
directly write a protocol.
|
contradictions and uncertainty, separate durable information from transient
|
||||||
|
discussion and prepare Canonical Meeting Knowledge.
|
||||||
|
|
||||||
## Renderers
|
## Renderers
|
||||||
|
|
||||||
|
|||||||
+34
-15
@@ -135,7 +135,15 @@ Deterministic
|
|||||||
|
|
||||||
## Current Status
|
## Current Status
|
||||||
|
|
||||||
Planned
|
Implemented as Canonicalizer V1.
|
||||||
|
|
||||||
|
CLI:
|
||||||
|
|
||||||
|
```text
|
||||||
|
PYTHONPATH=src .venv/bin/python -m meeting_lab.consolidation.canonicalize \
|
||||||
|
samples/whisper/meeting_speech_cleaned_chunks \
|
||||||
|
-o /tmp/canonicalized_extractions.json
|
||||||
|
```
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
@@ -231,7 +239,7 @@ LLM
|
|||||||
|
|
||||||
## Current Status
|
## Current Status
|
||||||
|
|
||||||
Planned
|
Implemented as Canonicalizer V1.
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
@@ -325,16 +333,23 @@ Canonicalized extraction objects.
|
|||||||
|
|
||||||
## Output
|
## Output
|
||||||
|
|
||||||
Canonical Meeting Knowledge.
|
Semantic Consolidator V0 output is `consolidated_extractions.json` with fact
|
||||||
|
groups and unchanged non-fact items.
|
||||||
|
|
||||||
|
Future broader semantic consolidation should produce Canonical Meeting
|
||||||
|
Knowledge.
|
||||||
|
|
||||||
## Responsibilities
|
## Responsibilities
|
||||||
|
|
||||||
- Merge semantically equivalent statements
|
- V0: merge semantically equivalent fact items only
|
||||||
- Group content by topic
|
- V0: preserve all non-fact categories unchanged
|
||||||
|
- V0: validate that every source fact ID appears exactly once
|
||||||
|
- Future: merge semantically equivalent statements across categories
|
||||||
|
- Future: group content by topic
|
||||||
- Preserve evidence from all contributing chunks
|
- Preserve evidence from all contributing chunks
|
||||||
- Mark contradictions and uncertainty
|
- Future: mark contradictions and uncertainty
|
||||||
- Separate durable information from transient discussion
|
- Future: separate durable information from transient discussion
|
||||||
- Reconcile category shifts where supported by evidence
|
- Future: reconcile category shifts where supported by evidence
|
||||||
|
|
||||||
## Must Not
|
## Must Not
|
||||||
|
|
||||||
@@ -348,7 +363,9 @@ Local LLM, with deterministic pre/post-processing where useful.
|
|||||||
|
|
||||||
## Current Status
|
## Current Status
|
||||||
|
|
||||||
Planned
|
Semantic Consolidator V0 is implemented and experimentally validated for
|
||||||
|
facts-only conservative duplicate detection. Broader semantic consolidation and
|
||||||
|
Canonical Meeting Knowledge generation remain planned.
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
@@ -507,16 +524,18 @@ A processing stage may be replaced by another implementation as long as it prese
|
|||||||
|
|
||||||
⬜ Specialized Extraction
|
⬜ Specialized Extraction
|
||||||
|
|
||||||
⬜ Deterministic Canonicalization
|
✔ Deterministic Canonicalization
|
||||||
|
|
||||||
⬜ Semantic Consolidation
|
✅ Semantic Consolidation V0 - facts-only duplicate detection
|
||||||
|
|
||||||
⬜ Canonical Meeting Knowledge
|
⬜ Canonical Meeting Knowledge
|
||||||
|
|
||||||
⬜ Output View Rendering
|
⬜ Output View Rendering
|
||||||
```
|
```
|
||||||
|
|
||||||
The immediate architecture focus is the Deterministic Canonicalizer followed by
|
The immediate evaluation focus is using the Semantic Consolidator V0 output as
|
||||||
the Semantic Consolidator. These stages preserve source evidence, recover
|
input for the unchanged Working Protocol renderer. Canonicalizer V1 now
|
||||||
global context from independent chunk extractions and prepare Canonical Meeting
|
preserves source evidence, and Semantic Consolidator V0 conservatively merges
|
||||||
Knowledge for parallel Output View rendering.
|
semantically equivalent fact items. Broader semantic consolidation should later
|
||||||
|
recover global context from independent chunk extractions and prepare Canonical
|
||||||
|
Meeting Knowledge for parallel Output View rendering.
|
||||||
|
|||||||
@@ -0,0 +1,52 @@
|
|||||||
|
You consolidate fact items from Meeting Lab.
|
||||||
|
|
||||||
|
Return only valid JSON.
|
||||||
|
|
||||||
|
Semantic equivalence is stricter than topical similarity.
|
||||||
|
|
||||||
|
Merge fact items only when they express the same core factual proposition.
|
||||||
|
Related facts are not enough. Facts from the same topic are not enough.
|
||||||
|
|
||||||
|
Preserve distinctions in:
|
||||||
|
|
||||||
|
- actor
|
||||||
|
- scope
|
||||||
|
- timing
|
||||||
|
- condition
|
||||||
|
- certainty
|
||||||
|
- responsibility
|
||||||
|
- current state versus future intention
|
||||||
|
|
||||||
|
Do not merge:
|
||||||
|
|
||||||
|
- broader and narrower statements
|
||||||
|
- cause and effect
|
||||||
|
- process and responsibility
|
||||||
|
- fact and interpretation
|
||||||
|
- related statements from the same topic
|
||||||
|
- statements that differ in actor, scope, timing, condition, or certainty
|
||||||
|
|
||||||
|
Use only the provided fact items.
|
||||||
|
Do not invent information.
|
||||||
|
Do not rewrite evidence.
|
||||||
|
Do not add source item IDs that were not provided.
|
||||||
|
Every provided fact item ID must appear exactly once.
|
||||||
|
When uncertain, keep items separate.
|
||||||
|
False negatives are preferable to false-positive merges.
|
||||||
|
|
||||||
|
Return JSON with this exact top-level shape:
|
||||||
|
|
||||||
|
{
|
||||||
|
"groups": [
|
||||||
|
{
|
||||||
|
"canonical_text": "One factual statement preserving the shared meaning",
|
||||||
|
"source_item_ids": ["fact_0001"],
|
||||||
|
"merge_reason": "Why these items are equivalent, or why this singleton remains separate"
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
|
||||||
|
For singleton groups, use the original fact text as canonical_text.
|
||||||
|
For merged groups, canonical_text must contain only information shared by all
|
||||||
|
source facts in the group.
|
||||||
|
|
||||||
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,168 @@
|
|||||||
|
# Canonicalizer V1 Evaluation Run
|
||||||
|
|
||||||
|
## Scope
|
||||||
|
|
||||||
|
This run evaluates Canonicalizer V1 as deterministic input preparation for a
|
||||||
|
future semantic consolidator. It does not evaluate protocol prose and does not
|
||||||
|
run the semantic consolidator.
|
||||||
|
|
||||||
|
Input files:
|
||||||
|
|
||||||
|
- `samples/whisper/meeting_speech_cleaned_chunks/chunk_01_extraction.json`
|
||||||
|
- `samples/whisper/meeting_speech_cleaned_chunks/chunk_02_extraction.json`
|
||||||
|
- `samples/whisper/meeting_speech_cleaned_chunks/chunk_03_extraction.json`
|
||||||
|
- `samples/whisper/meeting_speech_cleaned_chunks/chunk_04_extraction.json`
|
||||||
|
- `samples/whisper/meeting_speech_cleaned_chunks/chunk_05_extraction.json`
|
||||||
|
- `samples/whisper/meeting_speech_cleaned_chunks/chunk_06_extraction.json`
|
||||||
|
- `samples/whisper/meeting_speech_cleaned_chunks/chunk_07_extraction.json`
|
||||||
|
- `samples/whisper/meeting_speech_cleaned_chunks/chunk_08_extraction.json`
|
||||||
|
- `samples/whisper/meeting_speech_cleaned_chunks/chunk_09_extraction.json`
|
||||||
|
|
||||||
|
Output:
|
||||||
|
|
||||||
|
- `samples/benchmarks/canonicalizer_v1/canonicalized_extractions.json`
|
||||||
|
|
||||||
|
## Run Metrics
|
||||||
|
|
||||||
|
- Runtime: 0.04 seconds
|
||||||
|
- Input file count: 9
|
||||||
|
- Input byte size: 22,695 bytes
|
||||||
|
- Output byte size: 115,924 bytes
|
||||||
|
- Exact duplicate count: 0
|
||||||
|
- Merged exact duplicate count: 0
|
||||||
|
- Malformed or unparsed items: 0
|
||||||
|
- Source-reference mismatches found during comparison: 0
|
||||||
|
|
||||||
|
Input item count by category:
|
||||||
|
|
||||||
|
- fact: 33
|
||||||
|
- decision: 14
|
||||||
|
- action_item: 16
|
||||||
|
- open_question: 10
|
||||||
|
- position: 0
|
||||||
|
- technical_detail: 14
|
||||||
|
|
||||||
|
Output item count by category:
|
||||||
|
|
||||||
|
- fact: 33
|
||||||
|
- decision: 14
|
||||||
|
- action_item: 16
|
||||||
|
- open_question: 10
|
||||||
|
- position: 0
|
||||||
|
- technical_detail: 14
|
||||||
|
|
||||||
|
## Representative Inspection
|
||||||
|
|
||||||
|
### Fact
|
||||||
|
|
||||||
|
- Item: `fact_0001`
|
||||||
|
- Source: `chunk_01_extraction.json`, index 0
|
||||||
|
- Text: `Das ist das aktuelle Projektdeckblatt, was wir in der F&E benutzen.`
|
||||||
|
- Speaker: `Martin`
|
||||||
|
- Status: `clear`
|
||||||
|
- Evidence preserved from the original extraction item.
|
||||||
|
|
||||||
|
Assessment: original information and evidence were preserved, and the source
|
||||||
|
reference points back to the correct original item.
|
||||||
|
|
||||||
|
### Decision
|
||||||
|
|
||||||
|
- Item: `decision_0001`
|
||||||
|
- Source: `chunk_02_extraction.json`, index 0
|
||||||
|
- Text: `Der bestehende Prozess wird grundsätzlich auch für Business Development (BD)-Projekte genutzt, wobei die spezifischen Dokumente je nach Projekttyp angepasst werden.`
|
||||||
|
- Evidence preserved from the original extraction item.
|
||||||
|
|
||||||
|
Assessment: decision text and evidence were preserved without semantic
|
||||||
|
rewriting.
|
||||||
|
|
||||||
|
### Action Item
|
||||||
|
|
||||||
|
- Item: `action_item_0001`
|
||||||
|
- Source: `chunk_01_extraction.json`, index 0
|
||||||
|
- Text: `Giovanna stellt die Kriterien zusammen und erstellt einen ersten Entwurf für den Auswahlkatalog, der EDD-, Marketing- und PM-Kriterien integriert.`
|
||||||
|
- Responsible: `Giovanna`
|
||||||
|
- Deadline: `null`
|
||||||
|
- Evidence preserved from the original extraction item.
|
||||||
|
|
||||||
|
Assessment: responsibility was parsed from the legacy string and no deadline
|
||||||
|
was invented.
|
||||||
|
|
||||||
|
### Open Question
|
||||||
|
|
||||||
|
- Item: `open_question_0001`
|
||||||
|
- Source: `chunk_01_extraction.json`, index 0
|
||||||
|
- Text: `Wie werden spezifische Auswahlkriterien für digitale Produkte (z.B. Portal) versus Realprodukte definiert und integriert?`
|
||||||
|
- Evidence preserved from the original extraction item.
|
||||||
|
|
||||||
|
Assessment: question text and evidence were preserved.
|
||||||
|
|
||||||
|
### Position
|
||||||
|
|
||||||
|
The input extraction files contain zero `positions` items. The canonicalized
|
||||||
|
output correctly reports zero `position` items.
|
||||||
|
|
||||||
|
### Technical Detail
|
||||||
|
|
||||||
|
- Item: `technical_detail_0001`
|
||||||
|
- Source: `chunk_01_extraction.json`, index 0
|
||||||
|
- Subject: `Projektdeckblatt Struktur`
|
||||||
|
- Text: `Das ist eine kurze Projektidee, Ziel, Gegenüber, Entwicklungshemmnisse (Patente), Budgetabfrage.`
|
||||||
|
- Status: `clear`
|
||||||
|
- Evidence preserved from the original extraction item.
|
||||||
|
|
||||||
|
Assessment: subject, statement, status and evidence were parsed without adding
|
||||||
|
technical interpretation.
|
||||||
|
|
||||||
|
## Comparison Findings
|
||||||
|
|
||||||
|
- Original information is preserved. All output items retain `original_value`.
|
||||||
|
- Source references are correct. A comparison against the nine original JSON
|
||||||
|
files found zero source-reference mismatches.
|
||||||
|
- No semantic merges occurred. Output item counts match input item counts in
|
||||||
|
every category.
|
||||||
|
- Names and responsibilities were not invented. Missing optional values remain
|
||||||
|
`null`; for example, `action_item_0001` has `deadline: null`.
|
||||||
|
- Similar but non-identical items remain separate. For example, `fact_0025`
|
||||||
|
and `fact_0031` both discuss the F&E lead maintaining the project list, but
|
||||||
|
they have different wording and remain separate items.
|
||||||
|
- The output is suitable as deterministic input for a later LLM consolidator:
|
||||||
|
it has stable IDs, normalized categories, source references, evidence and
|
||||||
|
original values.
|
||||||
|
|
||||||
|
## Conclusion
|
||||||
|
|
||||||
|
1. Is the deterministic representation lossless enough?
|
||||||
|
|
||||||
|
Yes for the current extraction format. The canonicalized output preserves the
|
||||||
|
original value, parsed text fields, evidence and source references for each
|
||||||
|
item. The output is larger than the input because it adds deterministic
|
||||||
|
metadata and source-reference structure.
|
||||||
|
|
||||||
|
2. Are exact duplicates handled correctly?
|
||||||
|
|
||||||
|
Yes for this run. No exact duplicates were present, so no items were merged.
|
||||||
|
The input and output category counts are identical.
|
||||||
|
|
||||||
|
3. Are legacy extraction strings parsed reliably?
|
||||||
|
|
||||||
|
Yes for the inspected current files. Legacy pipe-delimited strings were parsed
|
||||||
|
into category-specific fields such as `speaker`, `status`, `responsible`,
|
||||||
|
`deadline` and `subject` where safely available. No malformed or unparsed items
|
||||||
|
were detected.
|
||||||
|
|
||||||
|
4. Does the result reduce avoidable work for the semantic consolidator?
|
||||||
|
|
||||||
|
Yes. The future consolidator can consume normalized categories, stable item
|
||||||
|
IDs, source references, evidence and parsed optional fields instead of
|
||||||
|
re-reading heterogeneous legacy strings directly.
|
||||||
|
|
||||||
|
5. What unresolved transformations must remain an LLM task?
|
||||||
|
|
||||||
|
- Merging semantically equivalent but differently worded statements.
|
||||||
|
- Grouping items into coherent topics.
|
||||||
|
- Reconciling facts, decisions, action items and questions across chunks.
|
||||||
|
- Detecting contradictions and uncertainty.
|
||||||
|
- Separating durable organizational knowledge from transient discussion.
|
||||||
|
- Deciding whether similar items such as `fact_0025` and `fact_0031` should be
|
||||||
|
merged, related or kept separate.
|
||||||
|
|
||||||
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,24 @@
|
|||||||
|
# Semantic Consolidator V0 Report
|
||||||
|
|
||||||
|
- Scope: facts only
|
||||||
|
- Model: `qwen3.5:9B`
|
||||||
|
- LLM call count: 1
|
||||||
|
- Runtime: 390.119 seconds
|
||||||
|
- Fact item count: 33
|
||||||
|
- Prompt characters: 15393
|
||||||
|
- Estimated prompt tokens: 3849
|
||||||
|
- Prompt eval count: 4295
|
||||||
|
- Eval count: 2344
|
||||||
|
- Merged fact groups: 1
|
||||||
|
- Source facts involved in merges: 2
|
||||||
|
- Singleton fact groups: 31
|
||||||
|
- Validation: passed
|
||||||
|
- Output path: `samples/benchmarks/semantic_consolidator_v0/consolidated_extractions.json`
|
||||||
|
|
||||||
|
## Actual Merges
|
||||||
|
|
||||||
|
### Der Leiter F&E führt die Projektliste auf dem zweiwöchentlichen Schnittstellen-Stand-Up.
|
||||||
|
|
||||||
|
- Source fact IDs: fact_0025, fact_0031
|
||||||
|
- Merge reason: Semantic equivalence: Both items state that the Head of R&D leads the project list at the bi-weekly interface stand-up meeting. The difference in wording ('auf den' vs 'auf dem') and minor elaboration on defining project types does not change the core factual proposition regarding who performs which action when.
|
||||||
|
- Review decision: accepted as correct.
|
||||||
@@ -0,0 +1,456 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""Canonicalize chunk extraction JSON without semantic merging."""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import json
|
||||||
|
import re
|
||||||
|
import sys
|
||||||
|
from collections import Counter
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
|
||||||
|
SCHEMA_VERSION = "1"
|
||||||
|
|
||||||
|
REQUIRED_CATEGORIES = (
|
||||||
|
"facts",
|
||||||
|
"decisions",
|
||||||
|
"todos",
|
||||||
|
"questions",
|
||||||
|
"positions",
|
||||||
|
"technical",
|
||||||
|
)
|
||||||
|
|
||||||
|
CATEGORY_NAMES = {
|
||||||
|
"facts": "fact",
|
||||||
|
"decisions": "decision",
|
||||||
|
"todos": "action_item",
|
||||||
|
"questions": "open_question",
|
||||||
|
"positions": "position",
|
||||||
|
"technical": "technical_detail",
|
||||||
|
}
|
||||||
|
|
||||||
|
CANONICAL_CATEGORIES = tuple(CATEGORY_NAMES.values())
|
||||||
|
|
||||||
|
CHUNK_EXTRACTION_RE = re.compile(r"^chunk_(\d+)_extraction\.json$")
|
||||||
|
WHITESPACE_RE = re.compile(r"\s+")
|
||||||
|
|
||||||
|
|
||||||
|
def parse_args() -> argparse.Namespace:
|
||||||
|
parser = argparse.ArgumentParser(
|
||||||
|
description="Canonicalize chunk extraction JSON files deterministically."
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"input_dir",
|
||||||
|
type=Path,
|
||||||
|
help="Directory containing chunk_XX_extraction.json files.",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"-o",
|
||||||
|
"--output",
|
||||||
|
type=Path,
|
||||||
|
default=Path("canonicalized_extractions.json"),
|
||||||
|
help="Output JSON path (default: canonicalized_extractions.json).",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--no-merge-exact-duplicates",
|
||||||
|
action="store_true",
|
||||||
|
help="Preserve exact duplicate items instead of merging them.",
|
||||||
|
)
|
||||||
|
return parser.parse_args()
|
||||||
|
|
||||||
|
|
||||||
|
def chunk_sort_key(path: Path) -> tuple[int, str]:
|
||||||
|
match = CHUNK_EXTRACTION_RE.match(path.name)
|
||||||
|
if not match:
|
||||||
|
return (sys.maxsize, path.name)
|
||||||
|
return (int(match.group(1)), path.name)
|
||||||
|
|
||||||
|
|
||||||
|
def find_extraction_files(input_dir: Path) -> list[Path]:
|
||||||
|
if not input_dir.is_dir():
|
||||||
|
raise FileNotFoundError(f"Input directory not found: {input_dir}")
|
||||||
|
return sorted(input_dir.glob("chunk_*_extraction.json"), key=chunk_sort_key)
|
||||||
|
|
||||||
|
|
||||||
|
def load_json_object(path: Path) -> dict[str, Any]:
|
||||||
|
try:
|
||||||
|
data = json.loads(path.read_text(encoding="utf-8-sig"))
|
||||||
|
except json.JSONDecodeError as exc:
|
||||||
|
raise ValueError(f"Invalid JSON in {path}: {exc}") from exc
|
||||||
|
if not isinstance(data, dict):
|
||||||
|
raise ValueError(f"Extraction file must contain a JSON object: {path}")
|
||||||
|
return data
|
||||||
|
|
||||||
|
|
||||||
|
def validate_required_categories(data: dict[str, Any], path: Path) -> None:
|
||||||
|
missing = [category for category in REQUIRED_CATEGORIES if category not in data]
|
||||||
|
if missing:
|
||||||
|
raise ValueError(
|
||||||
|
f"Missing required categories in {path}: {', '.join(missing)}"
|
||||||
|
)
|
||||||
|
invalid = [
|
||||||
|
category
|
||||||
|
for category in REQUIRED_CATEGORIES
|
||||||
|
if not isinstance(data.get(category), list)
|
||||||
|
]
|
||||||
|
if invalid:
|
||||||
|
raise ValueError(
|
||||||
|
f"Required categories must be lists in {path}: {', '.join(invalid)}"
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def clean_text(value: Any) -> str | None:
|
||||||
|
if value is None:
|
||||||
|
return None
|
||||||
|
text = WHITESPACE_RE.sub(" ", str(value)).strip()
|
||||||
|
return text if text else None
|
||||||
|
|
||||||
|
|
||||||
|
def split_legacy_string(value: str) -> list[str]:
|
||||||
|
return [part.strip() for part in value.split("|")]
|
||||||
|
|
||||||
|
|
||||||
|
def first_present(data: dict[str, Any], keys: tuple[str, ...]) -> str | None:
|
||||||
|
for key in keys:
|
||||||
|
text = clean_text(data.get(key))
|
||||||
|
if text:
|
||||||
|
return text
|
||||||
|
return None
|
||||||
|
|
||||||
|
|
||||||
|
def parse_fact(value: Any) -> dict[str, Any]:
|
||||||
|
if isinstance(value, dict):
|
||||||
|
return {
|
||||||
|
"speaker": clean_text(value.get("speaker")),
|
||||||
|
"text": first_present(value, ("statement", "fact", "text")),
|
||||||
|
"status": clean_text(value.get("status")),
|
||||||
|
"evidence": clean_text(value.get("evidence")),
|
||||||
|
}
|
||||||
|
|
||||||
|
text = clean_text(value)
|
||||||
|
parts = split_legacy_string(text or "")
|
||||||
|
if len(parts) >= 4:
|
||||||
|
return {
|
||||||
|
"speaker": clean_text(parts[0]),
|
||||||
|
"text": clean_text(parts[1]),
|
||||||
|
"status": clean_text(parts[2]),
|
||||||
|
"evidence": clean_text(" | ".join(parts[3:])),
|
||||||
|
}
|
||||||
|
if len(parts) >= 3:
|
||||||
|
return {
|
||||||
|
"speaker": None,
|
||||||
|
"text": clean_text(parts[0]),
|
||||||
|
"status": clean_text(parts[1]),
|
||||||
|
"evidence": clean_text(" | ".join(parts[2:])),
|
||||||
|
}
|
||||||
|
return {"speaker": None, "text": text, "status": None, "evidence": None}
|
||||||
|
|
||||||
|
|
||||||
|
def parse_decision(value: Any) -> dict[str, Any]:
|
||||||
|
if isinstance(value, dict):
|
||||||
|
return {
|
||||||
|
"text": first_present(value, ("decision", "text")),
|
||||||
|
"evidence": clean_text(value.get("evidence")),
|
||||||
|
}
|
||||||
|
|
||||||
|
text = clean_text(value)
|
||||||
|
parts = split_legacy_string(text or "")
|
||||||
|
if len(parts) >= 2:
|
||||||
|
return {
|
||||||
|
"text": clean_text(parts[0]),
|
||||||
|
"evidence": clean_text(" | ".join(parts[1:])),
|
||||||
|
}
|
||||||
|
return {"text": text, "evidence": None}
|
||||||
|
|
||||||
|
|
||||||
|
def parse_action_item(value: Any) -> dict[str, Any]:
|
||||||
|
if isinstance(value, dict):
|
||||||
|
return {
|
||||||
|
"text": first_present(value, ("task", "todo", "text")),
|
||||||
|
"responsible": first_present(value, ("responsible", "owner")),
|
||||||
|
"deadline": clean_text(value.get("deadline")),
|
||||||
|
"evidence": clean_text(value.get("evidence")),
|
||||||
|
}
|
||||||
|
|
||||||
|
text = clean_text(value)
|
||||||
|
parts = split_legacy_string(text or "")
|
||||||
|
if len(parts) >= 4:
|
||||||
|
return {
|
||||||
|
"text": clean_text(parts[0]),
|
||||||
|
"responsible": clean_text(parts[1]),
|
||||||
|
"deadline": clean_text(parts[2]),
|
||||||
|
"evidence": clean_text(" | ".join(parts[3:])),
|
||||||
|
}
|
||||||
|
if len(parts) == 3:
|
||||||
|
return {
|
||||||
|
"text": clean_text(parts[0]),
|
||||||
|
"responsible": clean_text(parts[1]),
|
||||||
|
"deadline": None,
|
||||||
|
"evidence": clean_text(parts[2]),
|
||||||
|
}
|
||||||
|
if len(parts) == 2:
|
||||||
|
return {
|
||||||
|
"text": clean_text(parts[0]),
|
||||||
|
"responsible": None,
|
||||||
|
"deadline": None,
|
||||||
|
"evidence": clean_text(parts[1]),
|
||||||
|
}
|
||||||
|
return {
|
||||||
|
"text": text,
|
||||||
|
"responsible": None,
|
||||||
|
"deadline": None,
|
||||||
|
"evidence": None,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def parse_question(value: Any) -> dict[str, Any]:
|
||||||
|
if isinstance(value, dict):
|
||||||
|
return {
|
||||||
|
"text": first_present(value, ("question", "text")),
|
||||||
|
"evidence": clean_text(value.get("evidence")),
|
||||||
|
}
|
||||||
|
|
||||||
|
text = clean_text(value)
|
||||||
|
parts = split_legacy_string(text or "")
|
||||||
|
if len(parts) >= 2:
|
||||||
|
return {
|
||||||
|
"text": clean_text(parts[0]),
|
||||||
|
"evidence": clean_text(" | ".join(parts[1:])),
|
||||||
|
}
|
||||||
|
return {"text": text, "evidence": None}
|
||||||
|
|
||||||
|
|
||||||
|
def parse_position(value: Any) -> dict[str, Any]:
|
||||||
|
if isinstance(value, dict):
|
||||||
|
return {
|
||||||
|
"speaker": clean_text(value.get("speaker")),
|
||||||
|
"text": first_present(value, ("position", "statement", "text")),
|
||||||
|
"evidence": clean_text(value.get("evidence")),
|
||||||
|
}
|
||||||
|
|
||||||
|
text = clean_text(value)
|
||||||
|
parts = split_legacy_string(text or "")
|
||||||
|
if len(parts) >= 3:
|
||||||
|
return {
|
||||||
|
"speaker": clean_text(parts[0]),
|
||||||
|
"text": clean_text(parts[1]),
|
||||||
|
"evidence": clean_text(" | ".join(parts[2:])),
|
||||||
|
}
|
||||||
|
if len(parts) == 2:
|
||||||
|
return {
|
||||||
|
"speaker": None,
|
||||||
|
"text": clean_text(parts[0]),
|
||||||
|
"evidence": clean_text(parts[1]),
|
||||||
|
}
|
||||||
|
return {"speaker": None, "text": text, "evidence": None}
|
||||||
|
|
||||||
|
|
||||||
|
def parse_technical_detail(value: Any) -> dict[str, Any]:
|
||||||
|
if isinstance(value, dict):
|
||||||
|
return {
|
||||||
|
"subject": clean_text(value.get("subject")),
|
||||||
|
"text": first_present(value, ("statement", "technical", "text")),
|
||||||
|
"status": clean_text(value.get("status")),
|
||||||
|
"evidence": clean_text(value.get("evidence")),
|
||||||
|
}
|
||||||
|
|
||||||
|
text = clean_text(value)
|
||||||
|
parts = split_legacy_string(text or "")
|
||||||
|
if len(parts) >= 4:
|
||||||
|
return {
|
||||||
|
"subject": clean_text(parts[0]),
|
||||||
|
"text": clean_text(parts[1]),
|
||||||
|
"status": clean_text(parts[2]),
|
||||||
|
"evidence": clean_text(" | ".join(parts[3:])),
|
||||||
|
}
|
||||||
|
if len(parts) >= 3:
|
||||||
|
return {
|
||||||
|
"subject": None,
|
||||||
|
"text": clean_text(parts[0]),
|
||||||
|
"status": clean_text(parts[1]),
|
||||||
|
"evidence": clean_text(" | ".join(parts[2:])),
|
||||||
|
}
|
||||||
|
return {"subject": None, "text": text, "status": None, "evidence": None}
|
||||||
|
|
||||||
|
|
||||||
|
PARSERS = {
|
||||||
|
"fact": parse_fact,
|
||||||
|
"decision": parse_decision,
|
||||||
|
"action_item": parse_action_item,
|
||||||
|
"open_question": parse_question,
|
||||||
|
"position": parse_position,
|
||||||
|
"technical_detail": parse_technical_detail,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def source_reference(
|
||||||
|
source_file: str,
|
||||||
|
source_index: int,
|
||||||
|
original_value: Any,
|
||||||
|
evidence: str | None,
|
||||||
|
) -> dict[str, Any]:
|
||||||
|
return {
|
||||||
|
"source_file": source_file,
|
||||||
|
"source_index": source_index,
|
||||||
|
"evidence": evidence,
|
||||||
|
"original_value": original_value,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def semantic_key(item: dict[str, Any]) -> str:
|
||||||
|
ignored = {
|
||||||
|
"item_id",
|
||||||
|
"source_file",
|
||||||
|
"source_index",
|
||||||
|
"original_value",
|
||||||
|
"source_references",
|
||||||
|
"duplicate_count",
|
||||||
|
}
|
||||||
|
comparable = {key: value for key, value in item.items() if key not in ignored}
|
||||||
|
return json.dumps(comparable, ensure_ascii=False, sort_keys=True)
|
||||||
|
|
||||||
|
|
||||||
|
def item_id_for(category: str, counts: Counter[str]) -> str:
|
||||||
|
counts[category] += 1
|
||||||
|
return f"{category}_{counts[category]:04d}"
|
||||||
|
|
||||||
|
|
||||||
|
def canonicalize_value(
|
||||||
|
category: str,
|
||||||
|
value: Any,
|
||||||
|
source_file: str,
|
||||||
|
source_index: int,
|
||||||
|
counts: Counter[str],
|
||||||
|
) -> dict[str, Any]:
|
||||||
|
parsed = PARSERS[category](value)
|
||||||
|
item: dict[str, Any] = {
|
||||||
|
"item_id": item_id_for(category, counts),
|
||||||
|
"category": category,
|
||||||
|
"text": parsed.pop("text", None),
|
||||||
|
"evidence": parsed.pop("evidence", None),
|
||||||
|
"source_file": source_file,
|
||||||
|
"source_index": source_index,
|
||||||
|
"original_value": value,
|
||||||
|
}
|
||||||
|
for key, parsed_value in parsed.items():
|
||||||
|
item[key] = parsed_value
|
||||||
|
item["source_references"] = [
|
||||||
|
source_reference(source_file, source_index, value, item["evidence"])
|
||||||
|
]
|
||||||
|
return item
|
||||||
|
|
||||||
|
|
||||||
|
def merge_exact_duplicates(items: list[dict[str, Any]]) -> tuple[list[dict[str, Any]], int]:
|
||||||
|
merged: list[dict[str, Any]] = []
|
||||||
|
seen: dict[str, dict[str, Any]] = {}
|
||||||
|
duplicates = 0
|
||||||
|
|
||||||
|
for item in items:
|
||||||
|
key = semantic_key(item)
|
||||||
|
existing = seen.get(key)
|
||||||
|
if existing is None:
|
||||||
|
item["duplicate_count"] = 1
|
||||||
|
seen[key] = item
|
||||||
|
merged.append(item)
|
||||||
|
continue
|
||||||
|
|
||||||
|
duplicates += 1
|
||||||
|
existing["duplicate_count"] += 1
|
||||||
|
existing["source_references"].extend(item["source_references"])
|
||||||
|
|
||||||
|
return merged, duplicates
|
||||||
|
|
||||||
|
|
||||||
|
def canonicalize_extractions(
|
||||||
|
input_dir: Path,
|
||||||
|
merge_duplicates: bool = True,
|
||||||
|
) -> dict[str, Any]:
|
||||||
|
files = find_extraction_files(input_dir)
|
||||||
|
if not files:
|
||||||
|
raise ValueError(f"No chunk extraction JSON files found in {input_dir}")
|
||||||
|
|
||||||
|
counts: Counter[str] = Counter()
|
||||||
|
input_counts: Counter[str] = Counter()
|
||||||
|
items: list[dict[str, Any]] = []
|
||||||
|
|
||||||
|
for path in files:
|
||||||
|
data = load_json_object(path)
|
||||||
|
validate_required_categories(data, path)
|
||||||
|
for raw_category in REQUIRED_CATEGORIES:
|
||||||
|
category = CATEGORY_NAMES[raw_category]
|
||||||
|
values = data[raw_category]
|
||||||
|
input_counts[category] += len(values)
|
||||||
|
for source_index, value in enumerate(values):
|
||||||
|
items.append(
|
||||||
|
canonicalize_value(
|
||||||
|
category=category,
|
||||||
|
value=value,
|
||||||
|
source_file=path.name,
|
||||||
|
source_index=source_index,
|
||||||
|
counts=counts,
|
||||||
|
)
|
||||||
|
)
|
||||||
|
|
||||||
|
exact_duplicates = 0
|
||||||
|
if merge_duplicates:
|
||||||
|
items, exact_duplicates = merge_exact_duplicates(items)
|
||||||
|
|
||||||
|
output_counts = Counter(item["category"] for item in items)
|
||||||
|
input_counts_by_category = {
|
||||||
|
category: input_counts[category] for category in CANONICAL_CATEGORIES
|
||||||
|
}
|
||||||
|
output_counts_by_category = {
|
||||||
|
category: output_counts[category] for category in CANONICAL_CATEGORIES
|
||||||
|
}
|
||||||
|
|
||||||
|
return {
|
||||||
|
"schema_version": SCHEMA_VERSION,
|
||||||
|
"source_files": [path.name for path in files],
|
||||||
|
"stats": {
|
||||||
|
"input_item_count": sum(input_counts.values()),
|
||||||
|
"output_item_count": len(items),
|
||||||
|
"input_item_count_by_category": input_counts_by_category,
|
||||||
|
"output_item_count_by_category": output_counts_by_category,
|
||||||
|
"exact_duplicates_merged": exact_duplicates,
|
||||||
|
},
|
||||||
|
"items": items,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def write_canonicalized(output: dict[str, Any], output_path: Path) -> Path:
|
||||||
|
output_path.parent.mkdir(parents=True, exist_ok=True)
|
||||||
|
output_path.write_text(
|
||||||
|
json.dumps(output, ensure_ascii=False, indent=2) + "\n",
|
||||||
|
encoding="utf-8",
|
||||||
|
)
|
||||||
|
return output_path
|
||||||
|
|
||||||
|
|
||||||
|
def main() -> int:
|
||||||
|
args = parse_args()
|
||||||
|
try:
|
||||||
|
output = canonicalize_extractions(
|
||||||
|
args.input_dir,
|
||||||
|
merge_duplicates=not args.no_merge_exact_duplicates,
|
||||||
|
)
|
||||||
|
output_path = write_canonicalized(output, args.output)
|
||||||
|
except (OSError, UnicodeError, ValueError) as exc:
|
||||||
|
print(f"Error: {exc}", file=sys.stderr)
|
||||||
|
return 1
|
||||||
|
|
||||||
|
stats = output["stats"]
|
||||||
|
print(f"Input directory: {args.input_dir}")
|
||||||
|
print(f"Input files processed: {len(output['source_files'])}")
|
||||||
|
print(f"Input item count by category: {stats['input_item_count_by_category']}")
|
||||||
|
print(f"Output item count by category: {stats['output_item_count_by_category']}")
|
||||||
|
print(f"Exact duplicates merged: {stats['exact_duplicates_merged']}")
|
||||||
|
print(f"Output: {output_path}")
|
||||||
|
print(f"Output JSON size: {output_path.stat().st_size} bytes")
|
||||||
|
return 0
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
raise SystemExit(main())
|
||||||
@@ -0,0 +1,603 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""LLM-backed V0 semantic consolidation for fact items only."""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import json
|
||||||
|
import sys
|
||||||
|
import time
|
||||||
|
from collections import Counter
|
||||||
|
from datetime import datetime
|
||||||
|
from pathlib import Path
|
||||||
|
from queue import Empty, Queue
|
||||||
|
from threading import Thread
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
import requests
|
||||||
|
|
||||||
|
try:
|
||||||
|
from meeting_lab.llm.prompts import load_prompt
|
||||||
|
except ModuleNotFoundError: # pragma: no cover - used by repository-root tests.
|
||||||
|
from src.meeting_lab.llm.prompts import load_prompt
|
||||||
|
|
||||||
|
|
||||||
|
DEFAULT_MODEL = "qwen3.5:9b"
|
||||||
|
DEFAULT_ENDPOINT = "http://127.0.0.1:11434/api/generate"
|
||||||
|
DEFAULT_NUM_CTX = 32768
|
||||||
|
DEFAULT_NUM_PREDICT = 4096
|
||||||
|
DEFAULT_PROGRESS_INTERVAL = 30
|
||||||
|
PROMPT_NAME = "consolidate_facts.md"
|
||||||
|
|
||||||
|
|
||||||
|
class ConsolidationValidationError(ValueError):
|
||||||
|
"""Raised when model consolidation output violates strict invariants."""
|
||||||
|
|
||||||
|
|
||||||
|
def parse_args() -> argparse.Namespace:
|
||||||
|
parser = argparse.ArgumentParser(
|
||||||
|
description="Consolidate semantically equivalent fact items only."
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"canonicalized_input",
|
||||||
|
type=Path,
|
||||||
|
help="Canonicalizer V1 JSON file.",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"-o",
|
||||||
|
"--output-dir",
|
||||||
|
type=Path,
|
||||||
|
required=True,
|
||||||
|
help="Directory for consolidated_extractions.json, report.md and raw response.",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--model",
|
||||||
|
default=DEFAULT_MODEL,
|
||||||
|
help=f"Ollama model name (default: {DEFAULT_MODEL}).",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--endpoint",
|
||||||
|
default=DEFAULT_ENDPOINT,
|
||||||
|
help=f"Ollama generate endpoint (default: {DEFAULT_ENDPOINT}).",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--timeout",
|
||||||
|
type=int,
|
||||||
|
default=1800,
|
||||||
|
help="HTTP timeout in seconds (default: 1800).",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--num-ctx",
|
||||||
|
type=int,
|
||||||
|
default=DEFAULT_NUM_CTX,
|
||||||
|
help=f"Context window tokens (default: {DEFAULT_NUM_CTX}).",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--num-predict",
|
||||||
|
type=int,
|
||||||
|
default=DEFAULT_NUM_PREDICT,
|
||||||
|
help=(
|
||||||
|
"Maximum generated tokens. The default is bounded for the expected "
|
||||||
|
f"fact-group JSON while leaving truncation headroom (default: {DEFAULT_NUM_PREDICT})."
|
||||||
|
),
|
||||||
|
)
|
||||||
|
thinking = parser.add_mutually_exclusive_group()
|
||||||
|
thinking.add_argument(
|
||||||
|
"--think",
|
||||||
|
dest="think",
|
||||||
|
action="store_true",
|
||||||
|
help="Enable Ollama thinking output when the selected model supports it.",
|
||||||
|
)
|
||||||
|
thinking.add_argument(
|
||||||
|
"--no-think",
|
||||||
|
dest="think",
|
||||||
|
action="store_false",
|
||||||
|
help="Disable Ollama thinking output for structured JSON consolidation.",
|
||||||
|
)
|
||||||
|
parser.set_defaults(think=False)
|
||||||
|
parser.add_argument(
|
||||||
|
"--progress-interval",
|
||||||
|
type=int,
|
||||||
|
default=DEFAULT_PROGRESS_INTERVAL,
|
||||||
|
help=(
|
||||||
|
"Seconds between waiting-status messages while the non-streaming "
|
||||||
|
f"Ollama request is in flight (default: {DEFAULT_PROGRESS_INTERVAL})."
|
||||||
|
),
|
||||||
|
)
|
||||||
|
return parser.parse_args()
|
||||||
|
|
||||||
|
|
||||||
|
def load_json_object(path: Path) -> dict[str, Any]:
|
||||||
|
try:
|
||||||
|
data = json.loads(path.read_text(encoding="utf-8-sig"))
|
||||||
|
except json.JSONDecodeError as exc:
|
||||||
|
raise ValueError(f"Invalid JSON in {path}: {exc}") from exc
|
||||||
|
if not isinstance(data, dict):
|
||||||
|
raise ValueError(f"JSON file must contain an object: {path}")
|
||||||
|
return data
|
||||||
|
|
||||||
|
|
||||||
|
def fact_items(canonicalized: dict[str, Any]) -> list[dict[str, Any]]:
|
||||||
|
items = canonicalized.get("items")
|
||||||
|
if not isinstance(items, list):
|
||||||
|
raise ValueError("Canonicalized input must contain an items list.")
|
||||||
|
return [item for item in items if item.get("category") == "fact"]
|
||||||
|
|
||||||
|
|
||||||
|
def non_fact_items(canonicalized: dict[str, Any]) -> list[dict[str, Any]]:
|
||||||
|
items = canonicalized.get("items")
|
||||||
|
if not isinstance(items, list):
|
||||||
|
raise ValueError("Canonicalized input must contain an items list.")
|
||||||
|
return [item for item in items if item.get("category") != "fact"]
|
||||||
|
|
||||||
|
|
||||||
|
def model_fact_payload(facts: list[dict[str, Any]]) -> list[dict[str, Any]]:
|
||||||
|
payload: list[dict[str, Any]] = []
|
||||||
|
for item in facts:
|
||||||
|
payload.append(
|
||||||
|
{
|
||||||
|
"item_id": item.get("item_id"),
|
||||||
|
"text": item.get("text"),
|
||||||
|
"evidence": item.get("evidence"),
|
||||||
|
"speaker": item.get("speaker"),
|
||||||
|
"status": item.get("status"),
|
||||||
|
"source_file": item.get("source_file"),
|
||||||
|
"source_index": item.get("source_index"),
|
||||||
|
}
|
||||||
|
)
|
||||||
|
return payload
|
||||||
|
|
||||||
|
|
||||||
|
def build_consolidation_prompt(facts: list[dict[str, Any]]) -> str:
|
||||||
|
task_prompt = load_prompt(PROMPT_NAME)
|
||||||
|
payload = json.dumps(
|
||||||
|
{"fact_items": model_fact_payload(facts)},
|
||||||
|
ensure_ascii=False,
|
||||||
|
indent=2,
|
||||||
|
)
|
||||||
|
return f"{task_prompt}\n\nFACT ITEMS:\n{payload}\n"
|
||||||
|
|
||||||
|
|
||||||
|
def response_text_from_ollama_data(data: dict[str, Any]) -> str | None:
|
||||||
|
text = data.get("response")
|
||||||
|
if isinstance(text, str) and text.strip():
|
||||||
|
return text
|
||||||
|
message = data.get("message")
|
||||||
|
if isinstance(message, dict):
|
||||||
|
content = message.get("content")
|
||||||
|
if isinstance(content, str) and content.strip():
|
||||||
|
return content
|
||||||
|
return text if isinstance(text, str) else None
|
||||||
|
|
||||||
|
|
||||||
|
def build_ollama_payload(
|
||||||
|
model: str,
|
||||||
|
prompt: str,
|
||||||
|
num_ctx: int,
|
||||||
|
num_predict: int,
|
||||||
|
think: bool,
|
||||||
|
) -> dict[str, Any]:
|
||||||
|
return {
|
||||||
|
"model": model,
|
||||||
|
"prompt": prompt,
|
||||||
|
"think": think,
|
||||||
|
"stream": False,
|
||||||
|
"format": "json",
|
||||||
|
"options": {
|
||||||
|
"temperature": 0.0,
|
||||||
|
"num_ctx": num_ctx,
|
||||||
|
"num_predict": num_predict,
|
||||||
|
},
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def print_response_metadata(data: dict[str, Any]) -> None:
|
||||||
|
fields = [
|
||||||
|
"total_duration",
|
||||||
|
"load_duration",
|
||||||
|
"prompt_eval_count",
|
||||||
|
"prompt_eval_duration",
|
||||||
|
"eval_count",
|
||||||
|
"eval_duration",
|
||||||
|
]
|
||||||
|
present = [(field, data.get(field)) for field in fields if field in data]
|
||||||
|
if not present:
|
||||||
|
print("Ollama response metadata: unavailable")
|
||||||
|
return
|
||||||
|
print("Ollama response metadata:")
|
||||||
|
for field, value in present:
|
||||||
|
print(f" {field}: {value}")
|
||||||
|
|
||||||
|
|
||||||
|
def post_with_progress(
|
||||||
|
endpoint: str,
|
||||||
|
payload: dict[str, Any],
|
||||||
|
timeout: int,
|
||||||
|
progress_interval: int,
|
||||||
|
) -> tuple[requests.Response, float]:
|
||||||
|
results: Queue[tuple[str, requests.Response | BaseException, float]] = Queue()
|
||||||
|
|
||||||
|
def worker() -> None:
|
||||||
|
started = time.perf_counter()
|
||||||
|
try:
|
||||||
|
response = requests.post(endpoint, json=payload, timeout=timeout)
|
||||||
|
except BaseException as exc: # noqa: BLE001 - forwarded to main thread.
|
||||||
|
results.put(("error", exc, time.perf_counter() - started))
|
||||||
|
else:
|
||||||
|
results.put(("response", response, time.perf_counter() - started))
|
||||||
|
|
||||||
|
started_wait = time.perf_counter()
|
||||||
|
thread = Thread(target=worker, daemon=True)
|
||||||
|
thread.start()
|
||||||
|
while True:
|
||||||
|
try:
|
||||||
|
status, result, elapsed = results.get(timeout=max(progress_interval, 1))
|
||||||
|
except Empty:
|
||||||
|
print(
|
||||||
|
"Waiting for Ollama response: "
|
||||||
|
f"{time.perf_counter() - started_wait:.1f} seconds elapsed",
|
||||||
|
flush=True,
|
||||||
|
)
|
||||||
|
continue
|
||||||
|
if status == "error":
|
||||||
|
raise result
|
||||||
|
return result, elapsed
|
||||||
|
|
||||||
|
|
||||||
|
def call_ollama(
|
||||||
|
endpoint: str,
|
||||||
|
model: str,
|
||||||
|
prompt: str,
|
||||||
|
timeout: int,
|
||||||
|
num_ctx: int,
|
||||||
|
num_predict: int,
|
||||||
|
think: bool,
|
||||||
|
progress_interval: int,
|
||||||
|
) -> tuple[str, dict[str, Any], float]:
|
||||||
|
payload = build_ollama_payload(
|
||||||
|
model=model,
|
||||||
|
prompt=prompt,
|
||||||
|
num_ctx=num_ctx,
|
||||||
|
num_predict=num_predict,
|
||||||
|
think=think,
|
||||||
|
)
|
||||||
|
print(f"Ollama request start: {datetime.now().isoformat(timespec='seconds')}")
|
||||||
|
print(f"Ollama endpoint: {endpoint}")
|
||||||
|
print(f"Ollama model: {model}")
|
||||||
|
print(f"Ollama stream: {payload['stream']}")
|
||||||
|
print(f"Ollama think: {payload['think']}")
|
||||||
|
print(f"Ollama timeout seconds: {timeout}")
|
||||||
|
print(f"Ollama options: {json.dumps(payload['options'], sort_keys=True)}")
|
||||||
|
response, elapsed = post_with_progress(
|
||||||
|
endpoint=endpoint,
|
||||||
|
payload=payload,
|
||||||
|
timeout=timeout,
|
||||||
|
progress_interval=progress_interval,
|
||||||
|
)
|
||||||
|
response.raise_for_status()
|
||||||
|
data = response.json()
|
||||||
|
print_response_metadata(data)
|
||||||
|
text = response_text_from_ollama_data(data)
|
||||||
|
if not isinstance(text, str) or not text.strip():
|
||||||
|
raise ValueError("Ollama returned no usable response text.")
|
||||||
|
return text.strip(), data, elapsed
|
||||||
|
|
||||||
|
|
||||||
|
def parse_model_json(text: str) -> dict[str, Any]:
|
||||||
|
try:
|
||||||
|
parsed = json.loads(text)
|
||||||
|
except json.JSONDecodeError as exc:
|
||||||
|
raise ConsolidationValidationError(f"Invalid model JSON: {exc}") from exc
|
||||||
|
if not isinstance(parsed, dict):
|
||||||
|
raise ConsolidationValidationError("Model JSON must be an object.")
|
||||||
|
return parsed
|
||||||
|
|
||||||
|
|
||||||
|
def validate_model_groups(
|
||||||
|
model_output: dict[str, Any],
|
||||||
|
expected_fact_ids: set[str],
|
||||||
|
) -> list[dict[str, Any]]:
|
||||||
|
groups = model_output.get("groups")
|
||||||
|
if not isinstance(groups, list):
|
||||||
|
raise ConsolidationValidationError("Model output must contain a groups list.")
|
||||||
|
|
||||||
|
seen: list[str] = []
|
||||||
|
validated: list[dict[str, Any]] = []
|
||||||
|
for index, group in enumerate(groups):
|
||||||
|
if not isinstance(group, dict):
|
||||||
|
raise ConsolidationValidationError(f"Group {index} must be an object.")
|
||||||
|
source_ids = group.get("source_item_ids")
|
||||||
|
if not isinstance(source_ids, list) or not source_ids:
|
||||||
|
raise ConsolidationValidationError(
|
||||||
|
f"Group {index} must contain source_item_ids."
|
||||||
|
)
|
||||||
|
if not all(isinstance(item_id, str) for item_id in source_ids):
|
||||||
|
raise ConsolidationValidationError(
|
||||||
|
f"Group {index} source_item_ids must be strings."
|
||||||
|
)
|
||||||
|
unknown = sorted(set(source_ids) - expected_fact_ids)
|
||||||
|
if unknown:
|
||||||
|
raise ConsolidationValidationError(
|
||||||
|
f"Group {index} contains unknown source item IDs: {unknown}"
|
||||||
|
)
|
||||||
|
duplicates_in_group = [
|
||||||
|
item_id for item_id, count in Counter(source_ids).items() if count > 1
|
||||||
|
]
|
||||||
|
if duplicates_in_group:
|
||||||
|
raise ConsolidationValidationError(
|
||||||
|
f"Group {index} repeats source item IDs: {duplicates_in_group}"
|
||||||
|
)
|
||||||
|
canonical_text = group.get("canonical_text")
|
||||||
|
if not isinstance(canonical_text, str) or not canonical_text.strip():
|
||||||
|
raise ConsolidationValidationError(
|
||||||
|
f"Group {index} must contain canonical_text."
|
||||||
|
)
|
||||||
|
merge_reason = group.get("merge_reason")
|
||||||
|
if not isinstance(merge_reason, str) or not merge_reason.strip():
|
||||||
|
raise ConsolidationValidationError(
|
||||||
|
f"Group {index} must contain merge_reason."
|
||||||
|
)
|
||||||
|
seen.extend(source_ids)
|
||||||
|
validated.append(
|
||||||
|
{
|
||||||
|
"canonical_text": canonical_text.strip(),
|
||||||
|
"source_item_ids": source_ids,
|
||||||
|
"merge_reason": merge_reason.strip(),
|
||||||
|
}
|
||||||
|
)
|
||||||
|
|
||||||
|
seen_counts = Counter(seen)
|
||||||
|
duplicated = sorted(item_id for item_id, count in seen_counts.items() if count > 1)
|
||||||
|
if duplicated:
|
||||||
|
raise ConsolidationValidationError(
|
||||||
|
f"Source item IDs appear in multiple groups: {duplicated}"
|
||||||
|
)
|
||||||
|
missing = sorted(expected_fact_ids - set(seen))
|
||||||
|
if missing:
|
||||||
|
raise ConsolidationValidationError(f"Missing source item IDs: {missing}")
|
||||||
|
|
||||||
|
return validated
|
||||||
|
|
||||||
|
|
||||||
|
def validate_group_shapes(groups: list[dict[str, Any]]) -> None:
|
||||||
|
for group in groups:
|
||||||
|
source_count = len(group["source_item_ids"])
|
||||||
|
if source_count < 1:
|
||||||
|
raise ConsolidationValidationError("Groups must not be empty.")
|
||||||
|
if source_count == 1:
|
||||||
|
continue
|
||||||
|
if source_count < 2:
|
||||||
|
raise ConsolidationValidationError("Merged groups need at least two IDs.")
|
||||||
|
|
||||||
|
|
||||||
|
def build_consolidated_fact_item(
|
||||||
|
group: dict[str, Any],
|
||||||
|
fact_by_id: dict[str, dict[str, Any]],
|
||||||
|
sequence: int,
|
||||||
|
) -> dict[str, Any]:
|
||||||
|
source_ids = group["source_item_ids"]
|
||||||
|
source_items = [fact_by_id[item_id] for item_id in source_ids]
|
||||||
|
source_references: list[dict[str, Any]] = []
|
||||||
|
evidence: list[str] = []
|
||||||
|
for item in source_items:
|
||||||
|
source_references.extend(item.get("source_references", []))
|
||||||
|
item_evidence = item.get("evidence")
|
||||||
|
if isinstance(item_evidence, str):
|
||||||
|
evidence.append(item_evidence)
|
||||||
|
return {
|
||||||
|
"consolidated_id": f"fact_group_{sequence:04d}",
|
||||||
|
"category": "fact",
|
||||||
|
"canonical_text": group["canonical_text"],
|
||||||
|
"source_item_ids": source_ids,
|
||||||
|
"source_references": source_references,
|
||||||
|
"evidence": evidence,
|
||||||
|
"merge_reason": group["merge_reason"],
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def build_consolidated_output(
|
||||||
|
canonicalized: dict[str, Any],
|
||||||
|
groups: list[dict[str, Any]],
|
||||||
|
) -> dict[str, Any]:
|
||||||
|
facts = fact_items(canonicalized)
|
||||||
|
fact_by_id = {item["item_id"]: item for item in facts}
|
||||||
|
consolidated_facts = [
|
||||||
|
build_consolidated_fact_item(group, fact_by_id, index)
|
||||||
|
for index, group in enumerate(groups, start=1)
|
||||||
|
]
|
||||||
|
output_items = consolidated_facts + non_fact_items(canonicalized)
|
||||||
|
merged_groups = [item for item in consolidated_facts if len(item["source_item_ids"]) > 1]
|
||||||
|
singleton_groups = [
|
||||||
|
item for item in consolidated_facts if len(item["source_item_ids"]) == 1
|
||||||
|
]
|
||||||
|
|
||||||
|
output = dict(canonicalized)
|
||||||
|
output["schema_version"] = "semantic_consolidator_v0"
|
||||||
|
output["items"] = output_items
|
||||||
|
output["semantic_consolidation"] = {
|
||||||
|
"scope": "facts_only",
|
||||||
|
"merged_fact_group_count": len(merged_groups),
|
||||||
|
"source_facts_in_merged_groups": sum(
|
||||||
|
len(item["source_item_ids"]) for item in merged_groups
|
||||||
|
),
|
||||||
|
"singleton_fact_group_count": len(singleton_groups),
|
||||||
|
}
|
||||||
|
return output
|
||||||
|
|
||||||
|
|
||||||
|
def validate_consolidated_output(
|
||||||
|
canonicalized: dict[str, Any],
|
||||||
|
output: dict[str, Any],
|
||||||
|
) -> None:
|
||||||
|
original_facts = fact_items(canonicalized)
|
||||||
|
expected_fact_ids = {item["item_id"] for item in original_facts}
|
||||||
|
output_items = output.get("items")
|
||||||
|
if not isinstance(output_items, list):
|
||||||
|
raise ConsolidationValidationError("Output items must be a list.")
|
||||||
|
output_fact_groups = [item for item in output_items if item.get("category") == "fact"]
|
||||||
|
seen: list[str] = []
|
||||||
|
for item in output_fact_groups:
|
||||||
|
source_ids = item.get("source_item_ids")
|
||||||
|
if not isinstance(source_ids, list):
|
||||||
|
raise ConsolidationValidationError("Fact groups need source_item_ids.")
|
||||||
|
if len(source_ids) == 0:
|
||||||
|
raise ConsolidationValidationError("Fact groups must not be empty.")
|
||||||
|
if len(source_ids) > 1 and not item.get("merge_reason"):
|
||||||
|
raise ConsolidationValidationError("Merged fact groups need merge_reason.")
|
||||||
|
seen.extend(source_ids)
|
||||||
|
source_refs = item.get("source_references")
|
||||||
|
if not isinstance(source_refs, list) or not source_refs:
|
||||||
|
raise ConsolidationValidationError("Fact groups need source references.")
|
||||||
|
seen_counts = Counter(seen)
|
||||||
|
duplicated = sorted(item_id for item_id, count in seen_counts.items() if count > 1)
|
||||||
|
if duplicated:
|
||||||
|
raise ConsolidationValidationError(
|
||||||
|
f"Output duplicates source fact IDs: {duplicated}"
|
||||||
|
)
|
||||||
|
missing = sorted(expected_fact_ids - set(seen))
|
||||||
|
if missing:
|
||||||
|
raise ConsolidationValidationError(f"Output misses source fact IDs: {missing}")
|
||||||
|
unknown = sorted(set(seen) - expected_fact_ids)
|
||||||
|
if unknown:
|
||||||
|
raise ConsolidationValidationError(f"Output has unknown source IDs: {unknown}")
|
||||||
|
|
||||||
|
original_non_facts = non_fact_items(canonicalized)
|
||||||
|
output_non_facts = [item for item in output_items if item.get("category") != "fact"]
|
||||||
|
if output_non_facts != original_non_facts:
|
||||||
|
raise ConsolidationValidationError("Non-fact categories changed.")
|
||||||
|
|
||||||
|
|
||||||
|
def write_json(path: Path, data: Any) -> None:
|
||||||
|
path.parent.mkdir(parents=True, exist_ok=True)
|
||||||
|
path.write_text(
|
||||||
|
json.dumps(data, ensure_ascii=False, indent=2) + "\n",
|
||||||
|
encoding="utf-8",
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def write_report(
|
||||||
|
path: Path,
|
||||||
|
model: str,
|
||||||
|
runtime: float,
|
||||||
|
prompt_chars: int,
|
||||||
|
prompt_token_estimate: int,
|
||||||
|
fact_count: int,
|
||||||
|
groups: list[dict[str, Any]],
|
||||||
|
output_path: Path,
|
||||||
|
) -> None:
|
||||||
|
merged = [group for group in groups if len(group["source_item_ids"]) > 1]
|
||||||
|
singletons = [group for group in groups if len(group["source_item_ids"]) == 1]
|
||||||
|
lines = [
|
||||||
|
"# Semantic Consolidator V0 Report",
|
||||||
|
"",
|
||||||
|
"- Scope: facts only",
|
||||||
|
f"- Model: `{model}`",
|
||||||
|
"- LLM call count: 1",
|
||||||
|
f"- Runtime: {runtime:.3f} seconds",
|
||||||
|
f"- Fact item count: {fact_count}",
|
||||||
|
f"- Prompt characters: {prompt_chars}",
|
||||||
|
f"- Estimated prompt tokens: {prompt_token_estimate}",
|
||||||
|
f"- Merged fact groups: {len(merged)}",
|
||||||
|
f"- Source facts involved in merges: {sum(len(group['source_item_ids']) for group in merged)}",
|
||||||
|
f"- Singleton fact groups: {len(singletons)}",
|
||||||
|
f"- Output path: `{output_path}`",
|
||||||
|
"",
|
||||||
|
"## Actual Merges",
|
||||||
|
"",
|
||||||
|
]
|
||||||
|
if not merged:
|
||||||
|
lines.append("- None.")
|
||||||
|
else:
|
||||||
|
for group in merged:
|
||||||
|
lines.extend(
|
||||||
|
[
|
||||||
|
f"### {group['canonical_text']}",
|
||||||
|
"",
|
||||||
|
f"- Source fact IDs: {', '.join(group['source_item_ids'])}",
|
||||||
|
f"- Merge reason: {group['merge_reason']}",
|
||||||
|
"",
|
||||||
|
]
|
||||||
|
)
|
||||||
|
path.write_text("\n".join(lines).rstrip() + "\n", encoding="utf-8")
|
||||||
|
|
||||||
|
|
||||||
|
def main() -> int:
|
||||||
|
args = parse_args()
|
||||||
|
args.output_dir.mkdir(parents=True, exist_ok=True)
|
||||||
|
raw_response_path = args.output_dir / "raw_model_response.txt"
|
||||||
|
output_path = args.output_dir / "consolidated_extractions.json"
|
||||||
|
report_path = args.output_dir / "report.md"
|
||||||
|
|
||||||
|
try:
|
||||||
|
canonicalized = load_json_object(args.canonicalized_input)
|
||||||
|
facts = fact_items(canonicalized)
|
||||||
|
prompt = build_consolidation_prompt(facts)
|
||||||
|
prompt_chars = len(prompt)
|
||||||
|
prompt_token_estimate = (prompt_chars + 3) // 4
|
||||||
|
print(f"Fact item count: {len(facts)}")
|
||||||
|
print(f"Estimated prompt size chars: {prompt_chars}")
|
||||||
|
print(f"Estimated prompt tokens: {prompt_token_estimate}")
|
||||||
|
print("Expected LLM call count: 1")
|
||||||
|
print("Expected runtime: 5-10 minutes on current local benchmark basis")
|
||||||
|
|
||||||
|
raw_text, _response_data, runtime = call_ollama(
|
||||||
|
endpoint=args.endpoint,
|
||||||
|
model=args.model,
|
||||||
|
prompt=prompt,
|
||||||
|
timeout=args.timeout,
|
||||||
|
num_ctx=args.num_ctx,
|
||||||
|
num_predict=args.num_predict,
|
||||||
|
think=args.think,
|
||||||
|
progress_interval=args.progress_interval,
|
||||||
|
)
|
||||||
|
raw_response_path.write_text(raw_text + "\n", encoding="utf-8")
|
||||||
|
model_output = parse_model_json(raw_text)
|
||||||
|
expected_fact_ids = {item["item_id"] for item in facts}
|
||||||
|
groups = validate_model_groups(model_output, expected_fact_ids)
|
||||||
|
validate_group_shapes(groups)
|
||||||
|
output = build_consolidated_output(canonicalized, groups)
|
||||||
|
validate_consolidated_output(canonicalized, output)
|
||||||
|
write_json(output_path, output)
|
||||||
|
write_report(
|
||||||
|
path=report_path,
|
||||||
|
model=args.model,
|
||||||
|
runtime=runtime,
|
||||||
|
prompt_chars=prompt_chars,
|
||||||
|
prompt_token_estimate=prompt_token_estimate,
|
||||||
|
fact_count=len(facts),
|
||||||
|
groups=groups,
|
||||||
|
output_path=output_path,
|
||||||
|
)
|
||||||
|
except requests.ConnectionError as exc:
|
||||||
|
print(f"Error: Ollama is not reachable at {args.endpoint}: {exc}", file=sys.stderr)
|
||||||
|
return 1
|
||||||
|
except requests.Timeout as exc:
|
||||||
|
print(
|
||||||
|
f"Error: Ollama request timed out after {args.timeout} seconds: {exc}",
|
||||||
|
file=sys.stderr,
|
||||||
|
)
|
||||||
|
return 1
|
||||||
|
except requests.HTTPError as exc:
|
||||||
|
print(f"Error: Ollama returned an HTTP error: {exc}", file=sys.stderr)
|
||||||
|
return 1
|
||||||
|
except (OSError, UnicodeError, ValueError) as exc:
|
||||||
|
print(f"Error: {exc}", file=sys.stderr)
|
||||||
|
return 1
|
||||||
|
|
||||||
|
merged = [group for group in groups if len(group["source_item_ids"]) > 1]
|
||||||
|
print(f"Runtime seconds: {runtime:.3f}")
|
||||||
|
print(f"Merged fact groups: {len(merged)}")
|
||||||
|
print(
|
||||||
|
"Source facts involved in merges: "
|
||||||
|
f"{sum(len(group['source_item_ids']) for group in merged)}"
|
||||||
|
)
|
||||||
|
print(f"Singleton fact groups: {len(groups) - len(merged)}")
|
||||||
|
print("Validation result: passed")
|
||||||
|
print(f"Output: {output_path}")
|
||||||
|
print(f"Report: {report_path}")
|
||||||
|
print(f"Raw model response: {raw_response_path}")
|
||||||
|
return 0
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
raise SystemExit(main())
|
||||||
@@ -0,0 +1,169 @@
|
|||||||
|
import json
|
||||||
|
import tempfile
|
||||||
|
import unittest
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
from src.meeting_lab.consolidation.canonicalize import (
|
||||||
|
canonicalize_extractions,
|
||||||
|
find_extraction_files,
|
||||||
|
load_json_object,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
EMPTY_EXTRACTION = {
|
||||||
|
"facts": [],
|
||||||
|
"decisions": [],
|
||||||
|
"todos": [],
|
||||||
|
"questions": [],
|
||||||
|
"positions": [],
|
||||||
|
"technical": [],
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def write_extraction(directory: Path, name: str, data: dict) -> Path:
|
||||||
|
path = directory / name
|
||||||
|
path.write_text(json.dumps(data), encoding="utf-8")
|
||||||
|
return path
|
||||||
|
|
||||||
|
|
||||||
|
class CanonicalizeTests(unittest.TestCase):
|
||||||
|
def test_stable_file_ordering(self) -> None:
|
||||||
|
with tempfile.TemporaryDirectory() as directory:
|
||||||
|
root = Path(directory)
|
||||||
|
write_extraction(root, "chunk_10_extraction.json", EMPTY_EXTRACTION)
|
||||||
|
write_extraction(root, "chunk_02_extraction.json", EMPTY_EXTRACTION)
|
||||||
|
write_extraction(root, "chunk_01_extraction.json", EMPTY_EXTRACTION)
|
||||||
|
|
||||||
|
self.assertEqual(
|
||||||
|
[path.name for path in find_extraction_files(root)],
|
||||||
|
[
|
||||||
|
"chunk_01_extraction.json",
|
||||||
|
"chunk_02_extraction.json",
|
||||||
|
"chunk_10_extraction.json",
|
||||||
|
],
|
||||||
|
)
|
||||||
|
|
||||||
|
def test_required_category_validation(self) -> None:
|
||||||
|
with tempfile.TemporaryDirectory() as directory:
|
||||||
|
root = Path(directory)
|
||||||
|
data = dict(EMPTY_EXTRACTION)
|
||||||
|
data.pop("technical")
|
||||||
|
write_extraction(root, "chunk_01_extraction.json", data)
|
||||||
|
|
||||||
|
with self.assertRaisesRegex(ValueError, "technical"):
|
||||||
|
canonicalize_extractions(root)
|
||||||
|
|
||||||
|
def test_stable_item_ids_and_string_parsing(self) -> None:
|
||||||
|
with tempfile.TemporaryDirectory() as directory:
|
||||||
|
root = Path(directory)
|
||||||
|
data = dict(EMPTY_EXTRACTION)
|
||||||
|
data["facts"] = ["Ada | Revenue increased | clear | Revenue increased"]
|
||||||
|
data["decisions"] = ["Ship the patch | Agreed to ship"]
|
||||||
|
data["todos"] = ["Update docs | Mira | Friday | I will update docs"]
|
||||||
|
data["questions"] = ["Which plan? | Which plan should we use?"]
|
||||||
|
data["positions"] = ["Kai | Prefer option B | I prefer option B"]
|
||||||
|
data["technical"] = ["API | Uses v2 mapping | clear | API uses v2"]
|
||||||
|
write_extraction(root, "chunk_01_extraction.json", data)
|
||||||
|
|
||||||
|
output = canonicalize_extractions(root)
|
||||||
|
items = output["items"]
|
||||||
|
|
||||||
|
self.assertEqual(
|
||||||
|
[item["item_id"] for item in items],
|
||||||
|
[
|
||||||
|
"fact_0001",
|
||||||
|
"decision_0001",
|
||||||
|
"action_item_0001",
|
||||||
|
"open_question_0001",
|
||||||
|
"position_0001",
|
||||||
|
"technical_detail_0001",
|
||||||
|
],
|
||||||
|
)
|
||||||
|
fact = items[0]
|
||||||
|
self.assertEqual(fact["speaker"], "Ada")
|
||||||
|
self.assertEqual(fact["text"], "Revenue increased")
|
||||||
|
self.assertEqual(fact["status"], "clear")
|
||||||
|
todo = items[2]
|
||||||
|
self.assertEqual(todo["responsible"], "Mira")
|
||||||
|
self.assertEqual(todo["deadline"], "Friday")
|
||||||
|
self.assertEqual(todo["evidence"], "I will update docs")
|
||||||
|
|
||||||
|
def test_source_reference_preservation(self) -> None:
|
||||||
|
with tempfile.TemporaryDirectory() as directory:
|
||||||
|
root = Path(directory)
|
||||||
|
data = dict(EMPTY_EXTRACTION)
|
||||||
|
data["decisions"] = ["Decision one | Evidence one"]
|
||||||
|
write_extraction(root, "chunk_01_extraction.json", data)
|
||||||
|
|
||||||
|
item = canonicalize_extractions(root)["items"][0]
|
||||||
|
|
||||||
|
self.assertEqual(item["source_file"], "chunk_01_extraction.json")
|
||||||
|
self.assertEqual(item["source_index"], 0)
|
||||||
|
self.assertEqual(item["original_value"], "Decision one | Evidence one")
|
||||||
|
self.assertEqual(
|
||||||
|
item["source_references"],
|
||||||
|
[
|
||||||
|
{
|
||||||
|
"source_file": "chunk_01_extraction.json",
|
||||||
|
"source_index": 0,
|
||||||
|
"evidence": "Evidence one",
|
||||||
|
"original_value": "Decision one | Evidence one",
|
||||||
|
}
|
||||||
|
],
|
||||||
|
)
|
||||||
|
|
||||||
|
def test_exact_duplicate_handling(self) -> None:
|
||||||
|
with tempfile.TemporaryDirectory() as directory:
|
||||||
|
root = Path(directory)
|
||||||
|
first = dict(EMPTY_EXTRACTION)
|
||||||
|
second = dict(EMPTY_EXTRACTION)
|
||||||
|
first["decisions"] = ["Same decision | Same evidence"]
|
||||||
|
second["decisions"] = ["Same decision | Same evidence"]
|
||||||
|
write_extraction(root, "chunk_01_extraction.json", first)
|
||||||
|
write_extraction(root, "chunk_02_extraction.json", second)
|
||||||
|
|
||||||
|
output = canonicalize_extractions(root)
|
||||||
|
|
||||||
|
self.assertEqual(len(output["items"]), 1)
|
||||||
|
self.assertEqual(output["stats"]["exact_duplicates_merged"], 1)
|
||||||
|
self.assertEqual(output["items"][0]["duplicate_count"], 2)
|
||||||
|
self.assertEqual(len(output["items"][0]["source_references"]), 2)
|
||||||
|
|
||||||
|
def test_no_merging_of_merely_similar_statements(self) -> None:
|
||||||
|
with tempfile.TemporaryDirectory() as directory:
|
||||||
|
root = Path(directory)
|
||||||
|
data = dict(EMPTY_EXTRACTION)
|
||||||
|
data["decisions"] = [
|
||||||
|
"Ship the patch Friday | Agreed to ship Friday",
|
||||||
|
"Ship the patch next week | Agreed to ship next week",
|
||||||
|
]
|
||||||
|
write_extraction(root, "chunk_01_extraction.json", data)
|
||||||
|
|
||||||
|
output = canonicalize_extractions(root)
|
||||||
|
|
||||||
|
self.assertEqual(len(output["items"]), 2)
|
||||||
|
self.assertEqual(output["stats"]["exact_duplicates_merged"], 0)
|
||||||
|
|
||||||
|
def test_invalid_json_handling(self) -> None:
|
||||||
|
with tempfile.TemporaryDirectory() as directory:
|
||||||
|
root = Path(directory)
|
||||||
|
(root / "chunk_01_extraction.json").write_text("{invalid", encoding="utf-8")
|
||||||
|
|
||||||
|
with self.assertRaisesRegex(ValueError, "Invalid JSON"):
|
||||||
|
load_json_object(root / "chunk_01_extraction.json")
|
||||||
|
|
||||||
|
def test_empty_categories(self) -> None:
|
||||||
|
with tempfile.TemporaryDirectory() as directory:
|
||||||
|
root = Path(directory)
|
||||||
|
write_extraction(root, "chunk_01_extraction.json", EMPTY_EXTRACTION)
|
||||||
|
|
||||||
|
output = canonicalize_extractions(root)
|
||||||
|
|
||||||
|
self.assertEqual(output["items"], [])
|
||||||
|
self.assertEqual(output["stats"]["input_item_count"], 0)
|
||||||
|
self.assertEqual(output["stats"]["output_item_count"], 0)
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
unittest.main()
|
||||||
|
|
||||||
@@ -0,0 +1,215 @@
|
|||||||
|
import unittest
|
||||||
|
|
||||||
|
from src.meeting_lab.consolidation.consolidate_facts import (
|
||||||
|
ConsolidationValidationError,
|
||||||
|
DEFAULT_NUM_PREDICT,
|
||||||
|
build_consolidated_output,
|
||||||
|
build_ollama_payload,
|
||||||
|
parse_model_json,
|
||||||
|
validate_consolidated_output,
|
||||||
|
validate_model_groups,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def canonicalized_fixture():
|
||||||
|
fact_one = {
|
||||||
|
"item_id": "fact_0001",
|
||||||
|
"category": "fact",
|
||||||
|
"text": "The lead maintains the list.",
|
||||||
|
"evidence": "lead maintains the list",
|
||||||
|
"source_file": "chunk_01_extraction.json",
|
||||||
|
"source_index": 0,
|
||||||
|
"original_value": "The lead maintains the list. | evidence",
|
||||||
|
"source_references": [
|
||||||
|
{
|
||||||
|
"source_file": "chunk_01_extraction.json",
|
||||||
|
"source_index": 0,
|
||||||
|
"evidence": "lead maintains the list",
|
||||||
|
"original_value": "The lead maintains the list. | evidence",
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"duplicate_count": 1,
|
||||||
|
}
|
||||||
|
fact_two = {
|
||||||
|
"item_id": "fact_0002",
|
||||||
|
"category": "fact",
|
||||||
|
"text": "The head maintains the project list.",
|
||||||
|
"evidence": "head maintains the project list",
|
||||||
|
"source_file": "chunk_02_extraction.json",
|
||||||
|
"source_index": 0,
|
||||||
|
"original_value": "The head maintains the project list. | evidence",
|
||||||
|
"source_references": [
|
||||||
|
{
|
||||||
|
"source_file": "chunk_02_extraction.json",
|
||||||
|
"source_index": 0,
|
||||||
|
"evidence": "head maintains the project list",
|
||||||
|
"original_value": "The head maintains the project list. | evidence",
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"duplicate_count": 1,
|
||||||
|
}
|
||||||
|
decision = {
|
||||||
|
"item_id": "decision_0001",
|
||||||
|
"category": "decision",
|
||||||
|
"text": "Ship it.",
|
||||||
|
"evidence": "Agreed.",
|
||||||
|
"source_file": "chunk_01_extraction.json",
|
||||||
|
"source_index": 0,
|
||||||
|
"original_value": "Ship it. | Agreed.",
|
||||||
|
"source_references": [],
|
||||||
|
"duplicate_count": 1,
|
||||||
|
}
|
||||||
|
return {
|
||||||
|
"schema_version": "1",
|
||||||
|
"source_files": ["chunk_01_extraction.json", "chunk_02_extraction.json"],
|
||||||
|
"items": [fact_one, fact_two, decision],
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
class ConsolidateFactsTests(unittest.TestCase):
|
||||||
|
def test_payload_construction_disables_streaming_and_thinking_by_default(self):
|
||||||
|
payload = build_ollama_payload(
|
||||||
|
model="qwen3.5:9B",
|
||||||
|
prompt="prompt",
|
||||||
|
num_ctx=32768,
|
||||||
|
num_predict=DEFAULT_NUM_PREDICT,
|
||||||
|
think=False,
|
||||||
|
)
|
||||||
|
|
||||||
|
self.assertEqual(payload["model"], "qwen3.5:9B")
|
||||||
|
self.assertEqual(payload["prompt"], "prompt")
|
||||||
|
self.assertIs(payload["stream"], False)
|
||||||
|
self.assertIs(payload["think"], False)
|
||||||
|
self.assertEqual(payload["format"], "json")
|
||||||
|
self.assertEqual(payload["options"]["temperature"], 0.0)
|
||||||
|
self.assertEqual(payload["options"]["num_ctx"], 32768)
|
||||||
|
self.assertEqual(payload["options"]["num_predict"], DEFAULT_NUM_PREDICT)
|
||||||
|
|
||||||
|
def test_payload_construction_can_enable_thinking_explicitly(self):
|
||||||
|
payload = build_ollama_payload(
|
||||||
|
model="qwen3.5:9B",
|
||||||
|
prompt="prompt",
|
||||||
|
num_ctx=16384,
|
||||||
|
num_predict=1024,
|
||||||
|
think=True,
|
||||||
|
)
|
||||||
|
|
||||||
|
self.assertIs(payload["think"], True)
|
||||||
|
self.assertEqual(payload["options"]["num_ctx"], 16384)
|
||||||
|
self.assertEqual(payload["options"]["num_predict"], 1024)
|
||||||
|
|
||||||
|
def test_grouping_validation_accepts_complete_singletons(self):
|
||||||
|
groups = validate_model_groups(
|
||||||
|
{
|
||||||
|
"groups": [
|
||||||
|
{
|
||||||
|
"canonical_text": "The lead maintains the list.",
|
||||||
|
"source_item_ids": ["fact_0001"],
|
||||||
|
"merge_reason": "Singleton.",
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"canonical_text": "The head maintains the project list.",
|
||||||
|
"source_item_ids": ["fact_0002"],
|
||||||
|
"merge_reason": "Singleton.",
|
||||||
|
},
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{"fact_0001", "fact_0002"},
|
||||||
|
)
|
||||||
|
|
||||||
|
self.assertEqual(len(groups), 2)
|
||||||
|
|
||||||
|
def test_no_missing_source_ids(self):
|
||||||
|
with self.assertRaisesRegex(ConsolidationValidationError, "Missing"):
|
||||||
|
validate_model_groups(
|
||||||
|
{
|
||||||
|
"groups": [
|
||||||
|
{
|
||||||
|
"canonical_text": "Only one.",
|
||||||
|
"source_item_ids": ["fact_0001"],
|
||||||
|
"merge_reason": "Singleton.",
|
||||||
|
}
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{"fact_0001", "fact_0002"},
|
||||||
|
)
|
||||||
|
|
||||||
|
def test_no_duplicate_source_ids(self):
|
||||||
|
with self.assertRaisesRegex(ConsolidationValidationError, "multiple"):
|
||||||
|
validate_model_groups(
|
||||||
|
{
|
||||||
|
"groups": [
|
||||||
|
{
|
||||||
|
"canonical_text": "One.",
|
||||||
|
"source_item_ids": ["fact_0001"],
|
||||||
|
"merge_reason": "Singleton.",
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"canonical_text": "Again.",
|
||||||
|
"source_item_ids": ["fact_0001"],
|
||||||
|
"merge_reason": "Singleton.",
|
||||||
|
},
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{"fact_0001"},
|
||||||
|
)
|
||||||
|
|
||||||
|
def test_merged_group_validation(self):
|
||||||
|
groups = validate_model_groups(
|
||||||
|
{
|
||||||
|
"groups": [
|
||||||
|
{
|
||||||
|
"canonical_text": "The lead maintains the project list.",
|
||||||
|
"source_item_ids": ["fact_0001", "fact_0002"],
|
||||||
|
"merge_reason": "Same proposition.",
|
||||||
|
}
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{"fact_0001", "fact_0002"},
|
||||||
|
)
|
||||||
|
|
||||||
|
output = build_consolidated_output(canonicalized_fixture(), groups)
|
||||||
|
validate_consolidated_output(canonicalized_fixture(), output)
|
||||||
|
self.assertEqual(output["items"][0]["source_item_ids"], ["fact_0001", "fact_0002"])
|
||||||
|
|
||||||
|
def test_preservation_of_non_fact_categories(self):
|
||||||
|
fixture = canonicalized_fixture()
|
||||||
|
groups = validate_model_groups(
|
||||||
|
{
|
||||||
|
"groups": [
|
||||||
|
{
|
||||||
|
"canonical_text": "The lead maintains the project list.",
|
||||||
|
"source_item_ids": ["fact_0001", "fact_0002"],
|
||||||
|
"merge_reason": "Same proposition.",
|
||||||
|
}
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{"fact_0001", "fact_0002"},
|
||||||
|
)
|
||||||
|
|
||||||
|
output = build_consolidated_output(fixture, groups)
|
||||||
|
|
||||||
|
self.assertEqual(output["items"][1:], fixture["items"][2:])
|
||||||
|
|
||||||
|
def test_invalid_model_json(self):
|
||||||
|
with self.assertRaisesRegex(ConsolidationValidationError, "Invalid model JSON"):
|
||||||
|
parse_model_json("{invalid")
|
||||||
|
|
||||||
|
def test_unknown_source_item_ids(self):
|
||||||
|
with self.assertRaisesRegex(ConsolidationValidationError, "unknown"):
|
||||||
|
validate_model_groups(
|
||||||
|
{
|
||||||
|
"groups": [
|
||||||
|
{
|
||||||
|
"canonical_text": "Unknown.",
|
||||||
|
"source_item_ids": ["fact_9999"],
|
||||||
|
"merge_reason": "Bad ID.",
|
||||||
|
}
|
||||||
|
]
|
||||||
|
},
|
||||||
|
{"fact_0001"},
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
unittest.main()
|
||||||
Reference in New Issue
Block a user