Document canonicalization and consolidation milestone

- preserve Working Protocol Synthesizer V0 as comparison baseline
- introduce deterministic canonicalization stage
- define semantic consolidator responsibilities
- clarify Canonical Meeting Knowledge generation
- document source-language output policy
- align roadmap, architecture and experiment log
This commit is contained in:
2026-07-31 09:32:12 +02:00
parent 09d125e54a
commit 23bbc744f7
12 changed files with 512 additions and 83 deletions
+37 -8
View File
@@ -30,7 +30,8 @@ Experimental/prototype:
Planned:
- Consolidation of extraction results.
- Deterministic Canonicalizer for extraction results.
- Semantic Consolidator for evidence-preserving semantic merging.
- Canonical Meeting Knowledge implementation as the semantic source of truth.
- Final Working Protocol / Arbeitsprotokoll, Distribution Protocol /
Verteilerprotokoll and Knowledge Objects / Wissensdatenbankeintrag renderers.
@@ -40,7 +41,7 @@ Planned:
```text
src/meeting_lab/
chunking/ technical transcript chunking
consolidation/ planned merge/consolidation area
consolidation/ planned canonicalization/consolidation area
extraction/ current local LLM extraction flow
io/ lightweight file and JSON helpers
llm/ Ollama and prompt support
@@ -112,10 +113,33 @@ Current Prompt Version 2 decision baseline:
## Canonical Knowledge Architecture
The next documented pipeline milestone is:
```text
Chunk Extractions
-> Deterministic Canonicalizer
-> Semantic Consolidator
-> Canonical Meeting Knowledge
-> Output View Renderers
```
The Deterministic Canonicalizer is planned Python code with no LLM. It should
validate and normalize extraction objects, assign stable source references and
IDs, normalize category names and basic field structure, perform only safe
deterministic cleanup, optionally group exact duplicates, and preserve all
source evidence. It must not perform uncertain semantic merging.
The Semantic Consolidator is planned local-LLM work. It should merge
semantically equivalent statements, group content by topic, preserve evidence
from all contributing chunks, mark contradictions and uncertainty, separate
durable information from transient discussion, and produce Canonical Meeting
Knowledge. It does not directly write a protocol.
Canonical Meeting Knowledge is the planned semantic intermediate model and
future single source of truth. It should preserve topics, facts, decisions,
action items, open questions, positions, technical details, rationale,
uncertainty, contradictions and source evidence.
future single source of truth. It should be structured, preferably JSON, and
preserve topics, facts, decisions, action items, open questions, positions,
technical details, rationale, uncertainty, contradictions and source evidence.
It is not itself a prose protocol.
Output views are planned as independent renderings from that canonical model:
@@ -129,6 +153,10 @@ Output views are planned as independent renderings from that canonical model:
The current `meeting_protocol.md` builder is an interim technical validation
tool, not the final output-view architecture.
Rendered protocol output should normally use the dominant language of the
source transcript or consolidated meeting knowledge unless an explicit output
language is requested.
## Current Limitations
- Discussion Blocks are documented as a stable semantic unit but are not yet a
@@ -136,7 +164,8 @@ tool, not the final output-view architecture.
- Topic segmentation exists as prototype tooling, not a stable pipeline stage.
- Extraction is still a combined current flow, even though separate extractors
are the intended architecture.
- Consolidation is not implemented.
- Deterministic Canonicalizer is documented but not implemented.
- Semantic Consolidator is documented but not implemented.
- Canonical Meeting Knowledge is documented but not implemented.
- Final output views are documented but not implemented.
- Most prompt files are placeholders except the common and decision prompts.
@@ -147,5 +176,5 @@ tool, not the final output-view architecture.
Stabilize repeatable local extraction evaluation before broadening the pipeline:
expand Gold Standard coverage by category, keep one-chunk extraction as the
baseline, and use small prompt experiments with immediate non-regression checks.
After extraction behavior is stable enough, implement consolidation with
evidence retention as the next major pipeline stage.
After extraction behavior is stable enough, implement the Deterministic
Canonicalizer first, then the Semantic Consolidator with evidence retention.