Document canonicalization and consolidation milestone

- preserve Working Protocol Synthesizer V0 as comparison baseline
- introduce deterministic canonicalization stage
- define semantic consolidator responsibilities
- clarify Canonical Meeting Knowledge generation
- document source-language output policy
- align roadmap, architecture and experiment log
This commit is contained in:
2026-07-31 09:32:12 +02:00
parent 09d125e54a
commit 23bbc744f7
12 changed files with 512 additions and 83 deletions
+35 -21
View File
@@ -98,7 +98,9 @@ Topic Segmentation
↓
Specialized Extraction
↓
Consolidation
Deterministic Canonicalization
↓
Semantic Consolidation
↓
Canonical Meeting Knowledge
↓
@@ -174,14 +176,32 @@ Each extractor has exactly one task and one prompt.
## consolidation/
Merges information extracted from multiple discussion segments.
Planned area for canonicalization and consolidation.
Typical responsibilities:
The next milestone splits this into two stages.
- merge duplicates
- combine partial information
- distinguish positions from decisions
- detect contradictions
Deterministic Canonicalizer:
- implemented in Python
- uses no LLM
- validates and normalizes extraction objects
- assigns stable source references and IDs
- normalizes category names and basic field structure
- performs only safe deterministic cleanup
- may group exact duplicates
- preserves all source evidence
- must not perform uncertain semantic merging
Semantic Consolidator:
- uses the local LLM
- merges semantically equivalent statements
- groups content by topic
- preserves evidence from all contributing chunks
- marks contradictions and uncertainty
- separates durable information from transient discussion
- produces Canonical Meeting Knowledge
- does not directly write a protocol
---
@@ -207,6 +227,10 @@ The planned output products are:
These are parallel renderings of the same canonical semantic model, not
documents derived from one another.
Rendered protocol language should normally match the dominant language of the
source transcript or consolidated meeting knowledge unless an explicit output
language is requested.
---
# Repository Layout
@@ -257,20 +281,10 @@ Views documented in `output-views.md`.
# Next Milestone
The next development step is the implementation of **topic segmentation**.
Its only responsibility is to identify the thematic structure of a discussion.
It should answer questions such as:
- Where does a topic begin?
- Where does it end?
- When does another topic start?
- When is an earlier topic resumed?
No facts, decisions or todos should be extracted at this stage.
Only after reliable topic segmentation has been achieved will the specialized extraction modules be implemented.
The next architecture milestone is the implementation of a deterministic
canonicalization stage followed by a semantic consolidation stage. These stages
convert raw chunk extraction JSON into evidence-preserving Canonical Meeting
Knowledge before any Output View renderer writes a protocol.
---
+64
View File
@@ -242,11 +242,74 @@ This object feeds the Canonical Meeting Knowledge representation.
---
# Canonical Extraction Object
Planned deterministic intermediate object created from raw chunk extraction
JSON.
Example:
```json
{
"id": "fact.chunk_03.0001",
"category": "fact",
"text": "...",
"source_references": [
{
"chunk_id": "chunk_03",
"source_file": "chunk_03_extraction.json",
"evidence": "..."
}
]
}
```
The Deterministic Canonicalizer should create this kind of object without an
LLM. It validates and normalizes raw extraction objects, assigns stable IDs and
source references, normalizes category names and basic field structure,
performs only safe deterministic cleanup, may group exact duplicates and must
preserve all source evidence.
It must not perform uncertain semantic merging.
---
# Consolidated Topic
Planned semantic object produced by the Semantic Consolidator.
Example:
```json
{
"topic_id": "topic_001",
"title": "...",
"background": [],
"decisions": [],
"action_items": [],
"open_questions": [],
"durable_information": [],
"uncertainty": [],
"source_references": []
}
```
The Semantic Consolidator may use the local LLM to merge semantically
equivalent statements, group content by topic, preserve evidence from all
contributing chunks, mark contradictions and uncertainty and separate durable
information from transient discussion.
It produces Canonical Meeting Knowledge. It does not directly write a protocol.
---
# Canonical Meeting Knowledge
The canonical semantic representation of one meeting.
This representation is the single source of truth for all downstream outputs.
It is a structured representation, preferably JSON, and is not itself a prose
protocol.
```json
{
@@ -260,6 +323,7 @@ This representation is the single source of truth for all downstream outputs.
"questions": [],
"positions": [],
"technical_details": [],
"durable_information": [],
"rationale": [],
"uncertainty": [],
"source_references": []
+73 -5
View File
@@ -945,14 +945,18 @@ Result:
The accepted design is a planned Canonical Meeting Knowledge layer as the
semantic source of truth, with Working Protocol, Distribution Protocol and
Knowledge Objects as parallel output views. Consolidation must merge duplicates,
preserve evidence, reconcile category shifts and mark contradictions or
uncertainty.
Knowledge Objects as parallel output views. The next consolidation architecture
is split into a Deterministic Canonicalizer and a Semantic Consolidator. The
canonicalizer prepares validated evidence-bearing objects without uncertain
semantic merging. The consolidator then merges semantically equivalent
statements, preserves evidence, reconciles category shifts where supported and
marks contradictions or uncertainty.
Decision:
Consolidation is the next major engineering step after stable local extraction.
Canonical Meeting Knowledge and final output views are planned, not implemented.
Deterministic canonicalization is the next implementation step after stable
local extraction, followed by semantic consolidation. Canonical Meeting
Knowledge and final output views are planned, not implemented.
Lessons learned:
@@ -969,3 +973,67 @@ Evidence:
- `PROJECT_KNOWLEDGE.md`
- `ROADMAP.md`
- Commit `5c03ed7`
## EXP-0020 - Working Protocol Synthesizer V0
Status: Accepted
Date or period: 2026-07-31
Hypothesis:
The current local synthesis model may be able to generate a useful detailed
Working Protocol directly from the existing independent chunk extraction JSON
files, before canonicalization or semantic consolidation exists.
Setup:
One synthesis prompt was constructed from exactly nine chunk extraction JSON
files. The model was instructed to use only those extraction files, merge
duplicates, group related information into topics, preserve useful discussion
context and write a neutral technical Working Protocol.
Inputs:
- `chunk_01_extraction.json` through `chunk_09_extraction.json`.
- No original transcript, normalized chunks or Whisper output were used as
synthesis input.
Model / configuration:
- Model: `qwen3.5:9b`
- Prompt characters: 24,979
- Actual prompt eval tokens: 5,929
- Output tokens: 1,486
- Runtime: 294.204 seconds
Result:
The generated Working Protocol was readable, well structured and
topic-oriented. It was still based directly on raw chunk extractions, without a
separate deterministic canonicalization stage or semantic consolidation stage.
The output language was English even though the source meeting material was
German.
Decision:
Preserve this output as the Working Protocol Synthesizer V0 benchmark baseline
for later canonicalizer, consolidator and renderer comparisons. This selected
generated artifact is intentionally versioned even though generated runtime
artifacts are normally ignored.
Lessons learned:
Direct synthesis from chunk extractions can create a useful recall-oriented
draft, but it does not replace Canonical Meeting Knowledge. The language
mismatch also establishes a default renderer rule: protocol output should
normally match the dominant source language unless an explicit output language
is requested.
Evidence:
- `samples/benchmarks/working_protocol_synthesizer_v0/README.md`
- `samples/benchmarks/working_protocol_synthesizer_v0/working_protocol.md`
- `PROJECT_KNOWLEDGE.md`
- `docs/output-views.md`
- See EXP-0017 and EXP-0019.
+20
View File
@@ -13,6 +13,10 @@ Meeting transcript
->
chunk extraction
->
Deterministic Canonicalizer
->
Semantic Consolidator
->
Canonical Meeting Knowledge
├── Working Protocol
├── Distribution Protocol
@@ -45,6 +49,18 @@ The detailed schema is future implementation work. The current implementation
still uses simple extraction JSON files and a basic Markdown protocol builder for
technical validation.
The next planned architecture stage before this representation is explicit:
- The Deterministic Canonicalizer validates and normalizes extraction objects,
assigns stable source references and IDs, performs only safe deterministic
cleanup and preserves all source evidence. It uses no LLM and must not make
uncertain semantic merges.
- The Semantic Consolidator uses the local LLM to merge semantically equivalent
statements, group content by topic, preserve evidence from all contributing
chunks, mark contradictions and uncertainty, separate durable information from
transient discussion and produce Canonical Meeting Knowledge. It does not
directly write a protocol.
## Renderers
Each Output View is produced by a renderer.
@@ -57,6 +73,10 @@ Depending on the implementation, rendering may be:
The architecture does not assume that every renderer must always use an LLM.
Rendered protocol language should normally match the dominant language of the
source transcript or consolidated meeting knowledge unless an explicit output
language is requested.
## Working Protocol
Suggested filename: `working_protocol.md`
+80 -19
View File
@@ -33,7 +33,10 @@ Topic Segmentation
Specialized Extraction
│
▼
Consolidation
Deterministic Canonicalization
│
▼
Semantic Consolidation
│
▼
Canonical Meeting Knowledge
@@ -267,35 +270,41 @@ Prototype exists as a combined extractor.
---
# Stage 6 – Consolidation
# Stage 6 – Deterministic Canonicalization
## Purpose
Merge analysis results originating from different discussion segments.
Normalize raw chunk extraction JSON into stable canonical extraction objects
without changing uncertain semantics.
## Input
Extraction results.
Chunk extraction JSON files.
## Output
Unified topic representation.
Validated canonical extraction objects with stable source references and IDs.
## Responsibilities
- Merge duplicates
- Merge complementary information
- Preserve contradictions
- Separate positions from decisions
- Combine related todos
- Validate extraction objects
- Normalize category names
- Normalize basic field structure
- Assign stable source references and IDs
- Preserve all source evidence
- Perform only safe deterministic cleanup
- Group exact duplicates where unambiguous
## Must Not
- Use an LLM
- Perform uncertain semantic merging
- Infer missing information
- Drop source evidence
## Processing Type
Hybrid
Deterministic wherever possible.
LLM support only if necessary.
Deterministic Python
## Current Status
@@ -303,13 +312,55 @@ Planned
---
# Stage 7 – Canonical Meeting Knowledge
# Stage 7 – Semantic Consolidation
## Purpose
Merge canonicalized extraction objects into an evidence-preserving semantic
meeting representation.
## Input
Canonicalized extraction objects.
## Output
Canonical Meeting Knowledge.
## Responsibilities
- Merge semantically equivalent statements
- Group content by topic
- Preserve evidence from all contributing chunks
- Mark contradictions and uncertainty
- Separate durable information from transient discussion
- Reconcile category shifts where supported by evidence
## Must Not
- Directly write a protocol
- Invent information
- Drop conflicting evidence silently
## Processing Type
Local LLM, with deterministic pre/post-processing where useful.
## Current Status
Planned
---
# Stage 8 – Canonical Meeting Knowledge
## Purpose
Produce the canonical semantic representation of one meeting.
This representation is the single source of truth for all downstream outputs.
It is a structured representation, preferably JSON, and is not itself a prose
protocol.
Example:
@@ -352,7 +403,7 @@ Planned
---
# Stage 8 – Output View Rendering
# Stage 9 – Output View Rendering
## Purpose
@@ -365,6 +416,7 @@ The planned output products are:
- Knowledge Objects, rendered as a Knowledge-base Entry (`knowledge_entry.md`)
and later stored in a structured format such as `knowledge_entry.json`
(Wissensdatenbankeintrag)
- Later additional views such as action lists
Output rendering never performs additional analysis.
@@ -383,6 +435,10 @@ Completeness differs by output:
- The Distribution Protocol optimizes for relevance and brevity.
- Knowledge Objects optimize for durability and reuse.
Rendered protocol language should normally match the dominant language of the
source transcript or consolidated meeting knowledge unless an explicit output
language is requested.
## Processing Type
LLM
@@ -451,11 +507,16 @@ A processing stage may be replaced by another implementation as long as it prese
⬜ Specialized Extraction
⬜ Consolidation
⬜ Deterministic Canonicalization
⬜ Semantic Consolidation
⬜ Canonical Meeting Knowledge
⬜ Output View Rendering
```
The immediate development focus is **Topic Segmentation**, as it provides the semantic structure on which all subsequent processing stages depend.
The immediate architecture focus is the Deterministic Canonicalizer followed by
the Semantic Consolidator. These stages preserve source evidence, recover
global context from independent chunk extractions and prepare Canonical Meeting
Knowledge for parallel Output View rendering.