Document canonicalization and consolidation milestone

- preserve Working Protocol Synthesizer V0 as comparison baseline
- introduce deterministic canonicalization stage
- define semantic consolidator responsibilities
- clarify Canonical Meeting Knowledge generation
- document source-language output policy
- align roadmap, architecture and experiment log
This commit is contained in:
2026-07-31 09:32:12 +02:00
parent 09d125e54a
commit 23bbc744f7
12 changed files with 512 additions and 83 deletions
+80 -19
View File
@@ -33,7 +33,10 @@ Topic Segmentation
Specialized Extraction
│
▼
Consolidation
Deterministic Canonicalization
│
▼
Semantic Consolidation
│
▼
Canonical Meeting Knowledge
@@ -267,35 +270,41 @@ Prototype exists as a combined extractor.
---
# Stage 6 – Consolidation
# Stage 6 – Deterministic Canonicalization
## Purpose
Merge analysis results originating from different discussion segments.
Normalize raw chunk extraction JSON into stable canonical extraction objects
without changing uncertain semantics.
## Input
Extraction results.
Chunk extraction JSON files.
## Output
Unified topic representation.
Validated canonical extraction objects with stable source references and IDs.
## Responsibilities
- Merge duplicates
- Merge complementary information
- Preserve contradictions
- Separate positions from decisions
- Combine related todos
- Validate extraction objects
- Normalize category names
- Normalize basic field structure
- Assign stable source references and IDs
- Preserve all source evidence
- Perform only safe deterministic cleanup
- Group exact duplicates where unambiguous
## Must Not
- Use an LLM
- Perform uncertain semantic merging
- Infer missing information
- Drop source evidence
## Processing Type
Hybrid
Deterministic wherever possible.
LLM support only if necessary.
Deterministic Python
## Current Status
@@ -303,13 +312,55 @@ Planned
---
# Stage 7 – Canonical Meeting Knowledge
# Stage 7 – Semantic Consolidation
## Purpose
Merge canonicalized extraction objects into an evidence-preserving semantic
meeting representation.
## Input
Canonicalized extraction objects.
## Output
Canonical Meeting Knowledge.
## Responsibilities
- Merge semantically equivalent statements
- Group content by topic
- Preserve evidence from all contributing chunks
- Mark contradictions and uncertainty
- Separate durable information from transient discussion
- Reconcile category shifts where supported by evidence
## Must Not
- Directly write a protocol
- Invent information
- Drop conflicting evidence silently
## Processing Type
Local LLM, with deterministic pre/post-processing where useful.
## Current Status
Planned
---
# Stage 8 – Canonical Meeting Knowledge
## Purpose
Produce the canonical semantic representation of one meeting.
This representation is the single source of truth for all downstream outputs.
It is a structured representation, preferably JSON, and is not itself a prose
protocol.
Example:
@@ -352,7 +403,7 @@ Planned
---
# Stage 8 – Output View Rendering
# Stage 9 – Output View Rendering
## Purpose
@@ -365,6 +416,7 @@ The planned output products are:
- Knowledge Objects, rendered as a Knowledge-base Entry (`knowledge_entry.md`)
and later stored in a structured format such as `knowledge_entry.json`
(Wissensdatenbankeintrag)
- Later additional views such as action lists
Output rendering never performs additional analysis.
@@ -383,6 +435,10 @@ Completeness differs by output:
- The Distribution Protocol optimizes for relevance and brevity.
- Knowledge Objects optimize for durability and reuse.
Rendered protocol language should normally match the dominant language of the
source transcript or consolidated meeting knowledge unless an explicit output
language is requested.
## Processing Type
LLM
@@ -451,11 +507,16 @@ A processing stage may be replaced by another implementation as long as it prese
⬜ Specialized Extraction
⬜ Consolidation
⬜ Deterministic Canonicalization
⬜ Semantic Consolidation
⬜ Canonical Meeting Knowledge
⬜ Output View Rendering
```
The immediate development focus is **Topic Segmentation**, as it provides the semantic structure on which all subsequent processing stages depend.
The immediate architecture focus is the Deterministic Canonicalizer followed by
the Semantic Consolidator. These stages preserve source evidence, recover
global context from independent chunk extractions and prepare Canonical Meeting
Knowledge for parallel Output View rendering.