Document canonicalization and consolidation milestone
- preserve Working Protocol Synthesizer V0 as comparison baseline - introduce deterministic canonicalization stage - define semantic consolidator responsibilities - clarify Canonical Meeting Knowledge generation - document source-language output policy - align roadmap, architecture and experiment log
This commit is contained in:
+35
-21
@@ -98,7 +98,9 @@ Topic Segmentation
|
||||
↓
|
||||
Specialized Extraction
|
||||
↓
|
||||
Consolidation
|
||||
Deterministic Canonicalization
|
||||
↓
|
||||
Semantic Consolidation
|
||||
↓
|
||||
Canonical Meeting Knowledge
|
||||
↓
|
||||
@@ -174,14 +176,32 @@ Each extractor has exactly one task and one prompt.
|
||||
|
||||
## consolidation/
|
||||
|
||||
Merges information extracted from multiple discussion segments.
|
||||
Planned area for canonicalization and consolidation.
|
||||
|
||||
Typical responsibilities:
|
||||
The next milestone splits this into two stages.
|
||||
|
||||
- merge duplicates
|
||||
- combine partial information
|
||||
- distinguish positions from decisions
|
||||
- detect contradictions
|
||||
Deterministic Canonicalizer:
|
||||
|
||||
- implemented in Python
|
||||
- uses no LLM
|
||||
- validates and normalizes extraction objects
|
||||
- assigns stable source references and IDs
|
||||
- normalizes category names and basic field structure
|
||||
- performs only safe deterministic cleanup
|
||||
- may group exact duplicates
|
||||
- preserves all source evidence
|
||||
- must not perform uncertain semantic merging
|
||||
|
||||
Semantic Consolidator:
|
||||
|
||||
- uses the local LLM
|
||||
- merges semantically equivalent statements
|
||||
- groups content by topic
|
||||
- preserves evidence from all contributing chunks
|
||||
- marks contradictions and uncertainty
|
||||
- separates durable information from transient discussion
|
||||
- produces Canonical Meeting Knowledge
|
||||
- does not directly write a protocol
|
||||
|
||||
---
|
||||
|
||||
@@ -207,6 +227,10 @@ The planned output products are:
|
||||
These are parallel renderings of the same canonical semantic model, not
|
||||
documents derived from one another.
|
||||
|
||||
Rendered protocol language should normally match the dominant language of the
|
||||
source transcript or consolidated meeting knowledge unless an explicit output
|
||||
language is requested.
|
||||
|
||||
---
|
||||
|
||||
# Repository Layout
|
||||
@@ -257,20 +281,10 @@ Views documented in `output-views.md`.
|
||||
|
||||
# Next Milestone
|
||||
|
||||
The next development step is the implementation of **topic segmentation**.
|
||||
|
||||
Its only responsibility is to identify the thematic structure of a discussion.
|
||||
|
||||
It should answer questions such as:
|
||||
|
||||
- Where does a topic begin?
|
||||
- Where does it end?
|
||||
- When does another topic start?
|
||||
- When is an earlier topic resumed?
|
||||
|
||||
No facts, decisions or todos should be extracted at this stage.
|
||||
|
||||
Only after reliable topic segmentation has been achieved will the specialized extraction modules be implemented.
|
||||
The next architecture milestone is the implementation of a deterministic
|
||||
canonicalization stage followed by a semantic consolidation stage. These stages
|
||||
convert raw chunk extraction JSON into evidence-preserving Canonical Meeting
|
||||
Knowledge before any Output View renderer writes a protocol.
|
||||
|
||||
---
|
||||
|
||||
|
||||
@@ -242,11 +242,74 @@ This object feeds the Canonical Meeting Knowledge representation.
|
||||
|
||||
---
|
||||
|
||||
# Canonical Extraction Object
|
||||
|
||||
Planned deterministic intermediate object created from raw chunk extraction
|
||||
JSON.
|
||||
|
||||
Example:
|
||||
|
||||
```json
|
||||
{
|
||||
"id": "fact.chunk_03.0001",
|
||||
"category": "fact",
|
||||
"text": "...",
|
||||
"source_references": [
|
||||
{
|
||||
"chunk_id": "chunk_03",
|
||||
"source_file": "chunk_03_extraction.json",
|
||||
"evidence": "..."
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
The Deterministic Canonicalizer should create this kind of object without an
|
||||
LLM. It validates and normalizes raw extraction objects, assigns stable IDs and
|
||||
source references, normalizes category names and basic field structure,
|
||||
performs only safe deterministic cleanup, may group exact duplicates and must
|
||||
preserve all source evidence.
|
||||
|
||||
It must not perform uncertain semantic merging.
|
||||
|
||||
---
|
||||
|
||||
# Consolidated Topic
|
||||
|
||||
Planned semantic object produced by the Semantic Consolidator.
|
||||
|
||||
Example:
|
||||
|
||||
```json
|
||||
{
|
||||
"topic_id": "topic_001",
|
||||
"title": "...",
|
||||
"background": [],
|
||||
"decisions": [],
|
||||
"action_items": [],
|
||||
"open_questions": [],
|
||||
"durable_information": [],
|
||||
"uncertainty": [],
|
||||
"source_references": []
|
||||
}
|
||||
```
|
||||
|
||||
The Semantic Consolidator may use the local LLM to merge semantically
|
||||
equivalent statements, group content by topic, preserve evidence from all
|
||||
contributing chunks, mark contradictions and uncertainty and separate durable
|
||||
information from transient discussion.
|
||||
|
||||
It produces Canonical Meeting Knowledge. It does not directly write a protocol.
|
||||
|
||||
---
|
||||
|
||||
# Canonical Meeting Knowledge
|
||||
|
||||
The canonical semantic representation of one meeting.
|
||||
|
||||
This representation is the single source of truth for all downstream outputs.
|
||||
It is a structured representation, preferably JSON, and is not itself a prose
|
||||
protocol.
|
||||
|
||||
```json
|
||||
{
|
||||
@@ -260,6 +323,7 @@ This representation is the single source of truth for all downstream outputs.
|
||||
"questions": [],
|
||||
"positions": [],
|
||||
"technical_details": [],
|
||||
"durable_information": [],
|
||||
"rationale": [],
|
||||
"uncertainty": [],
|
||||
"source_references": []
|
||||
|
||||
+73
-5
@@ -945,14 +945,18 @@ Result:
|
||||
|
||||
The accepted design is a planned Canonical Meeting Knowledge layer as the
|
||||
semantic source of truth, with Working Protocol, Distribution Protocol and
|
||||
Knowledge Objects as parallel output views. Consolidation must merge duplicates,
|
||||
preserve evidence, reconcile category shifts and mark contradictions or
|
||||
uncertainty.
|
||||
Knowledge Objects as parallel output views. The next consolidation architecture
|
||||
is split into a Deterministic Canonicalizer and a Semantic Consolidator. The
|
||||
canonicalizer prepares validated evidence-bearing objects without uncertain
|
||||
semantic merging. The consolidator then merges semantically equivalent
|
||||
statements, preserves evidence, reconciles category shifts where supported and
|
||||
marks contradictions or uncertainty.
|
||||
|
||||
Decision:
|
||||
|
||||
Consolidation is the next major engineering step after stable local extraction.
|
||||
Canonical Meeting Knowledge and final output views are planned, not implemented.
|
||||
Deterministic canonicalization is the next implementation step after stable
|
||||
local extraction, followed by semantic consolidation. Canonical Meeting
|
||||
Knowledge and final output views are planned, not implemented.
|
||||
|
||||
Lessons learned:
|
||||
|
||||
@@ -969,3 +973,67 @@ Evidence:
|
||||
- `PROJECT_KNOWLEDGE.md`
|
||||
- `ROADMAP.md`
|
||||
- Commit `5c03ed7`
|
||||
|
||||
## EXP-0020 - Working Protocol Synthesizer V0
|
||||
|
||||
Status: Accepted
|
||||
|
||||
Date or period: 2026-07-31
|
||||
|
||||
Hypothesis:
|
||||
|
||||
The current local synthesis model may be able to generate a useful detailed
|
||||
Working Protocol directly from the existing independent chunk extraction JSON
|
||||
files, before canonicalization or semantic consolidation exists.
|
||||
|
||||
Setup:
|
||||
|
||||
One synthesis prompt was constructed from exactly nine chunk extraction JSON
|
||||
files. The model was instructed to use only those extraction files, merge
|
||||
duplicates, group related information into topics, preserve useful discussion
|
||||
context and write a neutral technical Working Protocol.
|
||||
|
||||
Inputs:
|
||||
|
||||
- `chunk_01_extraction.json` through `chunk_09_extraction.json`.
|
||||
- No original transcript, normalized chunks or Whisper output were used as
|
||||
synthesis input.
|
||||
|
||||
Model / configuration:
|
||||
|
||||
- Model: `qwen3.5:9b`
|
||||
- Prompt characters: 24,979
|
||||
- Actual prompt eval tokens: 5,929
|
||||
- Output tokens: 1,486
|
||||
- Runtime: 294.204 seconds
|
||||
|
||||
Result:
|
||||
|
||||
The generated Working Protocol was readable, well structured and
|
||||
topic-oriented. It was still based directly on raw chunk extractions, without a
|
||||
separate deterministic canonicalization stage or semantic consolidation stage.
|
||||
The output language was English even though the source meeting material was
|
||||
German.
|
||||
|
||||
Decision:
|
||||
|
||||
Preserve this output as the Working Protocol Synthesizer V0 benchmark baseline
|
||||
for later canonicalizer, consolidator and renderer comparisons. This selected
|
||||
generated artifact is intentionally versioned even though generated runtime
|
||||
artifacts are normally ignored.
|
||||
|
||||
Lessons learned:
|
||||
|
||||
Direct synthesis from chunk extractions can create a useful recall-oriented
|
||||
draft, but it does not replace Canonical Meeting Knowledge. The language
|
||||
mismatch also establishes a default renderer rule: protocol output should
|
||||
normally match the dominant source language unless an explicit output language
|
||||
is requested.
|
||||
|
||||
Evidence:
|
||||
|
||||
- `samples/benchmarks/working_protocol_synthesizer_v0/README.md`
|
||||
- `samples/benchmarks/working_protocol_synthesizer_v0/working_protocol.md`
|
||||
- `PROJECT_KNOWLEDGE.md`
|
||||
- `docs/output-views.md`
|
||||
- See EXP-0017 and EXP-0019.
|
||||
|
||||
@@ -13,6 +13,10 @@ Meeting transcript
|
||||
->
|
||||
chunk extraction
|
||||
->
|
||||
Deterministic Canonicalizer
|
||||
->
|
||||
Semantic Consolidator
|
||||
->
|
||||
Canonical Meeting Knowledge
|
||||
├── Working Protocol
|
||||
├── Distribution Protocol
|
||||
@@ -45,6 +49,18 @@ The detailed schema is future implementation work. The current implementation
|
||||
still uses simple extraction JSON files and a basic Markdown protocol builder for
|
||||
technical validation.
|
||||
|
||||
The next planned architecture stage before this representation is explicit:
|
||||
|
||||
- The Deterministic Canonicalizer validates and normalizes extraction objects,
|
||||
assigns stable source references and IDs, performs only safe deterministic
|
||||
cleanup and preserves all source evidence. It uses no LLM and must not make
|
||||
uncertain semantic merges.
|
||||
- The Semantic Consolidator uses the local LLM to merge semantically equivalent
|
||||
statements, group content by topic, preserve evidence from all contributing
|
||||
chunks, mark contradictions and uncertainty, separate durable information from
|
||||
transient discussion and produce Canonical Meeting Knowledge. It does not
|
||||
directly write a protocol.
|
||||
|
||||
## Renderers
|
||||
|
||||
Each Output View is produced by a renderer.
|
||||
@@ -57,6 +73,10 @@ Depending on the implementation, rendering may be:
|
||||
|
||||
The architecture does not assume that every renderer must always use an LLM.
|
||||
|
||||
Rendered protocol language should normally match the dominant language of the
|
||||
source transcript or consolidated meeting knowledge unless an explicit output
|
||||
language is requested.
|
||||
|
||||
## Working Protocol
|
||||
|
||||
Suggested filename: `working_protocol.md`
|
||||
|
||||
+80
-19
@@ -33,7 +33,10 @@ Topic Segmentation
|
||||
Specialized Extraction
|
||||
│
|
||||
▼
|
||||
Consolidation
|
||||
Deterministic Canonicalization
|
||||
│
|
||||
▼
|
||||
Semantic Consolidation
|
||||
│
|
||||
▼
|
||||
Canonical Meeting Knowledge
|
||||
@@ -267,35 +270,41 @@ Prototype exists as a combined extractor.
|
||||
|
||||
---
|
||||
|
||||
# Stage 6 – Consolidation
|
||||
# Stage 6 – Deterministic Canonicalization
|
||||
|
||||
## Purpose
|
||||
|
||||
Merge analysis results originating from different discussion segments.
|
||||
Normalize raw chunk extraction JSON into stable canonical extraction objects
|
||||
without changing uncertain semantics.
|
||||
|
||||
## Input
|
||||
|
||||
Extraction results.
|
||||
Chunk extraction JSON files.
|
||||
|
||||
## Output
|
||||
|
||||
Unified topic representation.
|
||||
Validated canonical extraction objects with stable source references and IDs.
|
||||
|
||||
## Responsibilities
|
||||
|
||||
- Merge duplicates
|
||||
- Merge complementary information
|
||||
- Preserve contradictions
|
||||
- Separate positions from decisions
|
||||
- Combine related todos
|
||||
- Validate extraction objects
|
||||
- Normalize category names
|
||||
- Normalize basic field structure
|
||||
- Assign stable source references and IDs
|
||||
- Preserve all source evidence
|
||||
- Perform only safe deterministic cleanup
|
||||
- Group exact duplicates where unambiguous
|
||||
|
||||
## Must Not
|
||||
|
||||
- Use an LLM
|
||||
- Perform uncertain semantic merging
|
||||
- Infer missing information
|
||||
- Drop source evidence
|
||||
|
||||
## Processing Type
|
||||
|
||||
Hybrid
|
||||
|
||||
Deterministic wherever possible.
|
||||
|
||||
LLM support only if necessary.
|
||||
Deterministic Python
|
||||
|
||||
## Current Status
|
||||
|
||||
@@ -303,13 +312,55 @@ Planned
|
||||
|
||||
---
|
||||
|
||||
# Stage 7 – Canonical Meeting Knowledge
|
||||
# Stage 7 – Semantic Consolidation
|
||||
|
||||
## Purpose
|
||||
|
||||
Merge canonicalized extraction objects into an evidence-preserving semantic
|
||||
meeting representation.
|
||||
|
||||
## Input
|
||||
|
||||
Canonicalized extraction objects.
|
||||
|
||||
## Output
|
||||
|
||||
Canonical Meeting Knowledge.
|
||||
|
||||
## Responsibilities
|
||||
|
||||
- Merge semantically equivalent statements
|
||||
- Group content by topic
|
||||
- Preserve evidence from all contributing chunks
|
||||
- Mark contradictions and uncertainty
|
||||
- Separate durable information from transient discussion
|
||||
- Reconcile category shifts where supported by evidence
|
||||
|
||||
## Must Not
|
||||
|
||||
- Directly write a protocol
|
||||
- Invent information
|
||||
- Drop conflicting evidence silently
|
||||
|
||||
## Processing Type
|
||||
|
||||
Local LLM, with deterministic pre/post-processing where useful.
|
||||
|
||||
## Current Status
|
||||
|
||||
Planned
|
||||
|
||||
---
|
||||
|
||||
# Stage 8 – Canonical Meeting Knowledge
|
||||
|
||||
## Purpose
|
||||
|
||||
Produce the canonical semantic representation of one meeting.
|
||||
|
||||
This representation is the single source of truth for all downstream outputs.
|
||||
It is a structured representation, preferably JSON, and is not itself a prose
|
||||
protocol.
|
||||
|
||||
Example:
|
||||
|
||||
@@ -352,7 +403,7 @@ Planned
|
||||
|
||||
---
|
||||
|
||||
# Stage 8 – Output View Rendering
|
||||
# Stage 9 – Output View Rendering
|
||||
|
||||
## Purpose
|
||||
|
||||
@@ -365,6 +416,7 @@ The planned output products are:
|
||||
- Knowledge Objects, rendered as a Knowledge-base Entry (`knowledge_entry.md`)
|
||||
and later stored in a structured format such as `knowledge_entry.json`
|
||||
(Wissensdatenbankeintrag)
|
||||
- Later additional views such as action lists
|
||||
|
||||
Output rendering never performs additional analysis.
|
||||
|
||||
@@ -383,6 +435,10 @@ Completeness differs by output:
|
||||
- The Distribution Protocol optimizes for relevance and brevity.
|
||||
- Knowledge Objects optimize for durability and reuse.
|
||||
|
||||
Rendered protocol language should normally match the dominant language of the
|
||||
source transcript or consolidated meeting knowledge unless an explicit output
|
||||
language is requested.
|
||||
|
||||
## Processing Type
|
||||
|
||||
LLM
|
||||
@@ -451,11 +507,16 @@ A processing stage may be replaced by another implementation as long as it prese
|
||||
|
||||
⬜ Specialized Extraction
|
||||
|
||||
⬜ Consolidation
|
||||
⬜ Deterministic Canonicalization
|
||||
|
||||
⬜ Semantic Consolidation
|
||||
|
||||
⬜ Canonical Meeting Knowledge
|
||||
|
||||
⬜ Output View Rendering
|
||||
```
|
||||
|
||||
The immediate development focus is **Topic Segmentation**, as it provides the semantic structure on which all subsequent processing stages depend.
|
||||
The immediate architecture focus is the Deterministic Canonicalizer followed by
|
||||
the Semantic Consolidator. These stages preserve source evidence, recover
|
||||
global context from independent chunk extractions and prepare Canonical Meeting
|
||||
Knowledge for parallel Output View rendering.
|
||||
|
||||
Reference in New Issue
Block a user