Document canonicalization and consolidation milestone
- preserve Working Protocol Synthesizer V0 as comparison baseline - introduce deterministic canonicalization stage - define semantic consolidator responsibilities - clarify Canonical Meeting Knowledge generation - document source-language output policy - align roadmap, architecture and experiment log
This commit is contained in:
+35
-21
@@ -98,7 +98,9 @@ Topic Segmentation
|
||||
↓
|
||||
Specialized Extraction
|
||||
↓
|
||||
Consolidation
|
||||
Deterministic Canonicalization
|
||||
↓
|
||||
Semantic Consolidation
|
||||
↓
|
||||
Canonical Meeting Knowledge
|
||||
↓
|
||||
@@ -174,14 +176,32 @@ Each extractor has exactly one task and one prompt.
|
||||
|
||||
## consolidation/
|
||||
|
||||
Merges information extracted from multiple discussion segments.
|
||||
Planned area for canonicalization and consolidation.
|
||||
|
||||
Typical responsibilities:
|
||||
The next milestone splits this into two stages.
|
||||
|
||||
- merge duplicates
|
||||
- combine partial information
|
||||
- distinguish positions from decisions
|
||||
- detect contradictions
|
||||
Deterministic Canonicalizer:
|
||||
|
||||
- implemented in Python
|
||||
- uses no LLM
|
||||
- validates and normalizes extraction objects
|
||||
- assigns stable source references and IDs
|
||||
- normalizes category names and basic field structure
|
||||
- performs only safe deterministic cleanup
|
||||
- may group exact duplicates
|
||||
- preserves all source evidence
|
||||
- must not perform uncertain semantic merging
|
||||
|
||||
Semantic Consolidator:
|
||||
|
||||
- uses the local LLM
|
||||
- merges semantically equivalent statements
|
||||
- groups content by topic
|
||||
- preserves evidence from all contributing chunks
|
||||
- marks contradictions and uncertainty
|
||||
- separates durable information from transient discussion
|
||||
- produces Canonical Meeting Knowledge
|
||||
- does not directly write a protocol
|
||||
|
||||
---
|
||||
|
||||
@@ -207,6 +227,10 @@ The planned output products are:
|
||||
These are parallel renderings of the same canonical semantic model, not
|
||||
documents derived from one another.
|
||||
|
||||
Rendered protocol language should normally match the dominant language of the
|
||||
source transcript or consolidated meeting knowledge unless an explicit output
|
||||
language is requested.
|
||||
|
||||
---
|
||||
|
||||
# Repository Layout
|
||||
@@ -257,20 +281,10 @@ Views documented in `output-views.md`.
|
||||
|
||||
# Next Milestone
|
||||
|
||||
The next development step is the implementation of **topic segmentation**.
|
||||
|
||||
Its only responsibility is to identify the thematic structure of a discussion.
|
||||
|
||||
It should answer questions such as:
|
||||
|
||||
- Where does a topic begin?
|
||||
- Where does it end?
|
||||
- When does another topic start?
|
||||
- When is an earlier topic resumed?
|
||||
|
||||
No facts, decisions or todos should be extracted at this stage.
|
||||
|
||||
Only after reliable topic segmentation has been achieved will the specialized extraction modules be implemented.
|
||||
The next architecture milestone is the implementation of a deterministic
|
||||
canonicalization stage followed by a semantic consolidation stage. These stages
|
||||
convert raw chunk extraction JSON into evidence-preserving Canonical Meeting
|
||||
Knowledge before any Output View renderer writes a protocol.
|
||||
|
||||
---
|
||||
|
||||
|
||||
Reference in New Issue
Block a user