Document canonicalization and consolidation milestone

- preserve Working Protocol Synthesizer V0 as comparison baseline
- introduce deterministic canonicalization stage
- define semantic consolidator responsibilities
- clarify Canonical Meeting Knowledge generation
- document source-language output policy
- align roadmap, architecture and experiment log
This commit is contained in:
2026-07-31 09:32:12 +02:00
parent 09d125e54a
commit 23bbc744f7
12 changed files with 512 additions and 83 deletions
+21 -5
View File
@@ -21,7 +21,8 @@ Audio
-> normalization
-> chunking
-> local chunk extraction
-> consolidation
-> deterministic canonicalization
-> semantic consolidation
-> Canonical Meeting Knowledge
-> Output Views
```
@@ -31,8 +32,8 @@ Status:
- Implemented: Whisper JSON cleanup script, normalization, technical chunking,
local chunk extraction, interim Markdown protocol builder.
- Experimental/prototype: topic segmentation and review tooling.
- Planned: consolidation, Canonical Meeting Knowledge implementation, final
Output Views.
- Planned: Deterministic Canonicalizer, Semantic Consolidator, Canonical
Meeting Knowledge implementation, final Output Views.
## Architectural Principles
@@ -43,11 +44,26 @@ Status:
- Output views must not silently change meaning. They may select, condense or
render information for an audience, but not invent new semantics.
- Extraction, consolidation, synthesis and rendering are separate concerns.
- Deterministic canonicalization and semantic consolidation are separate
concerns.
- The Deterministic Canonicalizer is planned Python code. It validates and
normalizes extraction objects, assigns stable source references and IDs,
normalizes category names and basic field structure, performs only safe
deterministic cleanup, may group exact duplicates, and must preserve all
source evidence. It must not perform uncertain semantic merging.
- The Semantic Consolidator is planned local-LLM work. It merges semantically
equivalent statements, groups content by topic, preserves evidence from all
contributing chunks, marks contradictions and uncertainty, separates durable
information from transient discussion, and produces Canonical Meeting
Knowledge. It does not directly write a protocol.
- Prefer small, testable processing stages over one monolithic LLM prompt.
- Current extraction strategy is one normalized chunk per LLM call.
- Do not expand context windows or redesign the extraction strategy without an
explicit experiment.
- Deterministic stages should remain deterministic where possible.
- Rendered protocol output language should normally match the dominant language
of the source transcript or consolidated meeting knowledge unless an explicit
output language is requested.
## Prompt Engineering Rules
@@ -116,9 +132,9 @@ Current documented Prompt Version 2 decision baseline:
- `src/meeting_lab/segmentation/`: experimental topic segmentation tooling.
- `src/meeting_lab/extraction/`: local LLM extraction flow and category
extractor modules.
- `src/meeting_lab/consolidation/`: planned consolidation area.
- `src/meeting_lab/consolidation/`: planned deterministic canonicalization and
semantic consolidation area.
- `src/meeting_lab/protocol/`: interim protocol rendering.
- `src/meeting_lab/models/`: current lightweight data models.
- `samples/`: sample inputs and generated/experimental artifacts; do not treat
sample output as canonical source data.