Document canonicalization and consolidation milestone

- preserve Working Protocol Synthesizer V0 as comparison baseline
- introduce deterministic canonicalization stage
- define semantic consolidator responsibilities
- clarify Canonical Meeting Knowledge generation
- document source-language output policy
- align roadmap, architecture and experiment log
This commit is contained in:
2026-07-31 09:32:12 +02:00
parent 09d125e54a
commit 23bbc744f7
12 changed files with 512 additions and 83 deletions
+73 -5
View File
@@ -945,14 +945,18 @@ Result:
The accepted design is a planned Canonical Meeting Knowledge layer as the
semantic source of truth, with Working Protocol, Distribution Protocol and
Knowledge Objects as parallel output views. Consolidation must merge duplicates,
preserve evidence, reconcile category shifts and mark contradictions or
uncertainty.
Knowledge Objects as parallel output views. The next consolidation architecture
is split into a Deterministic Canonicalizer and a Semantic Consolidator. The
canonicalizer prepares validated evidence-bearing objects without uncertain
semantic merging. The consolidator then merges semantically equivalent
statements, preserves evidence, reconciles category shifts where supported and
marks contradictions or uncertainty.
Decision:
Consolidation is the next major engineering step after stable local extraction.
Canonical Meeting Knowledge and final output views are planned, not implemented.
Deterministic canonicalization is the next implementation step after stable
local extraction, followed by semantic consolidation. Canonical Meeting
Knowledge and final output views are planned, not implemented.
Lessons learned:
@@ -969,3 +973,67 @@ Evidence:
- `PROJECT_KNOWLEDGE.md`
- `ROADMAP.md`
- Commit `5c03ed7`
## EXP-0020 - Working Protocol Synthesizer V0
Status: Accepted
Date or period: 2026-07-31
Hypothesis:
The current local synthesis model may be able to generate a useful detailed
Working Protocol directly from the existing independent chunk extraction JSON
files, before canonicalization or semantic consolidation exists.
Setup:
One synthesis prompt was constructed from exactly nine chunk extraction JSON
files. The model was instructed to use only those extraction files, merge
duplicates, group related information into topics, preserve useful discussion
context and write a neutral technical Working Protocol.
Inputs:
- `chunk_01_extraction.json` through `chunk_09_extraction.json`.
- No original transcript, normalized chunks or Whisper output were used as
synthesis input.
Model / configuration:
- Model: `qwen3.5:9b`
- Prompt characters: 24,979
- Actual prompt eval tokens: 5,929
- Output tokens: 1,486
- Runtime: 294.204 seconds
Result:
The generated Working Protocol was readable, well structured and
topic-oriented. It was still based directly on raw chunk extractions, without a
separate deterministic canonicalization stage or semantic consolidation stage.
The output language was English even though the source meeting material was
German.
Decision:
Preserve this output as the Working Protocol Synthesizer V0 benchmark baseline
for later canonicalizer, consolidator and renderer comparisons. This selected
generated artifact is intentionally versioned even though generated runtime
artifacts are normally ignored.
Lessons learned:
Direct synthesis from chunk extractions can create a useful recall-oriented
draft, but it does not replace Canonical Meeting Knowledge. The language
mismatch also establishes a default renderer rule: protocol output should
normally match the dominant source language unless an explicit output language
is requested.
Evidence:
- `samples/benchmarks/working_protocol_synthesizer_v0/README.md`
- `samples/benchmarks/working_protocol_synthesizer_v0/working_protocol.md`
- `PROJECT_KNOWLEDGE.md`
- `docs/output-views.md`
- See EXP-0017 and EXP-0019.