Document canonicalization and consolidation milestone
- preserve Working Protocol Synthesizer V0 as comparison baseline - introduce deterministic canonicalization stage - define semantic consolidator responsibilities - clarify Canonical Meeting Knowledge generation - document source-language output policy - align roadmap, architecture and experiment log
This commit is contained in:
@@ -21,7 +21,8 @@ Audio
|
||||
-> normalization
|
||||
-> chunking
|
||||
-> local chunk extraction
|
||||
-> consolidation
|
||||
-> deterministic canonicalization
|
||||
-> semantic consolidation
|
||||
-> Canonical Meeting Knowledge
|
||||
-> Output Views
|
||||
```
|
||||
@@ -31,8 +32,8 @@ Status:
|
||||
- Implemented: Whisper JSON cleanup script, normalization, technical chunking,
|
||||
local chunk extraction, interim Markdown protocol builder.
|
||||
- Experimental/prototype: topic segmentation and review tooling.
|
||||
- Planned: consolidation, Canonical Meeting Knowledge implementation, final
|
||||
Output Views.
|
||||
- Planned: Deterministic Canonicalizer, Semantic Consolidator, Canonical
|
||||
Meeting Knowledge implementation, final Output Views.
|
||||
|
||||
## Architectural Principles
|
||||
|
||||
@@ -43,11 +44,26 @@ Status:
|
||||
- Output views must not silently change meaning. They may select, condense or
|
||||
render information for an audience, but not invent new semantics.
|
||||
- Extraction, consolidation, synthesis and rendering are separate concerns.
|
||||
- Deterministic canonicalization and semantic consolidation are separate
|
||||
concerns.
|
||||
- The Deterministic Canonicalizer is planned Python code. It validates and
|
||||
normalizes extraction objects, assigns stable source references and IDs,
|
||||
normalizes category names and basic field structure, performs only safe
|
||||
deterministic cleanup, may group exact duplicates, and must preserve all
|
||||
source evidence. It must not perform uncertain semantic merging.
|
||||
- The Semantic Consolidator is planned local-LLM work. It merges semantically
|
||||
equivalent statements, groups content by topic, preserves evidence from all
|
||||
contributing chunks, marks contradictions and uncertainty, separates durable
|
||||
information from transient discussion, and produces Canonical Meeting
|
||||
Knowledge. It does not directly write a protocol.
|
||||
- Prefer small, testable processing stages over one monolithic LLM prompt.
|
||||
- Current extraction strategy is one normalized chunk per LLM call.
|
||||
- Do not expand context windows or redesign the extraction strategy without an
|
||||
explicit experiment.
|
||||
- Deterministic stages should remain deterministic where possible.
|
||||
- Rendered protocol output language should normally match the dominant language
|
||||
of the source transcript or consolidated meeting knowledge unless an explicit
|
||||
output language is requested.
|
||||
|
||||
## Prompt Engineering Rules
|
||||
|
||||
@@ -116,9 +132,9 @@ Current documented Prompt Version 2 decision baseline:
|
||||
- `src/meeting_lab/segmentation/`: experimental topic segmentation tooling.
|
||||
- `src/meeting_lab/extraction/`: local LLM extraction flow and category
|
||||
extractor modules.
|
||||
- `src/meeting_lab/consolidation/`: planned consolidation area.
|
||||
- `src/meeting_lab/consolidation/`: planned deterministic canonicalization and
|
||||
semantic consolidation area.
|
||||
- `src/meeting_lab/protocol/`: interim protocol rendering.
|
||||
- `src/meeting_lab/models/`: current lightweight data models.
|
||||
- `samples/`: sample inputs and generated/experimental artifacts; do not treat
|
||||
sample output as canonical source data.
|
||||
|
||||
|
||||
Reference in New Issue
Block a user