Document canonicalization and consolidation milestone

- preserve Working Protocol Synthesizer V0 as comparison baseline
- introduce deterministic canonicalization stage
- define semantic consolidator responsibilities
- clarify Canonical Meeting Knowledge generation
- document source-language output policy
- align roadmap, architecture and experiment log
This commit is contained in:
2026-07-31 09:32:12 +02:00
parent 09d125e54a
commit 23bbc744f7
12 changed files with 512 additions and 83 deletions
+21 -5
View File
@@ -21,7 +21,8 @@ Audio
-> normalization
-> chunking
-> local chunk extraction
-> consolidation
-> deterministic canonicalization
-> semantic consolidation
-> Canonical Meeting Knowledge
-> Output Views
```
@@ -31,8 +32,8 @@ Status:
- Implemented: Whisper JSON cleanup script, normalization, technical chunking,
local chunk extraction, interim Markdown protocol builder.
- Experimental/prototype: topic segmentation and review tooling.
- Planned: consolidation, Canonical Meeting Knowledge implementation, final
Output Views.
- Planned: Deterministic Canonicalizer, Semantic Consolidator, Canonical
Meeting Knowledge implementation, final Output Views.
## Architectural Principles
@@ -43,11 +44,26 @@ Status:
- Output views must not silently change meaning. They may select, condense or
render information for an audience, but not invent new semantics.
- Extraction, consolidation, synthesis and rendering are separate concerns.
- Deterministic canonicalization and semantic consolidation are separate
concerns.
- The Deterministic Canonicalizer is planned Python code. It validates and
normalizes extraction objects, assigns stable source references and IDs,
normalizes category names and basic field structure, performs only safe
deterministic cleanup, may group exact duplicates, and must preserve all
source evidence. It must not perform uncertain semantic merging.
- The Semantic Consolidator is planned local-LLM work. It merges semantically
equivalent statements, groups content by topic, preserves evidence from all
contributing chunks, marks contradictions and uncertainty, separates durable
information from transient discussion, and produces Canonical Meeting
Knowledge. It does not directly write a protocol.
- Prefer small, testable processing stages over one monolithic LLM prompt.
- Current extraction strategy is one normalized chunk per LLM call.
- Do not expand context windows or redesign the extraction strategy without an
explicit experiment.
- Deterministic stages should remain deterministic where possible.
- Rendered protocol output language should normally match the dominant language
of the source transcript or consolidated meeting knowledge unless an explicit
output language is requested.
## Prompt Engineering Rules
@@ -116,9 +132,9 @@ Current documented Prompt Version 2 decision baseline:
- `src/meeting_lab/segmentation/`: experimental topic segmentation tooling.
- `src/meeting_lab/extraction/`: local LLM extraction flow and category
extractor modules.
- `src/meeting_lab/consolidation/`: planned consolidation area.
- `src/meeting_lab/consolidation/`: planned deterministic canonicalization and
semantic consolidation area.
- `src/meeting_lab/protocol/`: interim protocol rendering.
- `src/meeting_lab/models/`: current lightweight data models.
- `samples/`: sample inputs and generated/experimental artifacts; do not treat
sample output as canonical source data.
+7 -1
View File
@@ -17,6 +17,8 @@
- Output-view architecture documentation for Canonical Meeting Knowledge,
Working Protocol / Arbeitsprotokoll, Distribution Protocol /
Verteilerprotokoll and Knowledge Objects / Wissensdatenbankeintrag.
- Working Protocol Synthesizer V0 benchmark artifact for future
canonicalizer/consolidator comparisons.
### Changed
@@ -24,6 +26,11 @@
Meeting Knowledge as the planned semantic source of truth.
- Clarified that Working Protocol, Distribution Protocol and Knowledge Objects
are parallel renderings, not derived from one another.
- Documented the planned Deterministic Canonicalizer and Semantic Consolidator
stages before Canonical Meeting Knowledge.
- Documented that rendered protocol language should normally match the source
transcript or consolidated meeting knowledge unless explicitly requested
otherwise.
- Updated decision extraction semantics to include explicit process decisions
and deferrals.
- Simplified extraction prompt assembly around prompt files.
@@ -44,4 +51,3 @@
- Added Gold Standard prompt-engineering methodology.
- Added formal decision-definition documentation.
- Added scenario README files for the Gold Standard corpus.
+37 -8
View File
@@ -30,7 +30,8 @@ Experimental/prototype:
Planned:
- Consolidation of extraction results.
- Deterministic Canonicalizer for extraction results.
- Semantic Consolidator for evidence-preserving semantic merging.
- Canonical Meeting Knowledge implementation as the semantic source of truth.
- Final Working Protocol / Arbeitsprotokoll, Distribution Protocol /
Verteilerprotokoll and Knowledge Objects / Wissensdatenbankeintrag renderers.
@@ -40,7 +41,7 @@ Planned:
```text
src/meeting_lab/
chunking/ technical transcript chunking
consolidation/ planned merge/consolidation area
consolidation/ planned canonicalization/consolidation area
extraction/ current local LLM extraction flow
io/ lightweight file and JSON helpers
llm/ Ollama and prompt support
@@ -112,10 +113,33 @@ Current Prompt Version 2 decision baseline:
## Canonical Knowledge Architecture
The next documented pipeline milestone is:
```text
Chunk Extractions
-> Deterministic Canonicalizer
-> Semantic Consolidator
-> Canonical Meeting Knowledge
-> Output View Renderers
```
The Deterministic Canonicalizer is planned Python code with no LLM. It should
validate and normalize extraction objects, assign stable source references and
IDs, normalize category names and basic field structure, perform only safe
deterministic cleanup, optionally group exact duplicates, and preserve all
source evidence. It must not perform uncertain semantic merging.
The Semantic Consolidator is planned local-LLM work. It should merge
semantically equivalent statements, group content by topic, preserve evidence
from all contributing chunks, mark contradictions and uncertainty, separate
durable information from transient discussion, and produce Canonical Meeting
Knowledge. It does not directly write a protocol.
Canonical Meeting Knowledge is the planned semantic intermediate model and
future single source of truth. It should preserve topics, facts, decisions,
action items, open questions, positions, technical details, rationale,
uncertainty, contradictions and source evidence.
future single source of truth. It should be structured, preferably JSON, and
preserve topics, facts, decisions, action items, open questions, positions,
technical details, rationale, uncertainty, contradictions and source evidence.
It is not itself a prose protocol.
Output views are planned as independent renderings from that canonical model:
@@ -129,6 +153,10 @@ Output views are planned as independent renderings from that canonical model:
The current `meeting_protocol.md` builder is an interim technical validation
tool, not the final output-view architecture.
Rendered protocol output should normally use the dominant language of the
source transcript or consolidated meeting knowledge unless an explicit output
language is requested.
## Current Limitations
- Discussion Blocks are documented as a stable semantic unit but are not yet a
@@ -136,7 +164,8 @@ tool, not the final output-view architecture.
- Topic segmentation exists as prototype tooling, not a stable pipeline stage.
- Extraction is still a combined current flow, even though separate extractors
are the intended architecture.
- Consolidation is not implemented.
- Deterministic Canonicalizer is documented but not implemented.
- Semantic Consolidator is documented but not implemented.
- Canonical Meeting Knowledge is documented but not implemented.
- Final output views are documented but not implemented.
- Most prompt files are placeholders except the common and decision prompts.
@@ -147,5 +176,5 @@ tool, not the final output-view architecture.
Stabilize repeatable local extraction evaluation before broadening the pipeline:
expand Gold Standard coverage by category, keep one-chunk extraction as the
baseline, and use small prompt experiments with immediate non-regression checks.
After extraction behavior is stable enough, implement consolidation with
evidence retention as the next major pipeline stage.
After extraction behavior is stable enough, implement the Deterministic
Canonicalizer first, then the Semantic Consolidator with evidence retention.
+22 -3
View File
@@ -41,7 +41,9 @@ Themensegmentierung
↓
Extraktion
↓
Konsolidierung
Deterministic Canonicalizer
↓
Semantic Consolidator
↓
Canonical Meeting Knowledge
↓
@@ -91,12 +93,21 @@ chunk_transcript.py
↓
Extraktoren
↓
Deterministic Canonicalizer
↓
Semantic Consolidator
↓
Canonical Meeting Knowledge
↓
Output-Ansichten
```
Die Themensegmentierung bildet den nächsten großen Entwicklungsschritt.
Der nächste Architekturmeilenstein ist die Trennung zwischen deterministischer
Kanonisierung der Chunk-Extraktionen und semantischer Konsolidierung.
Die Kanonisierung validiert und normalisiert Extraktionsobjekte ohne LLM. Die
semantische Konsolidierung nutzt das lokale LLM, um gleichbedeutende Aussagen
zusammenzuführen, Evidenz zu erhalten und die Canonical Meeting Knowledge zu
erzeugen.
Das Meeting Lab behandelt "das Protokoll" nicht mehr als ein einzelnes
Endprodukt. Das konsolidierte Meeting-Wissen ist die **Canonical Meeting
@@ -118,11 +129,19 @@ Diese Ausgaben sind parallele Renderings desselben semantischen Modells. Das
Arbeitsprotokoll ist nicht die Quelle des Verteilerprotokolls, und das
Verteilerprotokoll ist nicht die Quelle der Knowledge Objects.
Gerenderte Protokolle sollen normalerweise in der dominanten Sprache des
Quelltranskripts beziehungsweise der konsolidierten Meeting Knowledge erstellt
werden, sofern keine explizite Ausgabesprache angefordert wurde.
---
## Projektstatus
Aktuell liegt der Schwerpunkt auf der Entwicklung eines modularen Diskussionsanalyzers.
Aktuell liegt der Schwerpunkt auf der Entwicklung eines modularen
Diskussionsanalyzers. Implementiert sind Vorverarbeitung, technische Chunking-
und lokale Chunk-Extraktionsschritte. Deterministic Canonicalizer, Semantic
Consolidator, Canonical Meeting Knowledge und finale Output-View-Renderer sind
geplante nächste Schritte.
Die eigentliche Ausgabeerzeugung ist bewusst der letzte Verarbeitungsschritt.
+52 -21
View File
@@ -25,36 +25,65 @@ Out of scope:
- Full-transcript LLM extraction.
- Larger context-window strategy changes without an explicit experiment.
- Consolidation or final protocol rendering.
- Canonicalization, semantic consolidation or final protocol rendering.
## Phase 2 - Consolidation
## Phase 2 - Deterministic Canonicalization
Goal:
- Merge independent extraction results into a coherent meeting-level
representation without losing evidence.
- Convert independent chunk extraction JSON into a validated, normalized,
evidence-bearing intermediate representation without semantic guessing.
Deliverables:
- Duplicate merging.
- Evidence retention.
- Category-shift reconciliation, especially facts versus positions and
positions versus decisions.
- Contradiction and uncertainty markers.
- Consolidated meeting representation.
- Deterministic Canonicalizer implemented in Python.
- Stable source references and IDs.
- Normalized category names and basic field structure.
- Safe deterministic cleanup.
- Exact duplicate grouping where unambiguous.
- Preservation of all source evidence.
Prerequisites:
- Stable local extraction baseline.
- Gold tests that expose cross-chunk duplication and category shifts.
- Agreement on the extraction object shape that should be canonicalized.
Out of scope:
- Final Canonical Meeting Knowledge schema.
- User-facing protocol polish.
- Uncertain semantic merging.
- Topic synthesis.
- Protocol writing.
- LLM calls.
## Phase 3 - Semantic Consolidation
Goal:
- Merge canonicalized extraction objects into a coherent semantic meeting
representation while preserving evidence and uncertainty.
Deliverables:
- Semantic Consolidator using the local LLM.
- Semantically equivalent statement merging.
- Topic grouping.
- Evidence preserved from all contributing chunks.
- Contradiction and uncertainty markers.
- Separation of durable information from transient discussion.
- Canonical Meeting Knowledge output.
Prerequisites:
- Deterministic Canonicalizer output with stable IDs and source references.
- Gold or benchmark cases that expose duplication and category shifts.
Out of scope:
- Direct protocol writing.
- Deriving output views from one another.
- Retrieval or RAG integration.
## Phase 3 - Canonical Meeting Knowledge
## Phase 4 - Canonical Meeting Knowledge
Goal:
@@ -70,7 +99,7 @@ Deliverables:
Prerequisites:
- Consolidation behavior that preserves evidence and uncertainty.
- Semantic consolidation behavior that preserves evidence and uncertainty.
- Agreement on required semantic categories.
Out of scope:
@@ -79,7 +108,7 @@ Out of scope:
- Export formats beyond those needed to validate the model.
- Knowledge-system storage design.
## Phase 4 - Output Views
## Phase 5 - Output Views
Goal:
@@ -91,8 +120,11 @@ Deliverables:
- Working Protocol / Arbeitsprotokoll renderer.
- Distribution Protocol / Verteilerprotokoll renderer.
- Knowledge Objects / Wissensdatenbankeintrag renderer or structured export.
- Later additional views such as action lists.
- Tests or checks showing that output views are parallel renderings of the same
canonical model.
- Default output-language policy: rendered protocols normally match the
dominant source language unless explicitly requested otherwise.
Prerequisites:
@@ -105,7 +137,7 @@ Out of scope:
- Deriving one output view from another.
- Retrieval integration.
## Phase 5 - Review and Quality Control
## Phase 6 - Review and Quality Control
Goal:
@@ -122,7 +154,7 @@ Deliverables:
Prerequisites:
- Stable extraction, consolidation and canonical model.
- Stable extraction, semantic consolidation and canonical model.
- Representative test meetings.
Out of scope:
@@ -131,7 +163,7 @@ Out of scope:
- Product UI work.
- Cloud deployment.
## Phase 6 - Productization
## Phase 7 - Productization
Goal:
@@ -158,7 +190,7 @@ Out of scope:
- Future Meeting Assistant integration beyond export contracts.
- Cloud-first architecture.
## Phase 7 - Knowledge-System Integration
## Phase 8 - Knowledge-System Integration
Goal:
@@ -182,4 +214,3 @@ Out of scope:
- Building a full enterprise search product inside Meeting Lab.
- Treating raw transcripts or generated protocols as the knowledge source of
truth.
+35 -21
View File
@@ -98,7 +98,9 @@ Topic Segmentation
↓
Specialized Extraction
↓
Consolidation
Deterministic Canonicalization
↓
Semantic Consolidation
↓
Canonical Meeting Knowledge
↓
@@ -174,14 +176,32 @@ Each extractor has exactly one task and one prompt.
## consolidation/
Merges information extracted from multiple discussion segments.
Planned area for canonicalization and consolidation.
Typical responsibilities:
The next milestone splits this into two stages.
- merge duplicates
- combine partial information
- distinguish positions from decisions
- detect contradictions
Deterministic Canonicalizer:
- implemented in Python
- uses no LLM
- validates and normalizes extraction objects
- assigns stable source references and IDs
- normalizes category names and basic field structure
- performs only safe deterministic cleanup
- may group exact duplicates
- preserves all source evidence
- must not perform uncertain semantic merging
Semantic Consolidator:
- uses the local LLM
- merges semantically equivalent statements
- groups content by topic
- preserves evidence from all contributing chunks
- marks contradictions and uncertainty
- separates durable information from transient discussion
- produces Canonical Meeting Knowledge
- does not directly write a protocol
---
@@ -207,6 +227,10 @@ The planned output products are:
These are parallel renderings of the same canonical semantic model, not
documents derived from one another.
Rendered protocol language should normally match the dominant language of the
source transcript or consolidated meeting knowledge unless an explicit output
language is requested.
---
# Repository Layout
@@ -257,20 +281,10 @@ Views documented in `output-views.md`.
# Next Milestone
The next development step is the implementation of **topic segmentation**.
Its only responsibility is to identify the thematic structure of a discussion.
It should answer questions such as:
- Where does a topic begin?
- Where does it end?
- When does another topic start?
- When is an earlier topic resumed?
No facts, decisions or todos should be extracted at this stage.
Only after reliable topic segmentation has been achieved will the specialized extraction modules be implemented.
The next architecture milestone is the implementation of a deterministic
canonicalization stage followed by a semantic consolidation stage. These stages
convert raw chunk extraction JSON into evidence-preserving Canonical Meeting
Knowledge before any Output View renderer writes a protocol.
---
+64
View File
@@ -242,11 +242,74 @@ This object feeds the Canonical Meeting Knowledge representation.
---
# Canonical Extraction Object
Planned deterministic intermediate object created from raw chunk extraction
JSON.
Example:
```json
{
"id": "fact.chunk_03.0001",
"category": "fact",
"text": "...",
"source_references": [
{
"chunk_id": "chunk_03",
"source_file": "chunk_03_extraction.json",
"evidence": "..."
}
]
}
```
The Deterministic Canonicalizer should create this kind of object without an
LLM. It validates and normalizes raw extraction objects, assigns stable IDs and
source references, normalizes category names and basic field structure,
performs only safe deterministic cleanup, may group exact duplicates and must
preserve all source evidence.
It must not perform uncertain semantic merging.
---
# Consolidated Topic
Planned semantic object produced by the Semantic Consolidator.
Example:
```json
{
"topic_id": "topic_001",
"title": "...",
"background": [],
"decisions": [],
"action_items": [],
"open_questions": [],
"durable_information": [],
"uncertainty": [],
"source_references": []
}
```
The Semantic Consolidator may use the local LLM to merge semantically
equivalent statements, group content by topic, preserve evidence from all
contributing chunks, mark contradictions and uncertainty and separate durable
information from transient discussion.
It produces Canonical Meeting Knowledge. It does not directly write a protocol.
---
# Canonical Meeting Knowledge
The canonical semantic representation of one meeting.
This representation is the single source of truth for all downstream outputs.
It is a structured representation, preferably JSON, and is not itself a prose
protocol.
```json
{
@@ -260,6 +323,7 @@ This representation is the single source of truth for all downstream outputs.
"questions": [],
"positions": [],
"technical_details": [],
"durable_information": [],
"rationale": [],
"uncertainty": [],
"source_references": []
+73 -5
View File
@@ -945,14 +945,18 @@ Result:
The accepted design is a planned Canonical Meeting Knowledge layer as the
semantic source of truth, with Working Protocol, Distribution Protocol and
Knowledge Objects as parallel output views. Consolidation must merge duplicates,
preserve evidence, reconcile category shifts and mark contradictions or
uncertainty.
Knowledge Objects as parallel output views. The next consolidation architecture
is split into a Deterministic Canonicalizer and a Semantic Consolidator. The
canonicalizer prepares validated evidence-bearing objects without uncertain
semantic merging. The consolidator then merges semantically equivalent
statements, preserves evidence, reconciles category shifts where supported and
marks contradictions or uncertainty.
Decision:
Consolidation is the next major engineering step after stable local extraction.
Canonical Meeting Knowledge and final output views are planned, not implemented.
Deterministic canonicalization is the next implementation step after stable
local extraction, followed by semantic consolidation. Canonical Meeting
Knowledge and final output views are planned, not implemented.
Lessons learned:
@@ -969,3 +973,67 @@ Evidence:
- `PROJECT_KNOWLEDGE.md`
- `ROADMAP.md`
- Commit `5c03ed7`
## EXP-0020 - Working Protocol Synthesizer V0
Status: Accepted
Date or period: 2026-07-31
Hypothesis:
The current local synthesis model may be able to generate a useful detailed
Working Protocol directly from the existing independent chunk extraction JSON
files, before canonicalization or semantic consolidation exists.
Setup:
One synthesis prompt was constructed from exactly nine chunk extraction JSON
files. The model was instructed to use only those extraction files, merge
duplicates, group related information into topics, preserve useful discussion
context and write a neutral technical Working Protocol.
Inputs:
- `chunk_01_extraction.json` through `chunk_09_extraction.json`.
- No original transcript, normalized chunks or Whisper output were used as
synthesis input.
Model / configuration:
- Model: `qwen3.5:9b`
- Prompt characters: 24,979
- Actual prompt eval tokens: 5,929
- Output tokens: 1,486
- Runtime: 294.204 seconds
Result:
The generated Working Protocol was readable, well structured and
topic-oriented. It was still based directly on raw chunk extractions, without a
separate deterministic canonicalization stage or semantic consolidation stage.
The output language was English even though the source meeting material was
German.
Decision:
Preserve this output as the Working Protocol Synthesizer V0 benchmark baseline
for later canonicalizer, consolidator and renderer comparisons. This selected
generated artifact is intentionally versioned even though generated runtime
artifacts are normally ignored.
Lessons learned:
Direct synthesis from chunk extractions can create a useful recall-oriented
draft, but it does not replace Canonical Meeting Knowledge. The language
mismatch also establishes a default renderer rule: protocol output should
normally match the dominant source language unless an explicit output language
is requested.
Evidence:
- `samples/benchmarks/working_protocol_synthesizer_v0/README.md`
- `samples/benchmarks/working_protocol_synthesizer_v0/working_protocol.md`
- `PROJECT_KNOWLEDGE.md`
- `docs/output-views.md`
- See EXP-0017 and EXP-0019.
+20
View File
@@ -13,6 +13,10 @@ Meeting transcript
->
chunk extraction
->
Deterministic Canonicalizer
->
Semantic Consolidator
->
Canonical Meeting Knowledge
├── Working Protocol
├── Distribution Protocol
@@ -45,6 +49,18 @@ The detailed schema is future implementation work. The current implementation
still uses simple extraction JSON files and a basic Markdown protocol builder for
technical validation.
The next planned architecture stage before this representation is explicit:
- The Deterministic Canonicalizer validates and normalizes extraction objects,
assigns stable source references and IDs, performs only safe deterministic
cleanup and preserves all source evidence. It uses no LLM and must not make
uncertain semantic merges.
- The Semantic Consolidator uses the local LLM to merge semantically equivalent
statements, group content by topic, preserve evidence from all contributing
chunks, mark contradictions and uncertainty, separate durable information from
transient discussion and produce Canonical Meeting Knowledge. It does not
directly write a protocol.
## Renderers
Each Output View is produced by a renderer.
@@ -57,6 +73,10 @@ Depending on the implementation, rendering may be:
The architecture does not assume that every renderer must always use an LLM.
Rendered protocol language should normally match the dominant language of the
source transcript or consolidated meeting knowledge unless an explicit output
language is requested.
## Working Protocol
Suggested filename: `working_protocol.md`
+80 -19
View File
@@ -33,7 +33,10 @@ Topic Segmentation
Specialized Extraction
│
▼
Consolidation
Deterministic Canonicalization
│
▼
Semantic Consolidation
│
▼
Canonical Meeting Knowledge
@@ -267,35 +270,41 @@ Prototype exists as a combined extractor.
---
# Stage 6 – Consolidation
# Stage 6 – Deterministic Canonicalization
## Purpose
Merge analysis results originating from different discussion segments.
Normalize raw chunk extraction JSON into stable canonical extraction objects
without changing uncertain semantics.
## Input
Extraction results.
Chunk extraction JSON files.
## Output
Unified topic representation.
Validated canonical extraction objects with stable source references and IDs.
## Responsibilities
- Merge duplicates
- Merge complementary information
- Preserve contradictions
- Separate positions from decisions
- Combine related todos
- Validate extraction objects
- Normalize category names
- Normalize basic field structure
- Assign stable source references and IDs
- Preserve all source evidence
- Perform only safe deterministic cleanup
- Group exact duplicates where unambiguous
## Must Not
- Use an LLM
- Perform uncertain semantic merging
- Infer missing information
- Drop source evidence
## Processing Type
Hybrid
Deterministic wherever possible.
LLM support only if necessary.
Deterministic Python
## Current Status
@@ -303,13 +312,55 @@ Planned
---
# Stage 7 – Canonical Meeting Knowledge
# Stage 7 – Semantic Consolidation
## Purpose
Merge canonicalized extraction objects into an evidence-preserving semantic
meeting representation.
## Input
Canonicalized extraction objects.
## Output
Canonical Meeting Knowledge.
## Responsibilities
- Merge semantically equivalent statements
- Group content by topic
- Preserve evidence from all contributing chunks
- Mark contradictions and uncertainty
- Separate durable information from transient discussion
- Reconcile category shifts where supported by evidence
## Must Not
- Directly write a protocol
- Invent information
- Drop conflicting evidence silently
## Processing Type
Local LLM, with deterministic pre/post-processing where useful.
## Current Status
Planned
---
# Stage 8 – Canonical Meeting Knowledge
## Purpose
Produce the canonical semantic representation of one meeting.
This representation is the single source of truth for all downstream outputs.
It is a structured representation, preferably JSON, and is not itself a prose
protocol.
Example:
@@ -352,7 +403,7 @@ Planned
---
# Stage 8 – Output View Rendering
# Stage 9 – Output View Rendering
## Purpose
@@ -365,6 +416,7 @@ The planned output products are:
- Knowledge Objects, rendered as a Knowledge-base Entry (`knowledge_entry.md`)
and later stored in a structured format such as `knowledge_entry.json`
(Wissensdatenbankeintrag)
- Later additional views such as action lists
Output rendering never performs additional analysis.
@@ -383,6 +435,10 @@ Completeness differs by output:
- The Distribution Protocol optimizes for relevance and brevity.
- Knowledge Objects optimize for durability and reuse.
Rendered protocol language should normally match the dominant language of the
source transcript or consolidated meeting knowledge unless an explicit output
language is requested.
## Processing Type
LLM
@@ -451,11 +507,16 @@ A processing stage may be replaced by another implementation as long as it prese
⬜ Specialized Extraction
⬜ Consolidation
⬜ Deterministic Canonicalization
⬜ Semantic Consolidation
⬜ Canonical Meeting Knowledge
⬜ Output View Rendering
```
The immediate development focus is **Topic Segmentation**, as it provides the semantic structure on which all subsequent processing stages depend.
The immediate architecture focus is the Deterministic Canonicalizer followed by
the Semantic Consolidator. These stages preserve source evidence, recover
global context from independent chunk extractions and prepare Canonical Meeting
Knowledge for parallel Output View rendering.
@@ -0,0 +1,32 @@
# Working Protocol Synthesizer V0
## Summary
- Experiment name: Working Protocol Synthesizer V0
- Source: nine independent chunk extraction JSON files
- Model: `qwen3.5:9b`
- Prompt characters: 24,979
- Actual prompt eval tokens: 5,929
- Output tokens: 1,486
- Runtime: 294.204 seconds
- Output language: English
- Purpose: baseline for later canonicalizer/consolidator comparisons
## Known Characteristics
- Readable and well structured.
- Topic-oriented synthesis.
- Still based directly on raw chunk extractions.
- No separate canonicalization or semantic consolidation.
- Output language did not match the German source material.
## Versioning Decision
Generated runtime artifacts are normally ignored in this repository. This
selected output is intentionally versioned as a deliberate benchmark reference
artifact so future deterministic canonicalizer, semantic consolidator and
renderer experiments can be compared against the pre-consolidation baseline.
Do not treat this Markdown file as canonical source data. It is a preserved
experiment result.
@@ -0,0 +1,69 @@
# Working Protocol: Project Idea Intake and Process Adaptation
## Overview
The meeting focused on adapting the existing R&D (F&E) project process to accommodate Business Development (BD), Marketing, and Product Management projects. The primary objective was to define a flexible framework that avoids over-engineering for a mid-sized company while establishing clear criteria for filtering ideas before they enter formal development pools. Key outcomes include defining minimum requirements for new ideas, assigning responsibility for initial screening, and agreeing on the structure of project documentation (Project Deck Sheet vs. Lastenheft) based on project type.
## Topics
### 1. Project Intake Requirements and Documentation
**Background:**
The current process is historically an R&D-focused workflow involving a "Project Deck Sheet" used in F&E departments. It was clarified that this document is not intended to be a filter system but rather a concise summary of the project idea, goal, counterpart, development hurdles (e.g., patents), and budget inquiry. The use of Gantt diagrams on every deck sheet is deemed unnecessary for all projects; they are only required when specifically needed during search or planning phases.
**Decisions:**
* **Minimum Requirements for New Ideas:** A new idea must minimally answer: Title, Project Idea (a few sentences), and Project Goal. Optionally, underlying assumptions (e.g., expected market volume) may be listed but are not mandatory initially.
* **Documentation Adaptation by Type:** The general process flow remains consistent across project types (F&E, PM, Marketing, BD). However, specific documents vary:
* F&E Projects use the standard Project Deck Sheet and Lastenheft structure.
* Business Development (BD) projects require a Project Deck Sheet and a Lastenheft defining market targets.
* The process steps for different types are left open in Swimlane diagrams to accommodate various perspectives, rather than being formally rigidly defined at this stage.
**Action Items:**
* **Giovanna:** Compile criteria and create the first draft of the selection catalog integrating F&E, Marketing, and PM criteria.
* **Björn:** Define specific criteria for Business Development projects (e.g., digital affinity, cultural hurdles) from a marketing perspective and integrate these into the filter.
**Open Questions:**
* How will specific selection criteria be defined and integrated for digital products (e.g., Portals) versus physical real products? Criteria are acknowledged as never being general but always specific to the context.
* Exactly how should requirements for the Lastenheft in non-F&E projects be defined, allowing departments autonomy if they have different needs?
### 2. Filtering and Screening Process
**Background:**
Ideas must undergo pre-processing before F&E or other units consider them. It is noted that secretariats are often dismissive of vague requests (e.g., just a keyword), so ideas need to be enriched with minimal requirements beforehand. The network connection at municipalities is currently very slow, which may impact communication speeds but does not alter the process logic.
**Decisions:**
* **Pre-Filtering Responsibility:** There was discussion regarding whether pre-filtering should be centralized or decentralized. While some hesitation exists about placing this responsibility solely within F&E (specifically David and James), there is agreement that someone must handle initial sorting and data collection formally acting as a "Gatekeeper." Björn expressed willingness to let others take over the preliminary filtering if desired, provided it happens.
* **Filter Criteria Definition:** The group agreed not to plan everything in advance but rather adapt the existing framework based on needs observed after 2-4 meetings or 10-15 projects. Filter criteria must be defined for different cases (new products/applications vs. existing markets).
**Action Items:**
* **Björn & PM/BD:** Discuss specific project examples (e.g., Papatikus-Anker) and route them to the respective specialist departments.
* **Gatekeepers (James/David):** Formally process a few items and gather information at the entry point of the pool, though full resolution is not expected immediately.
**Open Questions:**
* At which stage does a department decide independently versus when it must refer an idea to the central "pot" or management? It remains unclear exactly where projects enter this pot and where they do not.
* Are we convinced that specific market studies should be initiated based on these filtered ideas?
### 3. Project Classification and Management Levels
**Background:**
The process distinguishes between tasks handled at the team/department level versus those requiring cross-departmental or strategic relevance/budget approval. The R&D "pot" is jointly managed and not fragmented. It was clarified that F&E leadership maintains an overview of projects during bi-weekly interface stand-ups to categorize them as Research (F), Development, PM, or BD projects at the time they are presented on this list.
**Decisions:**
* **Project Categorization:** At the point where a project is listed and categorized by F&E leadership, it has already passed certain pre-filters. The processing of these projects must occur exclusively within the specialist departments (Fachabteilungen).
* **Management Scope:** Management does not decide on the success or failure of specific marketing projects but maintains oversight of the overall portfolio.
**Action Items:**
* Define a few "checkpoints" (Prüfsteine) to determine if an idea is cross-departmental/strategic, requires budget, or can be handled internally by a team/unit. If it does not meet these criteria, it returns to the department level.
### 4. Process Implementation and Governance
**Background:**
The goal is to establish a process where filters are securely queried without creating complex additional processes for a mid-sized company. The implementation of this new workflow relies on execution rather than perfect upfront planning. A Miro board containing relevant links (e.g., from Malte) was discussed as a resource for the team, specifically shared with Henning and others via chat.
**Decisions:**
* **Process Flexibility:** Agree to utilize the existing process framework and adapt it to needs without creating complex additional processes suitable only for larger enterprises. The group agreed on this approach collectively.
* **Filter Nature:** The first filter is defined as a decision proposal (Entscheidungsvorlage) rather than a final binding decision, allowing flexibility before formal review by the committee or leadership.
**Action Items:**
* **Malte:** Copy and distribute the Miro link to all relevant persons in the chat.
* **Giovanna & Björn:** Await their feedback on the proposed criteria catalog; once they agree, the discussion is considered concluded for now.
* **Raik:** Must ensure ideas are entered into the pool with proper context (not just keywords) when contacted or instructed to do so by others.
**Open Questions:**
* Which checkpoints are relevant specifically for the specialist departments independent of the main process?
* Who presents which project list in which committee, and who is responsible for setting up the necessary checks or creating a project plan at that specific term?