Files
meeting-lab/PROJECT_KNOWLEDGE.md
T
admin 23bbc744f7 Document canonicalization and consolidation milestone
- preserve Working Protocol Synthesizer V0 as comparison baseline
- introduce deterministic canonicalization stage
- define semantic consolidator responsibilities
- clarify Canonical Meeting Knowledge generation
- document source-language output policy
- align roadmap, architecture and experiment log
2026-07-31 09:32:12 +02:00

181 lines
7.5 KiB
Markdown

# Project Knowledge
This is a compact operational summary of the current Meeting Lab state.
## Objective
Meeting Lab develops and evaluates local methods for extracting structured
organizational knowledge from real meeting recordings and transcripts. The
project is a research and validation environment for a future Meeting
Assistant, not a finished product.
## Implemented Pipeline Stages
Implemented:
- Whisper JSON cleanup via `scripts/clean_whisper_json.py`.
- Transcript normalization in `src/meeting_lab/normalization/`.
- Technical chunking in `src/meeting_lab/chunking/`.
- Local per-chunk extraction in `src/meeting_lab/extraction/extract_chunks.py`.
- Prompt loading from `src/meeting_lab/llm/prompts.py`.
- Interim Markdown protocol generation in `src/meeting_lab/protocol/`.
- Non-LLM unit tests for chunking, extraction helpers, protocol rendering and
gold-test runner validation.
Experimental/prototype:
- Topic segmentation in `src/meeting_lab/segmentation/`.
- Windowed segmentation and review output in `samples/chunks/`.
- Gold Standard extraction corpus under `tests/gold/`.
Planned:
- Deterministic Canonicalizer for extraction results.
- Semantic Consolidator for evidence-preserving semantic merging.
- Canonical Meeting Knowledge implementation as the semantic source of truth.
- Final Working Protocol / Arbeitsprotokoll, Distribution Protocol /
Verteilerprotokoll and Knowledge Objects / Wissensdatenbankeintrag renderers.
## Source Tree
```text
src/meeting_lab/
chunking/ technical transcript chunking
consolidation/ planned canonicalization/consolidation area
extraction/ current local LLM extraction flow
io/ lightweight file and JSON helpers
llm/ Ollama and prompt support
models/ current lightweight model definitions
normalization/ deterministic transcript cleanup
protocol/ interim Markdown protocol builder
segmentation/ experimental topic segmentation tooling
```
Supporting areas:
- `docs/`: architecture, pipeline, data models and output-view concepts.
- `prompts/`: active prompt files. Only `common.md` and `decisions.md` contain
substantive extraction prompt text in the current tree.
- `tests/gold/`: semantic gold tests and prompt-engineering methodology.
- `samples/`: sample inputs and generated or experimental artifacts.
- `scripts/`: operational scripts for cleanup and gold-test execution.
## Current Model Strategy
The current extraction strategy is one normalized chunk per LLM call. This is
preferred over expanding context windows or asking one model call to analyze a
full meeting.
Known working models from current project notes and experiment practice:
- `qwen3:1.7b`: useful for smoke tests.
- `qwen3.5:9b`: useful for meaningful extraction and segmentation work.
LLM calls use Ollama locally. The current extractor defaults to `qwen3:8b`, but
validated work may specify another model explicitly.
## Important Findings
- Whisper JSON chunking must use `segments[*].text`, not only the top-level
`text` field.
- Independent chunk extraction is currently preferred.
- Larger context windows can change classification behavior and increase
instability.
- Extraction and consolidation are separate problems.
- Generation limits can truncate JSON.
- Qwen thinking may be returned separately by the Ollama API.
- Gold Standard tests are also a formal specification of meeting semantics.
- Raw model responses should be preserved when diagnosing parser or truncation
failures.
## Decision Taxonomy
Accepted decision semantics:
- A decision is an explicit agreement that creates a binding change in action,
process, responsibility, approval status, timing or next step.
- Included: substantive decisions, organizational decisions, process decisions,
approvals, rejections, deferrals, explicit agreement not to decide yet, and
explicit agreement to gather more information before deciding.
- Excluded: opinions, preferences, proposals without agreement, open questions,
current-state descriptions and explanations without commitment.
- A process decision to defer a substantive decision is still a decision.
- "No decision was reached" is different from "the group decided to defer the
decision."
Current Prompt Version 2 decision baseline:
- `decision_simple`: passing.
- `decision_deferred`: passing.
- `decision_none`: passing.
- Prompt Version 2 explicitly supports process decisions where the group agrees
to defer a substantive decision until more information is available.
## Canonical Knowledge Architecture
The next documented pipeline milestone is:
```text
Chunk Extractions
-> Deterministic Canonicalizer
-> Semantic Consolidator
-> Canonical Meeting Knowledge
-> Output View Renderers
```
The Deterministic Canonicalizer is planned Python code with no LLM. It should
validate and normalize extraction objects, assign stable source references and
IDs, normalize category names and basic field structure, perform only safe
deterministic cleanup, optionally group exact duplicates, and preserve all
source evidence. It must not perform uncertain semantic merging.
The Semantic Consolidator is planned local-LLM work. It should merge
semantically equivalent statements, group content by topic, preserve evidence
from all contributing chunks, mark contradictions and uncertainty, separate
durable information from transient discussion, and produce Canonical Meeting
Knowledge. It does not directly write a protocol.
Canonical Meeting Knowledge is the planned semantic intermediate model and
future single source of truth. It should be structured, preferably JSON, and
preserve topics, facts, decisions, action items, open questions, positions,
technical details, rationale, uncertainty, contradictions and source evidence.
It is not itself a prose protocol.
Output views are planned as independent renderings from that canonical model:
- Working Protocol / Arbeitsprotokoll: relatively complete, optimized for
recall and traceability.
- Distribution Protocol / Verteilerprotokoll: concise and outcome-oriented,
optimized for circulation.
- Knowledge Objects / Wissensdatenbankeintrag: durable organizational knowledge
optimized for reuse.
The current `meeting_protocol.md` builder is an interim technical validation
tool, not the final output-view architecture.
Rendered protocol output should normally use the dominant language of the
source transcript or consolidated meeting knowledge unless an explicit output
language is requested.
## Current Limitations
- Discussion Blocks are documented as a stable semantic unit but are not yet a
separate implemented pipeline artifact.
- Topic segmentation exists as prototype tooling, not a stable pipeline stage.
- Extraction is still a combined current flow, even though separate extractors
are the intended architecture.
- Deterministic Canonicalizer is documented but not implemented.
- Semantic Consolidator is documented but not implemented.
- Canonical Meeting Knowledge is documented but not implemented.
- Final output views are documented but not implemented.
- Most prompt files are placeholders except the common and decision prompts.
- Gold tests currently emphasize extraction semantics, especially decisions.
## Next Recommended Engineering Step
Stabilize repeatable local extraction evaluation before broadening the pipeline:
expand Gold Standard coverage by category, keep one-chunk extraction as the
baseline, and use small prompt experiments with immediate non-regression checks.
After extraction behavior is stable enough, implement the Deterministic
Canonicalizer first, then the Semantic Consolidator with evidence retention.