Files
meeting-lab/PROJECT_KNOWLEDGE.md
T
admin 90aa34d5d0 Implement Semantic Consolidator V0
- add deterministic canonicalization support for extraction items
- add facts-only semantic consolidation using local Ollama
- preserve source evidence and validate complete fact coverage
- add conservative merge rules and non-LLM tests
- record the first validated real-life consolidation benchmark
- document current scope, limitations and next evaluation step
2026-07-31 11:25:46 +02:00

8.6 KiB

Project Knowledge

This is a compact operational summary of the current Meeting Lab state.

Objective

Meeting Lab develops and evaluates local methods for extracting structured organizational knowledge from real meeting recordings and transcripts. The project is a research and validation environment for a future Meeting Assistant, not a finished product.

Implemented Pipeline Stages

Implemented:

  • Whisper JSON cleanup via scripts/clean_whisper_json.py.
  • Transcript normalization in src/meeting_lab/normalization/.
  • Technical chunking in src/meeting_lab/chunking/.
  • Local per-chunk extraction in src/meeting_lab/extraction/extract_chunks.py.
  • Deterministic Canonicalizer V1 in src/meeting_lab/consolidation/canonicalize.py.
  • Semantic Consolidator V0 in src/meeting_lab/consolidation/consolidate_facts.py for facts-only semantic duplicate detection.
  • Prompt loading from src/meeting_lab/llm/prompts.py.
  • Interim Markdown protocol generation in src/meeting_lab/protocol/.
  • Non-LLM unit tests for chunking, extraction helpers, protocol rendering and gold-test runner validation.

Experimental/prototype:

  • Topic segmentation in src/meeting_lab/segmentation/.
  • Windowed segmentation and review output in samples/chunks/.
  • Gold Standard extraction corpus under tests/gold/.

Planned:

  • Broader semantic consolidation for topic grouping, contradiction handling, uncertainty marking and durable/transient separation.
  • Canonical Meeting Knowledge implementation as the semantic source of truth.
  • Final Working Protocol / Arbeitsprotokoll, Distribution Protocol / Verteilerprotokoll and Knowledge Objects / Wissensdatenbankeintrag renderers.

Source Tree

src/meeting_lab/
  chunking/        technical transcript chunking
  consolidation/   planned canonicalization/consolidation area
  extraction/      current local LLM extraction flow
  io/              lightweight file and JSON helpers
  llm/             Ollama and prompt support
  models/          current lightweight model definitions
  normalization/   deterministic transcript cleanup
  protocol/        interim Markdown protocol builder
  segmentation/    experimental topic segmentation tooling

Supporting areas:

  • docs/: architecture, pipeline, data models and output-view concepts.
  • prompts/: active prompt files. Only common.md and decisions.md contain substantive extraction prompt text in the current tree.
  • tests/gold/: semantic gold tests and prompt-engineering methodology.
  • samples/: sample inputs and generated or experimental artifacts.
  • scripts/: operational scripts for cleanup and gold-test execution.

Current Model Strategy

The current extraction strategy is one normalized chunk per LLM call. This is preferred over expanding context windows or asking one model call to analyze a full meeting.

Known working models from current project notes and experiment practice:

  • qwen3:1.7b: useful for smoke tests.
  • qwen3.5:9b: useful for meaningful extraction and segmentation work.

LLM calls use Ollama locally. The current extractor defaults to qwen3:8b, but validated work may specify another model explicitly.

Important Findings

  • Whisper JSON chunking must use segments[*].text, not only the top-level text field.
  • Independent chunk extraction is currently preferred.
  • Larger context windows can change classification behavior and increase instability.
  • Extraction and consolidation are separate problems.
  • Generation limits can truncate JSON.
  • Qwen thinking may be returned separately by the Ollama API.
  • Gold Standard tests are also a formal specification of meeting semantics.
  • Raw model responses should be preserved when diagnosing parser or truncation failures.

Decision Taxonomy

Accepted decision semantics:

  • A decision is an explicit agreement that creates a binding change in action, process, responsibility, approval status, timing or next step.
  • Included: substantive decisions, organizational decisions, process decisions, approvals, rejections, deferrals, explicit agreement not to decide yet, and explicit agreement to gather more information before deciding.
  • Excluded: opinions, preferences, proposals without agreement, open questions, current-state descriptions and explanations without commitment.
  • A process decision to defer a substantive decision is still a decision.
  • "No decision was reached" is different from "the group decided to defer the decision."

Current Prompt Version 2 decision baseline:

  • decision_simple: passing.
  • decision_deferred: passing.
  • decision_none: passing.
  • Prompt Version 2 explicitly supports process decisions where the group agrees to defer a substantive decision until more information is available.

Canonical Knowledge Architecture

The next documented pipeline milestone is:

Chunk Extractions
  -> Deterministic Canonicalizer
  -> Semantic Consolidator
  -> Canonical Meeting Knowledge
  -> Output View Renderers

Canonicalizer V1 is implemented Python code with no LLM. It validates and normalizes extraction objects, assigns stable source references and IDs, normalizes category names and basic field structure, performs only safe deterministic cleanup, optionally groups exact duplicates, and preserves all source evidence. It must not perform uncertain semantic merging.

CLI:

PYTHONPATH=src .venv/bin/python -m meeting_lab.consolidation.canonicalize \
  samples/whisper/meeting_speech_cleaned_chunks \
  -o /tmp/canonicalized_extractions.json

Semantic Consolidator V0 is implemented as narrow local-LLM work for fact items only. It merges semantically equivalent fact statements, preserves source references and evidence, and prefers false negatives over false-positive merges. It is not a summarizer, topic grouper, protocol renderer or complete Canonical Meeting Knowledge stage.

The first accepted V0 benchmark used qwen3.5:9B in one Ollama call over 33 fact items. Runtime on the current machine was 390.119 seconds. One correct merge was accepted, involving fact_0025 and fact_0031; 31 facts remained singletons, validation passed, no source fact was lost or duplicated, and non-fact categories remained unchanged. This is a local benchmark, not a general hardware claim.

Broader semantic consolidation remains planned. It should group content by topic, mark contradictions and uncertainty, separate durable information from transient discussion, and produce Canonical Meeting Knowledge. It does not directly write a protocol.

Canonical Meeting Knowledge is the planned semantic intermediate model and future single source of truth. It should be structured, preferably JSON, and preserve topics, facts, decisions, action items, open questions, positions, technical details, rationale, uncertainty, contradictions and source evidence. It is not itself a prose protocol.

Output views are planned as independent renderings from that canonical model:

  • Working Protocol / Arbeitsprotokoll: relatively complete, optimized for recall and traceability.
  • Distribution Protocol / Verteilerprotokoll: concise and outcome-oriented, optimized for circulation.
  • Knowledge Objects / Wissensdatenbankeintrag: durable organizational knowledge optimized for reuse.

The current meeting_protocol.md builder is an interim technical validation tool, not the final output-view architecture.

Rendered protocol output should normally use the dominant language of the source transcript or consolidated meeting knowledge unless an explicit output language is requested.

Current Limitations

  • Discussion Blocks are documented as a stable semantic unit but are not yet a separate implemented pipeline artifact.
  • Topic segmentation exists as prototype tooling, not a stable pipeline stage.
  • Extraction is still a combined current flow, even though separate extractors are the intended architecture.
  • Semantic Consolidator V0 is implemented only for facts-only duplicate detection.
  • Canonical Meeting Knowledge is documented but not implemented.
  • Final output views are documented but not implemented.
  • Most prompt files are placeholders except the common and decision prompts.
  • Gold tests currently emphasize extraction semantics, especially decisions.

Stabilize repeatable local extraction evaluation before broadening the pipeline: expand Gold Standard coverage by category, keep one-chunk extraction as the baseline, and use small prompt experiments with immediate non-regression checks. The next recommended evaluation step is to use the consolidated V0 result as input for the unchanged Working Protocol renderer and compare that output against the Working Protocol Synthesizer V0 baseline and the human reference protocol.