Files
meeting-lab/docs/architecture.md
T
admin 90aa34d5d0 Implement Semantic Consolidator V0
- add deterministic canonicalization support for extraction items
- add facts-only semantic consolidation using local Ollama
- preserve source evidence and validate complete fact coverage
- add conservative merge rules and non-LLM tests
- record the first validated real-life consolidation benchmark
- document current scope, limitations and next evaluation step
2026-07-31 11:25:46 +02:00

7.0 KiB

Architecture

Purpose

The Meeting Lab is an experimental environment for developing and evaluating methods to extract structured knowledge from real meeting transcripts.

Its purpose is not to build a complete meeting assistant, but to answer a single question:

How can knowledge be extracted from real discussions as reliably as possible?

Successful approaches will later be integrated into the Meeting Assistant project.


Design Goals

The architecture follows a small set of guiding principles.

Modular Pipeline

Complex problems are divided into small, well-defined processing steps.

Each module has exactly one responsibility.

Deterministic where possible

Tasks that can be solved reliably without an LLM should use deterministic algorithms.

Examples include:

  • transcript normalization
  • whitespace cleanup
  • duplicate removal
  • chunk generation

LLMs are only used where semantic understanding is required.

Preserve Information

The pipeline should never remove or rewrite information unless it is certain that the content is merely noise.

Losing information is considered worse than keeping harmless redundancy.

Explainable Results

Every processing step should be understandable.

Intermediate results should remain inspectable throughout the pipeline.

Reproducible Experiments

Experiments must be repeatable.

Given the same input, prompt, model and parameters, another developer should be able to reproduce the result.

Local First

The complete pipeline should run locally.

Cloud services may be supported in the future but are not a design requirement.


Core Idea

Traditional meeting summarization attempts to solve everything in one step.

Transcript
    ↓
LLM
    ↓
Summary

Real discussions do not work that way.

Topics are introduced, interrupted, resumed later, expanded, questioned and finally concluded.

Instead of building a better summarizer, the Meeting Lab develops a Discussion Analyzer.

The analyzer gradually transforms an unstructured discussion into structured knowledge.


High-Level Pipeline

Transcript
    ↓
Normalization
    ↓
Discussion Blocks
    ↓
Technical Chunking
    ↓
Topic Segmentation
    ↓
Specialized Extraction
    ↓
Deterministic Canonicalization
    ↓
Semantic Consolidation
    ↓
Canonical Meeting Knowledge
    ↓
Output View Rendering
    ↓
Working Protocol / Distribution Protocol / Knowledge Objects

Each stage solves one clearly defined problem.

No module should perform multiple semantic tasks simultaneously.


Module Overview

The current architecture consists of the following processing stages.

normalization/

Deterministic transcript cleanup.

Responsibilities:

  • remove filler words
  • remove immediate repetitions
  • whitespace cleanup
  • generate change log

chunking/

Creates model-sized chunks.

Chunking is purely technical.

It does not recognize discussion topics.


segmentation/

Identifies discussion topics.

Responsibilities:

  • detect topic start
  • detect topic end
  • detect topic switches
  • recognize resumed topics

This is the next major development milestone.


extraction/

Contains specialized LLM modules.

Planned extractors include:

  • facts
  • questions
  • positions
  • decisions
  • todos
  • technical information

Each extractor has exactly one task and one prompt.


consolidation/

Planned area for canonicalization and consolidation.

The next milestone splits this into two stages.

Deterministic Canonicalizer:

  • implemented in Python
  • uses no LLM
  • validates and normalizes extraction objects
  • assigns stable source references and IDs
  • normalizes category names and basic field structure
  • performs only safe deterministic cleanup
  • may group exact duplicates
  • preserves all source evidence
  • must not perform uncertain semantic merging

Semantic Consolidator:

  • uses the local LLM
  • V0 is implemented for facts-only semantic duplicate detection
  • V0 merges semantically equivalent fact items conservatively
  • V0 preserves source references and evidence
  • V0 validates that every source fact appears exactly once
  • V0 does not process non-fact categories semantically
  • later versions should group content by topic, mark contradictions and uncertainty, separate durable information from transient discussion and prepare Canonical Meeting Knowledge
  • does not directly write a protocol

protocol/

Generates output views from Canonical Meeting Knowledge.

Output generation never invents information.

It only reformulates the analysis results for a specific audience and purpose. Depending on the output and maturity of the implementation, a renderer may be deterministic, template-based or LLM-assisted.

The planned output products are:

  • Working Protocol (working_protocol.md, Arbeitsprotokoll)
  • Distribution Protocol (distribution_protocol.md, Verteilerprotokoll)
  • Knowledge Objects, which may be rendered as a Knowledge-base Entry (knowledge_entry.md) and later stored in a structured format such as knowledge_entry.json (Wissensdatenbankeintrag)

These are parallel renderings of the same canonical semantic model, not documents derived from one another.

Rendered protocol language should normally match the dominant language of the source transcript or consolidated meeting knowledge unless an explicit output language is requested.


Repository Layout

meeting-lab/
│
├── src/
├── prompts/
├── experiments/
├── samples/
├── tests/
└── docs/

Additional documentation is intentionally split into focused documents.

Examples:

  • pipeline.md
  • segmentation.md
  • prompts.md
  • experiments.md
  • output-views.md

The architecture document only describes the overall system.


Current State

Implemented:

  • Transcript normalization
  • Technical chunk generation
  • Experimental LLM-based information extraction
  • Canonicalizer V1 deterministic extraction canonicalization

The current extraction step still performs multiple tasks simultaneously.

This was sufficient as a proof of concept but does not reflect the intended long-term architecture.

The current protocol builder is also an interim implementation. It concatenates extraction results into meeting_protocol.md for technical validation. The planned architecture separates Canonical Meeting Knowledge from the final Output Views documented in output-views.md.


Next Milestone

The next architecture milestone is the implementation of a deterministic canonicalization stage followed by a semantic consolidation stage. These stages convert raw chunk extraction JSON into evidence-preserving Canonical Meeting Knowledge before any Output View renderer writes a protocol.


Guiding Principle

The Meeting Lab assumes that the greatest improvement in transcript quality will not come from increasingly powerful language models.

Instead, quality is expected to emerge from a pipeline that decomposes a complex problem into many small, clearly defined and independently testable processing steps.