410 lines
12 KiB
Markdown
410 lines
12 KiB
Markdown
# Architecture
|
|
|
|
## Purpose
|
|
|
|
The **Meeting Lab** is an experimental environment for developing and evaluating methods to extract structured knowledge from real meeting transcripts.
|
|
|
|
Its purpose is not to build a complete meeting assistant, but to answer a single question:
|
|
|
|
> **How can knowledge be extracted from real discussions as reliably as possible?**
|
|
|
|
Successful approaches will later be integrated into the Meeting Assistant project.
|
|
|
|
---
|
|
|
|
# Design Goals
|
|
|
|
The architecture follows a small set of guiding principles.
|
|
|
|
## Modular Pipeline
|
|
|
|
Complex problems are divided into small, well-defined processing steps.
|
|
|
|
Each module has exactly one responsibility.
|
|
|
|
## Deterministic where possible
|
|
|
|
Tasks that can be solved reliably without an LLM should use deterministic algorithms.
|
|
|
|
Examples include:
|
|
|
|
- transcript normalization
|
|
- whitespace cleanup
|
|
- duplicate removal
|
|
- chunk generation
|
|
|
|
LLMs are only used where semantic understanding is required.
|
|
|
|
## Preserve Information
|
|
|
|
The pipeline should never remove or rewrite information unless it is certain that the content is merely noise.
|
|
|
|
Losing information is considered worse than keeping harmless redundancy.
|
|
|
|
## Explainable Results
|
|
|
|
Every processing step should be understandable.
|
|
|
|
Intermediate results should remain inspectable throughout the pipeline.
|
|
|
|
## Responsibility Attribution Integrity
|
|
|
|
Responsibility, ownership, organizational roles and action-item assignments may
|
|
be recorded only when meeting evidence explicitly assigns, accepts or confirms
|
|
them.
|
|
|
|
The system must not infer responsibility from thematic proximity,
|
|
participation in a discussion, mentioning a task, commenting on another
|
|
department, organizational assumptions, likely job roles, speaker adjacency or
|
|
model world knowledge.
|
|
|
|
When evidence is incomplete or ambiguous, the responsible person remains unset
|
|
or unclear and the supporting evidence is preserved.
|
|
|
|
## Reproducible Experiments
|
|
|
|
Experiments must be repeatable.
|
|
|
|
Given the same input, prompt, model and parameters, another developer should be able to reproduce the result.
|
|
|
|
## Local First
|
|
|
|
The complete pipeline should run locally.
|
|
|
|
Cloud services may be supported in the future but are not a design requirement.
|
|
|
|
---
|
|
|
|
# Core Idea
|
|
|
|
Traditional meeting summarization attempts to solve everything in one step.
|
|
|
|
```text
|
|
Transcript
|
|
↓
|
|
LLM
|
|
↓
|
|
Summary
|
|
```
|
|
|
|
Real discussions do not work that way.
|
|
|
|
Topics are introduced, interrupted, resumed later, expanded, questioned and finally concluded.
|
|
|
|
Instead of building a better summarizer, the Meeting Lab develops a **Discussion Analyzer**.
|
|
|
|
The analyzer gradually transforms an unstructured discussion into structured knowledge.
|
|
|
|
---
|
|
|
|
# High-Level Pipeline
|
|
|
|
Current implemented and intended analysis flow:
|
|
|
|
```text
|
|
Transcript
|
|
↓
|
|
Normalization
|
|
↓
|
|
Discussion Blocks
|
|
↓
|
|
Technical Chunking
|
|
↓
|
|
Topic Segmentation
|
|
↓
|
|
Specialized Extraction
|
|
↓
|
|
Deterministic Canonicalization
|
|
↓
|
|
Semantic Consolidation
|
|
↓
|
|
Canonical Meeting Knowledge
|
|
↓
|
|
Output View Rendering
|
|
↓
|
|
Working Protocol / Distribution Protocol / Knowledge Objects
|
|
```
|
|
|
|
Accepted future Meeting Context V2 preparation flow:
|
|
|
|
```text
|
|
Whisper
|
|
↓
|
|
Entity Detection
|
|
↓
|
|
User Confirmation
|
|
↓
|
|
Entity Registry Update
|
|
↓
|
|
Meeting Context Builder
|
|
↓
|
|
meeting_context.yaml
|
|
↓
|
|
Extraction Pipeline
|
|
```
|
|
|
|
Each stage solves one clearly defined problem.
|
|
|
|
No module should perform multiple semantic tasks simultaneously.
|
|
|
|
---
|
|
|
|
# Module Overview
|
|
|
|
The current architecture consists of the following processing stages.
|
|
|
|
## normalization/
|
|
|
|
Deterministic transcript cleanup.
|
|
|
|
Responsibilities:
|
|
|
|
- remove filler words
|
|
- remove immediate repetitions
|
|
- whitespace cleanup
|
|
- generate change log
|
|
|
|
---
|
|
|
|
## chunking/
|
|
|
|
Creates model-sized chunks.
|
|
|
|
Chunking is purely technical.
|
|
|
|
It does **not** recognize discussion topics.
|
|
|
|
---
|
|
|
|
## segmentation/
|
|
|
|
Identifies discussion topics.
|
|
|
|
Responsibilities:
|
|
|
|
- detect topic start
|
|
- detect topic end
|
|
- detect topic switches
|
|
- recognize resumed topics
|
|
|
|
This is the next major development milestone.
|
|
|
|
---
|
|
|
|
## extraction/
|
|
|
|
Contains specialized LLM modules.
|
|
|
|
Planned extractors include:
|
|
|
|
- facts
|
|
- questions
|
|
- positions
|
|
- decisions
|
|
- todos
|
|
- technical information
|
|
|
|
Each extractor has exactly one task and one prompt.
|
|
|
|
---
|
|
|
|
## Meeting Context
|
|
|
|
Meeting Context V1 is a manually maintained YAML scaffold for reliable meeting
|
|
metadata such as title, language, participants, aliases, departments,
|
|
abbreviations and known entities.
|
|
|
|
It is documented in `docs/meeting-context.md` and templated at
|
|
`samples/templates/meeting_context.template.yaml`. It is implemented for
|
|
loading, validation and optional injection into chunk extraction prompts.
|
|
Extraction results record only minimal context provenance. It is not yet
|
|
connected to consolidation, Canonical Meeting Knowledge or output rendering.
|
|
|
|
The context can help prevent non-participants from being interpreted as
|
|
attendees and can normalize known aliases for extraction. It must not infer
|
|
roles, departments, responsibilities or decisions.
|
|
|
|
Accepted future direction:
|
|
|
|
Meeting Context V2 should be generated or assisted from an interactive entity
|
|
confirmation workflow and a persistent Entity Registry. The Entity Registry is
|
|
the persistent cross-meeting knowledge source for confirmed people,
|
|
organizations, departments, products, projects, locations, aliases and
|
|
organizational metadata under stable internal IDs. Display names may change,
|
|
but internal IDs remain stable. Aliases are first-class data.
|
|
|
|
The registry never learns automatically. It may propose matches and aliases,
|
|
including spelling variants, Whisper transcription variants, umlaut variants
|
|
and OCR-like mistakes, but only explicit user confirmation updates registry
|
|
state. Unknown names should be presented to the user as meeting participant,
|
|
mentioned person, external person, transcription error or ignore.
|
|
|
|
In this architecture, `meeting_context.yaml` remains the authoritative
|
|
meeting-specific Point of Truth and reproducible input artifact consumed by the
|
|
extraction pipeline. It is a meeting-specific snapshot built from the Entity
|
|
Registry, user confirmations and meeting metadata. The Registry must not
|
|
override explicit meeting-specific confirmations, and Registry changes after a
|
|
meeting run must not silently change the historical Meeting Context used for
|
|
that run.
|
|
|
|
Status: Accepted Architecture; implementation deferred. See
|
|
`docs/adr-meeting-context-v2-entity-registry.md`.
|
|
|
|
---
|
|
|
|
## consolidation/
|
|
|
|
Planned area for canonicalization and consolidation.
|
|
|
|
The next milestone splits this into two stages.
|
|
|
|
Deterministic Canonicalizer:
|
|
|
|
- implemented in Python
|
|
- uses no LLM
|
|
- validates and normalizes extraction objects
|
|
- assigns stable source references and IDs
|
|
- normalizes category names and basic field structure
|
|
- validates and normalizes action-item responsible fields against Meeting
|
|
Context when available: known participant and mentioned-person aliases are
|
|
normalized to canonical display names, while dates, locations, projects,
|
|
products, technical terms, generic process words and unknown free text are
|
|
cleared with a structured validation record
|
|
- performs only safe deterministic cleanup
|
|
- may group exact duplicates
|
|
- preserves all source evidence
|
|
- must not perform uncertain semantic merging
|
|
|
|
Semantic Consolidator:
|
|
|
|
- uses the local LLM
|
|
- V0 is implemented for facts-only semantic duplicate detection
|
|
- V0 merges semantically equivalent fact items conservatively
|
|
- V0 preserves source references and evidence
|
|
- V0 validates that every source fact appears exactly once
|
|
- V0 sizes its Ollama output budget from the actual fact payload instead of
|
|
using a fixed response cap for every meeting
|
|
- V0 may apply deterministic source-coverage repair after valid model JSON is
|
|
parsed: duplicate source IDs are removed after their first occurrence, empty
|
|
groups are removed and missing source facts are restored as singleton groups
|
|
from canonicalized input before strict validation runs
|
|
- V0 does not process non-fact categories semantically
|
|
- later versions should group content by topic, mark contradictions and
|
|
uncertainty, separate durable information from transient discussion and
|
|
prepare Canonical Meeting Knowledge
|
|
- does not directly write a protocol
|
|
|
|
---
|
|
|
|
## protocol/
|
|
|
|
Generates output views from Canonical Meeting Knowledge.
|
|
|
|
Output generation never invents information.
|
|
|
|
It only reformulates the analysis results for a specific audience and purpose.
|
|
Depending on the output and maturity of the implementation, a renderer may be
|
|
deterministic, template-based or LLM-assisted.
|
|
|
|
LLM-assisted renderers preserve raw model output separately and write the final
|
|
output artifact only after deterministic contract validation succeeds. Renderer
|
|
post-processing may remove non-semantic wrapper text, but must not fabricate
|
|
missing semantic sections or relabel an invalid summary as a valid output view.
|
|
|
|
The planned output products are:
|
|
|
|
- Working Protocol (`working_protocol.md`, Arbeitsprotokoll)
|
|
- Distribution Protocol (`distribution_protocol.md`,
|
|
Verteilerprotokoll)
|
|
- Knowledge Objects, which may be rendered as a Knowledge-base Entry
|
|
(`knowledge_entry.md`) and later stored in a structured format such as
|
|
`knowledge_entry.json` (Wissensdatenbankeintrag)
|
|
|
|
These are parallel renderings of the same canonical semantic model, not
|
|
documents derived from one another.
|
|
|
|
Rendered protocol language should normally match the dominant language of the
|
|
source transcript or consolidated meeting knowledge unless an explicit output
|
|
language is requested.
|
|
|
|
---
|
|
|
|
# Repository Layout
|
|
|
|
```text
|
|
meeting-lab/
|
|
│
|
|
├── src/
|
|
├── prompts/
|
|
├── experiments/
|
|
├── samples/
|
|
├── tests/
|
|
└── docs/
|
|
```
|
|
|
|
Additional documentation is intentionally split into focused documents.
|
|
|
|
Examples:
|
|
|
|
- pipeline.md
|
|
- segmentation.md
|
|
- prompts.md
|
|
- experiments.md
|
|
- output-views.md
|
|
|
|
The architecture document only describes the overall system.
|
|
|
|
---
|
|
|
|
# Current State
|
|
|
|
Implemented:
|
|
|
|
- Transcript normalization
|
|
- Technical chunk generation
|
|
- Experimental LLM-based information extraction
|
|
- Meeting Context V1 loading, validation and extraction prompt integration
|
|
- Canonicalizer V1 deterministic extraction canonicalization
|
|
|
|
The current extraction step still performs multiple tasks simultaneously.
|
|
|
|
This was sufficient as a proof of concept but does not reflect the intended long-term architecture.
|
|
|
|
The current protocol builder is also an interim implementation. It concatenates
|
|
extraction results into `meeting_protocol.md` for technical validation. The
|
|
planned architecture separates Canonical Meeting Knowledge from the final Output
|
|
Views documented in `output-views.md`.
|
|
|
|
## Semantic Consolidator failure handling
|
|
|
|
Semantic Consolidator V0 preserves every raw model response before parsing.
|
|
Its normal path accepts parseable grouping JSON and leaves duplicate source-ID
|
|
and missing-source-ID correction to the existing deterministic coverage
|
|
repair.
|
|
|
|
An invalid or truncated response is not retried by default. One controlled
|
|
retry is allowed only when deterministic inspection finds at least three
|
|
consecutive complete groups with an identical structural signature:
|
|
`canonical_text`, ordered `source_item_ids` and `merge_reason`. The retry keeps
|
|
the same model, temperature, context window and generation limit and adds only
|
|
an instruction not to emit an identical group more than once. Both attempts
|
|
and the detected repetition metadata are preserved. If the retry also fails,
|
|
the stage fails normally; it does not make another LLM call.
|
|
|
|
---
|
|
|
|
# Next Milestone
|
|
|
|
The next architecture milestone is the implementation of a deterministic
|
|
canonicalization stage followed by a semantic consolidation stage. These stages
|
|
convert raw chunk extraction JSON into evidence-preserving Canonical Meeting
|
|
Knowledge before any Output View renderer writes a protocol.
|
|
|
|
---
|
|
|
|
# Guiding Principle
|
|
|
|
The Meeting Lab assumes that the greatest improvement in transcript quality will not come from increasingly powerful language models.
|
|
|
|
Instead, quality is expected to emerge from a pipeline that decomposes a complex problem into many small, clearly defined and independently testable processing steps.
|