Introduce Meeting Context V1 with YAML schema, validation and template. Support optional --meeting-context during chunk extraction. Inject authoritative Meeting Context into extraction prompts. Record Meeting Context provenance in extraction output. Activate todos.md in shared prompt assembly. Strengthen responsibility attribution and decision/todo boundaries. Add focused Gold scenarios and validation tests. Update architecture and pipeline documentation.
332 lines
8.3 KiB
Markdown
332 lines
8.3 KiB
Markdown
# Architecture
|
|
|
|
## Purpose
|
|
|
|
The **Meeting Lab** is an experimental environment for developing and evaluating methods to extract structured knowledge from real meeting transcripts.
|
|
|
|
Its purpose is not to build a complete meeting assistant, but to answer a single question:
|
|
|
|
> **How can knowledge be extracted from real discussions as reliably as possible?**
|
|
|
|
Successful approaches will later be integrated into the Meeting Assistant project.
|
|
|
|
---
|
|
|
|
# Design Goals
|
|
|
|
The architecture follows a small set of guiding principles.
|
|
|
|
## Modular Pipeline
|
|
|
|
Complex problems are divided into small, well-defined processing steps.
|
|
|
|
Each module has exactly one responsibility.
|
|
|
|
## Deterministic where possible
|
|
|
|
Tasks that can be solved reliably without an LLM should use deterministic algorithms.
|
|
|
|
Examples include:
|
|
|
|
- transcript normalization
|
|
- whitespace cleanup
|
|
- duplicate removal
|
|
- chunk generation
|
|
|
|
LLMs are only used where semantic understanding is required.
|
|
|
|
## Preserve Information
|
|
|
|
The pipeline should never remove or rewrite information unless it is certain that the content is merely noise.
|
|
|
|
Losing information is considered worse than keeping harmless redundancy.
|
|
|
|
## Explainable Results
|
|
|
|
Every processing step should be understandable.
|
|
|
|
Intermediate results should remain inspectable throughout the pipeline.
|
|
|
|
## Responsibility Attribution Integrity
|
|
|
|
Responsibility, ownership, organizational roles and action-item assignments may
|
|
be recorded only when meeting evidence explicitly assigns, accepts or confirms
|
|
them.
|
|
|
|
The system must not infer responsibility from thematic proximity,
|
|
participation in a discussion, mentioning a task, commenting on another
|
|
department, organizational assumptions, likely job roles, speaker adjacency or
|
|
model world knowledge.
|
|
|
|
When evidence is incomplete or ambiguous, the responsible person remains unset
|
|
or unclear and the supporting evidence is preserved.
|
|
|
|
## Reproducible Experiments
|
|
|
|
Experiments must be repeatable.
|
|
|
|
Given the same input, prompt, model and parameters, another developer should be able to reproduce the result.
|
|
|
|
## Local First
|
|
|
|
The complete pipeline should run locally.
|
|
|
|
Cloud services may be supported in the future but are not a design requirement.
|
|
|
|
---
|
|
|
|
# Core Idea
|
|
|
|
Traditional meeting summarization attempts to solve everything in one step.
|
|
|
|
```text
|
|
Transcript
|
|
↓
|
|
LLM
|
|
↓
|
|
Summary
|
|
```
|
|
|
|
Real discussions do not work that way.
|
|
|
|
Topics are introduced, interrupted, resumed later, expanded, questioned and finally concluded.
|
|
|
|
Instead of building a better summarizer, the Meeting Lab develops a **Discussion Analyzer**.
|
|
|
|
The analyzer gradually transforms an unstructured discussion into structured knowledge.
|
|
|
|
---
|
|
|
|
# High-Level Pipeline
|
|
|
|
```text
|
|
Transcript
|
|
↓
|
|
Normalization
|
|
↓
|
|
Discussion Blocks
|
|
↓
|
|
Technical Chunking
|
|
↓
|
|
Topic Segmentation
|
|
↓
|
|
Specialized Extraction
|
|
↓
|
|
Deterministic Canonicalization
|
|
↓
|
|
Semantic Consolidation
|
|
↓
|
|
Canonical Meeting Knowledge
|
|
↓
|
|
Output View Rendering
|
|
↓
|
|
Working Protocol / Distribution Protocol / Knowledge Objects
|
|
```
|
|
|
|
Each stage solves one clearly defined problem.
|
|
|
|
No module should perform multiple semantic tasks simultaneously.
|
|
|
|
---
|
|
|
|
# Module Overview
|
|
|
|
The current architecture consists of the following processing stages.
|
|
|
|
## normalization/
|
|
|
|
Deterministic transcript cleanup.
|
|
|
|
Responsibilities:
|
|
|
|
- remove filler words
|
|
- remove immediate repetitions
|
|
- whitespace cleanup
|
|
- generate change log
|
|
|
|
---
|
|
|
|
## chunking/
|
|
|
|
Creates model-sized chunks.
|
|
|
|
Chunking is purely technical.
|
|
|
|
It does **not** recognize discussion topics.
|
|
|
|
---
|
|
|
|
## segmentation/
|
|
|
|
Identifies discussion topics.
|
|
|
|
Responsibilities:
|
|
|
|
- detect topic start
|
|
- detect topic end
|
|
- detect topic switches
|
|
- recognize resumed topics
|
|
|
|
This is the next major development milestone.
|
|
|
|
---
|
|
|
|
## extraction/
|
|
|
|
Contains specialized LLM modules.
|
|
|
|
Planned extractors include:
|
|
|
|
- facts
|
|
- questions
|
|
- positions
|
|
- decisions
|
|
- todos
|
|
- technical information
|
|
|
|
Each extractor has exactly one task and one prompt.
|
|
|
|
---
|
|
|
|
## Meeting Context
|
|
|
|
Meeting Context V1 is a manually maintained YAML scaffold for reliable meeting
|
|
metadata such as title, language, participants, aliases, departments,
|
|
abbreviations and known entities.
|
|
|
|
It is documented in `docs/meeting-context.md` and templated at
|
|
`samples/templates/meeting_context.template.yaml`. It is implemented for
|
|
loading, validation and optional injection into chunk extraction prompts.
|
|
Extraction results record only minimal context provenance. It is not yet
|
|
connected to consolidation, Canonical Meeting Knowledge or output rendering.
|
|
|
|
The context can help prevent non-participants from being interpreted as
|
|
attendees and can normalize known aliases for extraction. It must not infer
|
|
roles, departments, responsibilities or decisions.
|
|
|
|
---
|
|
|
|
## consolidation/
|
|
|
|
Planned area for canonicalization and consolidation.
|
|
|
|
The next milestone splits this into two stages.
|
|
|
|
Deterministic Canonicalizer:
|
|
|
|
- implemented in Python
|
|
- uses no LLM
|
|
- validates and normalizes extraction objects
|
|
- assigns stable source references and IDs
|
|
- normalizes category names and basic field structure
|
|
- performs only safe deterministic cleanup
|
|
- may group exact duplicates
|
|
- preserves all source evidence
|
|
- must not perform uncertain semantic merging
|
|
|
|
Semantic Consolidator:
|
|
|
|
- uses the local LLM
|
|
- V0 is implemented for facts-only semantic duplicate detection
|
|
- V0 merges semantically equivalent fact items conservatively
|
|
- V0 preserves source references and evidence
|
|
- V0 validates that every source fact appears exactly once
|
|
- V0 does not process non-fact categories semantically
|
|
- later versions should group content by topic, mark contradictions and
|
|
uncertainty, separate durable information from transient discussion and
|
|
prepare Canonical Meeting Knowledge
|
|
- does not directly write a protocol
|
|
|
|
---
|
|
|
|
## protocol/
|
|
|
|
Generates output views from Canonical Meeting Knowledge.
|
|
|
|
Output generation never invents information.
|
|
|
|
It only reformulates the analysis results for a specific audience and purpose.
|
|
Depending on the output and maturity of the implementation, a renderer may be
|
|
deterministic, template-based or LLM-assisted.
|
|
|
|
The planned output products are:
|
|
|
|
- Working Protocol (`working_protocol.md`, Arbeitsprotokoll)
|
|
- Distribution Protocol (`distribution_protocol.md`,
|
|
Verteilerprotokoll)
|
|
- Knowledge Objects, which may be rendered as a Knowledge-base Entry
|
|
(`knowledge_entry.md`) and later stored in a structured format such as
|
|
`knowledge_entry.json` (Wissensdatenbankeintrag)
|
|
|
|
These are parallel renderings of the same canonical semantic model, not
|
|
documents derived from one another.
|
|
|
|
Rendered protocol language should normally match the dominant language of the
|
|
source transcript or consolidated meeting knowledge unless an explicit output
|
|
language is requested.
|
|
|
|
---
|
|
|
|
# Repository Layout
|
|
|
|
```text
|
|
meeting-lab/
|
|
│
|
|
├── src/
|
|
├── prompts/
|
|
├── experiments/
|
|
├── samples/
|
|
├── tests/
|
|
└── docs/
|
|
```
|
|
|
|
Additional documentation is intentionally split into focused documents.
|
|
|
|
Examples:
|
|
|
|
- pipeline.md
|
|
- segmentation.md
|
|
- prompts.md
|
|
- experiments.md
|
|
- output-views.md
|
|
|
|
The architecture document only describes the overall system.
|
|
|
|
---
|
|
|
|
# Current State
|
|
|
|
Implemented:
|
|
|
|
- Transcript normalization
|
|
- Technical chunk generation
|
|
- Experimental LLM-based information extraction
|
|
- Meeting Context V1 loading, validation and extraction prompt integration
|
|
- Canonicalizer V1 deterministic extraction canonicalization
|
|
|
|
The current extraction step still performs multiple tasks simultaneously.
|
|
|
|
This was sufficient as a proof of concept but does not reflect the intended long-term architecture.
|
|
|
|
The current protocol builder is also an interim implementation. It concatenates
|
|
extraction results into `meeting_protocol.md` for technical validation. The
|
|
planned architecture separates Canonical Meeting Knowledge from the final Output
|
|
Views documented in `output-views.md`.
|
|
|
|
---
|
|
|
|
# Next Milestone
|
|
|
|
The next architecture milestone is the implementation of a deterministic
|
|
canonicalization stage followed by a semantic consolidation stage. These stages
|
|
convert raw chunk extraction JSON into evidence-preserving Canonical Meeting
|
|
Knowledge before any Output View renderer writes a protocol.
|
|
|
|
---
|
|
|
|
# Guiding Principle
|
|
|
|
The Meeting Lab assumes that the greatest improvement in transcript quality will not come from increasingly powerful language models.
|
|
|
|
Instead, quality is expected to emerge from a pipeline that decomposes a complex problem into many small, clearly defined and independently testable processing steps.
|