259 lines
5.1 KiB
Markdown
259 lines
5.1 KiB
Markdown
# Architecture
|
|
|
|
## Purpose
|
|
|
|
The **Meeting Lab** is an experimental environment for developing and evaluating methods to extract structured knowledge from real meeting transcripts.
|
|
|
|
Its purpose is not to build a complete meeting assistant, but to answer a single question:
|
|
|
|
> **How can knowledge be extracted from real discussions as reliably as possible?**
|
|
|
|
Successful approaches will later be integrated into the Meeting Assistant project.
|
|
|
|
---
|
|
|
|
# Design Goals
|
|
|
|
The architecture follows a small set of guiding principles.
|
|
|
|
## Modular Pipeline
|
|
|
|
Complex problems are divided into small, well-defined processing steps.
|
|
|
|
Each module has exactly one responsibility.
|
|
|
|
## Deterministic where possible
|
|
|
|
Tasks that can be solved reliably without an LLM should use deterministic algorithms.
|
|
|
|
Examples include:
|
|
|
|
- transcript normalization
|
|
- whitespace cleanup
|
|
- duplicate removal
|
|
- chunk generation
|
|
|
|
LLMs are only used where semantic understanding is required.
|
|
|
|
## Preserve Information
|
|
|
|
The pipeline should never remove or rewrite information unless it is certain that the content is merely noise.
|
|
|
|
Losing information is considered worse than keeping harmless redundancy.
|
|
|
|
## Explainable Results
|
|
|
|
Every processing step should be understandable.
|
|
|
|
Intermediate results should remain inspectable throughout the pipeline.
|
|
|
|
## Reproducible Experiments
|
|
|
|
Experiments must be repeatable.
|
|
|
|
Given the same input, prompt, model and parameters, another developer should be able to reproduce the result.
|
|
|
|
## Local First
|
|
|
|
The complete pipeline should run locally.
|
|
|
|
Cloud services may be supported in the future but are not a design requirement.
|
|
|
|
---
|
|
|
|
# Core Idea
|
|
|
|
Traditional meeting summarization attempts to solve everything in one step.
|
|
|
|
```text
|
|
Transcript
|
|
↓
|
|
LLM
|
|
↓
|
|
Summary
|
|
```
|
|
|
|
Real discussions do not work that way.
|
|
|
|
Topics are introduced, interrupted, resumed later, expanded, questioned and finally concluded.
|
|
|
|
Instead of building a better summarizer, the Meeting Lab develops a **Discussion Analyzer**.
|
|
|
|
The analyzer gradually transforms an unstructured discussion into structured knowledge.
|
|
|
|
---
|
|
|
|
# High-Level Pipeline
|
|
|
|
```text
|
|
Transcript
|
|
↓
|
|
Normalization
|
|
↓
|
|
Discussion Blocks
|
|
↓
|
|
Technical Chunking
|
|
↓
|
|
Topic Segmentation
|
|
↓
|
|
Specialized Extraction
|
|
↓
|
|
Consolidation
|
|
↓
|
|
Structured Meeting Data
|
|
↓
|
|
Protocol Generation
|
|
```
|
|
|
|
Each stage solves one clearly defined problem.
|
|
|
|
No module should perform multiple semantic tasks simultaneously.
|
|
|
|
---
|
|
|
|
# Module Overview
|
|
|
|
The current architecture consists of the following processing stages.
|
|
|
|
## normalization/
|
|
|
|
Deterministic transcript cleanup.
|
|
|
|
Responsibilities:
|
|
|
|
- remove filler words
|
|
- remove immediate repetitions
|
|
- whitespace cleanup
|
|
- generate change log
|
|
|
|
---
|
|
|
|
## chunking/
|
|
|
|
Creates model-sized chunks.
|
|
|
|
Chunking is purely technical.
|
|
|
|
It does **not** recognize discussion topics.
|
|
|
|
---
|
|
|
|
## segmentation/
|
|
|
|
Identifies discussion topics.
|
|
|
|
Responsibilities:
|
|
|
|
- detect topic start
|
|
- detect topic end
|
|
- detect topic switches
|
|
- recognize resumed topics
|
|
|
|
This is the next major development milestone.
|
|
|
|
---
|
|
|
|
## extraction/
|
|
|
|
Contains specialized LLM modules.
|
|
|
|
Planned extractors include:
|
|
|
|
- facts
|
|
- questions
|
|
- positions
|
|
- decisions
|
|
- todos
|
|
- technical information
|
|
|
|
Each extractor has exactly one task and one prompt.
|
|
|
|
---
|
|
|
|
## consolidation/
|
|
|
|
Merges information extracted from multiple discussion segments.
|
|
|
|
Typical responsibilities:
|
|
|
|
- merge duplicates
|
|
- combine partial information
|
|
- distinguish positions from decisions
|
|
- detect contradictions
|
|
|
|
---
|
|
|
|
## protocol/
|
|
|
|
Generates human-readable output from structured meeting data.
|
|
|
|
Protocol generation never invents information.
|
|
|
|
It only reformulates the analysis results.
|
|
|
|
---
|
|
|
|
# Repository Layout
|
|
|
|
```text
|
|
meeting-lab/
|
|
│
|
|
├── src/
|
|
├── prompts/
|
|
├── experiments/
|
|
├── samples/
|
|
├── tests/
|
|
└── docs/
|
|
```
|
|
|
|
Additional documentation is intentionally split into focused documents.
|
|
|
|
Examples:
|
|
|
|
- pipeline.md
|
|
- segmentation.md
|
|
- prompts.md
|
|
- experiments.md
|
|
|
|
The architecture document only describes the overall system.
|
|
|
|
---
|
|
|
|
# Current State
|
|
|
|
Implemented:
|
|
|
|
- Transcript normalization
|
|
- Technical chunk generation
|
|
- Experimental LLM-based information extraction
|
|
|
|
The current extraction step still performs multiple tasks simultaneously.
|
|
|
|
This was sufficient as a proof of concept but does not reflect the intended long-term architecture.
|
|
|
|
---
|
|
|
|
# Next Milestone
|
|
|
|
The next development step is the implementation of **topic segmentation**.
|
|
|
|
Its only responsibility is to identify the thematic structure of a discussion.
|
|
|
|
It should answer questions such as:
|
|
|
|
- Where does a topic begin?
|
|
- Where does it end?
|
|
- When does another topic start?
|
|
- When is an earlier topic resumed?
|
|
|
|
No facts, decisions or todos should be extracted at this stage.
|
|
|
|
Only after reliable topic segmentation has been achieved will the specialized extraction modules be implemented.
|
|
|
|
---
|
|
|
|
# Guiding Principle
|
|
|
|
The Meeting Lab assumes that the greatest improvement in transcript quality will not come from increasingly powerful language models.
|
|
|
|
Instead, quality is expected to emerge from a pipeline that decomposes a complex problem into many small, clearly defined and independently testable processing steps. |