# Architecture ## Purpose The **Meeting Lab** is an experimental environment for developing and evaluating methods to extract structured knowledge from real meeting transcripts. Its purpose is not to build a complete meeting assistant, but to answer a single question: > **How can knowledge be extracted from real discussions as reliably as possible?** Successful approaches will later be integrated into the Meeting Assistant project. --- # Design Goals The architecture follows a small set of guiding principles. ## Modular Pipeline Complex problems are divided into small, well-defined processing steps. Each module has exactly one responsibility. ## Deterministic where possible Tasks that can be solved reliably without an LLM should use deterministic algorithms. Examples include: - transcript normalization - whitespace cleanup - duplicate removal - chunk generation LLMs are only used where semantic understanding is required. ## Preserve Information The pipeline should never remove or rewrite information unless it is certain that the content is merely noise. Losing information is considered worse than keeping harmless redundancy. ## Explainable Results Every processing step should be understandable. Intermediate results should remain inspectable throughout the pipeline. ## Reproducible Experiments Experiments must be repeatable. Given the same input, prompt, model and parameters, another developer should be able to reproduce the result. ## Local First The complete pipeline should run locally. Cloud services may be supported in the future but are not a design requirement. --- # Core Idea Traditional meeting summarization attempts to solve everything in one step. ```text Transcript ↓ LLM ↓ Summary ``` Real discussions do not work that way. Topics are introduced, interrupted, resumed later, expanded, questioned and finally concluded. Instead of building a better summarizer, the Meeting Lab develops a **Discussion Analyzer**. The analyzer gradually transforms an unstructured discussion into structured knowledge. --- # High-Level Pipeline ```text Transcript ↓ Normalization ↓ Discussion Blocks ↓ Technical Chunking ↓ Topic Segmentation ↓ Specialized Extraction ↓ Consolidation ↓ Structured Meeting Data ↓ Protocol Generation ``` Each stage solves one clearly defined problem. No module should perform multiple semantic tasks simultaneously. --- # Module Overview The current architecture consists of the following processing stages. ## normalization/ Deterministic transcript cleanup. Responsibilities: - remove filler words - remove immediate repetitions - whitespace cleanup - generate change log --- ## chunking/ Creates model-sized chunks. Chunking is purely technical. It does **not** recognize discussion topics. --- ## segmentation/ Identifies discussion topics. Responsibilities: - detect topic start - detect topic end - detect topic switches - recognize resumed topics This is the next major development milestone. --- ## extraction/ Contains specialized LLM modules. Planned extractors include: - facts - questions - positions - decisions - todos - technical information Each extractor has exactly one task and one prompt. --- ## consolidation/ Merges information extracted from multiple discussion segments. Typical responsibilities: - merge duplicates - combine partial information - distinguish positions from decisions - detect contradictions --- ## protocol/ Generates human-readable output from structured meeting data. Protocol generation never invents information. It only reformulates the analysis results. --- # Repository Layout ```text meeting-lab/ │ ├── src/ ├── prompts/ ├── experiments/ ├── samples/ ├── tests/ └── docs/ ``` Additional documentation is intentionally split into focused documents. Examples: - pipeline.md - segmentation.md - prompts.md - experiments.md The architecture document only describes the overall system. --- # Current State Implemented: - Transcript normalization - Technical chunk generation - Experimental LLM-based information extraction The current extraction step still performs multiple tasks simultaneously. This was sufficient as a proof of concept but does not reflect the intended long-term architecture. --- # Next Milestone The next development step is the implementation of **topic segmentation**. Its only responsibility is to identify the thematic structure of a discussion. It should answer questions such as: - Where does a topic begin? - Where does it end? - When does another topic start? - When is an earlier topic resumed? No facts, decisions or todos should be extracted at this stage. Only after reliable topic segmentation has been achieved will the specialized extraction modules be implemented. --- # Guiding Principle The Meeting Lab assumes that the greatest improvement in transcript quality will not come from increasingly powerful language models. Instead, quality is expected to emerge from a pipeline that decomposes a complex problem into many small, clearly defined and independently testable processing steps.