# Architecture ## Purpose The **Meeting Lab** is an experimental environment for developing and evaluating methods to extract structured knowledge from real meeting transcripts. Its purpose is not to build a complete meeting assistant, but to answer a single question: > **How can knowledge be extracted from real discussions as reliably as possible?** Successful approaches will later be integrated into the Meeting Assistant project. --- # Design Goals The architecture follows a small set of guiding principles. ## Modular Pipeline Complex problems are divided into small, well-defined processing steps. Each module has exactly one responsibility. ## Deterministic where possible Tasks that can be solved reliably without an LLM should use deterministic algorithms. Examples include: - transcript normalization - whitespace cleanup - duplicate removal - chunk generation LLMs are only used where semantic understanding is required. ## Preserve Information The pipeline should never remove or rewrite information unless it is certain that the content is merely noise. Losing information is considered worse than keeping harmless redundancy. ## Explainable Results Every processing step should be understandable. Intermediate results should remain inspectable throughout the pipeline. ## Responsibility Attribution Integrity Responsibility, ownership, organizational roles and action-item assignments may be recorded only when meeting evidence explicitly assigns, accepts or confirms them. The system must not infer responsibility from thematic proximity, participation in a discussion, mentioning a task, commenting on another department, organizational assumptions, likely job roles, speaker adjacency or model world knowledge. When evidence is incomplete or ambiguous, the responsible person remains unset or unclear and the supporting evidence is preserved. ## Reproducible Experiments Experiments must be repeatable. Given the same input, prompt, model and parameters, another developer should be able to reproduce the result. ## Local First The complete pipeline should run locally. Cloud services may be supported in the future but are not a design requirement. --- # Core Idea Traditional meeting summarization attempts to solve everything in one step. ```text Transcript ↓ LLM ↓ Summary ``` Real discussions do not work that way. Topics are introduced, interrupted, resumed later, expanded, questioned and finally concluded. Instead of building a better summarizer, the Meeting Lab develops a **Discussion Analyzer**. The analyzer gradually transforms an unstructured discussion into structured knowledge. --- # High-Level Pipeline Current implemented and intended analysis flow: ```text Transcript ↓ Normalization ↓ Discussion Blocks ↓ Technical Chunking ↓ Topic Segmentation ↓ Specialized Extraction ↓ Deterministic Canonicalization ↓ Semantic Consolidation ↓ Canonical Meeting Knowledge ↓ Output View Rendering ↓ Working Protocol / Distribution Protocol / Knowledge Objects ``` Accepted future Meeting Context V2 preparation flow: ```text Whisper ↓ Entity Detection ↓ User Confirmation ↓ Entity Registry Update ↓ Meeting Context Builder ↓ meeting_context.yaml ↓ Extraction Pipeline ``` Each stage solves one clearly defined problem. No module should perform multiple semantic tasks simultaneously. --- # Module Overview The current architecture consists of the following processing stages. ## normalization/ Deterministic transcript cleanup. Responsibilities: - remove filler words - remove immediate repetitions - whitespace cleanup - generate change log --- ## chunking/ Creates model-sized chunks. Chunking is purely technical. It does **not** recognize discussion topics. --- ## segmentation/ Identifies discussion topics. Responsibilities: - detect topic start - detect topic end - detect topic switches - recognize resumed topics This is the next major development milestone. --- ## extraction/ Contains specialized LLM modules. Planned extractors include: - facts - questions - positions - decisions - todos - technical information Each extractor has exactly one task and one prompt. --- ## Meeting Context Meeting Context V1 is a manually maintained YAML scaffold for reliable meeting metadata such as title, language, participants, aliases, departments, abbreviations and known entities. It is documented in `docs/meeting-context.md` and templated at `samples/templates/meeting_context.template.yaml`. It is implemented for loading, validation and optional injection into chunk extraction prompts. Extraction results record only minimal context provenance. It is not yet connected to consolidation, Canonical Meeting Knowledge or output rendering. The context can help prevent non-participants from being interpreted as attendees and can normalize known aliases for extraction. It must not infer roles, departments, responsibilities or decisions. Accepted future direction: Meeting Context V2 should be generated or assisted from an interactive entity confirmation workflow and a persistent Entity Registry. The Entity Registry is the persistent cross-meeting knowledge source for confirmed people, organizations, departments, products, projects, locations, aliases and organizational metadata under stable internal IDs. Display names may change, but internal IDs remain stable. Aliases are first-class data. The registry never learns automatically. It may propose matches and aliases, including spelling variants, Whisper transcription variants, umlaut variants and OCR-like mistakes, but only explicit user confirmation updates registry state. Unknown names should be presented to the user as meeting participant, mentioned person, external person, transcription error or ignore. In this architecture, `meeting_context.yaml` remains the authoritative meeting-specific Point of Truth and reproducible input artifact consumed by the extraction pipeline. It is a meeting-specific snapshot built from the Entity Registry, user confirmations and meeting metadata. The Registry must not override explicit meeting-specific confirmations, and Registry changes after a meeting run must not silently change the historical Meeting Context used for that run. Status: Accepted Architecture; implementation deferred. See `docs/adr-meeting-context-v2-entity-registry.md`. --- ## consolidation/ Planned area for canonicalization and consolidation. The next milestone splits this into two stages. Deterministic Canonicalizer: - implemented in Python - uses no LLM - validates and normalizes extraction objects - assigns stable source references and IDs - normalizes category names and basic field structure - validates and normalizes action-item responsible fields against Meeting Context when available: known participant and mentioned-person aliases are normalized to canonical display names, while dates, locations, projects, products, technical terms, generic process words and unknown free text are cleared with a structured validation record - performs only safe deterministic cleanup - may group exact duplicates - preserves all source evidence - must not perform uncertain semantic merging Semantic Consolidator: - uses the local LLM - V0 is implemented for facts-only semantic duplicate detection - V0 merges semantically equivalent fact items conservatively - V0 preserves source references and evidence - V0 validates that every source fact appears exactly once - V0 sizes its Ollama output budget from the actual fact payload instead of using a fixed response cap for every meeting - V0 may apply deterministic source-coverage repair after valid model JSON is parsed: duplicate source IDs are removed after their first occurrence, empty groups are removed and missing source facts are restored as singleton groups from canonicalized input before strict validation runs - V0 does not process non-fact categories semantically - later versions should group content by topic, mark contradictions and uncertainty, separate durable information from transient discussion and prepare Canonical Meeting Knowledge - does not directly write a protocol --- ## protocol/ Generates output views from Canonical Meeting Knowledge. Output generation never invents information. It only reformulates the analysis results for a specific audience and purpose. Depending on the output and maturity of the implementation, a renderer may be deterministic, template-based or LLM-assisted. LLM-assisted renderers preserve raw model output separately and write the final output artifact only after deterministic contract validation succeeds. Renderer post-processing may remove non-semantic wrapper text, but must not fabricate missing semantic sections or relabel an invalid summary as a valid output view. The planned output products are: - Working Protocol (`working_protocol.md`, Arbeitsprotokoll) - Distribution Protocol (`distribution_protocol.md`, Verteilerprotokoll) - Knowledge Objects, which may be rendered as a Knowledge-base Entry (`knowledge_entry.md`) and later stored in a structured format such as `knowledge_entry.json` (Wissensdatenbankeintrag) These are parallel renderings of the same canonical semantic model, not documents derived from one another. Rendered protocol language should normally match the dominant language of the source transcript or consolidated meeting knowledge unless an explicit output language is requested. --- # Repository Layout ```text meeting-lab/ │ ├── src/ ├── prompts/ ├── experiments/ ├── samples/ ├── tests/ └── docs/ ``` Additional documentation is intentionally split into focused documents. Examples: - pipeline.md - segmentation.md - prompts.md - experiments.md - output-views.md The architecture document only describes the overall system. --- # Current State Implemented: - Transcript normalization - Technical chunk generation - Experimental LLM-based information extraction - Meeting Context V1 loading, validation and extraction prompt integration - Canonicalizer V1 deterministic extraction canonicalization The current extraction step still performs multiple tasks simultaneously. This was sufficient as a proof of concept but does not reflect the intended long-term architecture. The current protocol builder is also an interim implementation. It concatenates extraction results into `meeting_protocol.md` for technical validation. The planned architecture separates Canonical Meeting Knowledge from the final Output Views documented in `output-views.md`. ## Semantic Consolidator failure handling Semantic Consolidator V0 preserves every raw model response before parsing. Its normal path accepts parseable grouping JSON and leaves duplicate source-ID and missing-source-ID correction to the existing deterministic coverage repair. An invalid or truncated response is not retried by default. One controlled retry is allowed only when deterministic inspection finds at least three consecutive complete groups with an identical structural signature: `canonical_text`, ordered `source_item_ids` and `merge_reason`. The retry keeps the same model, temperature, context window and generation limit and adds only an instruction not to emit an identical group more than once. Both attempts and the detected repetition metadata are preserved. If the retry also fails, the stage fails normally; it does not make another LLM call. --- # Next Milestone The next architecture milestone is the implementation of a deterministic canonicalization stage followed by a semantic consolidation stage. These stages convert raw chunk extraction JSON into evidence-preserving Canonical Meeting Knowledge before any Output View renderer writes a protocol. --- # Guiding Principle The Meeting Lab assumes that the greatest improvement in transcript quality will not come from increasingly powerful language models. Instead, quality is expected to emerge from a pipeline that decomposes a complex problem into many small, clearly defined and independently testable processing steps.