# Pipeline ## Purpose This document describes the processing pipeline of the Meeting Lab. Unlike `architecture.md`, which describes the overall system, this document focuses on the individual processing stages, their responsibilities and the data flowing between them. The guiding principle is simple: > **Each processing stage has exactly one responsibility.** --- # Pipeline Overview ```text Whisper Transcript │ ▼ Normalization │ ▼ Discussion Blocks │ ▼ Technical Chunking │ ▼ Topic Segmentation │ ▼ Specialized Extraction │ ▼ Consolidation │ ▼ Canonical Meeting Knowledge │ ▼ Output View Rendering │ ├── Working Protocol ├── Distribution Protocol └── Knowledge Objects ``` Each stage receives a well-defined input and produces a well-defined output. --- # Stage 1 – Normalization ## Purpose Remove transcription artifacts without changing the meaning of the discussion. ## Input Raw transcript generated by Whisper. ## Output Normalized transcript. Change log containing every modification. ## Processing Type Deterministic ## Responsibilities - Remove filler words - Remove immediate duplicate words - Remove immediate duplicate short phrases - Normalize whitespace - Preserve all semantic content ## Must Not - Rephrase text - Summarize - Interpret statements - Correct factual content ## Current Status Implemented --- # Stage 2 – Discussion Blocks ## Purpose Convert the transcript into stable processing units. Discussion blocks are the smallest semantic unit used throughout the pipeline. ## Input Normalized transcript. ## Output Ordered list of discussion blocks. Example: ```json { "block_id": 42, "speaker": "A", "start": 351.2, "end": 367.8, "text": "..." } ``` ## Processing Type Deterministic ## Responsibilities - Create stable identifiers - Preserve ordering - Preserve timestamps - Preserve speaker information where available ## Current Status Planned --- # Stage 3 – Technical Chunking ## Purpose Split large meetings into model-sized chunks. Chunking exists only because language models have limited context windows. ## Input Discussion blocks. ## Output Chunk manifest and chunk files. ## Processing Type Deterministic ## Responsibilities - Respect block boundaries - Keep chunk size below model limits - Optionally create overlapping context ## Must Not - Detect discussion topics - Merge discussion content - Interpret meaning ## Current Status Implemented --- # Stage 4 – Topic Segmentation ## Purpose Identify the thematic structure of the discussion. This is considered the central research problem of the Meeting Lab. ## Input Discussion blocks or technical chunks. ## Output Topics consisting of one or more discussion segments. Example: ```json { "topic": "Ventilation", "segments": [ { "start_block": 40, "end_block": 152 }, { "start_block": 1618, "end_block": 1697 } ] } ``` ## Processing Type LLM ## Responsibilities - Detect topic start - Detect topic end - Detect topic changes - Detect resumed topics - Associate discussion blocks with topics ## Must Not - Extract facts - Detect todos - Generate summaries ## Current Status Planned --- # Stage 5 – Specialized Extraction ## Purpose Extract one specific type of information from each topic. Every extractor performs exactly one task. ## Planned Extractors ```text extract_facts.py extract_questions.py extract_positions.py extract_decisions.py extract_todos.py extract_technical.py ``` Each extractor has: - one prompt - one responsibility - one output schema ## Processing Type LLM ## Current Status Prototype exists as a combined extractor. --- # Stage 6 – Consolidation ## Purpose Merge analysis results originating from different discussion segments. ## Input Extraction results. ## Output Unified topic representation. ## Responsibilities - Merge duplicates - Merge complementary information - Preserve contradictions - Separate positions from decisions - Combine related todos ## Processing Type Hybrid Deterministic wherever possible. LLM support only if necessary. ## Current Status Planned --- # Stage 7 – Canonical Meeting Knowledge ## Purpose Produce the canonical semantic representation of one meeting. This representation is the single source of truth for all downstream outputs. Example: ```json { "topics": [ { "title": "...", "facts": [], "questions": [], "positions": [], "decisions": [], "todos": [] } ] } ``` The exact schema will evolve during development. Conceptually, the Canonical Meeting Knowledge should include: - meeting metadata - topics - facts - decisions - action items - open questions - positions - technical information - rationale and discussion context - contradictions or uncertainty - source references and evidence The detailed schema remains future implementation work. ## Current Status Planned --- # Stage 8 – Output View Rendering ## Purpose Render purpose-specific outputs from Canonical Meeting Knowledge. The planned output products are: - Working Protocol (`working_protocol.md`, Arbeitsprotokoll) - Distribution Protocol (`distribution_protocol.md`, Verteilerprotokoll) - Knowledge Objects, rendered as a Knowledge-base Entry (`knowledge_entry.md`) and later stored in a structured format such as `knowledge_entry.json` (Wissensdatenbankeintrag) Output rendering never performs additional analysis. It only transforms existing structured information into the required view. The outputs are rendered in parallel from the canonical representation. The Distribution Protocol is not derived from the Working Protocol, and Knowledge Objects are not derived from either protocol. Rendering may be deterministic, template-based or LLM-assisted depending on the output and implementation maturity. Completeness differs by output: - The Working Protocol optimizes for recall and traceability. - The Distribution Protocol optimizes for relevance and brevity. - Knowledge Objects optimize for durability and reuse. ## Processing Type LLM ## Current Status Planned --- # Data Flow Each stage consumes only the output of the previous stage. ```text Stage N │ Structured Output │ ▼ Stage N + 1 ``` Intermediate results remain available for inspection, testing and experimentation. --- # Guiding Principles Every pipeline stage should satisfy the following rules. ## Single Responsibility One module. One task. ## Explicit Input Every stage expects a clearly defined input format. ## Explicit Output Every stage produces a clearly defined output format. ## Independent Evaluation Each stage should be testable without executing the entire pipeline. ## Replaceable Components A processing stage may be replaced by another implementation as long as it preserves the same interface. --- # Current Development Roadmap ```text ✔ Normalization ⬜ Discussion Blocks ✔ Technical Chunking ⬜ Topic Segmentation ⬜ Specialized Extraction ⬜ Consolidation ⬜ Canonical Meeting Knowledge ⬜ Output View Rendering ``` The immediate development focus is **Topic Segmentation**, as it provides the semantic structure on which all subsequent processing stages depend.