# Pipeline ## Purpose This document describes the processing pipeline of the Meeting Lab. Unlike `architecture.md`, which describes the overall system, this document focuses on the individual processing stages, their responsibilities and the data flowing between them. The guiding principle is simple: > **Each processing stage has exactly one responsibility.** --- # Pipeline Overview ```text Whisper Transcript │ ▼ Normalization │ ▼ Discussion Blocks │ ▼ Technical Chunking │ ▼ Topic Segmentation │ ▼ Specialized Extraction │ ▼ Consolidation │ ▼ Structured Meeting │ ▼ Protocol Generation ``` Each stage receives a well-defined input and produces a well-defined output. --- # Stage 1 – Normalization ## Purpose Remove transcription artifacts without changing the meaning of the discussion. ## Input Raw transcript generated by Whisper. ## Output Normalized transcript. Change log containing every modification. ## Processing Type Deterministic ## Responsibilities - Remove filler words - Remove immediate duplicate words - Remove immediate duplicate short phrases - Normalize whitespace - Preserve all semantic content ## Must Not - Rephrase text - Summarize - Interpret statements - Correct factual content ## Current Status Implemented --- # Stage 2 – Discussion Blocks ## Purpose Convert the transcript into stable processing units. Discussion blocks are the smallest semantic unit used throughout the pipeline. ## Input Normalized transcript. ## Output Ordered list of discussion blocks. Example: ```json { "block_id": 42, "speaker": "A", "start": 351.2, "end": 367.8, "text": "..." } ``` ## Processing Type Deterministic ## Responsibilities - Create stable identifiers - Preserve ordering - Preserve timestamps - Preserve speaker information where available ## Current Status Planned --- # Stage 3 – Technical Chunking ## Purpose Split large meetings into model-sized chunks. Chunking exists only because language models have limited context windows. ## Input Discussion blocks. ## Output Chunk manifest and chunk files. ## Processing Type Deterministic ## Responsibilities - Respect block boundaries - Keep chunk size below model limits - Optionally create overlapping context ## Must Not - Detect discussion topics - Merge discussion content - Interpret meaning ## Current Status Implemented --- # Stage 4 – Topic Segmentation ## Purpose Identify the thematic structure of the discussion. This is considered the central research problem of the Meeting Lab. ## Input Discussion blocks or technical chunks. ## Output Topics consisting of one or more discussion segments. Example: ```json { "topic": "Ventilation", "segments": [ { "start_block": 40, "end_block": 152 }, { "start_block": 1618, "end_block": 1697 } ] } ``` ## Processing Type LLM ## Responsibilities - Detect topic start - Detect topic end - Detect topic changes - Detect resumed topics - Associate discussion blocks with topics ## Must Not - Extract facts - Detect todos - Generate summaries ## Current Status Planned --- # Stage 5 – Specialized Extraction ## Purpose Extract one specific type of information from each topic. Every extractor performs exactly one task. ## Planned Extractors ```text extract_facts.py extract_questions.py extract_positions.py extract_decisions.py extract_todos.py extract_technical.py ``` Each extractor has: - one prompt - one responsibility - one output schema ## Processing Type LLM ## Current Status Prototype exists as a combined extractor. --- # Stage 6 – Consolidation ## Purpose Merge analysis results originating from different discussion segments. ## Input Extraction results. ## Output Unified topic representation. ## Responsibilities - Merge duplicates - Merge complementary information - Preserve contradictions - Separate positions from decisions - Combine related todos ## Processing Type Hybrid Deterministic wherever possible. LLM support only if necessary. ## Current Status Planned --- # Stage 7 – Structured Meeting ## Purpose Produce a complete machine-readable representation of the meeting. This is the primary output of the analysis pipeline. Example: ```json { "topics": [ { "title": "...", "facts": [], "questions": [], "positions": [], "decisions": [], "todos": [] } ] } ``` The exact schema will evolve during development. ## Current Status Planned --- # Stage 8 – Protocol Generation ## Purpose Generate human-readable documents from structured meeting data. Possible outputs include: - Full protocol - Executive summary - Action list - Decision log - Technical report Protocol generation never performs additional analysis. It only transforms existing structured information into readable text. ## Processing Type LLM ## Current Status Planned --- # Data Flow Each stage consumes only the output of the previous stage. ```text Stage N │ Structured Output │ ▼ Stage N + 1 ``` Intermediate results remain available for inspection, testing and experimentation. --- # Guiding Principles Every pipeline stage should satisfy the following rules. ## Single Responsibility One module. One task. ## Explicit Input Every stage expects a clearly defined input format. ## Explicit Output Every stage produces a clearly defined output format. ## Independent Evaluation Each stage should be testable without executing the entire pipeline. ## Replaceable Components A processing stage may be replaced by another implementation as long as it preserves the same interface. --- # Current Development Roadmap ```text ✔ Normalization ⬜ Discussion Blocks ✔ Technical Chunking ⬜ Topic Segmentation ⬜ Specialized Extraction ⬜ Consolidation ⬜ Structured Meeting ⬜ Protocol Generation ``` The immediate development focus is **Topic Segmentation**, as it provides the semantic structure on which all subsequent processing stages depend.