Files
meeting-lab/docs/pipeline.md
T
2026-07-21 16:32:49 +02:00

428 lines
6.2 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Pipeline
## Purpose
This document describes the processing pipeline of the Meeting Lab.
Unlike `architecture.md`, which describes the overall system, this document focuses on the individual processing stages, their responsibilities and the data flowing between them.
The guiding principle is simple:
> **Each processing stage has exactly one responsibility.**
---
# Pipeline Overview
```text
Whisper Transcript
│
▼
Normalization
│
▼
Discussion Blocks
│
▼
Technical Chunking
│
▼
Topic Segmentation
│
▼
Specialized Extraction
│
▼
Consolidation
│
▼
Structured Meeting
│
▼
Protocol Generation
```
Each stage receives a well-defined input and produces a well-defined output.
---
# Stage 1 – Normalization
## Purpose
Remove transcription artifacts without changing the meaning of the discussion.
## Input
Raw transcript generated by Whisper.
## Output
Normalized transcript.
Change log containing every modification.
## Processing Type
Deterministic
## Responsibilities
- Remove filler words
- Remove immediate duplicate words
- Remove immediate duplicate short phrases
- Normalize whitespace
- Preserve all semantic content
## Must Not
- Rephrase text
- Summarize
- Interpret statements
- Correct factual content
## Current Status
Implemented
---
# Stage 2 – Discussion Blocks
## Purpose
Convert the transcript into stable processing units.
Discussion blocks are the smallest semantic unit used throughout the pipeline.
## Input
Normalized transcript.
## Output
Ordered list of discussion blocks.
Example:
```json
{
"block_id": 42,
"speaker": "A",
"start": 351.2,
"end": 367.8,
"text": "..."
}
```
## Processing Type
Deterministic
## Responsibilities
- Create stable identifiers
- Preserve ordering
- Preserve timestamps
- Preserve speaker information where available
## Current Status
Planned
---
# Stage 3 – Technical Chunking
## Purpose
Split large meetings into model-sized chunks.
Chunking exists only because language models have limited context windows.
## Input
Discussion blocks.
## Output
Chunk manifest and chunk files.
## Processing Type
Deterministic
## Responsibilities
- Respect block boundaries
- Keep chunk size below model limits
- Optionally create overlapping context
## Must Not
- Detect discussion topics
- Merge discussion content
- Interpret meaning
## Current Status
Implemented
---
# Stage 4 – Topic Segmentation
## Purpose
Identify the thematic structure of the discussion.
This is considered the central research problem of the Meeting Lab.
## Input
Discussion blocks or technical chunks.
## Output
Topics consisting of one or more discussion segments.
Example:
```json
{
"topic": "Ventilation",
"segments": [
{
"start_block": 40,
"end_block": 152
},
{
"start_block": 1618,
"end_block": 1697
}
]
}
```
## Processing Type
LLM
## Responsibilities
- Detect topic start
- Detect topic end
- Detect topic changes
- Detect resumed topics
- Associate discussion blocks with topics
## Must Not
- Extract facts
- Detect todos
- Generate summaries
## Current Status
Planned
---
# Stage 5 – Specialized Extraction
## Purpose
Extract one specific type of information from each topic.
Every extractor performs exactly one task.
## Planned Extractors
```text
extract_facts.py
extract_questions.py
extract_positions.py
extract_decisions.py
extract_todos.py
extract_technical.py
```
Each extractor has:
- one prompt
- one responsibility
- one output schema
## Processing Type
LLM
## Current Status
Prototype exists as a combined extractor.
---
# Stage 6 – Consolidation
## Purpose
Merge analysis results originating from different discussion segments.
## Input
Extraction results.
## Output
Unified topic representation.
## Responsibilities
- Merge duplicates
- Merge complementary information
- Preserve contradictions
- Separate positions from decisions
- Combine related todos
## Processing Type
Hybrid
Deterministic wherever possible.
LLM support only if necessary.
## Current Status
Planned
---
# Stage 7 – Structured Meeting
## Purpose
Produce a complete machine-readable representation of the meeting.
This is the primary output of the analysis pipeline.
Example:
```json
{
"topics": [
{
"title": "...",
"facts": [],
"questions": [],
"positions": [],
"decisions": [],
"todos": []
}
]
}
```
The exact schema will evolve during development.
## Current Status
Planned
---
# Stage 8 – Protocol Generation
## Purpose
Generate human-readable documents from structured meeting data.
Possible outputs include:
- Full protocol
- Executive summary
- Action list
- Decision log
- Technical report
Protocol generation never performs additional analysis.
It only transforms existing structured information into readable text.
## Processing Type
LLM
## Current Status
Planned
---
# Data Flow
Each stage consumes only the output of the previous stage.
```text
Stage N
│
Structured Output
│
▼
Stage N + 1
```
Intermediate results remain available for inspection, testing and experimentation.
---
# Guiding Principles
Every pipeline stage should satisfy the following rules.
## Single Responsibility
One module.
One task.
## Explicit Input
Every stage expects a clearly defined input format.
## Explicit Output
Every stage produces a clearly defined output format.
## Independent Evaluation
Each stage should be testable without executing the entire pipeline.
## Replaceable Components
A processing stage may be replaced by another implementation as long as it preserves the same interface.
---
# Current Development Roadmap
```text
✔ Normalization
⬜ Discussion Blocks
✔ Technical Chunking
⬜ Topic Segmentation
⬜ Specialized Extraction
⬜ Consolidation
⬜ Structured Meeting
⬜ Protocol Generation
```
The immediate development focus is **Topic Segmentation**, as it provides the semantic structure on which all subsequent processing stages depend.