428 lines
6.2 KiB
Markdown
428 lines
6.2 KiB
Markdown
# Pipeline
|
||
|
||
## Purpose
|
||
|
||
This document describes the processing pipeline of the Meeting Lab.
|
||
|
||
Unlike `architecture.md`, which describes the overall system, this document focuses on the individual processing stages, their responsibilities and the data flowing between them.
|
||
|
||
The guiding principle is simple:
|
||
|
||
> **Each processing stage has exactly one responsibility.**
|
||
|
||
---
|
||
|
||
# Pipeline Overview
|
||
|
||
```text
|
||
Whisper Transcript
|
||
│
|
||
▼
|
||
Normalization
|
||
│
|
||
▼
|
||
Discussion Blocks
|
||
│
|
||
▼
|
||
Technical Chunking
|
||
│
|
||
▼
|
||
Topic Segmentation
|
||
│
|
||
▼
|
||
Specialized Extraction
|
||
│
|
||
▼
|
||
Consolidation
|
||
│
|
||
▼
|
||
Structured Meeting
|
||
│
|
||
▼
|
||
Protocol Generation
|
||
```
|
||
|
||
Each stage receives a well-defined input and produces a well-defined output.
|
||
|
||
---
|
||
|
||
# Stage 1 – Normalization
|
||
|
||
## Purpose
|
||
|
||
Remove transcription artifacts without changing the meaning of the discussion.
|
||
|
||
## Input
|
||
|
||
Raw transcript generated by Whisper.
|
||
|
||
## Output
|
||
|
||
Normalized transcript.
|
||
|
||
Change log containing every modification.
|
||
|
||
## Processing Type
|
||
|
||
Deterministic
|
||
|
||
## Responsibilities
|
||
|
||
- Remove filler words
|
||
- Remove immediate duplicate words
|
||
- Remove immediate duplicate short phrases
|
||
- Normalize whitespace
|
||
- Preserve all semantic content
|
||
|
||
## Must Not
|
||
|
||
- Rephrase text
|
||
- Summarize
|
||
- Interpret statements
|
||
- Correct factual content
|
||
|
||
## Current Status
|
||
|
||
Implemented
|
||
|
||
---
|
||
|
||
# Stage 2 – Discussion Blocks
|
||
|
||
## Purpose
|
||
|
||
Convert the transcript into stable processing units.
|
||
|
||
Discussion blocks are the smallest semantic unit used throughout the pipeline.
|
||
|
||
## Input
|
||
|
||
Normalized transcript.
|
||
|
||
## Output
|
||
|
||
Ordered list of discussion blocks.
|
||
|
||
Example:
|
||
|
||
```json
|
||
{
|
||
"block_id": 42,
|
||
"speaker": "A",
|
||
"start": 351.2,
|
||
"end": 367.8,
|
||
"text": "..."
|
||
}
|
||
```
|
||
|
||
## Processing Type
|
||
|
||
Deterministic
|
||
|
||
## Responsibilities
|
||
|
||
- Create stable identifiers
|
||
- Preserve ordering
|
||
- Preserve timestamps
|
||
- Preserve speaker information where available
|
||
|
||
## Current Status
|
||
|
||
Planned
|
||
|
||
---
|
||
|
||
# Stage 3 – Technical Chunking
|
||
|
||
## Purpose
|
||
|
||
Split large meetings into model-sized chunks.
|
||
|
||
Chunking exists only because language models have limited context windows.
|
||
|
||
## Input
|
||
|
||
Discussion blocks.
|
||
|
||
## Output
|
||
|
||
Chunk manifest and chunk files.
|
||
|
||
## Processing Type
|
||
|
||
Deterministic
|
||
|
||
## Responsibilities
|
||
|
||
- Respect block boundaries
|
||
- Keep chunk size below model limits
|
||
- Optionally create overlapping context
|
||
|
||
## Must Not
|
||
|
||
- Detect discussion topics
|
||
- Merge discussion content
|
||
- Interpret meaning
|
||
|
||
## Current Status
|
||
|
||
Implemented
|
||
|
||
---
|
||
|
||
# Stage 4 – Topic Segmentation
|
||
|
||
## Purpose
|
||
|
||
Identify the thematic structure of the discussion.
|
||
|
||
This is considered the central research problem of the Meeting Lab.
|
||
|
||
## Input
|
||
|
||
Discussion blocks or technical chunks.
|
||
|
||
## Output
|
||
|
||
Topics consisting of one or more discussion segments.
|
||
|
||
Example:
|
||
|
||
```json
|
||
{
|
||
"topic": "Ventilation",
|
||
"segments": [
|
||
{
|
||
"start_block": 40,
|
||
"end_block": 152
|
||
},
|
||
{
|
||
"start_block": 1618,
|
||
"end_block": 1697
|
||
}
|
||
]
|
||
}
|
||
```
|
||
|
||
## Processing Type
|
||
|
||
LLM
|
||
|
||
## Responsibilities
|
||
|
||
- Detect topic start
|
||
- Detect topic end
|
||
- Detect topic changes
|
||
- Detect resumed topics
|
||
- Associate discussion blocks with topics
|
||
|
||
## Must Not
|
||
|
||
- Extract facts
|
||
- Detect todos
|
||
- Generate summaries
|
||
|
||
## Current Status
|
||
|
||
Planned
|
||
|
||
---
|
||
|
||
# Stage 5 – Specialized Extraction
|
||
|
||
## Purpose
|
||
|
||
Extract one specific type of information from each topic.
|
||
|
||
Every extractor performs exactly one task.
|
||
|
||
## Planned Extractors
|
||
|
||
```text
|
||
extract_facts.py
|
||
extract_questions.py
|
||
extract_positions.py
|
||
extract_decisions.py
|
||
extract_todos.py
|
||
extract_technical.py
|
||
```
|
||
|
||
Each extractor has:
|
||
|
||
- one prompt
|
||
- one responsibility
|
||
- one output schema
|
||
|
||
## Processing Type
|
||
|
||
LLM
|
||
|
||
## Current Status
|
||
|
||
Prototype exists as a combined extractor.
|
||
|
||
---
|
||
|
||
# Stage 6 – Consolidation
|
||
|
||
## Purpose
|
||
|
||
Merge analysis results originating from different discussion segments.
|
||
|
||
## Input
|
||
|
||
Extraction results.
|
||
|
||
## Output
|
||
|
||
Unified topic representation.
|
||
|
||
## Responsibilities
|
||
|
||
- Merge duplicates
|
||
- Merge complementary information
|
||
- Preserve contradictions
|
||
- Separate positions from decisions
|
||
- Combine related todos
|
||
|
||
## Processing Type
|
||
|
||
Hybrid
|
||
|
||
Deterministic wherever possible.
|
||
|
||
LLM support only if necessary.
|
||
|
||
## Current Status
|
||
|
||
Planned
|
||
|
||
---
|
||
|
||
# Stage 7 – Structured Meeting
|
||
|
||
## Purpose
|
||
|
||
Produce a complete machine-readable representation of the meeting.
|
||
|
||
This is the primary output of the analysis pipeline.
|
||
|
||
Example:
|
||
|
||
```json
|
||
{
|
||
"topics": [
|
||
{
|
||
"title": "...",
|
||
"facts": [],
|
||
"questions": [],
|
||
"positions": [],
|
||
"decisions": [],
|
||
"todos": []
|
||
}
|
||
]
|
||
}
|
||
```
|
||
|
||
The exact schema will evolve during development.
|
||
|
||
## Current Status
|
||
|
||
Planned
|
||
|
||
---
|
||
|
||
# Stage 8 – Protocol Generation
|
||
|
||
## Purpose
|
||
|
||
Generate human-readable documents from structured meeting data.
|
||
|
||
Possible outputs include:
|
||
|
||
- Full protocol
|
||
- Executive summary
|
||
- Action list
|
||
- Decision log
|
||
- Technical report
|
||
|
||
Protocol generation never performs additional analysis.
|
||
|
||
It only transforms existing structured information into readable text.
|
||
|
||
## Processing Type
|
||
|
||
LLM
|
||
|
||
## Current Status
|
||
|
||
Planned
|
||
|
||
---
|
||
|
||
# Data Flow
|
||
|
||
Each stage consumes only the output of the previous stage.
|
||
|
||
```text
|
||
Stage N
|
||
│
|
||
Structured Output
|
||
│
|
||
▼
|
||
Stage N + 1
|
||
```
|
||
|
||
Intermediate results remain available for inspection, testing and experimentation.
|
||
|
||
---
|
||
|
||
# Guiding Principles
|
||
|
||
Every pipeline stage should satisfy the following rules.
|
||
|
||
## Single Responsibility
|
||
|
||
One module.
|
||
|
||
One task.
|
||
|
||
## Explicit Input
|
||
|
||
Every stage expects a clearly defined input format.
|
||
|
||
## Explicit Output
|
||
|
||
Every stage produces a clearly defined output format.
|
||
|
||
## Independent Evaluation
|
||
|
||
Each stage should be testable without executing the entire pipeline.
|
||
|
||
## Replaceable Components
|
||
|
||
A processing stage may be replaced by another implementation as long as it preserves the same interface.
|
||
|
||
---
|
||
|
||
# Current Development Roadmap
|
||
|
||
```text
|
||
✔ Normalization
|
||
|
||
⬜ Discussion Blocks
|
||
|
||
✔ Technical Chunking
|
||
|
||
⬜ Topic Segmentation
|
||
|
||
⬜ Specialized Extraction
|
||
|
||
⬜ Consolidation
|
||
|
||
⬜ Structured Meeting
|
||
|
||
⬜ Protocol Generation
|
||
```
|
||
|
||
The immediate development focus is **Topic Segmentation**, as it provides the semantic structure on which all subsequent processing stages depend. |