Initial project structure

This commit is contained in:
2026-07-21 16:32:49 +02:00
commit b2c195c365
45 changed files with 999 additions and 0 deletions
+428
View File
@@ -0,0 +1,428 @@
# Pipeline
## Purpose
This document describes the processing pipeline of the Meeting Lab.
Unlike `architecture.md`, which describes the overall system, this document focuses on the individual processing stages, their responsibilities and the data flowing between them.
The guiding principle is simple:
> **Each processing stage has exactly one responsibility.**
---
# Pipeline Overview
```text
Whisper Transcript
│
▼
Normalization
│
▼
Discussion Blocks
│
▼
Technical Chunking
│
▼
Topic Segmentation
│
▼
Specialized Extraction
│
▼
Consolidation
│
▼
Structured Meeting
│
▼
Protocol Generation
```
Each stage receives a well-defined input and produces a well-defined output.
---
# Stage 1 – Normalization
## Purpose
Remove transcription artifacts without changing the meaning of the discussion.
## Input
Raw transcript generated by Whisper.
## Output
Normalized transcript.
Change log containing every modification.
## Processing Type
Deterministic
## Responsibilities
- Remove filler words
- Remove immediate duplicate words
- Remove immediate duplicate short phrases
- Normalize whitespace
- Preserve all semantic content
## Must Not
- Rephrase text
- Summarize
- Interpret statements
- Correct factual content
## Current Status
Implemented
---
# Stage 2 – Discussion Blocks
## Purpose
Convert the transcript into stable processing units.
Discussion blocks are the smallest semantic unit used throughout the pipeline.
## Input
Normalized transcript.
## Output
Ordered list of discussion blocks.
Example:
```json
{
"block_id": 42,
"speaker": "A",
"start": 351.2,
"end": 367.8,
"text": "..."
}
```
## Processing Type
Deterministic
## Responsibilities
- Create stable identifiers
- Preserve ordering
- Preserve timestamps
- Preserve speaker information where available
## Current Status
Planned
---
# Stage 3 – Technical Chunking
## Purpose
Split large meetings into model-sized chunks.
Chunking exists only because language models have limited context windows.
## Input
Discussion blocks.
## Output
Chunk manifest and chunk files.
## Processing Type
Deterministic
## Responsibilities
- Respect block boundaries
- Keep chunk size below model limits
- Optionally create overlapping context
## Must Not
- Detect discussion topics
- Merge discussion content
- Interpret meaning
## Current Status
Implemented
---
# Stage 4 – Topic Segmentation
## Purpose
Identify the thematic structure of the discussion.
This is considered the central research problem of the Meeting Lab.
## Input
Discussion blocks or technical chunks.
## Output
Topics consisting of one or more discussion segments.
Example:
```json
{
"topic": "Ventilation",
"segments": [
{
"start_block": 40,
"end_block": 152
},
{
"start_block": 1618,
"end_block": 1697
}
]
}
```
## Processing Type
LLM
## Responsibilities
- Detect topic start
- Detect topic end
- Detect topic changes
- Detect resumed topics
- Associate discussion blocks with topics
## Must Not
- Extract facts
- Detect todos
- Generate summaries
## Current Status
Planned
---
# Stage 5 – Specialized Extraction
## Purpose
Extract one specific type of information from each topic.
Every extractor performs exactly one task.
## Planned Extractors
```text
extract_facts.py
extract_questions.py
extract_positions.py
extract_decisions.py
extract_todos.py
extract_technical.py
```
Each extractor has:
- one prompt
- one responsibility
- one output schema
## Processing Type
LLM
## Current Status
Prototype exists as a combined extractor.
---
# Stage 6 – Consolidation
## Purpose
Merge analysis results originating from different discussion segments.
## Input
Extraction results.
## Output
Unified topic representation.
## Responsibilities
- Merge duplicates
- Merge complementary information
- Preserve contradictions
- Separate positions from decisions
- Combine related todos
## Processing Type
Hybrid
Deterministic wherever possible.
LLM support only if necessary.
## Current Status
Planned
---
# Stage 7 – Structured Meeting
## Purpose
Produce a complete machine-readable representation of the meeting.
This is the primary output of the analysis pipeline.
Example:
```json
{
"topics": [
{
"title": "...",
"facts": [],
"questions": [],
"positions": [],
"decisions": [],
"todos": []
}
]
}
```
The exact schema will evolve during development.
## Current Status
Planned
---
# Stage 8 – Protocol Generation
## Purpose
Generate human-readable documents from structured meeting data.
Possible outputs include:
- Full protocol
- Executive summary
- Action list
- Decision log
- Technical report
Protocol generation never performs additional analysis.
It only transforms existing structured information into readable text.
## Processing Type
LLM
## Current Status
Planned
---
# Data Flow
Each stage consumes only the output of the previous stage.
```text
Stage N
│
Structured Output
│
▼
Stage N + 1
```
Intermediate results remain available for inspection, testing and experimentation.
---
# Guiding Principles
Every pipeline stage should satisfy the following rules.
## Single Responsibility
One module.
One task.
## Explicit Input
Every stage expects a clearly defined input format.
## Explicit Output
Every stage produces a clearly defined output format.
## Independent Evaluation
Each stage should be testable without executing the entire pipeline.
## Replaceable Components
A processing stage may be replaced by another implementation as long as it preserves the same interface.
---
# Current Development Roadmap
```text
✔ Normalization
⬜ Discussion Blocks
✔ Technical Chunking
⬜ Topic Segmentation
⬜ Specialized Extraction
⬜ Consolidation
⬜ Structured Meeting
⬜ Protocol Generation
```
The immediate development focus is **Topic Segmentation**, as it provides the semantic structure on which all subsequent processing stages depend.