Initial project structure
This commit is contained in:
@@ -0,0 +1,428 @@
|
||||
# Pipeline
|
||||
|
||||
## Purpose
|
||||
|
||||
This document describes the processing pipeline of the Meeting Lab.
|
||||
|
||||
Unlike `architecture.md`, which describes the overall system, this document focuses on the individual processing stages, their responsibilities and the data flowing between them.
|
||||
|
||||
The guiding principle is simple:
|
||||
|
||||
> **Each processing stage has exactly one responsibility.**
|
||||
|
||||
---
|
||||
|
||||
# Pipeline Overview
|
||||
|
||||
```text
|
||||
Whisper Transcript
|
||||
│
|
||||
▼
|
||||
Normalization
|
||||
│
|
||||
▼
|
||||
Discussion Blocks
|
||||
│
|
||||
▼
|
||||
Technical Chunking
|
||||
│
|
||||
▼
|
||||
Topic Segmentation
|
||||
│
|
||||
▼
|
||||
Specialized Extraction
|
||||
│
|
||||
▼
|
||||
Consolidation
|
||||
│
|
||||
▼
|
||||
Structured Meeting
|
||||
│
|
||||
▼
|
||||
Protocol Generation
|
||||
```
|
||||
|
||||
Each stage receives a well-defined input and produces a well-defined output.
|
||||
|
||||
---
|
||||
|
||||
# Stage 1 – Normalization
|
||||
|
||||
## Purpose
|
||||
|
||||
Remove transcription artifacts without changing the meaning of the discussion.
|
||||
|
||||
## Input
|
||||
|
||||
Raw transcript generated by Whisper.
|
||||
|
||||
## Output
|
||||
|
||||
Normalized transcript.
|
||||
|
||||
Change log containing every modification.
|
||||
|
||||
## Processing Type
|
||||
|
||||
Deterministic
|
||||
|
||||
## Responsibilities
|
||||
|
||||
- Remove filler words
|
||||
- Remove immediate duplicate words
|
||||
- Remove immediate duplicate short phrases
|
||||
- Normalize whitespace
|
||||
- Preserve all semantic content
|
||||
|
||||
## Must Not
|
||||
|
||||
- Rephrase text
|
||||
- Summarize
|
||||
- Interpret statements
|
||||
- Correct factual content
|
||||
|
||||
## Current Status
|
||||
|
||||
Implemented
|
||||
|
||||
---
|
||||
|
||||
# Stage 2 – Discussion Blocks
|
||||
|
||||
## Purpose
|
||||
|
||||
Convert the transcript into stable processing units.
|
||||
|
||||
Discussion blocks are the smallest semantic unit used throughout the pipeline.
|
||||
|
||||
## Input
|
||||
|
||||
Normalized transcript.
|
||||
|
||||
## Output
|
||||
|
||||
Ordered list of discussion blocks.
|
||||
|
||||
Example:
|
||||
|
||||
```json
|
||||
{
|
||||
"block_id": 42,
|
||||
"speaker": "A",
|
||||
"start": 351.2,
|
||||
"end": 367.8,
|
||||
"text": "..."
|
||||
}
|
||||
```
|
||||
|
||||
## Processing Type
|
||||
|
||||
Deterministic
|
||||
|
||||
## Responsibilities
|
||||
|
||||
- Create stable identifiers
|
||||
- Preserve ordering
|
||||
- Preserve timestamps
|
||||
- Preserve speaker information where available
|
||||
|
||||
## Current Status
|
||||
|
||||
Planned
|
||||
|
||||
---
|
||||
|
||||
# Stage 3 – Technical Chunking
|
||||
|
||||
## Purpose
|
||||
|
||||
Split large meetings into model-sized chunks.
|
||||
|
||||
Chunking exists only because language models have limited context windows.
|
||||
|
||||
## Input
|
||||
|
||||
Discussion blocks.
|
||||
|
||||
## Output
|
||||
|
||||
Chunk manifest and chunk files.
|
||||
|
||||
## Processing Type
|
||||
|
||||
Deterministic
|
||||
|
||||
## Responsibilities
|
||||
|
||||
- Respect block boundaries
|
||||
- Keep chunk size below model limits
|
||||
- Optionally create overlapping context
|
||||
|
||||
## Must Not
|
||||
|
||||
- Detect discussion topics
|
||||
- Merge discussion content
|
||||
- Interpret meaning
|
||||
|
||||
## Current Status
|
||||
|
||||
Implemented
|
||||
|
||||
---
|
||||
|
||||
# Stage 4 – Topic Segmentation
|
||||
|
||||
## Purpose
|
||||
|
||||
Identify the thematic structure of the discussion.
|
||||
|
||||
This is considered the central research problem of the Meeting Lab.
|
||||
|
||||
## Input
|
||||
|
||||
Discussion blocks or technical chunks.
|
||||
|
||||
## Output
|
||||
|
||||
Topics consisting of one or more discussion segments.
|
||||
|
||||
Example:
|
||||
|
||||
```json
|
||||
{
|
||||
"topic": "Ventilation",
|
||||
"segments": [
|
||||
{
|
||||
"start_block": 40,
|
||||
"end_block": 152
|
||||
},
|
||||
{
|
||||
"start_block": 1618,
|
||||
"end_block": 1697
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
## Processing Type
|
||||
|
||||
LLM
|
||||
|
||||
## Responsibilities
|
||||
|
||||
- Detect topic start
|
||||
- Detect topic end
|
||||
- Detect topic changes
|
||||
- Detect resumed topics
|
||||
- Associate discussion blocks with topics
|
||||
|
||||
## Must Not
|
||||
|
||||
- Extract facts
|
||||
- Detect todos
|
||||
- Generate summaries
|
||||
|
||||
## Current Status
|
||||
|
||||
Planned
|
||||
|
||||
---
|
||||
|
||||
# Stage 5 – Specialized Extraction
|
||||
|
||||
## Purpose
|
||||
|
||||
Extract one specific type of information from each topic.
|
||||
|
||||
Every extractor performs exactly one task.
|
||||
|
||||
## Planned Extractors
|
||||
|
||||
```text
|
||||
extract_facts.py
|
||||
extract_questions.py
|
||||
extract_positions.py
|
||||
extract_decisions.py
|
||||
extract_todos.py
|
||||
extract_technical.py
|
||||
```
|
||||
|
||||
Each extractor has:
|
||||
|
||||
- one prompt
|
||||
- one responsibility
|
||||
- one output schema
|
||||
|
||||
## Processing Type
|
||||
|
||||
LLM
|
||||
|
||||
## Current Status
|
||||
|
||||
Prototype exists as a combined extractor.
|
||||
|
||||
---
|
||||
|
||||
# Stage 6 – Consolidation
|
||||
|
||||
## Purpose
|
||||
|
||||
Merge analysis results originating from different discussion segments.
|
||||
|
||||
## Input
|
||||
|
||||
Extraction results.
|
||||
|
||||
## Output
|
||||
|
||||
Unified topic representation.
|
||||
|
||||
## Responsibilities
|
||||
|
||||
- Merge duplicates
|
||||
- Merge complementary information
|
||||
- Preserve contradictions
|
||||
- Separate positions from decisions
|
||||
- Combine related todos
|
||||
|
||||
## Processing Type
|
||||
|
||||
Hybrid
|
||||
|
||||
Deterministic wherever possible.
|
||||
|
||||
LLM support only if necessary.
|
||||
|
||||
## Current Status
|
||||
|
||||
Planned
|
||||
|
||||
---
|
||||
|
||||
# Stage 7 – Structured Meeting
|
||||
|
||||
## Purpose
|
||||
|
||||
Produce a complete machine-readable representation of the meeting.
|
||||
|
||||
This is the primary output of the analysis pipeline.
|
||||
|
||||
Example:
|
||||
|
||||
```json
|
||||
{
|
||||
"topics": [
|
||||
{
|
||||
"title": "...",
|
||||
"facts": [],
|
||||
"questions": [],
|
||||
"positions": [],
|
||||
"decisions": [],
|
||||
"todos": []
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
The exact schema will evolve during development.
|
||||
|
||||
## Current Status
|
||||
|
||||
Planned
|
||||
|
||||
---
|
||||
|
||||
# Stage 8 – Protocol Generation
|
||||
|
||||
## Purpose
|
||||
|
||||
Generate human-readable documents from structured meeting data.
|
||||
|
||||
Possible outputs include:
|
||||
|
||||
- Full protocol
|
||||
- Executive summary
|
||||
- Action list
|
||||
- Decision log
|
||||
- Technical report
|
||||
|
||||
Protocol generation never performs additional analysis.
|
||||
|
||||
It only transforms existing structured information into readable text.
|
||||
|
||||
## Processing Type
|
||||
|
||||
LLM
|
||||
|
||||
## Current Status
|
||||
|
||||
Planned
|
||||
|
||||
---
|
||||
|
||||
# Data Flow
|
||||
|
||||
Each stage consumes only the output of the previous stage.
|
||||
|
||||
```text
|
||||
Stage N
|
||||
│
|
||||
Structured Output
|
||||
│
|
||||
▼
|
||||
Stage N + 1
|
||||
```
|
||||
|
||||
Intermediate results remain available for inspection, testing and experimentation.
|
||||
|
||||
---
|
||||
|
||||
# Guiding Principles
|
||||
|
||||
Every pipeline stage should satisfy the following rules.
|
||||
|
||||
## Single Responsibility
|
||||
|
||||
One module.
|
||||
|
||||
One task.
|
||||
|
||||
## Explicit Input
|
||||
|
||||
Every stage expects a clearly defined input format.
|
||||
|
||||
## Explicit Output
|
||||
|
||||
Every stage produces a clearly defined output format.
|
||||
|
||||
## Independent Evaluation
|
||||
|
||||
Each stage should be testable without executing the entire pipeline.
|
||||
|
||||
## Replaceable Components
|
||||
|
||||
A processing stage may be replaced by another implementation as long as it preserves the same interface.
|
||||
|
||||
---
|
||||
|
||||
# Current Development Roadmap
|
||||
|
||||
```text
|
||||
✔ Normalization
|
||||
|
||||
⬜ Discussion Blocks
|
||||
|
||||
✔ Technical Chunking
|
||||
|
||||
⬜ Topic Segmentation
|
||||
|
||||
⬜ Specialized Extraction
|
||||
|
||||
⬜ Consolidation
|
||||
|
||||
⬜ Structured Meeting
|
||||
|
||||
⬜ Protocol Generation
|
||||
```
|
||||
|
||||
The immediate development focus is **Topic Segmentation**, as it provides the semantic structure on which all subsequent processing stages depend.
|
||||
Reference in New Issue
Block a user