- establish Canonical Meeting Knowledge as the semantic source of truth - introduce Output View Rendering architecture - define Working Protocol, Distribution Protocol and Knowledge Objects as parallel renderers - document renderer responsibilities and terminology - clarify future Knowledge Object architecture - document long-term reuse for enterprise knowledge systems
462 lines
7.4 KiB
Markdown
462 lines
7.4 KiB
Markdown
# Pipeline
|
||
|
||
## Purpose
|
||
|
||
This document describes the processing pipeline of the Meeting Lab.
|
||
|
||
Unlike `architecture.md`, which describes the overall system, this document focuses on the individual processing stages, their responsibilities and the data flowing between them.
|
||
|
||
The guiding principle is simple:
|
||
|
||
> **Each processing stage has exactly one responsibility.**
|
||
|
||
---
|
||
|
||
# Pipeline Overview
|
||
|
||
```text
|
||
Whisper Transcript
|
||
│
|
||
▼
|
||
Normalization
|
||
│
|
||
▼
|
||
Discussion Blocks
|
||
│
|
||
▼
|
||
Technical Chunking
|
||
│
|
||
▼
|
||
Topic Segmentation
|
||
│
|
||
▼
|
||
Specialized Extraction
|
||
│
|
||
▼
|
||
Consolidation
|
||
│
|
||
▼
|
||
Canonical Meeting Knowledge
|
||
│
|
||
▼
|
||
Output View Rendering
|
||
│
|
||
├── Working Protocol
|
||
├── Distribution Protocol
|
||
└── Knowledge Objects
|
||
```
|
||
|
||
Each stage receives a well-defined input and produces a well-defined output.
|
||
|
||
---
|
||
|
||
# Stage 1 – Normalization
|
||
|
||
## Purpose
|
||
|
||
Remove transcription artifacts without changing the meaning of the discussion.
|
||
|
||
## Input
|
||
|
||
Raw transcript generated by Whisper.
|
||
|
||
## Output
|
||
|
||
Normalized transcript.
|
||
|
||
Change log containing every modification.
|
||
|
||
## Processing Type
|
||
|
||
Deterministic
|
||
|
||
## Responsibilities
|
||
|
||
- Remove filler words
|
||
- Remove immediate duplicate words
|
||
- Remove immediate duplicate short phrases
|
||
- Normalize whitespace
|
||
- Preserve all semantic content
|
||
|
||
## Must Not
|
||
|
||
- Rephrase text
|
||
- Summarize
|
||
- Interpret statements
|
||
- Correct factual content
|
||
|
||
## Current Status
|
||
|
||
Implemented
|
||
|
||
---
|
||
|
||
# Stage 2 – Discussion Blocks
|
||
|
||
## Purpose
|
||
|
||
Convert the transcript into stable processing units.
|
||
|
||
Discussion blocks are the smallest semantic unit used throughout the pipeline.
|
||
|
||
## Input
|
||
|
||
Normalized transcript.
|
||
|
||
## Output
|
||
|
||
Ordered list of discussion blocks.
|
||
|
||
Example:
|
||
|
||
```json
|
||
{
|
||
"block_id": 42,
|
||
"speaker": "A",
|
||
"start": 351.2,
|
||
"end": 367.8,
|
||
"text": "..."
|
||
}
|
||
```
|
||
|
||
## Processing Type
|
||
|
||
Deterministic
|
||
|
||
## Responsibilities
|
||
|
||
- Create stable identifiers
|
||
- Preserve ordering
|
||
- Preserve timestamps
|
||
- Preserve speaker information where available
|
||
|
||
## Current Status
|
||
|
||
Planned
|
||
|
||
---
|
||
|
||
# Stage 3 – Technical Chunking
|
||
|
||
## Purpose
|
||
|
||
Split large meetings into model-sized chunks.
|
||
|
||
Chunking exists only because language models have limited context windows.
|
||
|
||
## Input
|
||
|
||
Discussion blocks.
|
||
|
||
## Output
|
||
|
||
Chunk manifest and chunk files.
|
||
|
||
## Processing Type
|
||
|
||
Deterministic
|
||
|
||
## Responsibilities
|
||
|
||
- Respect block boundaries
|
||
- Keep chunk size below model limits
|
||
- Optionally create overlapping context
|
||
|
||
## Must Not
|
||
|
||
- Detect discussion topics
|
||
- Merge discussion content
|
||
- Interpret meaning
|
||
|
||
## Current Status
|
||
|
||
Implemented
|
||
|
||
---
|
||
|
||
# Stage 4 – Topic Segmentation
|
||
|
||
## Purpose
|
||
|
||
Identify the thematic structure of the discussion.
|
||
|
||
This is considered the central research problem of the Meeting Lab.
|
||
|
||
## Input
|
||
|
||
Discussion blocks or technical chunks.
|
||
|
||
## Output
|
||
|
||
Topics consisting of one or more discussion segments.
|
||
|
||
Example:
|
||
|
||
```json
|
||
{
|
||
"topic": "Ventilation",
|
||
"segments": [
|
||
{
|
||
"start_block": 40,
|
||
"end_block": 152
|
||
},
|
||
{
|
||
"start_block": 1618,
|
||
"end_block": 1697
|
||
}
|
||
]
|
||
}
|
||
```
|
||
|
||
## Processing Type
|
||
|
||
LLM
|
||
|
||
## Responsibilities
|
||
|
||
- Detect topic start
|
||
- Detect topic end
|
||
- Detect topic changes
|
||
- Detect resumed topics
|
||
- Associate discussion blocks with topics
|
||
|
||
## Must Not
|
||
|
||
- Extract facts
|
||
- Detect todos
|
||
- Generate summaries
|
||
|
||
## Current Status
|
||
|
||
Planned
|
||
|
||
---
|
||
|
||
# Stage 5 – Specialized Extraction
|
||
|
||
## Purpose
|
||
|
||
Extract one specific type of information from each topic.
|
||
|
||
Every extractor performs exactly one task.
|
||
|
||
## Planned Extractors
|
||
|
||
```text
|
||
extract_facts.py
|
||
extract_questions.py
|
||
extract_positions.py
|
||
extract_decisions.py
|
||
extract_todos.py
|
||
extract_technical.py
|
||
```
|
||
|
||
Each extractor has:
|
||
|
||
- one prompt
|
||
- one responsibility
|
||
- one output schema
|
||
|
||
## Processing Type
|
||
|
||
LLM
|
||
|
||
## Current Status
|
||
|
||
Prototype exists as a combined extractor.
|
||
|
||
---
|
||
|
||
# Stage 6 – Consolidation
|
||
|
||
## Purpose
|
||
|
||
Merge analysis results originating from different discussion segments.
|
||
|
||
## Input
|
||
|
||
Extraction results.
|
||
|
||
## Output
|
||
|
||
Unified topic representation.
|
||
|
||
## Responsibilities
|
||
|
||
- Merge duplicates
|
||
- Merge complementary information
|
||
- Preserve contradictions
|
||
- Separate positions from decisions
|
||
- Combine related todos
|
||
|
||
## Processing Type
|
||
|
||
Hybrid
|
||
|
||
Deterministic wherever possible.
|
||
|
||
LLM support only if necessary.
|
||
|
||
## Current Status
|
||
|
||
Planned
|
||
|
||
---
|
||
|
||
# Stage 7 – Canonical Meeting Knowledge
|
||
|
||
## Purpose
|
||
|
||
Produce the canonical semantic representation of one meeting.
|
||
|
||
This representation is the single source of truth for all downstream outputs.
|
||
|
||
Example:
|
||
|
||
```json
|
||
{
|
||
"topics": [
|
||
{
|
||
"title": "...",
|
||
"facts": [],
|
||
"questions": [],
|
||
"positions": [],
|
||
"decisions": [],
|
||
"todos": []
|
||
}
|
||
]
|
||
}
|
||
```
|
||
|
||
The exact schema will evolve during development.
|
||
|
||
Conceptually, the Canonical Meeting Knowledge should include:
|
||
|
||
- meeting metadata
|
||
- topics
|
||
- facts
|
||
- decisions
|
||
- action items
|
||
- open questions
|
||
- positions
|
||
- technical information
|
||
- rationale and discussion context
|
||
- contradictions or uncertainty
|
||
- source references and evidence
|
||
|
||
The detailed schema remains future implementation work.
|
||
|
||
## Current Status
|
||
|
||
Planned
|
||
|
||
---
|
||
|
||
# Stage 8 – Output View Rendering
|
||
|
||
## Purpose
|
||
|
||
Render purpose-specific outputs from Canonical Meeting Knowledge.
|
||
|
||
The planned output products are:
|
||
|
||
- Working Protocol (`working_protocol.md`, Arbeitsprotokoll)
|
||
- Distribution Protocol (`distribution_protocol.md`, Verteilerprotokoll)
|
||
- Knowledge Objects, rendered as a Knowledge-base Entry (`knowledge_entry.md`)
|
||
and later stored in a structured format such as `knowledge_entry.json`
|
||
(Wissensdatenbankeintrag)
|
||
|
||
Output rendering never performs additional analysis.
|
||
|
||
It only transforms existing structured information into the required view.
|
||
|
||
The outputs are rendered in parallel from the canonical representation. The
|
||
Distribution Protocol is not derived from the Working Protocol, and Knowledge
|
||
Objects are not derived from either protocol.
|
||
|
||
Rendering may be deterministic, template-based or LLM-assisted depending on the
|
||
output and implementation maturity.
|
||
|
||
Completeness differs by output:
|
||
|
||
- The Working Protocol optimizes for recall and traceability.
|
||
- The Distribution Protocol optimizes for relevance and brevity.
|
||
- Knowledge Objects optimize for durability and reuse.
|
||
|
||
## Processing Type
|
||
|
||
LLM
|
||
|
||
## Current Status
|
||
|
||
Planned
|
||
|
||
---
|
||
|
||
# Data Flow
|
||
|
||
Each stage consumes only the output of the previous stage.
|
||
|
||
```text
|
||
Stage N
|
||
│
|
||
Structured Output
|
||
│
|
||
▼
|
||
Stage N + 1
|
||
```
|
||
|
||
Intermediate results remain available for inspection, testing and experimentation.
|
||
|
||
---
|
||
|
||
# Guiding Principles
|
||
|
||
Every pipeline stage should satisfy the following rules.
|
||
|
||
## Single Responsibility
|
||
|
||
One module.
|
||
|
||
One task.
|
||
|
||
## Explicit Input
|
||
|
||
Every stage expects a clearly defined input format.
|
||
|
||
## Explicit Output
|
||
|
||
Every stage produces a clearly defined output format.
|
||
|
||
## Independent Evaluation
|
||
|
||
Each stage should be testable without executing the entire pipeline.
|
||
|
||
## Replaceable Components
|
||
|
||
A processing stage may be replaced by another implementation as long as it preserves the same interface.
|
||
|
||
---
|
||
|
||
# Current Development Roadmap
|
||
|
||
```text
|
||
✔ Normalization
|
||
|
||
⬜ Discussion Blocks
|
||
|
||
✔ Technical Chunking
|
||
|
||
⬜ Topic Segmentation
|
||
|
||
⬜ Specialized Extraction
|
||
|
||
⬜ Consolidation
|
||
|
||
⬜ Canonical Meeting Knowledge
|
||
|
||
⬜ Output View Rendering
|
||
```
|
||
|
||
The immediate development focus is **Topic Segmentation**, as it provides the semantic structure on which all subsequent processing stages depend.
|