- preserve Working Protocol Synthesizer V0 as comparison baseline - introduce deterministic canonicalization stage - define semantic consolidator responsibilities - clarify Canonical Meeting Knowledge generation - document source-language output policy - align roadmap, architecture and experiment log
9.0 KiB
Pipeline
Purpose
This document describes the processing pipeline of the Meeting Lab.
Unlike architecture.md, which describes the overall system, this document focuses on the individual processing stages, their responsibilities and the data flowing between them.
The guiding principle is simple:
Each processing stage has exactly one responsibility.
Pipeline Overview
Whisper Transcript
│
▼
Normalization
│
▼
Discussion Blocks
│
▼
Technical Chunking
│
▼
Topic Segmentation
│
▼
Specialized Extraction
│
▼
Deterministic Canonicalization
│
▼
Semantic Consolidation
│
▼
Canonical Meeting Knowledge
│
▼
Output View Rendering
│
├── Working Protocol
├── Distribution Protocol
└── Knowledge Objects
Each stage receives a well-defined input and produces a well-defined output.
Stage 1 – Normalization
Purpose
Remove transcription artifacts without changing the meaning of the discussion.
Input
Raw transcript generated by Whisper.
Output
Normalized transcript.
Change log containing every modification.
Processing Type
Deterministic
Responsibilities
- Remove filler words
- Remove immediate duplicate words
- Remove immediate duplicate short phrases
- Normalize whitespace
- Preserve all semantic content
Must Not
- Rephrase text
- Summarize
- Interpret statements
- Correct factual content
Current Status
Implemented
Stage 2 – Discussion Blocks
Purpose
Convert the transcript into stable processing units.
Discussion blocks are the smallest semantic unit used throughout the pipeline.
Input
Normalized transcript.
Output
Ordered list of discussion blocks.
Example:
{
"block_id": 42,
"speaker": "A",
"start": 351.2,
"end": 367.8,
"text": "..."
}
Processing Type
Deterministic
Responsibilities
- Create stable identifiers
- Preserve ordering
- Preserve timestamps
- Preserve speaker information where available
Current Status
Planned
Stage 3 – Technical Chunking
Purpose
Split large meetings into model-sized chunks.
Chunking exists only because language models have limited context windows.
Input
Discussion blocks.
Output
Chunk manifest and chunk files.
Processing Type
Deterministic
Responsibilities
- Respect block boundaries
- Keep chunk size below model limits
- Optionally create overlapping context
Must Not
- Detect discussion topics
- Merge discussion content
- Interpret meaning
Current Status
Implemented
Stage 4 – Topic Segmentation
Purpose
Identify the thematic structure of the discussion.
This is considered the central research problem of the Meeting Lab.
Input
Discussion blocks or technical chunks.
Output
Topics consisting of one or more discussion segments.
Example:
{
"topic": "Ventilation",
"segments": [
{
"start_block": 40,
"end_block": 152
},
{
"start_block": 1618,
"end_block": 1697
}
]
}
Processing Type
LLM
Responsibilities
- Detect topic start
- Detect topic end
- Detect topic changes
- Detect resumed topics
- Associate discussion blocks with topics
Must Not
- Extract facts
- Detect todos
- Generate summaries
Current Status
Planned
Stage 5 – Specialized Extraction
Purpose
Extract one specific type of information from each topic.
Every extractor performs exactly one task.
Planned Extractors
extract_facts.py
extract_questions.py
extract_positions.py
extract_decisions.py
extract_todos.py
extract_technical.py
Each extractor has:
- one prompt
- one responsibility
- one output schema
Processing Type
LLM
Current Status
Prototype exists as a combined extractor.
Stage 6 – Deterministic Canonicalization
Purpose
Normalize raw chunk extraction JSON into stable canonical extraction objects without changing uncertain semantics.
Input
Chunk extraction JSON files.
Output
Validated canonical extraction objects with stable source references and IDs.
Responsibilities
- Validate extraction objects
- Normalize category names
- Normalize basic field structure
- Assign stable source references and IDs
- Preserve all source evidence
- Perform only safe deterministic cleanup
- Group exact duplicates where unambiguous
Must Not
- Use an LLM
- Perform uncertain semantic merging
- Infer missing information
- Drop source evidence
Processing Type
Deterministic Python
Current Status
Planned
Stage 7 – Semantic Consolidation
Purpose
Merge canonicalized extraction objects into an evidence-preserving semantic meeting representation.
Input
Canonicalized extraction objects.
Output
Canonical Meeting Knowledge.
Responsibilities
- Merge semantically equivalent statements
- Group content by topic
- Preserve evidence from all contributing chunks
- Mark contradictions and uncertainty
- Separate durable information from transient discussion
- Reconcile category shifts where supported by evidence
Must Not
- Directly write a protocol
- Invent information
- Drop conflicting evidence silently
Processing Type
Local LLM, with deterministic pre/post-processing where useful.
Current Status
Planned
Stage 8 – Canonical Meeting Knowledge
Purpose
Produce the canonical semantic representation of one meeting.
This representation is the single source of truth for all downstream outputs. It is a structured representation, preferably JSON, and is not itself a prose protocol.
Example:
{
"topics": [
{
"title": "...",
"facts": [],
"questions": [],
"positions": [],
"decisions": [],
"todos": []
}
]
}
The exact schema will evolve during development.
Conceptually, the Canonical Meeting Knowledge should include:
- meeting metadata
- topics
- facts
- decisions
- action items
- open questions
- positions
- technical information
- rationale and discussion context
- contradictions or uncertainty
- source references and evidence
The detailed schema remains future implementation work.
Current Status
Planned
Stage 9 – Output View Rendering
Purpose
Render purpose-specific outputs from Canonical Meeting Knowledge.
The planned output products are:
- Working Protocol (
working_protocol.md, Arbeitsprotokoll) - Distribution Protocol (
distribution_protocol.md, Verteilerprotokoll) - Knowledge Objects, rendered as a Knowledge-base Entry (
knowledge_entry.md) and later stored in a structured format such asknowledge_entry.json(Wissensdatenbankeintrag) - Later additional views such as action lists
Output rendering never performs additional analysis.
It only transforms existing structured information into the required view.
The outputs are rendered in parallel from the canonical representation. The Distribution Protocol is not derived from the Working Protocol, and Knowledge Objects are not derived from either protocol.
Rendering may be deterministic, template-based or LLM-assisted depending on the output and implementation maturity.
Completeness differs by output:
- The Working Protocol optimizes for recall and traceability.
- The Distribution Protocol optimizes for relevance and brevity.
- Knowledge Objects optimize for durability and reuse.
Rendered protocol language should normally match the dominant language of the source transcript or consolidated meeting knowledge unless an explicit output language is requested.
Processing Type
LLM
Current Status
Planned
Data Flow
Each stage consumes only the output of the previous stage.
Stage N
│
Structured Output
│
▼
Stage N + 1
Intermediate results remain available for inspection, testing and experimentation.
Guiding Principles
Every pipeline stage should satisfy the following rules.
Single Responsibility
One module.
One task.
Explicit Input
Every stage expects a clearly defined input format.
Explicit Output
Every stage produces a clearly defined output format.
Independent Evaluation
Each stage should be testable without executing the entire pipeline.
Replaceable Components
A processing stage may be replaced by another implementation as long as it preserves the same interface.
Current Development Roadmap
✔ Normalization
⬜ Discussion Blocks
✔ Technical Chunking
⬜ Topic Segmentation
⬜ Specialized Extraction
⬜ Deterministic Canonicalization
⬜ Semantic Consolidation
⬜ Canonical Meeting Knowledge
⬜ Output View Rendering
The immediate architecture focus is the Deterministic Canonicalizer followed by the Semantic Consolidator. These stages preserve source evidence, recover global context from independent chunk extractions and prepare Canonical Meeting Knowledge for parallel Output View rendering.