Files
meeting-lab/docs/pipeline.md
T
admin 9446c6e0be Document validation architecture and renderer faithfulness findings
- document Entity Registry and Meeting Context V2 architecture
- preserve meeting_context.yaml as the authoritative meeting-specific input
- define immutable authoritative metadata across all pipeline stages
- restrict Constraint Repair to deterministic structured-data operations
- record BUG-003 root cause and deferred entity-verification resolution
- document BUG-005 attendance-consistency design
- add BUG-006 renderer faithfulness root-cause analysis
- distinguish Engineering Readiness from Practical Usability
- update the persistent regression bug tracker
2026-08-03 16:05:47 +02:00

622 lines
13 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Pipeline
## Purpose
This document describes the processing pipeline of the Meeting Lab.
Unlike `architecture.md`, which describes the overall system, this document focuses on the individual processing stages, their responsibilities and the data flowing between them.
The guiding principle is simple:
> **Each processing stage has exactly one responsibility.**
Cross-cutting invariant:
> **Responsibility attribution requires explicit evidence.**
A person, team or department may be recorded as responsible only when the
source material explicitly assigns, accepts or confirms that responsibility.
The pipeline must not infer ownership from thematic proximity, discussion
participation, mentioning a task, commenting on another department,
organizational assumptions, likely job roles, speaker adjacency or model world
knowledge.
When support is incomplete or ambiguous, leave the responsible person unset,
mark the item as unclear where supported, and preserve the attribution
evidence. This applies to extraction, canonicalization, semantic consolidation,
Canonical Meeting Knowledge and every Output View renderer.
---
# Pipeline Overview
Current analysis pipeline:
```text
Whisper Transcript
│
▼
Normalization
│
▼
Discussion Blocks
│
▼
Technical Chunking
│
▼
Topic Segmentation
│
▼
Specialized Extraction
│
▼
Deterministic Canonicalization
│
▼
Semantic Consolidation
│
▼
Canonical Meeting Knowledge
│
▼
Output View Rendering
│
├── Working Protocol
├── Distribution Protocol
└── Knowledge Objects
```
Each stage receives a well-defined input and produces a well-defined output.
Meeting Context V1 exists as a manually maintained YAML metadata scaffold. It
is implemented for validation, optional `--meeting-context` use during chunk
extraction, deterministic prompt injection and minimal extraction JSON
provenance. Later Canonicalizer, Semantic Consolidator, Canonical Meeting
Knowledge and renderer integration remains future work.
Accepted future Meeting Context V2 preparation flow:
```text
Whisper
│
▼
Entity Detection
│
▼
User Confirmation
│
▼
Entity Registry Update
│
▼
Meeting Context Builder
│
▼
meeting_context.yaml
│
▼
Extraction Pipeline
```
This preparation flow is not implemented. It is the accepted long-term
direction for reducing manual Meeting Context work while preserving explicit
user control. The Entity Registry is the persistent cross-meeting knowledge
source for confirmed entities, aliases and organizational metadata. It stores
stable internal IDs and confirmed aliases, and never updates itself
automatically.
For each meeting run, `meeting_context.yaml` remains the authoritative
meeting-specific Point of Truth and reproducible input artifact consumed by the
pipeline. V2 changes how that artifact is prepared: it may be generated or
assisted from Registry data, user confirmations and meeting metadata. The
Registry must not override explicit meeting-specific confirmations, and
Registry changes after a meeting run must not silently change the historical
Meeting Context used for that run. Similarity suggestions and unknown entity
classifications require explicit user confirmation.
---
# Stage 1 – Normalization
## Purpose
Remove transcription artifacts without changing the meaning of the discussion.
## Input
Raw transcript generated by Whisper.
## Output
Normalized transcript.
Change log containing every modification.
## Processing Type
Deterministic
## Responsibilities
- Remove filler words
- Remove immediate duplicate words
- Remove immediate duplicate short phrases
- Normalize whitespace
- Preserve all semantic content
## Must Not
- Rephrase text
- Summarize
- Interpret statements
- Correct factual content
## Current Status
Implemented
---
# Stage 2 – Discussion Blocks
## Purpose
Convert the transcript into stable processing units.
Discussion blocks are the smallest semantic unit used throughout the pipeline.
## Input
Normalized transcript.
## Output
Ordered list of discussion blocks.
Example:
```json
{
"block_id": 42,
"speaker": "A",
"start": 351.2,
"end": 367.8,
"text": "..."
}
```
## Processing Type
Deterministic
## Responsibilities
- Create stable identifiers
- Preserve ordering
- Preserve timestamps
- Preserve speaker information where available
## Current Status
Implemented as Canonicalizer V1.
CLI:
```text
PYTHONPATH=src .venv/bin/python -m meeting_lab.consolidation.canonicalize \
samples/whisper/meeting_speech_cleaned_chunks \
-o /tmp/canonicalized_extractions.json
```
---
# Stage 3 – Technical Chunking
## Purpose
Split large meetings into model-sized chunks.
Chunking exists only because language models have limited context windows.
## Input
Discussion blocks.
## Output
Chunk manifest and chunk files.
## Processing Type
Deterministic
## Responsibilities
- Respect block boundaries
- Keep chunk size below model limits
- Optionally create overlapping context
## Must Not
- Detect discussion topics
- Merge discussion content
- Interpret meaning
## Current Status
Implemented
---
# Stage 4 – Topic Segmentation
## Purpose
Identify the thematic structure of the discussion.
This is considered the central research problem of the Meeting Lab.
## Input
Discussion blocks or technical chunks.
## Output
Topics consisting of one or more discussion segments.
Example:
```json
{
"topic": "Ventilation",
"segments": [
{
"start_block": 40,
"end_block": 152
},
{
"start_block": 1618,
"end_block": 1697
}
]
}
```
## Processing Type
LLM
## Responsibilities
- Detect topic start
- Detect topic end
- Detect topic changes
- Detect resumed topics
- Associate discussion blocks with topics
## Must Not
- Extract facts
- Detect todos
- Generate summaries
## Current Status
Implemented as Canonicalizer V1.
---
# Stage 5 – Specialized Extraction
## Purpose
Extract one specific type of information from each topic.
Every extractor performs exactly one task.
## Planned Extractors
```text
extract_facts.py
extract_questions.py
extract_positions.py
extract_decisions.py
extract_todos.py
extract_technical.py
```
Each extractor has:
- one prompt
- one responsibility
- one output schema
The current shared extraction prompt is assembled from:
- `common.md`
- `decisions.md`
- `todos.md`
The todo prompt requires explicit assignment, volunteering or acceptance before
recording a named responsible person. Meeting Context may validate identity,
role, department and attendance, but never establishes responsibility.
The decision prompt treats a personal commitment to concrete future work as a
todo unless the group separately establishes a binding outcome, rule, approval,
rejection, deferral, selection, process state or responsibility policy. The
same proposition should not be duplicated under decisions and todos; extract
both only for semantically separate propositions.
## Processing Type
LLM
## Current Status
Prototype exists as a combined extractor.
---
# Stage 6 – Deterministic Canonicalization
## Purpose
Normalize raw chunk extraction JSON into stable canonical extraction objects
without changing uncertain semantics.
## Input
Chunk extraction JSON files.
## Output
Validated canonical extraction objects with stable source references and IDs.
## Responsibilities
- Validate extraction objects
- Normalize category names
- Normalize basic field structure
- Assign stable source references and IDs
- Preserve all source evidence
- Perform only safe deterministic cleanup
- Group exact duplicates where unambiguous
## Must Not
- Use an LLM
- Perform uncertain semantic merging
- Infer missing information
- Drop source evidence
## Processing Type
Deterministic Python
## Current Status
Planned
---
# Stage 7 – Semantic Consolidation
## Purpose
Merge canonicalized extraction objects into an evidence-preserving semantic
meeting representation.
## Input
Canonicalized extraction objects.
## Output
Semantic Consolidator V0 output is `consolidated_extractions.json` with fact
groups and unchanged non-fact items.
Future broader semantic consolidation should produce Canonical Meeting
Knowledge.
## Responsibilities
- V0: merge semantically equivalent fact items only
- V0: preserve all non-fact categories unchanged
- V0: validate that every source fact ID appears exactly once
- Future: merge semantically equivalent statements across categories
- Future: group content by topic
- Preserve evidence from all contributing chunks
- Future: mark contradictions and uncertainty
- Future: separate durable information from transient discussion
- Future: reconcile category shifts where supported by evidence
## Must Not
- Directly write a protocol
- Invent information
- Drop conflicting evidence silently
## Processing Type
Local LLM, with deterministic pre/post-processing where useful.
## Current Status
Semantic Consolidator V0 is implemented and experimentally validated for
facts-only conservative duplicate detection. Broader semantic consolidation and
Canonical Meeting Knowledge generation remain planned.
---
# Stage 8 – Canonical Meeting Knowledge
## Purpose
Produce the canonical semantic representation of one meeting.
This representation is the single source of truth for all downstream outputs.
It is a structured representation, preferably JSON, and is not itself a prose
protocol.
Example:
```json
{
"topics": [
{
"title": "...",
"facts": [],
"questions": [],
"positions": [],
"decisions": [],
"todos": []
}
]
}
```
The exact schema will evolve during development.
Conceptually, the Canonical Meeting Knowledge should include:
- meeting metadata
- topics
- facts
- decisions
- action items
- open questions
- positions
- technical information
- rationale and discussion context
- contradictions or uncertainty
- source references and evidence
The detailed schema remains future implementation work.
## Current Status
Planned
---
# Stage 9 – Output View Rendering
## Purpose
Render purpose-specific outputs from Canonical Meeting Knowledge.
The planned output products are:
- Working Protocol (`working_protocol.md`, Arbeitsprotokoll)
- Distribution Protocol (`distribution_protocol.md`, Verteilerprotokoll)
- Knowledge Objects, rendered as a Knowledge-base Entry (`knowledge_entry.md`)
and later stored in a structured format such as `knowledge_entry.json`
(Wissensdatenbankeintrag)
- Later additional views such as action lists
Output rendering never performs additional analysis.
It only transforms existing structured information into the required view.
The outputs are rendered in parallel from the canonical representation. The
Distribution Protocol is not derived from the Working Protocol, and Knowledge
Objects are not derived from either protocol.
Rendering may be deterministic, template-based or LLM-assisted depending on the
output and implementation maturity.
Completeness differs by output:
- The Working Protocol optimizes for recall and traceability.
- The Distribution Protocol optimizes for relevance and brevity.
- Knowledge Objects optimize for durability and reuse.
Rendered protocol language should normally match the dominant language of the
source transcript or consolidated meeting knowledge unless an explicit output
language is requested.
## Processing Type
LLM
## Current Status
Planned
---
# Data Flow
Each stage consumes only the output of the previous stage.
```text
Stage N
│
Structured Output
│
▼
Stage N + 1
```
Intermediate results remain available for inspection, testing and experimentation.
---
# Guiding Principles
Every pipeline stage should satisfy the following rules.
## Single Responsibility
One module.
One task.
## Explicit Input
Every stage expects a clearly defined input format.
## Explicit Output
Every stage produces a clearly defined output format.
## Independent Evaluation
Each stage should be testable without executing the entire pipeline.
## Replaceable Components
A processing stage may be replaced by another implementation as long as it preserves the same interface.
---
# Current Development Roadmap
```text
✔ Normalization
⬜ Discussion Blocks
✔ Technical Chunking
⬜ Topic Segmentation
⬜ Specialized Extraction
✔ Deterministic Canonicalization
✅ Semantic Consolidation V0 - facts-only duplicate detection
⬜ Canonical Meeting Knowledge
⬜ Output View Rendering
```
The immediate evaluation focus is using the Semantic Consolidator V0 output as
input for the unchanged Working Protocol renderer. Canonicalizer V1 now
preserves source evidence, and Semantic Consolidator V0 conservatively merges
semantically equivalent fact items. Broader semantic consolidation should later
recover global context from independent chunk extractions and prepare Canonical
Meeting Knowledge for parallel Output View rendering.