- document Entity Registry and Meeting Context V2 architecture - preserve meeting_context.yaml as the authoritative meeting-specific input - define immutable authoritative metadata across all pipeline stages - restrict Constraint Repair to deterministic structured-data operations - record BUG-003 root cause and deferred entity-verification resolution - document BUG-005 attendance-consistency design - add BUG-006 renderer faithfulness root-cause analysis - distinguish Engineering Readiness from Practical Usability - update the persistent regression bug tracker
622 lines
13 KiB
Markdown
622 lines
13 KiB
Markdown
# Pipeline
|
||
|
||
## Purpose
|
||
|
||
This document describes the processing pipeline of the Meeting Lab.
|
||
|
||
Unlike `architecture.md`, which describes the overall system, this document focuses on the individual processing stages, their responsibilities and the data flowing between them.
|
||
|
||
The guiding principle is simple:
|
||
|
||
> **Each processing stage has exactly one responsibility.**
|
||
|
||
Cross-cutting invariant:
|
||
|
||
> **Responsibility attribution requires explicit evidence.**
|
||
|
||
A person, team or department may be recorded as responsible only when the
|
||
source material explicitly assigns, accepts or confirms that responsibility.
|
||
The pipeline must not infer ownership from thematic proximity, discussion
|
||
participation, mentioning a task, commenting on another department,
|
||
organizational assumptions, likely job roles, speaker adjacency or model world
|
||
knowledge.
|
||
|
||
When support is incomplete or ambiguous, leave the responsible person unset,
|
||
mark the item as unclear where supported, and preserve the attribution
|
||
evidence. This applies to extraction, canonicalization, semantic consolidation,
|
||
Canonical Meeting Knowledge and every Output View renderer.
|
||
|
||
---
|
||
|
||
# Pipeline Overview
|
||
|
||
Current analysis pipeline:
|
||
|
||
```text
|
||
Whisper Transcript
|
||
│
|
||
▼
|
||
Normalization
|
||
│
|
||
▼
|
||
Discussion Blocks
|
||
│
|
||
▼
|
||
Technical Chunking
|
||
│
|
||
▼
|
||
Topic Segmentation
|
||
│
|
||
▼
|
||
Specialized Extraction
|
||
│
|
||
▼
|
||
Deterministic Canonicalization
|
||
│
|
||
▼
|
||
Semantic Consolidation
|
||
│
|
||
▼
|
||
Canonical Meeting Knowledge
|
||
│
|
||
▼
|
||
Output View Rendering
|
||
│
|
||
├── Working Protocol
|
||
├── Distribution Protocol
|
||
└── Knowledge Objects
|
||
```
|
||
|
||
Each stage receives a well-defined input and produces a well-defined output.
|
||
|
||
Meeting Context V1 exists as a manually maintained YAML metadata scaffold. It
|
||
is implemented for validation, optional `--meeting-context` use during chunk
|
||
extraction, deterministic prompt injection and minimal extraction JSON
|
||
provenance. Later Canonicalizer, Semantic Consolidator, Canonical Meeting
|
||
Knowledge and renderer integration remains future work.
|
||
|
||
Accepted future Meeting Context V2 preparation flow:
|
||
|
||
```text
|
||
Whisper
|
||
│
|
||
▼
|
||
Entity Detection
|
||
│
|
||
▼
|
||
User Confirmation
|
||
│
|
||
▼
|
||
Entity Registry Update
|
||
│
|
||
▼
|
||
Meeting Context Builder
|
||
│
|
||
▼
|
||
meeting_context.yaml
|
||
│
|
||
▼
|
||
Extraction Pipeline
|
||
```
|
||
|
||
This preparation flow is not implemented. It is the accepted long-term
|
||
direction for reducing manual Meeting Context work while preserving explicit
|
||
user control. The Entity Registry is the persistent cross-meeting knowledge
|
||
source for confirmed entities, aliases and organizational metadata. It stores
|
||
stable internal IDs and confirmed aliases, and never updates itself
|
||
automatically.
|
||
|
||
For each meeting run, `meeting_context.yaml` remains the authoritative
|
||
meeting-specific Point of Truth and reproducible input artifact consumed by the
|
||
pipeline. V2 changes how that artifact is prepared: it may be generated or
|
||
assisted from Registry data, user confirmations and meeting metadata. The
|
||
Registry must not override explicit meeting-specific confirmations, and
|
||
Registry changes after a meeting run must not silently change the historical
|
||
Meeting Context used for that run. Similarity suggestions and unknown entity
|
||
classifications require explicit user confirmation.
|
||
|
||
---
|
||
|
||
# Stage 1 – Normalization
|
||
|
||
## Purpose
|
||
|
||
Remove transcription artifacts without changing the meaning of the discussion.
|
||
|
||
## Input
|
||
|
||
Raw transcript generated by Whisper.
|
||
|
||
## Output
|
||
|
||
Normalized transcript.
|
||
|
||
Change log containing every modification.
|
||
|
||
## Processing Type
|
||
|
||
Deterministic
|
||
|
||
## Responsibilities
|
||
|
||
- Remove filler words
|
||
- Remove immediate duplicate words
|
||
- Remove immediate duplicate short phrases
|
||
- Normalize whitespace
|
||
- Preserve all semantic content
|
||
|
||
## Must Not
|
||
|
||
- Rephrase text
|
||
- Summarize
|
||
- Interpret statements
|
||
- Correct factual content
|
||
|
||
## Current Status
|
||
|
||
Implemented
|
||
|
||
---
|
||
|
||
# Stage 2 – Discussion Blocks
|
||
|
||
## Purpose
|
||
|
||
Convert the transcript into stable processing units.
|
||
|
||
Discussion blocks are the smallest semantic unit used throughout the pipeline.
|
||
|
||
## Input
|
||
|
||
Normalized transcript.
|
||
|
||
## Output
|
||
|
||
Ordered list of discussion blocks.
|
||
|
||
Example:
|
||
|
||
```json
|
||
{
|
||
"block_id": 42,
|
||
"speaker": "A",
|
||
"start": 351.2,
|
||
"end": 367.8,
|
||
"text": "..."
|
||
}
|
||
```
|
||
|
||
## Processing Type
|
||
|
||
Deterministic
|
||
|
||
## Responsibilities
|
||
|
||
- Create stable identifiers
|
||
- Preserve ordering
|
||
- Preserve timestamps
|
||
- Preserve speaker information where available
|
||
|
||
## Current Status
|
||
|
||
Implemented as Canonicalizer V1.
|
||
|
||
CLI:
|
||
|
||
```text
|
||
PYTHONPATH=src .venv/bin/python -m meeting_lab.consolidation.canonicalize \
|
||
samples/whisper/meeting_speech_cleaned_chunks \
|
||
-o /tmp/canonicalized_extractions.json
|
||
```
|
||
|
||
---
|
||
|
||
# Stage 3 – Technical Chunking
|
||
|
||
## Purpose
|
||
|
||
Split large meetings into model-sized chunks.
|
||
|
||
Chunking exists only because language models have limited context windows.
|
||
|
||
## Input
|
||
|
||
Discussion blocks.
|
||
|
||
## Output
|
||
|
||
Chunk manifest and chunk files.
|
||
|
||
## Processing Type
|
||
|
||
Deterministic
|
||
|
||
## Responsibilities
|
||
|
||
- Respect block boundaries
|
||
- Keep chunk size below model limits
|
||
- Optionally create overlapping context
|
||
|
||
## Must Not
|
||
|
||
- Detect discussion topics
|
||
- Merge discussion content
|
||
- Interpret meaning
|
||
|
||
## Current Status
|
||
|
||
Implemented
|
||
|
||
---
|
||
|
||
# Stage 4 – Topic Segmentation
|
||
|
||
## Purpose
|
||
|
||
Identify the thematic structure of the discussion.
|
||
|
||
This is considered the central research problem of the Meeting Lab.
|
||
|
||
## Input
|
||
|
||
Discussion blocks or technical chunks.
|
||
|
||
## Output
|
||
|
||
Topics consisting of one or more discussion segments.
|
||
|
||
Example:
|
||
|
||
```json
|
||
{
|
||
"topic": "Ventilation",
|
||
"segments": [
|
||
{
|
||
"start_block": 40,
|
||
"end_block": 152
|
||
},
|
||
{
|
||
"start_block": 1618,
|
||
"end_block": 1697
|
||
}
|
||
]
|
||
}
|
||
```
|
||
|
||
## Processing Type
|
||
|
||
LLM
|
||
|
||
## Responsibilities
|
||
|
||
- Detect topic start
|
||
- Detect topic end
|
||
- Detect topic changes
|
||
- Detect resumed topics
|
||
- Associate discussion blocks with topics
|
||
|
||
## Must Not
|
||
|
||
- Extract facts
|
||
- Detect todos
|
||
- Generate summaries
|
||
|
||
## Current Status
|
||
|
||
Implemented as Canonicalizer V1.
|
||
|
||
---
|
||
|
||
# Stage 5 – Specialized Extraction
|
||
|
||
## Purpose
|
||
|
||
Extract one specific type of information from each topic.
|
||
|
||
Every extractor performs exactly one task.
|
||
|
||
## Planned Extractors
|
||
|
||
```text
|
||
extract_facts.py
|
||
extract_questions.py
|
||
extract_positions.py
|
||
extract_decisions.py
|
||
extract_todos.py
|
||
extract_technical.py
|
||
```
|
||
|
||
Each extractor has:
|
||
|
||
- one prompt
|
||
- one responsibility
|
||
- one output schema
|
||
|
||
The current shared extraction prompt is assembled from:
|
||
|
||
- `common.md`
|
||
- `decisions.md`
|
||
- `todos.md`
|
||
|
||
The todo prompt requires explicit assignment, volunteering or acceptance before
|
||
recording a named responsible person. Meeting Context may validate identity,
|
||
role, department and attendance, but never establishes responsibility.
|
||
|
||
The decision prompt treats a personal commitment to concrete future work as a
|
||
todo unless the group separately establishes a binding outcome, rule, approval,
|
||
rejection, deferral, selection, process state or responsibility policy. The
|
||
same proposition should not be duplicated under decisions and todos; extract
|
||
both only for semantically separate propositions.
|
||
|
||
## Processing Type
|
||
|
||
LLM
|
||
|
||
## Current Status
|
||
|
||
Prototype exists as a combined extractor.
|
||
|
||
---
|
||
|
||
# Stage 6 – Deterministic Canonicalization
|
||
|
||
## Purpose
|
||
|
||
Normalize raw chunk extraction JSON into stable canonical extraction objects
|
||
without changing uncertain semantics.
|
||
|
||
## Input
|
||
|
||
Chunk extraction JSON files.
|
||
|
||
## Output
|
||
|
||
Validated canonical extraction objects with stable source references and IDs.
|
||
|
||
## Responsibilities
|
||
|
||
- Validate extraction objects
|
||
- Normalize category names
|
||
- Normalize basic field structure
|
||
- Assign stable source references and IDs
|
||
- Preserve all source evidence
|
||
- Perform only safe deterministic cleanup
|
||
- Group exact duplicates where unambiguous
|
||
|
||
## Must Not
|
||
|
||
- Use an LLM
|
||
- Perform uncertain semantic merging
|
||
- Infer missing information
|
||
- Drop source evidence
|
||
|
||
## Processing Type
|
||
|
||
Deterministic Python
|
||
|
||
## Current Status
|
||
|
||
Planned
|
||
|
||
---
|
||
|
||
# Stage 7 – Semantic Consolidation
|
||
|
||
## Purpose
|
||
|
||
Merge canonicalized extraction objects into an evidence-preserving semantic
|
||
meeting representation.
|
||
|
||
## Input
|
||
|
||
Canonicalized extraction objects.
|
||
|
||
## Output
|
||
|
||
Semantic Consolidator V0 output is `consolidated_extractions.json` with fact
|
||
groups and unchanged non-fact items.
|
||
|
||
Future broader semantic consolidation should produce Canonical Meeting
|
||
Knowledge.
|
||
|
||
## Responsibilities
|
||
|
||
- V0: merge semantically equivalent fact items only
|
||
- V0: preserve all non-fact categories unchanged
|
||
- V0: validate that every source fact ID appears exactly once
|
||
- Future: merge semantically equivalent statements across categories
|
||
- Future: group content by topic
|
||
- Preserve evidence from all contributing chunks
|
||
- Future: mark contradictions and uncertainty
|
||
- Future: separate durable information from transient discussion
|
||
- Future: reconcile category shifts where supported by evidence
|
||
|
||
## Must Not
|
||
|
||
- Directly write a protocol
|
||
- Invent information
|
||
- Drop conflicting evidence silently
|
||
|
||
## Processing Type
|
||
|
||
Local LLM, with deterministic pre/post-processing where useful.
|
||
|
||
## Current Status
|
||
|
||
Semantic Consolidator V0 is implemented and experimentally validated for
|
||
facts-only conservative duplicate detection. Broader semantic consolidation and
|
||
Canonical Meeting Knowledge generation remain planned.
|
||
|
||
---
|
||
|
||
# Stage 8 – Canonical Meeting Knowledge
|
||
|
||
## Purpose
|
||
|
||
Produce the canonical semantic representation of one meeting.
|
||
|
||
This representation is the single source of truth for all downstream outputs.
|
||
It is a structured representation, preferably JSON, and is not itself a prose
|
||
protocol.
|
||
|
||
Example:
|
||
|
||
```json
|
||
{
|
||
"topics": [
|
||
{
|
||
"title": "...",
|
||
"facts": [],
|
||
"questions": [],
|
||
"positions": [],
|
||
"decisions": [],
|
||
"todos": []
|
||
}
|
||
]
|
||
}
|
||
```
|
||
|
||
The exact schema will evolve during development.
|
||
|
||
Conceptually, the Canonical Meeting Knowledge should include:
|
||
|
||
- meeting metadata
|
||
- topics
|
||
- facts
|
||
- decisions
|
||
- action items
|
||
- open questions
|
||
- positions
|
||
- technical information
|
||
- rationale and discussion context
|
||
- contradictions or uncertainty
|
||
- source references and evidence
|
||
|
||
The detailed schema remains future implementation work.
|
||
|
||
## Current Status
|
||
|
||
Planned
|
||
|
||
---
|
||
|
||
# Stage 9 – Output View Rendering
|
||
|
||
## Purpose
|
||
|
||
Render purpose-specific outputs from Canonical Meeting Knowledge.
|
||
|
||
The planned output products are:
|
||
|
||
- Working Protocol (`working_protocol.md`, Arbeitsprotokoll)
|
||
- Distribution Protocol (`distribution_protocol.md`, Verteilerprotokoll)
|
||
- Knowledge Objects, rendered as a Knowledge-base Entry (`knowledge_entry.md`)
|
||
and later stored in a structured format such as `knowledge_entry.json`
|
||
(Wissensdatenbankeintrag)
|
||
- Later additional views such as action lists
|
||
|
||
Output rendering never performs additional analysis.
|
||
|
||
It only transforms existing structured information into the required view.
|
||
|
||
The outputs are rendered in parallel from the canonical representation. The
|
||
Distribution Protocol is not derived from the Working Protocol, and Knowledge
|
||
Objects are not derived from either protocol.
|
||
|
||
Rendering may be deterministic, template-based or LLM-assisted depending on the
|
||
output and implementation maturity.
|
||
|
||
Completeness differs by output:
|
||
|
||
- The Working Protocol optimizes for recall and traceability.
|
||
- The Distribution Protocol optimizes for relevance and brevity.
|
||
- Knowledge Objects optimize for durability and reuse.
|
||
|
||
Rendered protocol language should normally match the dominant language of the
|
||
source transcript or consolidated meeting knowledge unless an explicit output
|
||
language is requested.
|
||
|
||
## Processing Type
|
||
|
||
LLM
|
||
|
||
## Current Status
|
||
|
||
Planned
|
||
|
||
---
|
||
|
||
# Data Flow
|
||
|
||
Each stage consumes only the output of the previous stage.
|
||
|
||
```text
|
||
Stage N
|
||
│
|
||
Structured Output
|
||
│
|
||
▼
|
||
Stage N + 1
|
||
```
|
||
|
||
Intermediate results remain available for inspection, testing and experimentation.
|
||
|
||
---
|
||
|
||
# Guiding Principles
|
||
|
||
Every pipeline stage should satisfy the following rules.
|
||
|
||
## Single Responsibility
|
||
|
||
One module.
|
||
|
||
One task.
|
||
|
||
## Explicit Input
|
||
|
||
Every stage expects a clearly defined input format.
|
||
|
||
## Explicit Output
|
||
|
||
Every stage produces a clearly defined output format.
|
||
|
||
## Independent Evaluation
|
||
|
||
Each stage should be testable without executing the entire pipeline.
|
||
|
||
## Replaceable Components
|
||
|
||
A processing stage may be replaced by another implementation as long as it preserves the same interface.
|
||
|
||
---
|
||
|
||
# Current Development Roadmap
|
||
|
||
```text
|
||
✔ Normalization
|
||
|
||
⬜ Discussion Blocks
|
||
|
||
✔ Technical Chunking
|
||
|
||
⬜ Topic Segmentation
|
||
|
||
⬜ Specialized Extraction
|
||
|
||
✔ Deterministic Canonicalization
|
||
|
||
✅ Semantic Consolidation V0 - facts-only duplicate detection
|
||
|
||
⬜ Canonical Meeting Knowledge
|
||
|
||
⬜ Output View Rendering
|
||
```
|
||
|
||
The immediate evaluation focus is using the Semantic Consolidator V0 output as
|
||
input for the unchanged Working Protocol renderer. Canonicalizer V1 now
|
||
preserves source evidence, and Semantic Consolidator V0 conservatively merges
|
||
semantically equivalent fact items. Broader semantic consolidation should later
|
||
recover global context from independent chunk extractions and prepare Canonical
|
||
Meeting Knowledge for parallel Output View rendering.
|