Files
meeting-lab/docs/pipeline.md
T
admin 63075eaca9 Enforce explicit responsibility attribution
- document responsibility attribution as a project-wide invariant
- prevent inferred ownership in protocol rendering
- add negative gold regression for false responsibility assignment
- document future responsibility evidence and attribution model
- record the real-life benchmark finding
2026-07-31 13:01:11 +02:00

11 KiB
Raw Blame History

Pipeline

Purpose

This document describes the processing pipeline of the Meeting Lab.

Unlike architecture.md, which describes the overall system, this document focuses on the individual processing stages, their responsibilities and the data flowing between them.

The guiding principle is simple:

Each processing stage has exactly one responsibility.

Cross-cutting invariant:

Responsibility attribution requires explicit evidence.

A person, team or department may be recorded as responsible only when the source material explicitly assigns, accepts or confirms that responsibility. The pipeline must not infer ownership from thematic proximity, discussion participation, mentioning a task, commenting on another department, organizational assumptions, likely job roles, speaker adjacency or model world knowledge.

When support is incomplete or ambiguous, leave the responsible person unset, mark the item as unclear where supported, and preserve the attribution evidence. This applies to extraction, canonicalization, semantic consolidation, Canonical Meeting Knowledge and every Output View renderer.


Pipeline Overview

Whisper Transcript
        │
        ▼
Normalization
        │
        ▼
Discussion Blocks
        │
        ▼
Technical Chunking
        │
        ▼
Topic Segmentation
        │
        ▼
Specialized Extraction
        │
        ▼
Deterministic Canonicalization
        │
        ▼
Semantic Consolidation
        │
        ▼
Canonical Meeting Knowledge
        │
        ▼
Output View Rendering
        │
        ├── Working Protocol
        ├── Distribution Protocol
        └── Knowledge Objects

Each stage receives a well-defined input and produces a well-defined output.


Stage 1 – Normalization

Purpose

Remove transcription artifacts without changing the meaning of the discussion.

Input

Raw transcript generated by Whisper.

Output

Normalized transcript.

Change log containing every modification.

Processing Type

Deterministic

Responsibilities

  • Remove filler words
  • Remove immediate duplicate words
  • Remove immediate duplicate short phrases
  • Normalize whitespace
  • Preserve all semantic content

Must Not

  • Rephrase text
  • Summarize
  • Interpret statements
  • Correct factual content

Current Status

Implemented


Stage 2 – Discussion Blocks

Purpose

Convert the transcript into stable processing units.

Discussion blocks are the smallest semantic unit used throughout the pipeline.

Input

Normalized transcript.

Output

Ordered list of discussion blocks.

Example:

{
    "block_id": 42,
    "speaker": "A",
    "start": 351.2,
    "end": 367.8,
    "text": "..."
}

Processing Type

Deterministic

Responsibilities

  • Create stable identifiers
  • Preserve ordering
  • Preserve timestamps
  • Preserve speaker information where available

Current Status

Implemented as Canonicalizer V1.

CLI:

PYTHONPATH=src .venv/bin/python -m meeting_lab.consolidation.canonicalize \
  samples/whisper/meeting_speech_cleaned_chunks \
  -o /tmp/canonicalized_extractions.json

Stage 3 – Technical Chunking

Purpose

Split large meetings into model-sized chunks.

Chunking exists only because language models have limited context windows.

Input

Discussion blocks.

Output

Chunk manifest and chunk files.

Processing Type

Deterministic

Responsibilities

  • Respect block boundaries
  • Keep chunk size below model limits
  • Optionally create overlapping context

Must Not

  • Detect discussion topics
  • Merge discussion content
  • Interpret meaning

Current Status

Implemented


Stage 4 – Topic Segmentation

Purpose

Identify the thematic structure of the discussion.

This is considered the central research problem of the Meeting Lab.

Input

Discussion blocks or technical chunks.

Output

Topics consisting of one or more discussion segments.

Example:

{
    "topic": "Ventilation",
    "segments": [
        {
            "start_block": 40,
            "end_block": 152
        },
        {
            "start_block": 1618,
            "end_block": 1697
        }
    ]
}

Processing Type

LLM

Responsibilities

  • Detect topic start
  • Detect topic end
  • Detect topic changes
  • Detect resumed topics
  • Associate discussion blocks with topics

Must Not

  • Extract facts
  • Detect todos
  • Generate summaries

Current Status

Implemented as Canonicalizer V1.


Stage 5 – Specialized Extraction

Purpose

Extract one specific type of information from each topic.

Every extractor performs exactly one task.

Planned Extractors

extract_facts.py
extract_questions.py
extract_positions.py
extract_decisions.py
extract_todos.py
extract_technical.py

Each extractor has:

  • one prompt
  • one responsibility
  • one output schema

Processing Type

LLM

Current Status

Prototype exists as a combined extractor.


Stage 6 – Deterministic Canonicalization

Purpose

Normalize raw chunk extraction JSON into stable canonical extraction objects without changing uncertain semantics.

Input

Chunk extraction JSON files.

Output

Validated canonical extraction objects with stable source references and IDs.

Responsibilities

  • Validate extraction objects
  • Normalize category names
  • Normalize basic field structure
  • Assign stable source references and IDs
  • Preserve all source evidence
  • Perform only safe deterministic cleanup
  • Group exact duplicates where unambiguous

Must Not

  • Use an LLM
  • Perform uncertain semantic merging
  • Infer missing information
  • Drop source evidence

Processing Type

Deterministic Python

Current Status

Planned


Stage 7 – Semantic Consolidation

Purpose

Merge canonicalized extraction objects into an evidence-preserving semantic meeting representation.

Input

Canonicalized extraction objects.

Output

Semantic Consolidator V0 output is consolidated_extractions.json with fact groups and unchanged non-fact items.

Future broader semantic consolidation should produce Canonical Meeting Knowledge.

Responsibilities

  • V0: merge semantically equivalent fact items only
  • V0: preserve all non-fact categories unchanged
  • V0: validate that every source fact ID appears exactly once
  • Future: merge semantically equivalent statements across categories
  • Future: group content by topic
  • Preserve evidence from all contributing chunks
  • Future: mark contradictions and uncertainty
  • Future: separate durable information from transient discussion
  • Future: reconcile category shifts where supported by evidence

Must Not

  • Directly write a protocol
  • Invent information
  • Drop conflicting evidence silently

Processing Type

Local LLM, with deterministic pre/post-processing where useful.

Current Status

Semantic Consolidator V0 is implemented and experimentally validated for facts-only conservative duplicate detection. Broader semantic consolidation and Canonical Meeting Knowledge generation remain planned.


Stage 8 – Canonical Meeting Knowledge

Purpose

Produce the canonical semantic representation of one meeting.

This representation is the single source of truth for all downstream outputs. It is a structured representation, preferably JSON, and is not itself a prose protocol.

Example:

{
    "topics": [
        {
            "title": "...",
            "facts": [],
            "questions": [],
            "positions": [],
            "decisions": [],
            "todos": []
        }
    ]
}

The exact schema will evolve during development.

Conceptually, the Canonical Meeting Knowledge should include:

  • meeting metadata
  • topics
  • facts
  • decisions
  • action items
  • open questions
  • positions
  • technical information
  • rationale and discussion context
  • contradictions or uncertainty
  • source references and evidence

The detailed schema remains future implementation work.

Current Status

Planned


Stage 9 – Output View Rendering

Purpose

Render purpose-specific outputs from Canonical Meeting Knowledge.

The planned output products are:

  • Working Protocol (working_protocol.md, Arbeitsprotokoll)
  • Distribution Protocol (distribution_protocol.md, Verteilerprotokoll)
  • Knowledge Objects, rendered as a Knowledge-base Entry (knowledge_entry.md) and later stored in a structured format such as knowledge_entry.json (Wissensdatenbankeintrag)
  • Later additional views such as action lists

Output rendering never performs additional analysis.

It only transforms existing structured information into the required view.

The outputs are rendered in parallel from the canonical representation. The Distribution Protocol is not derived from the Working Protocol, and Knowledge Objects are not derived from either protocol.

Rendering may be deterministic, template-based or LLM-assisted depending on the output and implementation maturity.

Completeness differs by output:

  • The Working Protocol optimizes for recall and traceability.
  • The Distribution Protocol optimizes for relevance and brevity.
  • Knowledge Objects optimize for durability and reuse.

Rendered protocol language should normally match the dominant language of the source transcript or consolidated meeting knowledge unless an explicit output language is requested.

Processing Type

LLM

Current Status

Planned


Data Flow

Each stage consumes only the output of the previous stage.

Stage N
    │
Structured Output
    │
    ▼
Stage N + 1

Intermediate results remain available for inspection, testing and experimentation.


Guiding Principles

Every pipeline stage should satisfy the following rules.

Single Responsibility

One module.

One task.

Explicit Input

Every stage expects a clearly defined input format.

Explicit Output

Every stage produces a clearly defined output format.

Independent Evaluation

Each stage should be testable without executing the entire pipeline.

Replaceable Components

A processing stage may be replaced by another implementation as long as it preserves the same interface.


Current Development Roadmap

✔ Normalization

⬜ Discussion Blocks

✔ Technical Chunking

⬜ Topic Segmentation

⬜ Specialized Extraction

✔ Deterministic Canonicalization

✅ Semantic Consolidation V0 - facts-only duplicate detection

⬜ Canonical Meeting Knowledge

⬜ Output View Rendering

The immediate evaluation focus is using the Semantic Consolidator V0 output as input for the unchanged Working Protocol renderer. Canonicalizer V1 now preserves source evidence, and Semantic Consolidator V0 conservatively merges semantically equivalent fact items. Broader semantic consolidation should later recover global context from independent chunk extractions and prepare Canonical Meeting Knowledge for parallel Output View rendering.