Files
meeting-lab/docs/data-models.md
admin 9446c6e0be Document validation architecture and renderer faithfulness findings
- document Entity Registry and Meeting Context V2 architecture
- preserve meeting_context.yaml as the authoritative meeting-specific input
- define immutable authoritative metadata across all pipeline stages
- restrict Constraint Repair to deterministic structured-data operations
- record BUG-003 root cause and deferred entity-verification resolution
- document BUG-005 attendance-consistency design
- add BUG-006 renderer faithfulness root-cause analysis
- distinguish Engineering Readiness from Practical Usability
- update the persistent regression bug tracker
2026-08-03 16:05:47 +02:00

11 KiB

Data Models

Purpose

This document describes the logical data structures exchanged between the pipeline stages of the Meeting Lab.

The goal is not to define a final database schema.

Instead, these models represent stable interfaces between processing modules.

Models should evolve only when required by new functionality.


Design Principles

Keep Models Small

Only include fields that are currently required.

Avoid speculative attributes.

Bad:

{
  "priority": "...",
  "confidence": 0.93,
  "risk": "...",
  "category": "...",
  "importance": "...",
  "status": "..."
}

Good:

{
  "text": "...",
  "owner": "..."
}

New fields can always be added later.


Preserve Information

Models should preserve information rather than interpret it.

Interpretation belongs to processing modules.


Stable Interfaces

Modules communicate only through documented data models.

A module must never depend on another module's internal implementation.


Transcript

Represents the complete meeting transcript.

Example

{
    "meeting_id": "meeting_001",
    "language": "en",
    "blocks": []
}

Discussion Block

The discussion block is the fundamental processing unit.

{
    "block_id": 42,
    "speaker": "Speaker A",
    "start": 351.2,
    "end": 367.8,
    "text": "..."
}

Required fields

  • block_id
  • text

Optional fields

  • speaker
  • timestamps

Chunk

Technical processing unit.

{
    "chunk_id": 3,
    "blocks": [
        40,
        41,
        42
    ]
}

Chunks are implementation details.

They never represent discussion topics.


Topic

Represents one discussion topic.

{
    "topic_id": "topic_003",
    "title": "Ventilation",
    "segments": []
}

Topic Segment

A continuous part of a topic.

{
    "start_block": 40,
    "end_block": 152
}

One topic may contain multiple segments.


Fact

{
    "text": "..."
}

Question

{
    "text": "..."
}

Position

{
    "text": "...",
    "speaker": "..."
}

Decision

{
    "text": "..."
}

Todo

{
    "text": "...",
    "owner": "..."
}

Owner remains empty if unknown.


Technical Detail

{
    "text": "..."
}

Meeting Context V1

Manually maintained YAML metadata scaffold with an implemented Python loader, validator and deterministic prompt renderer for chunk extraction.

Top-level structure:

schema_version: "1"
meeting: {}
participants: []
mentioned_people: []
organization: {}
known_entities: {}
context_rules: {}

Meeting Context separates participants, mentioned people, transcript speakers, responsible people, roles and departments. It is authoritative only for explicitly supplied metadata. It must not be used to infer responsibilities, decisions or commitments.

Template:

  • samples/templates/meeting_context.template.yaml

Documentation:

  • docs/meeting-context.md

When --meeting-context is supplied to extraction, output JSON receives only minimal provenance:

{
    "context": {
        "meeting_id": "...",
        "source_file": "...",
        "schema_version": "1"
    }
}

Later Canonicalizer, Semantic Consolidator, Canonical Meeting Knowledge and renderer integration remains planned.


Entity Registry

Accepted Architecture. Implementation deferred.

The Entity Registry is the persistent cross-meeting knowledge source for confirmed entities, aliases and organizational metadata. It is independent from individual meetings and is the planned long-term source used to prepare Meeting Context V2.

Entity types include:

  • people
  • organizations
  • departments
  • products
  • projects
  • locations
  • abbreviations

Each entity has a stable internal identifier. The displayed name may change over time, but the internal identifier must remain stable.

Conceptual shape:

{
    "entity_id": "person_0001",
    "entity_type": "person",
    "display_name": "Jovana",
    "aliases": [
        "Jovana",
        "Giovanna",
        "Jovanna",
        "Giovana"
    ],
    "status": "confirmed"
}

The registry never learns automatically. It may propose matches, but only confirmed user actions update it. Similarity search may suggest spelling variants, Whisper transcription variants, umlaut variants or OCR-like mistakes, but suggestions require explicit confirmation.

Previously unseen names should be classified by the user as one of:

  • meeting participant
  • mentioned person
  • external person
  • transcription error
  • ignore

The Entity Registry must not infer responsibility, decisions, attendance or ownership.


Meeting Context V2

Accepted Architecture. Implementation deferred.

Meeting Context V2 is an authoritative meeting-specific YAML Point of Truth generated or assisted from:

  • Entity Registry
  • user confirmations
  • meeting metadata

The YAML remains the extraction pipeline interface and the authoritative meeting-specific Point of Truth for that meeting run. It is also a reproducible input artifact: changes to the Entity Registry after a meeting run must not silently change the historical Meeting Context used for that run.

The Entity Registry remains the persistent cross-meeting knowledge source. It must not override explicit meeting-specific confirmations.

Meeting Context V2 should reduce manual work, improve alias handling, detect transcription errors earlier and make Meeting Context quality scalable across many meetings.

See docs/adr-meeting-context-v2-entity-registry.md.


Topic Result

After extraction, every topic contains the collected information.

{
    "topic_id": "topic_003",
    "title": "Ventilation",

    "segments": [],

    "facts": [],
    "questions": [],
    "positions": [],
    "decisions": [],
    "todos": [],
    "technical_details": []
}

This object feeds the Canonical Meeting Knowledge representation.


Canonicalized Extractions

Implemented deterministic intermediate file created from raw chunk extraction JSON by Canonicalizer V1. This is not Canonical Meeting Knowledge.

Top-level structure:

{
    "schema_version": "1",
    "source_files": [],
    "stats": {},
    "items": []
}

Each item contains at least:

  • item_id
  • category
  • text
  • evidence
  • source_file
  • source_index
  • original_value
  • source_references

Action items also preserve deterministic fields such as responsible and deadline when present.

Responsibility attribution is stricter than mention or participation. A responsible value may be kept only when source evidence explicitly assigns, accepts or confirms responsibility. If evidence is incomplete, ambiguous or only based on a suggestion, objection, topic expertise or department mention, the field remains null or unset and the evidence is preserved.

Do not collapse these concepts into one field:

  • speaker: person who uttered the evidence.
  • mentioned_person: person named in the evidence.
  • participant: person present in the meeting.
  • responsible_person: person explicitly assigned to or accepting an action.
  • department: organizational unit discussed or represented.
  • owner: durable ownership of a process, system or knowledge object.
  • assignee: operational person or team assigned to a concrete task.

Future compatible fields may include:

  • responsibility_status: explicit, accepted, proposed or unclear.
  • attribution_evidence: source evidence supporting the assignment status.

Example:

{
    "id": "fact.chunk_03.0001",
    "category": "fact",
    "text": "...",
    "source_references": [
        {
            "chunk_id": "chunk_03",
            "source_file": "chunk_03_extraction.json",
            "evidence": "..."
        }
    ]
}

Canonicalizer V1 creates this kind of object without an LLM. It validates and normalizes raw extraction objects, assigns stable IDs and source references, normalizes category names and basic field structure, performs only safe deterministic cleanup, may group exact duplicates and must preserve all source evidence.

It must not perform uncertain semantic merging.


Semantic Fact Group

Implemented by Semantic Consolidator V0.

Example:

{
    "consolidated_id": "fact_group_0001",
    "category": "fact",
    "canonical_text": "...",
    "source_item_ids": ["fact_0001"],
    "source_references": [],
    "evidence": [],
    "merge_reason": "Singleton; no semantically equivalent fact found."
}

Semantic Consolidator V0 only processes fact items. It merges semantically equivalent facts conservatively, preserves source references and evidence, and validates that every source fact appears exactly once. Non-fact categories are copied unchanged. It is not a summarizer, topic grouper, protocol renderer or Canonical Meeting Knowledge generator.


Consolidated Topic

Planned semantic object produced by the Semantic Consolidator.

Example:

{
    "topic_id": "topic_001",
    "title": "...",
    "background": [],
    "decisions": [],
    "action_items": [],
    "open_questions": [],
    "durable_information": [],
    "uncertainty": [],
    "source_references": []
}

Future Semantic Consolidator versions may use the local LLM to merge semantically equivalent statements beyond facts, group content by topic, preserve evidence from all contributing chunks, mark contradictions and uncertainty and separate durable information from transient discussion.

It produces Canonical Meeting Knowledge. It does not directly write a protocol.


Canonical Meeting Knowledge

The canonical semantic representation of one meeting.

This representation is the single source of truth for all downstream outputs. It is a structured representation, preferably JSON, and is not itself a prose protocol.

{
    "meeting_id": "meeting_001",

    "metadata": {},
    "topics": [],
    "facts": [],
    "decisions": [],
    "todos": [],
    "questions": [],
    "positions": [],
    "technical_details": [],
    "durable_information": [],
    "rationale": [],
    "uncertainty": [],
    "source_references": []
}

This is the common intermediate representation for all final Output Views. The exact schema is not final and should be refined during future implementation work.


Output Views

The final outputs are independent renderings of the Canonical Meeting Knowledge.

Canonical Meeting Knowledge
    ├── Working Protocol
    ├── Distribution Protocol
    └── Knowledge Objects

The Working Protocol, Distribution Protocol and Knowledge Objects are not derived from one another. Each renderer reads the same canonical semantic model and selects the level of detail appropriate for its purpose.

Knowledge Objects represent durable organizational knowledge such as processes, definitions, responsibilities, rules, accepted practices and long-term decisions. They are independent of the original meeting wording. Markdown is one possible presentation, but JSON or another structured format is expected to become the canonical storage format later.


Future Extensions

Possible future additions include:

  • confidence values
  • evidence references
  • source blocks
  • priorities
  • deadlines
  • status tracking
  • semantic relationships

These fields will only be introduced when they provide measurable benefits.

The Meeting Lab intentionally avoids designing an overly complex schema in advance.