Files
meeting-lab/docs/data-models.md
T
admin 5c03ed7efd Refine canonical meeting knowledge architecture
- establish Canonical Meeting Knowledge as the semantic source of truth
- introduce Output View Rendering architecture
- define Working Protocol, Distribution Protocol and Knowledge Objects as parallel renderers
- document renderer responsibilities and terminology
- clarify future Knowledge Object architecture
- document long-term reuse for enterprise knowledge systems
2026-07-30 16:59:58 +02:00

4.4 KiB

Data Models

Purpose

This document describes the logical data structures exchanged between the pipeline stages of the Meeting Lab.

The goal is not to define a final database schema.

Instead, these models represent stable interfaces between processing modules.

Models should evolve only when required by new functionality.


Design Principles

Keep Models Small

Only include fields that are currently required.

Avoid speculative attributes.

Bad:

{
  "priority": "...",
  "confidence": 0.93,
  "risk": "...",
  "category": "...",
  "importance": "...",
  "status": "..."
}

Good:

{
  "text": "...",
  "owner": "..."
}

New fields can always be added later.


Preserve Information

Models should preserve information rather than interpret it.

Interpretation belongs to processing modules.


Stable Interfaces

Modules communicate only through documented data models.

A module must never depend on another module's internal implementation.


Transcript

Represents the complete meeting transcript.

Example

{
    "meeting_id": "meeting_001",
    "language": "en",
    "blocks": []
}

Discussion Block

The discussion block is the fundamental processing unit.

{
    "block_id": 42,
    "speaker": "Speaker A",
    "start": 351.2,
    "end": 367.8,
    "text": "..."
}

Required fields

  • block_id
  • text

Optional fields

  • speaker
  • timestamps

Chunk

Technical processing unit.

{
    "chunk_id": 3,
    "blocks": [
        40,
        41,
        42
    ]
}

Chunks are implementation details.

They never represent discussion topics.


Topic

Represents one discussion topic.

{
    "topic_id": "topic_003",
    "title": "Ventilation",
    "segments": []
}

Topic Segment

A continuous part of a topic.

{
    "start_block": 40,
    "end_block": 152
}

One topic may contain multiple segments.


Fact

{
    "text": "..."
}

Question

{
    "text": "..."
}

Position

{
    "text": "...",
    "speaker": "..."
}

Decision

{
    "text": "..."
}

Todo

{
    "text": "...",
    "owner": "..."
}

Owner remains empty if unknown.


Technical Detail

{
    "text": "..."
}

Topic Result

After extraction, every topic contains the collected information.

{
    "topic_id": "topic_003",
    "title": "Ventilation",

    "segments": [],

    "facts": [],
    "questions": [],
    "positions": [],
    "decisions": [],
    "todos": [],
    "technical_details": []
}

This object feeds the Canonical Meeting Knowledge representation.


Canonical Meeting Knowledge

The canonical semantic representation of one meeting.

This representation is the single source of truth for all downstream outputs.

{
    "meeting_id": "meeting_001",

    "metadata": {},
    "topics": [],
    "facts": [],
    "decisions": [],
    "todos": [],
    "questions": [],
    "positions": [],
    "technical_details": [],
    "rationale": [],
    "uncertainty": [],
    "source_references": []
}

This is the common intermediate representation for all final Output Views. The exact schema is not final and should be refined during future implementation work.


Output Views

The final outputs are independent renderings of the Canonical Meeting Knowledge.

Canonical Meeting Knowledge
    ├── Working Protocol
    ├── Distribution Protocol
    └── Knowledge Objects

The Working Protocol, Distribution Protocol and Knowledge Objects are not derived from one another. Each renderer reads the same canonical semantic model and selects the level of detail appropriate for its purpose.

Knowledge Objects represent durable organizational knowledge such as processes, definitions, responsibilities, rules, accepted practices and long-term decisions. They are independent of the original meeting wording. Markdown is one possible presentation, but JSON or another structured format is expected to become the canonical storage format later.


Future Extensions

Possible future additions include:

  • confidence values
  • evidence references
  • source blocks
  • priorities
  • deadlines
  • status tracking
  • semantic relationships

These fields will only be introduced when they provide measurable benefits.

The Meeting Lab intentionally avoids designing an overly complex schema in advance.