Files
meeting-lab/docs/data-models.md
T
admin 90aa34d5d0 Implement Semantic Consolidator V0
- add deterministic canonicalization support for extraction items
- add facts-only semantic consolidation using local Ollama
- preserve source evidence and validate complete fact coverage
- add conservative merge rules and non-LLM tests
- record the first validated real-life consolidation benchmark
- document current scope, limitations and next evaluation step
2026-07-31 11:25:46 +02:00

428 lines
7.0 KiB
Markdown

# Data Models
## Purpose
This document describes the logical data structures exchanged between the pipeline stages of the Meeting Lab.
The goal is **not** to define a final database schema.
Instead, these models represent stable interfaces between processing modules.
Models should evolve only when required by new functionality.
---
# Design Principles
## Keep Models Small
Only include fields that are currently required.
Avoid speculative attributes.
Bad:
```json
{
"priority": "...",
"confidence": 0.93,
"risk": "...",
"category": "...",
"importance": "...",
"status": "..."
}
```
Good:
```json
{
"text": "...",
"owner": "..."
}
```
New fields can always be added later.
---
## Preserve Information
Models should preserve information rather than interpret it.
Interpretation belongs to processing modules.
---
## Stable Interfaces
Modules communicate only through documented data models.
A module must never depend on another module's internal implementation.
---
# Transcript
Represents the complete meeting transcript.
Example
```json
{
"meeting_id": "meeting_001",
"language": "en",
"blocks": []
}
```
---
# Discussion Block
The discussion block is the fundamental processing unit.
```json
{
"block_id": 42,
"speaker": "Speaker A",
"start": 351.2,
"end": 367.8,
"text": "..."
}
```
Required fields
- block_id
- text
Optional fields
- speaker
- timestamps
---
# Chunk
Technical processing unit.
```json
{
"chunk_id": 3,
"blocks": [
40,
41,
42
]
}
```
Chunks are implementation details.
They never represent discussion topics.
---
# Topic
Represents one discussion topic.
```json
{
"topic_id": "topic_003",
"title": "Ventilation",
"segments": []
}
```
---
# Topic Segment
A continuous part of a topic.
```json
{
"start_block": 40,
"end_block": 152
}
```
One topic may contain multiple segments.
---
# Fact
```json
{
"text": "..."
}
```
---
# Question
```json
{
"text": "..."
}
```
---
# Position
```json
{
"text": "...",
"speaker": "..."
}
```
---
# Decision
```json
{
"text": "..."
}
```
---
# Todo
```json
{
"text": "...",
"owner": "..."
}
```
Owner remains empty if unknown.
---
# Technical Detail
```json
{
"text": "..."
}
```
---
# Topic Result
After extraction, every topic contains the collected information.
```json
{
"topic_id": "topic_003",
"title": "Ventilation",
"segments": [],
"facts": [],
"questions": [],
"positions": [],
"decisions": [],
"todos": [],
"technical_details": []
}
```
This object feeds the Canonical Meeting Knowledge representation.
---
# Canonicalized Extractions
Implemented deterministic intermediate file created from raw chunk extraction
JSON by Canonicalizer V1. This is not Canonical Meeting Knowledge.
Top-level structure:
```json
{
"schema_version": "1",
"source_files": [],
"stats": {},
"items": []
}
```
Each item contains at least:
- item_id
- category
- text
- evidence
- source_file
- source_index
- original_value
- source_references
Action items also preserve deterministic fields such as `responsible` and
`deadline` when present.
Example:
```json
{
"id": "fact.chunk_03.0001",
"category": "fact",
"text": "...",
"source_references": [
{
"chunk_id": "chunk_03",
"source_file": "chunk_03_extraction.json",
"evidence": "..."
}
]
}
```
Canonicalizer V1 creates this kind of object without an LLM. It validates and
normalizes raw extraction objects, assigns stable IDs and source references,
normalizes category names and basic field structure, performs only safe
deterministic cleanup, may group exact duplicates and must preserve all source
evidence.
It must not perform uncertain semantic merging.
---
# Semantic Fact Group
Implemented by Semantic Consolidator V0.
Example:
```json
{
"consolidated_id": "fact_group_0001",
"category": "fact",
"canonical_text": "...",
"source_item_ids": ["fact_0001"],
"source_references": [],
"evidence": [],
"merge_reason": "Singleton; no semantically equivalent fact found."
}
```
Semantic Consolidator V0 only processes fact items. It merges semantically
equivalent facts conservatively, preserves source references and evidence, and
validates that every source fact appears exactly once. Non-fact categories are
copied unchanged. It is not a summarizer, topic grouper, protocol renderer or
Canonical Meeting Knowledge generator.
---
# Consolidated Topic
Planned semantic object produced by the Semantic Consolidator.
Example:
```json
{
"topic_id": "topic_001",
"title": "...",
"background": [],
"decisions": [],
"action_items": [],
"open_questions": [],
"durable_information": [],
"uncertainty": [],
"source_references": []
}
```
Future Semantic Consolidator versions may use the local LLM to merge
semantically equivalent statements beyond facts, group content by topic,
preserve evidence from all contributing chunks, mark contradictions and
uncertainty and separate durable information from transient discussion.
It produces Canonical Meeting Knowledge. It does not directly write a protocol.
---
# Canonical Meeting Knowledge
The canonical semantic representation of one meeting.
This representation is the single source of truth for all downstream outputs.
It is a structured representation, preferably JSON, and is not itself a prose
protocol.
```json
{
"meeting_id": "meeting_001",
"metadata": {},
"topics": [],
"facts": [],
"decisions": [],
"todos": [],
"questions": [],
"positions": [],
"technical_details": [],
"durable_information": [],
"rationale": [],
"uncertainty": [],
"source_references": []
}
```
This is the common intermediate representation for all final Output Views.
The exact schema is not final and should be refined during future
implementation work.
---
# Output Views
The final outputs are independent renderings of the Canonical Meeting Knowledge.
```text
Canonical Meeting Knowledge
├── Working Protocol
├── Distribution Protocol
└── Knowledge Objects
```
The Working Protocol, Distribution Protocol and Knowledge Objects are not
derived from one another. Each renderer reads the same canonical semantic
model and selects the level of detail appropriate for its purpose.
Knowledge Objects represent durable organizational knowledge such as processes,
definitions, responsibilities, rules, accepted practices and long-term
decisions. They are independent of the original meeting wording. Markdown is one
possible presentation, but JSON or another structured format is expected to
become the canonical storage format later.
---
# Future Extensions
Possible future additions include:
- confidence values
- evidence references
- source blocks
- priorities
- deadlines
- status tracking
- semantic relationships
These fields will only be introduced when they provide measurable benefits.
The Meeting Lab intentionally avoids designing an overly complex schema in advance.