- add deterministic canonicalization support for extraction items - add facts-only semantic consolidation using local Ollama - preserve source evidence and validate complete fact coverage - add conservative merge rules and non-LLM tests - record the first validated real-life consolidation benchmark - document current scope, limitations and next evaluation step
428 lines
7.0 KiB
Markdown
428 lines
7.0 KiB
Markdown
# Data Models
|
|
|
|
## Purpose
|
|
|
|
This document describes the logical data structures exchanged between the pipeline stages of the Meeting Lab.
|
|
|
|
The goal is **not** to define a final database schema.
|
|
|
|
Instead, these models represent stable interfaces between processing modules.
|
|
|
|
Models should evolve only when required by new functionality.
|
|
|
|
---
|
|
|
|
# Design Principles
|
|
|
|
## Keep Models Small
|
|
|
|
Only include fields that are currently required.
|
|
|
|
Avoid speculative attributes.
|
|
|
|
Bad:
|
|
|
|
```json
|
|
{
|
|
"priority": "...",
|
|
"confidence": 0.93,
|
|
"risk": "...",
|
|
"category": "...",
|
|
"importance": "...",
|
|
"status": "..."
|
|
}
|
|
```
|
|
|
|
Good:
|
|
|
|
```json
|
|
{
|
|
"text": "...",
|
|
"owner": "..."
|
|
}
|
|
```
|
|
|
|
New fields can always be added later.
|
|
|
|
---
|
|
|
|
## Preserve Information
|
|
|
|
Models should preserve information rather than interpret it.
|
|
|
|
Interpretation belongs to processing modules.
|
|
|
|
---
|
|
|
|
## Stable Interfaces
|
|
|
|
Modules communicate only through documented data models.
|
|
|
|
A module must never depend on another module's internal implementation.
|
|
|
|
---
|
|
|
|
# Transcript
|
|
|
|
Represents the complete meeting transcript.
|
|
|
|
Example
|
|
|
|
```json
|
|
{
|
|
"meeting_id": "meeting_001",
|
|
"language": "en",
|
|
"blocks": []
|
|
}
|
|
```
|
|
|
|
---
|
|
|
|
# Discussion Block
|
|
|
|
The discussion block is the fundamental processing unit.
|
|
|
|
```json
|
|
{
|
|
"block_id": 42,
|
|
"speaker": "Speaker A",
|
|
"start": 351.2,
|
|
"end": 367.8,
|
|
"text": "..."
|
|
}
|
|
```
|
|
|
|
Required fields
|
|
|
|
- block_id
|
|
- text
|
|
|
|
Optional fields
|
|
|
|
- speaker
|
|
- timestamps
|
|
|
|
---
|
|
|
|
# Chunk
|
|
|
|
Technical processing unit.
|
|
|
|
```json
|
|
{
|
|
"chunk_id": 3,
|
|
"blocks": [
|
|
40,
|
|
41,
|
|
42
|
|
]
|
|
}
|
|
```
|
|
|
|
Chunks are implementation details.
|
|
|
|
They never represent discussion topics.
|
|
|
|
---
|
|
|
|
# Topic
|
|
|
|
Represents one discussion topic.
|
|
|
|
```json
|
|
{
|
|
"topic_id": "topic_003",
|
|
"title": "Ventilation",
|
|
"segments": []
|
|
}
|
|
```
|
|
|
|
---
|
|
|
|
# Topic Segment
|
|
|
|
A continuous part of a topic.
|
|
|
|
```json
|
|
{
|
|
"start_block": 40,
|
|
"end_block": 152
|
|
}
|
|
```
|
|
|
|
One topic may contain multiple segments.
|
|
|
|
---
|
|
|
|
# Fact
|
|
|
|
```json
|
|
{
|
|
"text": "..."
|
|
}
|
|
```
|
|
|
|
---
|
|
|
|
# Question
|
|
|
|
```json
|
|
{
|
|
"text": "..."
|
|
}
|
|
```
|
|
|
|
---
|
|
|
|
# Position
|
|
|
|
```json
|
|
{
|
|
"text": "...",
|
|
"speaker": "..."
|
|
}
|
|
```
|
|
|
|
---
|
|
|
|
# Decision
|
|
|
|
```json
|
|
{
|
|
"text": "..."
|
|
}
|
|
```
|
|
|
|
---
|
|
|
|
# Todo
|
|
|
|
```json
|
|
{
|
|
"text": "...",
|
|
"owner": "..."
|
|
}
|
|
```
|
|
|
|
Owner remains empty if unknown.
|
|
|
|
---
|
|
|
|
# Technical Detail
|
|
|
|
```json
|
|
{
|
|
"text": "..."
|
|
}
|
|
```
|
|
|
|
---
|
|
|
|
# Topic Result
|
|
|
|
After extraction, every topic contains the collected information.
|
|
|
|
```json
|
|
{
|
|
"topic_id": "topic_003",
|
|
"title": "Ventilation",
|
|
|
|
"segments": [],
|
|
|
|
"facts": [],
|
|
"questions": [],
|
|
"positions": [],
|
|
"decisions": [],
|
|
"todos": [],
|
|
"technical_details": []
|
|
}
|
|
```
|
|
|
|
This object feeds the Canonical Meeting Knowledge representation.
|
|
|
|
---
|
|
|
|
# Canonicalized Extractions
|
|
|
|
Implemented deterministic intermediate file created from raw chunk extraction
|
|
JSON by Canonicalizer V1. This is not Canonical Meeting Knowledge.
|
|
|
|
Top-level structure:
|
|
|
|
```json
|
|
{
|
|
"schema_version": "1",
|
|
"source_files": [],
|
|
"stats": {},
|
|
"items": []
|
|
}
|
|
```
|
|
|
|
Each item contains at least:
|
|
|
|
- item_id
|
|
- category
|
|
- text
|
|
- evidence
|
|
- source_file
|
|
- source_index
|
|
- original_value
|
|
- source_references
|
|
|
|
Action items also preserve deterministic fields such as `responsible` and
|
|
`deadline` when present.
|
|
|
|
Example:
|
|
|
|
```json
|
|
{
|
|
"id": "fact.chunk_03.0001",
|
|
"category": "fact",
|
|
"text": "...",
|
|
"source_references": [
|
|
{
|
|
"chunk_id": "chunk_03",
|
|
"source_file": "chunk_03_extraction.json",
|
|
"evidence": "..."
|
|
}
|
|
]
|
|
}
|
|
```
|
|
|
|
Canonicalizer V1 creates this kind of object without an LLM. It validates and
|
|
normalizes raw extraction objects, assigns stable IDs and source references,
|
|
normalizes category names and basic field structure, performs only safe
|
|
deterministic cleanup, may group exact duplicates and must preserve all source
|
|
evidence.
|
|
|
|
It must not perform uncertain semantic merging.
|
|
|
|
---
|
|
|
|
# Semantic Fact Group
|
|
|
|
Implemented by Semantic Consolidator V0.
|
|
|
|
Example:
|
|
|
|
```json
|
|
{
|
|
"consolidated_id": "fact_group_0001",
|
|
"category": "fact",
|
|
"canonical_text": "...",
|
|
"source_item_ids": ["fact_0001"],
|
|
"source_references": [],
|
|
"evidence": [],
|
|
"merge_reason": "Singleton; no semantically equivalent fact found."
|
|
}
|
|
```
|
|
|
|
Semantic Consolidator V0 only processes fact items. It merges semantically
|
|
equivalent facts conservatively, preserves source references and evidence, and
|
|
validates that every source fact appears exactly once. Non-fact categories are
|
|
copied unchanged. It is not a summarizer, topic grouper, protocol renderer or
|
|
Canonical Meeting Knowledge generator.
|
|
|
|
---
|
|
|
|
# Consolidated Topic
|
|
|
|
Planned semantic object produced by the Semantic Consolidator.
|
|
|
|
Example:
|
|
|
|
```json
|
|
{
|
|
"topic_id": "topic_001",
|
|
"title": "...",
|
|
"background": [],
|
|
"decisions": [],
|
|
"action_items": [],
|
|
"open_questions": [],
|
|
"durable_information": [],
|
|
"uncertainty": [],
|
|
"source_references": []
|
|
}
|
|
```
|
|
|
|
Future Semantic Consolidator versions may use the local LLM to merge
|
|
semantically equivalent statements beyond facts, group content by topic,
|
|
preserve evidence from all contributing chunks, mark contradictions and
|
|
uncertainty and separate durable information from transient discussion.
|
|
|
|
It produces Canonical Meeting Knowledge. It does not directly write a protocol.
|
|
|
|
---
|
|
|
|
# Canonical Meeting Knowledge
|
|
|
|
The canonical semantic representation of one meeting.
|
|
|
|
This representation is the single source of truth for all downstream outputs.
|
|
It is a structured representation, preferably JSON, and is not itself a prose
|
|
protocol.
|
|
|
|
```json
|
|
{
|
|
"meeting_id": "meeting_001",
|
|
|
|
"metadata": {},
|
|
"topics": [],
|
|
"facts": [],
|
|
"decisions": [],
|
|
"todos": [],
|
|
"questions": [],
|
|
"positions": [],
|
|
"technical_details": [],
|
|
"durable_information": [],
|
|
"rationale": [],
|
|
"uncertainty": [],
|
|
"source_references": []
|
|
}
|
|
```
|
|
|
|
This is the common intermediate representation for all final Output Views.
|
|
The exact schema is not final and should be refined during future
|
|
implementation work.
|
|
|
|
---
|
|
|
|
# Output Views
|
|
|
|
The final outputs are independent renderings of the Canonical Meeting Knowledge.
|
|
|
|
```text
|
|
Canonical Meeting Knowledge
|
|
├── Working Protocol
|
|
├── Distribution Protocol
|
|
└── Knowledge Objects
|
|
```
|
|
|
|
The Working Protocol, Distribution Protocol and Knowledge Objects are not
|
|
derived from one another. Each renderer reads the same canonical semantic
|
|
model and selects the level of detail appropriate for its purpose.
|
|
|
|
Knowledge Objects represent durable organizational knowledge such as processes,
|
|
definitions, responsibilities, rules, accepted practices and long-term
|
|
decisions. They are independent of the original meeting wording. Markdown is one
|
|
possible presentation, but JSON or another structured format is expected to
|
|
become the canonical storage format later.
|
|
|
|
---
|
|
|
|
# Future Extensions
|
|
|
|
Possible future additions include:
|
|
|
|
- confidence values
|
|
- evidence references
|
|
- source blocks
|
|
- priorities
|
|
- deadlines
|
|
- status tracking
|
|
- semantic relationships
|
|
|
|
These fields will only be introduced when they provide measurable benefits.
|
|
|
|
The Meeting Lab intentionally avoids designing an overly complex schema in advance.
|