Initial project structure

This commit is contained in:
2026-07-21 16:32:49 +02:00
commit b2c195c365
45 changed files with 999 additions and 0 deletions
+275
View File
@@ -0,0 +1,275 @@
# Data Models
## Purpose
This document describes the logical data structures exchanged between the pipeline stages of the Meeting Lab.
The goal is **not** to define a final database schema.
Instead, these models represent stable interfaces between processing modules.
Models should evolve only when required by new functionality.
---
# Design Principles
## Keep Models Small
Only include fields that are currently required.
Avoid speculative attributes.
Bad:
```json
{
"priority": "...",
"confidence": 0.93,
"risk": "...",
"category": "...",
"importance": "...",
"status": "..."
}
```
Good:
```json
{
"text": "...",
"owner": "..."
}
```
New fields can always be added later.
---
## Preserve Information
Models should preserve information rather than interpret it.
Interpretation belongs to processing modules.
---
## Stable Interfaces
Modules communicate only through documented data models.
A module must never depend on another module's internal implementation.
---
# Transcript
Represents the complete meeting transcript.
Example
```json
{
"meeting_id": "meeting_001",
"language": "en",
"blocks": []
}
```
---
# Discussion Block
The discussion block is the fundamental processing unit.
```json
{
"block_id": 42,
"speaker": "Speaker A",
"start": 351.2,
"end": 367.8,
"text": "..."
}
```
Required fields
- block_id
- text
Optional fields
- speaker
- timestamps
---
# Chunk
Technical processing unit.
```json
{
"chunk_id": 3,
"blocks": [
40,
41,
42
]
}
```
Chunks are implementation details.
They never represent discussion topics.
---
# Topic
Represents one discussion topic.
```json
{
"topic_id": "topic_003",
"title": "Ventilation",
"segments": []
}
```
---
# Topic Segment
A continuous part of a topic.
```json
{
"start_block": 40,
"end_block": 152
}
```
One topic may contain multiple segments.
---
# Fact
```json
{
"text": "..."
}
```
---
# Question
```json
{
"text": "..."
}
```
---
# Position
```json
{
"text": "...",
"speaker": "..."
}
```
---
# Decision
```json
{
"text": "..."
}
```
---
# Todo
```json
{
"text": "...",
"owner": "..."
}
```
Owner remains empty if unknown.
---
# Technical Detail
```json
{
"text": "..."
}
```
---
# Topic Result
After extraction, every topic contains the collected information.
```json
{
"topic_id": "topic_003",
"title": "Ventilation",
"segments": [],
"facts": [],
"questions": [],
"positions": [],
"decisions": [],
"todos": [],
"technical_details": []
}
```
This object represents the main output of the analysis pipeline.
---
# Meeting Result
The complete structured meeting.
```json
{
"meeting_id": "meeting_001",
"topics": []
}
```
Protocol generation operates exclusively on this structure.
---
# Future Extensions
Possible future additions include:
- confidence values
- evidence references
- source blocks
- priorities
- deadlines
- status tracking
- semantic relationships
These fields will only be introduced when they provide measurable benefits.
The Meeting Lab intentionally avoids designing an overly complex schema in advance.