Initial project structure
This commit is contained in:
@@ -0,0 +1,259 @@
|
||||
# Architecture
|
||||
|
||||
## Purpose
|
||||
|
||||
The **Meeting Lab** is an experimental environment for developing and evaluating methods to extract structured knowledge from real meeting transcripts.
|
||||
|
||||
Its purpose is not to build a complete meeting assistant, but to answer a single question:
|
||||
|
||||
> **How can knowledge be extracted from real discussions as reliably as possible?**
|
||||
|
||||
Successful approaches will later be integrated into the Meeting Assistant project.
|
||||
|
||||
---
|
||||
|
||||
# Design Goals
|
||||
|
||||
The architecture follows a small set of guiding principles.
|
||||
|
||||
## Modular Pipeline
|
||||
|
||||
Complex problems are divided into small, well-defined processing steps.
|
||||
|
||||
Each module has exactly one responsibility.
|
||||
|
||||
## Deterministic where possible
|
||||
|
||||
Tasks that can be solved reliably without an LLM should use deterministic algorithms.
|
||||
|
||||
Examples include:
|
||||
|
||||
- transcript normalization
|
||||
- whitespace cleanup
|
||||
- duplicate removal
|
||||
- chunk generation
|
||||
|
||||
LLMs are only used where semantic understanding is required.
|
||||
|
||||
## Preserve Information
|
||||
|
||||
The pipeline should never remove or rewrite information unless it is certain that the content is merely noise.
|
||||
|
||||
Losing information is considered worse than keeping harmless redundancy.
|
||||
|
||||
## Explainable Results
|
||||
|
||||
Every processing step should be understandable.
|
||||
|
||||
Intermediate results should remain inspectable throughout the pipeline.
|
||||
|
||||
## Reproducible Experiments
|
||||
|
||||
Experiments must be repeatable.
|
||||
|
||||
Given the same input, prompt, model and parameters, another developer should be able to reproduce the result.
|
||||
|
||||
## Local First
|
||||
|
||||
The complete pipeline should run locally.
|
||||
|
||||
Cloud services may be supported in the future but are not a design requirement.
|
||||
|
||||
---
|
||||
|
||||
# Core Idea
|
||||
|
||||
Traditional meeting summarization attempts to solve everything in one step.
|
||||
|
||||
```text
|
||||
Transcript
|
||||
↓
|
||||
LLM
|
||||
↓
|
||||
Summary
|
||||
```
|
||||
|
||||
Real discussions do not work that way.
|
||||
|
||||
Topics are introduced, interrupted, resumed later, expanded, questioned and finally concluded.
|
||||
|
||||
Instead of building a better summarizer, the Meeting Lab develops a **Discussion Analyzer**.
|
||||
|
||||
The analyzer gradually transforms an unstructured discussion into structured knowledge.
|
||||
|
||||
---
|
||||
|
||||
# High-Level Pipeline
|
||||
|
||||
```text
|
||||
Transcript
|
||||
↓
|
||||
Normalization
|
||||
↓
|
||||
Discussion Blocks
|
||||
↓
|
||||
Technical Chunking
|
||||
↓
|
||||
Topic Segmentation
|
||||
↓
|
||||
Specialized Extraction
|
||||
↓
|
||||
Consolidation
|
||||
↓
|
||||
Structured Meeting Data
|
||||
↓
|
||||
Protocol Generation
|
||||
```
|
||||
|
||||
Each stage solves one clearly defined problem.
|
||||
|
||||
No module should perform multiple semantic tasks simultaneously.
|
||||
|
||||
---
|
||||
|
||||
# Module Overview
|
||||
|
||||
The current architecture consists of the following processing stages.
|
||||
|
||||
## normalization/
|
||||
|
||||
Deterministic transcript cleanup.
|
||||
|
||||
Responsibilities:
|
||||
|
||||
- remove filler words
|
||||
- remove immediate repetitions
|
||||
- whitespace cleanup
|
||||
- generate change log
|
||||
|
||||
---
|
||||
|
||||
## chunking/
|
||||
|
||||
Creates model-sized chunks.
|
||||
|
||||
Chunking is purely technical.
|
||||
|
||||
It does **not** recognize discussion topics.
|
||||
|
||||
---
|
||||
|
||||
## segmentation/
|
||||
|
||||
Identifies discussion topics.
|
||||
|
||||
Responsibilities:
|
||||
|
||||
- detect topic start
|
||||
- detect topic end
|
||||
- detect topic switches
|
||||
- recognize resumed topics
|
||||
|
||||
This is the next major development milestone.
|
||||
|
||||
---
|
||||
|
||||
## extraction/
|
||||
|
||||
Contains specialized LLM modules.
|
||||
|
||||
Planned extractors include:
|
||||
|
||||
- facts
|
||||
- questions
|
||||
- positions
|
||||
- decisions
|
||||
- todos
|
||||
- technical information
|
||||
|
||||
Each extractor has exactly one task and one prompt.
|
||||
|
||||
---
|
||||
|
||||
## consolidation/
|
||||
|
||||
Merges information extracted from multiple discussion segments.
|
||||
|
||||
Typical responsibilities:
|
||||
|
||||
- merge duplicates
|
||||
- combine partial information
|
||||
- distinguish positions from decisions
|
||||
- detect contradictions
|
||||
|
||||
---
|
||||
|
||||
## protocol/
|
||||
|
||||
Generates human-readable output from structured meeting data.
|
||||
|
||||
Protocol generation never invents information.
|
||||
|
||||
It only reformulates the analysis results.
|
||||
|
||||
---
|
||||
|
||||
# Repository Layout
|
||||
|
||||
```text
|
||||
meeting-lab/
|
||||
│
|
||||
├── src/
|
||||
├── prompts/
|
||||
├── experiments/
|
||||
├── samples/
|
||||
├── tests/
|
||||
└── docs/
|
||||
```
|
||||
|
||||
Additional documentation is intentionally split into focused documents.
|
||||
|
||||
Examples:
|
||||
|
||||
- pipeline.md
|
||||
- segmentation.md
|
||||
- prompts.md
|
||||
- experiments.md
|
||||
|
||||
The architecture document only describes the overall system.
|
||||
|
||||
---
|
||||
|
||||
# Current State
|
||||
|
||||
Implemented:
|
||||
|
||||
- Transcript normalization
|
||||
- Technical chunk generation
|
||||
- Experimental LLM-based information extraction
|
||||
|
||||
The current extraction step still performs multiple tasks simultaneously.
|
||||
|
||||
This was sufficient as a proof of concept but does not reflect the intended long-term architecture.
|
||||
|
||||
---
|
||||
|
||||
# Next Milestone
|
||||
|
||||
The next development step is the implementation of **topic segmentation**.
|
||||
|
||||
Its only responsibility is to identify the thematic structure of a discussion.
|
||||
|
||||
It should answer questions such as:
|
||||
|
||||
- Where does a topic begin?
|
||||
- Where does it end?
|
||||
- When does another topic start?
|
||||
- When is an earlier topic resumed?
|
||||
|
||||
No facts, decisions or todos should be extracted at this stage.
|
||||
|
||||
Only after reliable topic segmentation has been achieved will the specialized extraction modules be implemented.
|
||||
|
||||
---
|
||||
|
||||
# Guiding Principle
|
||||
|
||||
The Meeting Lab assumes that the greatest improvement in transcript quality will not come from increasingly powerful language models.
|
||||
|
||||
Instead, quality is expected to emerge from a pipeline that decomposes a complex problem into many small, clearly defined and independently testable processing steps.
|
||||
Reference in New Issue
Block a user