Files

655 lines
23 KiB
Markdown

# Architecture
## Purpose
The **Meeting Lab** is the experimental R&D environment for developing and
evaluating methods to extract structured knowledge from real meeting
transcripts. It is the research platform, architecture playground, regression
framework, benchmark environment and prototype implementation for the future
Meeting Assistant.
Its purpose is not to build a complete meeting assistant, but to answer a single question:
> **How can knowledge be extracted from real discussions as reliably as possible?**
Its purpose is to validate ideas before they are promoted into the product.
Only sufficiently mature and verified components should migrate into Meeting
Assistant. Meeting Lab may intentionally contain experiments or development
branches that are rejected, remain inconclusive or never reach the Assistant.
## Meeting Lab and Meeting Assistant lifecycle
Meeting Lab and Meeting Assistant have different long-term responsibilities:
- **Meeting Lab** is the long-term innovation branch. It favors learning,
inspectable experiments, regression evidence, benchmarks and architectural
change.
- **Meeting Assistant** is the stable product branch. It favors a polished user
experience, installation, configuration and a production pipeline.
Architectural promotion follows an evidence-based lifecycle:
```text
Research idea
->
Meeting Lab experiment
->
Regression tests
->
Stable architecture
->
Meeting Assistant implementation
```
The expected release progression is:
```text
Meeting Lab Alpha
->
Meeting Lab Beta
->
Meeting Assistant Beta
->
Meeting Assistant Release
```
The first public Meeting Assistant beta should be based on a stable Meeting
Lab MVP. It should provide a polished user experience, an installer,
configuration and a production-quality pipeline. A GUI is optional.
Experimental features should not be enabled by default. Meeting Lab continues
to evolve independently after components have migrated; promotion does not
turn the Lab itself into the product branch.
---
# Design Goals
The architecture follows a small set of guiding principles.
## Modular Pipeline
Complex problems are divided into small, well-defined processing steps.
Each module has exactly one responsibility.
## Deterministic where possible
Tasks that can be solved reliably without an LLM should use deterministic algorithms.
Examples include:
- transcript normalization
- whitespace cleanup
- duplicate removal
- chunk generation
LLMs are only used where semantic understanding is required.
## Preserve Information
The pipeline should never remove or rewrite information unless it is certain that the content is merely noise.
Losing information is considered worse than keeping harmless redundancy.
## Explainable Results
Every processing step should be understandable.
Intermediate results should remain inspectable throughout the pipeline.
## Responsibility Attribution Integrity
Responsibility, ownership, organizational roles and action-item assignments may
be recorded only when meeting evidence explicitly assigns, accepts or confirms
them.
The system must not infer responsibility from thematic proximity,
participation in a discussion, mentioning a task, commenting on another
department, organizational assumptions, likely job roles, speaker adjacency or
model world knowledge.
When evidence is incomplete or ambiguous, the responsible person remains unset
or unclear and the supporting evidence is preserved.
## Reproducible Experiments
Experiments must be repeatable.
Given the same input, prompt, model and parameters, another developer should be able to reproduce the result.
## Local First
The complete pipeline should run locally.
Cloud services may be supported in the future but are not a design requirement.
---
# Core Idea
Traditional meeting summarization attempts to solve everything in one step.
```text
Transcript
↓
LLM
↓
Summary
```
Real discussions do not work that way.
Topics are introduced, interrupted, resumed later, expanded, questioned and finally concluded.
Instead of building a better summarizer, the Meeting Lab develops a **Discussion Analyzer**.
The analyzer gradually transforms an unstructured discussion into structured knowledge.
---
# High-Level Pipeline
Current implemented and intended analysis flow:
```text
Transcript
↓
Normalization
↓
Discussion Blocks
↓
Technical Chunking
↓
Topic Segmentation
↓
Specialized Extraction
↓
Deterministic Canonicalization
↓
Semantic Consolidation
↓
Canonical Meeting Knowledge
↓
Output View Rendering
↓
Working Protocol / Distribution Protocol / Knowledge Objects
```
Accepted future Meeting Context V2 preparation flow:
```text
Whisper
↓
Entity Detection
↓
User Confirmation
↓
Entity Registry Update
↓
Meeting Context Builder
↓
meeting_context.yaml
↓
Extraction Pipeline
```
Each stage solves one clearly defined problem.
No module should perform multiple semantic tasks simultaneously.
---
# Module Overview
The current architecture consists of the following processing stages.
## normalization/
Deterministic transcript cleanup.
Responsibilities:
- remove filler words
- remove immediate repetitions
- whitespace cleanup
- generate change log
---
## chunking/
Creates model-sized chunks.
Chunking is purely technical.
It does **not** recognize discussion topics.
---
## segmentation/
Identifies discussion topics.
Responsibilities:
- detect topic start
- detect topic end
- detect topic switches
- recognize resumed topics
This is the next major development milestone.
---
## extraction/
Contains specialized LLM modules.
Planned extractors include:
- facts
- questions
- positions
- decisions
- todos
- technical information
Each extractor has exactly one task and one prompt.
---
## Meeting Context
Meeting Context V1 is a manually maintained YAML scaffold for reliable meeting
metadata such as title, language, participants, aliases, departments,
abbreviations and known entities.
It is documented in `docs/meeting-context.md` and templated at
`samples/templates/meeting_context.template.yaml`. It is implemented for
loading, validation and optional injection into chunk extraction prompts.
Extraction results record only minimal context provenance. It is not yet
connected to consolidation, Canonical Meeting Knowledge or output rendering.
The context can help prevent non-participants from being interpreted as
attendees and can normalize known aliases for extraction. It must not infer
roles, departments, responsibilities or decisions.
Accepted future direction:
Meeting Context V2 should be generated or assisted from an interactive entity
confirmation workflow and a persistent Entity Registry. The Entity Registry is
the persistent cross-meeting knowledge source for confirmed people,
organizations, departments, products, projects, locations, aliases and
organizational metadata under stable internal IDs. Display names may change,
but internal IDs remain stable. Aliases are first-class data.
The registry never learns automatically. It may propose matches and aliases,
including spelling variants, Whisper transcription variants, umlaut variants
and OCR-like mistakes, but only explicit user confirmation updates registry
state. Unknown names should be presented to the user as meeting participant,
mentioned person, external person, transcription error or ignore.
In this architecture, `meeting_context.yaml` remains the authoritative
meeting-specific Point of Truth and reproducible input artifact consumed by the
extraction pipeline. It is a meeting-specific snapshot built from the Entity
Registry, user confirmations and meeting metadata. The Registry must not
override explicit meeting-specific confirmations, and Registry changes after a
meeting run must not silently change the historical Meeting Context used for
that run.
Status: Accepted Architecture; implementation deferred. See
`docs/adr-meeting-context-v2-entity-registry.md`.
---
## consolidation/
Planned area for canonicalization and consolidation.
The next milestone splits this into two stages.
Deterministic Canonicalizer:
- implemented in Python
- uses no LLM
- validates and normalizes extraction objects
- assigns stable source references and IDs
- normalizes category names and basic field structure
- validates and normalizes action-item responsible fields against Meeting
Context when available: known participant and mentioned-person aliases are
normalized to canonical display names, while dates, locations, projects,
products, technical terms, generic process words and unknown free text are
cleared with a structured validation record
- performs only safe deterministic cleanup
- may group exact duplicates
- preserves all source evidence
- must not perform uncertain semantic merging
Semantic Consolidator:
- uses the local LLM
- V0 is implemented for facts-only semantic duplicate detection
- V0 merges semantically equivalent fact items conservatively
- V0 preserves source references and evidence
- V0 validates that every source fact appears exactly once
- V0 sizes its Ollama output budget from the actual fact payload instead of
using a fixed response cap for every meeting
- V0 may apply deterministic source-coverage repair after valid model JSON is
parsed: duplicate source IDs are removed after their first occurrence, empty
groups are removed and missing source facts are restored as singleton groups
from canonicalized input before strict validation runs
- V0 does not process non-fact categories semantically
- later versions should group content by topic, mark contradictions and
uncertainty, separate durable information from transient discussion and
prepare Canonical Meeting Knowledge
- does not directly write a protocol
---
## protocol/
Generates output views from Canonical Meeting Knowledge.
Output generation never invents information.
It only reformulates the analysis results for a specific audience and purpose.
Depending on the output and maturity of the implementation, a renderer may be
deterministic, template-based or LLM-assisted.
LLM-assisted renderers preserve raw model output separately and write the final
output artifact only after deterministic contract validation succeeds. Renderer
post-processing may remove non-semantic wrapper text, but must not fabricate
missing semantic sections or relabel an invalid summary as a valid output view.
The planned output products are:
- Working Protocol (`working_protocol.md`, Arbeitsprotokoll)
- Distribution Protocol (`distribution_protocol.md`,
Verteilerprotokoll)
- Knowledge Objects, which may be rendered as a Knowledge-base Entry
(`knowledge_entry.md`) and later stored in a structured format such as
`knowledge_entry.json` (Wissensdatenbankeintrag)
These are parallel renderings of the same canonical semantic model, not
documents derived from one another.
Rendered protocol language should normally match the dominant language of the
source transcript or consolidated meeting knowledge unless an explicit output
language is requested.
---
# Repository Layout
```text
meeting-lab/
│
├── src/
├── prompts/
├── experiments/
├── samples/
├── tests/
└── docs/
```
Additional documentation is intentionally split into focused documents.
Examples:
- pipeline.md
- segmentation.md
- prompts.md
- experiments.md
- output-views.md
The architecture document only describes the overall system.
---
# Version 2 Accepted Architectural Direction
The following topics are accepted architectural goals for Version 2. They
record direction reached through the BUG-011 through BUG-015 investigations;
they are not descriptions of implemented behavior or authorization to change
the current pipeline.
## Speaker diarization before semantic analysis
Version 2 should determine **who is speaking before semantic analysis**.
Speaker identity contains evidence that cannot reliably be reconstructed from
text alone. It helps distinguish, for example, who answers a question, accepts
work, agrees with a proposal, or advances the discussion after another
speaker. It also preserves conversational flow that anonymous transcript text
can erase.
Diarization is therefore a semantic prerequisite in the intended Version 2
architecture, not merely a display enhancement. Its output should remain
traceable to transcript segments so later stages can preserve speaker and
source provenance.
## Persistent speaker identification
Version 2 should add a persistent speaker database and an interactive identity
workflow during import:
```text
Unknown speaker detected
->
Representative audio sample (approximately 20 seconds)
->
User selects an existing identity or creates a new identity
->
Known speaker available for future recognition
```
Automatic recognition may suggest an identity, but user confirmation governs
the persistent association. Over the long term, speaker embeddings rather
than raw meeting recordings should be the persistent recognition
representation. Representative raw audio is an import and confirmation aid,
not the intended durable identity store. Privacy, deletion and false-match
handling require separate design before implementation.
## Meeting Context as a probabilistic prior
Known meeting participants should influence semantic interpretation, but
Meeting Context is a **probabilistic prior**, not a deterministic semantic
rule. It can make one interpretation more plausible and help focus review; it
must never manufacture a commitment, decision or responsibility assignment.
For example, if Marleen is confirmed as present, “Marleen müsste sich mal
äußern” is more likely to be conversation management directed at a current
participant than future project work. Presence alone does not prove this
interpretation, and it does not establish an Action Item or responsibility.
Explicit meeting evidence remains authoritative. This extends, rather than
weakens, the responsibility attribution invariant.
The existing Meeting Context V2 entity direction is documented in
[`adr-meeting-context-v2-entity-registry.md`](adr-meeting-context-v2-entity-registry.md).
Speaker identities and meeting-specific participant confirmation should
eventually feed that context without turning registry metadata into semantic
facts.
## Conversation Management versus Meeting Content
BUG-015 reinforced that not every utterance is protocol-worthy content.
Version 2 should conceptually distinguish:
- **Conversation Management**: utterances that coordinate the meeting itself,
such as asking a present participant to speak, moderation, requesting a
slide, or asking someone to repeat something.
- **Meeting Content**: propositions that may contribute to the meeting's
durable knowledge, including facts, technical findings, decisions, action
items and open questions.
This is an architectural concept, not a currently implemented category or
filter. The distinction should prevent conversational coordination from being
promoted into project commitments while retaining sufficient provenance to
understand dialogue. Context and diarization can inform the distinction, but
neither should act as a deterministic keyword or participant rule.
## Evidence and commitment before protocol eligibility
The BUG-015 design study concludes that semantic state should be classified
before policy determines whether an item is eligible for a protocol. A binary
keep/reject verifier conflates evidence recognition with publication policy
and loses valid intermediate states.
The proposed decision progression is:
```text
idea -> option -> proposal -> preferred option -> tentative agreement -> decision
```
The proposed action progression is:
```text
possible next step -> recommendation -> requested action
-> established action -> ongoing work -> completed
```
Questions combine a communicative **kind** with an independent **resolution
state**, rather than treating every uncertainty or interrogative as an Open
Question. Responsibility remains an independent dimension and may be recorded
only when explicitly assigned, accepted or confirmed. Evidence strength is
also independent: it describes support for a semantic label, not semantic
maturity or protocol eligibility.
After semantic state classification, explicit policy should select Decisions,
Action Items and Open Questions for a particular output view. Renderers should
receive policy-selected semantic content and must not promote proposals or
conversation management into commitments. The complete taxonomy, trade-offs,
architecture interactions and migration questions are recorded in
[`design/evidence_commitment_model.md`](design/evidence_commitment_model.md).
## Topic-oriented primary protocol
One of the highest-level Version 2 requirements is:
> **The protocol is primarily a topic-oriented reconstruction of the meeting, not a category-oriented listing of extracted information.**
Semantic categories remain metadata and supporting structure within topics;
they must not dictate the main document structure. The conceptual target flow
is `Transcript -> Evidence Extraction -> Topic Reconstruction -> Semantic
Synthesis -> Protocol Rendering`. The primary renderer should eventually
receive topic-oriented semantic knowledge. Category-oriented Action Item,
Decision, Open Question and management views remain useful derived outputs.
The detailed rationale, example, architectural implications and explicit
deferral of a final Topic Reconstruction schema are documented in
[`design/evidence_commitment_model.md`](design/evidence_commitment_model.md#thematic-protocol-as-the-primary-structure).
Implementation is intentionally postponed until this architecture has been
reviewed. No Version 2 goal in this section changes current extraction,
canonicalization, consolidation or rendering behavior.
---
# Current State
Implemented:
- Transcript normalization
- Technical chunk generation
- Experimental LLM-based information extraction
- Meeting Context V1 loading, validation and extraction prompt integration
- Canonicalizer V1 deterministic extraction canonicalization
The current extraction step still performs multiple tasks simultaneously.
This was sufficient as a proof of concept but does not reflect the intended long-term architecture.
The current protocol builder is also an interim implementation. It concatenates
extraction results into `meeting_protocol.md` for technical validation. The
planned architecture separates Canonical Meeting Knowledge from the final Output
Views documented in `output-views.md`.
## Extraction classification contract
Decision, Action Item and Open Question extraction shares one evidence-oriented
classification contract. It is placed after the transcript so it remains the
final classification instruction in the single multi-category extraction call.
Decisions require a settled outcome; Action Items require established work;
Open Questions require a concrete unresolved need. Unsupported candidates must
not be moved into another category.
Action existence and responsibility attribution are separate checks. A valid
Action Item may have no known owner, while a named owner requires explicit
assignment, volunteering or acceptance. These are semantic LLM classifications;
deterministic validation must not guess intent from keywords.
### Classification Verifier
Before extraction output is normalized for Canonicalizer input, Decision,
Action Item and Open Question candidates pass through a semantic precision
gate. Facts and technical details pass through unchanged. Each candidate is
reviewed independently with its evidence and bounded local chunk context; the
verifier may only keep or reject the existing candidate. It cannot add or
rewrite semantic items.
Verifier output contains the stable candidate ID, category, `keep|reject`
verdict, evidence-based reason and `responsibility_supported`. A kept Action
Item with an unsupported named owner is retained with its responsibility
cleared. Malformed output fails the extraction verification substage closed,
after preserving candidate input and raw response. Per-candidate results and an
aggregate audit remain traceable before Canonicalizer input is written.
## Semantic Consolidator failure handling
Semantic Consolidator V0 preserves every raw model response before parsing.
Its normal path accepts parseable grouping JSON and leaves duplicate source-ID
unknown-source-ID and missing-source-ID correction to the deterministic
coverage repair. Unknown source IDs are removed without attempting numeric or
semantic remapping. A group retains its valid IDs and remains present when at
least one valid ID survives. A group containing only unknown IDs is removed
after it becomes empty. Existing coverage repair then restores every genuinely
missing canonical fact as a singleton. Strict validation runs against the
repaired result and still rejects any unknown ID that survives this process.
An invalid or truncated response is not retried by default. One controlled
retry is allowed only when deterministic inspection finds at least three
consecutive complete groups with an identical structural signature:
`canonical_text`, ordered `source_item_ids` and `merge_reason`. The retry keeps
the same model, temperature, context window and generation limit and adds only
an instruction not to emit an identical group more than once. Both attempts
and the detected repetition metadata are preserved. If the retry also fails,
the stage fails normally; it does not make another LLM call.
## Working Protocol V2 renderer contract
The renderer deterministically projects consolidated items to the semantic
fields required for presentation and omits bulky provenance fields from the
LLM request. Every renderable item remains represented; structurally empty
items are recorded separately rather than turned into invented prose. The
compact renderer input is preserved as an artifact.
The exact Markdown structure is generated from the same heading constants used
by the validator and appended after the renderer input so it remains visible
within the evaluated context. Decisions, action items and open questions carry
input-derived hidden coverage markers. Strict validation requires every such
renderable priority item exactly once in its matching section and rejects
missing, duplicate, wrong-section or invented markers. Facts and technical
details remain condensable as background.
Renderer output budgeting is adaptive to required priority content and prompt
size, while an explicit `num_predict` override remains authoritative. Raw model
output is always preserved, and `working_protocol.md` is written only after
strict structure and coverage validation passes. The renderer does not retry
automatically.
---
# Next Milestone
The next architecture milestone is the implementation of a deterministic
canonicalization stage followed by a semantic consolidation stage. These stages
convert raw chunk extraction JSON into evidence-preserving Canonical Meeting
Knowledge before any Output View renderer writes a protocol.
---
# Guiding Principle
The Meeting Lab assumes that the greatest improvement in transcript quality will not come from increasingly powerful language models.
Instead, quality is expected to emerge from a pipeline that decomposes a complex problem into many small, clearly defined and independently testable processing steps.