655 lines
23 KiB
Markdown
655 lines
23 KiB
Markdown
# Architecture
|
|
|
|
## Purpose
|
|
|
|
The **Meeting Lab** is the experimental R&D environment for developing and
|
|
evaluating methods to extract structured knowledge from real meeting
|
|
transcripts. It is the research platform, architecture playground, regression
|
|
framework, benchmark environment and prototype implementation for the future
|
|
Meeting Assistant.
|
|
|
|
Its purpose is not to build a complete meeting assistant, but to answer a single question:
|
|
|
|
> **How can knowledge be extracted from real discussions as reliably as possible?**
|
|
|
|
Its purpose is to validate ideas before they are promoted into the product.
|
|
Only sufficiently mature and verified components should migrate into Meeting
|
|
Assistant. Meeting Lab may intentionally contain experiments or development
|
|
branches that are rejected, remain inconclusive or never reach the Assistant.
|
|
|
|
## Meeting Lab and Meeting Assistant lifecycle
|
|
|
|
Meeting Lab and Meeting Assistant have different long-term responsibilities:
|
|
|
|
- **Meeting Lab** is the long-term innovation branch. It favors learning,
|
|
inspectable experiments, regression evidence, benchmarks and architectural
|
|
change.
|
|
- **Meeting Assistant** is the stable product branch. It favors a polished user
|
|
experience, installation, configuration and a production pipeline.
|
|
|
|
Architectural promotion follows an evidence-based lifecycle:
|
|
|
|
```text
|
|
Research idea
|
|
->
|
|
Meeting Lab experiment
|
|
->
|
|
Regression tests
|
|
->
|
|
Stable architecture
|
|
->
|
|
Meeting Assistant implementation
|
|
```
|
|
|
|
The expected release progression is:
|
|
|
|
```text
|
|
Meeting Lab Alpha
|
|
->
|
|
Meeting Lab Beta
|
|
->
|
|
Meeting Assistant Beta
|
|
->
|
|
Meeting Assistant Release
|
|
```
|
|
|
|
The first public Meeting Assistant beta should be based on a stable Meeting
|
|
Lab MVP. It should provide a polished user experience, an installer,
|
|
configuration and a production-quality pipeline. A GUI is optional.
|
|
Experimental features should not be enabled by default. Meeting Lab continues
|
|
to evolve independently after components have migrated; promotion does not
|
|
turn the Lab itself into the product branch.
|
|
|
|
---
|
|
|
|
# Design Goals
|
|
|
|
The architecture follows a small set of guiding principles.
|
|
|
|
## Modular Pipeline
|
|
|
|
Complex problems are divided into small, well-defined processing steps.
|
|
|
|
Each module has exactly one responsibility.
|
|
|
|
## Deterministic where possible
|
|
|
|
Tasks that can be solved reliably without an LLM should use deterministic algorithms.
|
|
|
|
Examples include:
|
|
|
|
- transcript normalization
|
|
- whitespace cleanup
|
|
- duplicate removal
|
|
- chunk generation
|
|
|
|
LLMs are only used where semantic understanding is required.
|
|
|
|
## Preserve Information
|
|
|
|
The pipeline should never remove or rewrite information unless it is certain that the content is merely noise.
|
|
|
|
Losing information is considered worse than keeping harmless redundancy.
|
|
|
|
## Explainable Results
|
|
|
|
Every processing step should be understandable.
|
|
|
|
Intermediate results should remain inspectable throughout the pipeline.
|
|
|
|
## Responsibility Attribution Integrity
|
|
|
|
Responsibility, ownership, organizational roles and action-item assignments may
|
|
be recorded only when meeting evidence explicitly assigns, accepts or confirms
|
|
them.
|
|
|
|
The system must not infer responsibility from thematic proximity,
|
|
participation in a discussion, mentioning a task, commenting on another
|
|
department, organizational assumptions, likely job roles, speaker adjacency or
|
|
model world knowledge.
|
|
|
|
When evidence is incomplete or ambiguous, the responsible person remains unset
|
|
or unclear and the supporting evidence is preserved.
|
|
|
|
## Reproducible Experiments
|
|
|
|
Experiments must be repeatable.
|
|
|
|
Given the same input, prompt, model and parameters, another developer should be able to reproduce the result.
|
|
|
|
## Local First
|
|
|
|
The complete pipeline should run locally.
|
|
|
|
Cloud services may be supported in the future but are not a design requirement.
|
|
|
|
---
|
|
|
|
# Core Idea
|
|
|
|
Traditional meeting summarization attempts to solve everything in one step.
|
|
|
|
```text
|
|
Transcript
|
|
↓
|
|
LLM
|
|
↓
|
|
Summary
|
|
```
|
|
|
|
Real discussions do not work that way.
|
|
|
|
Topics are introduced, interrupted, resumed later, expanded, questioned and finally concluded.
|
|
|
|
Instead of building a better summarizer, the Meeting Lab develops a **Discussion Analyzer**.
|
|
|
|
The analyzer gradually transforms an unstructured discussion into structured knowledge.
|
|
|
|
---
|
|
|
|
# High-Level Pipeline
|
|
|
|
Current implemented and intended analysis flow:
|
|
|
|
```text
|
|
Transcript
|
|
↓
|
|
Normalization
|
|
↓
|
|
Discussion Blocks
|
|
↓
|
|
Technical Chunking
|
|
↓
|
|
Topic Segmentation
|
|
↓
|
|
Specialized Extraction
|
|
↓
|
|
Deterministic Canonicalization
|
|
↓
|
|
Semantic Consolidation
|
|
↓
|
|
Canonical Meeting Knowledge
|
|
↓
|
|
Output View Rendering
|
|
↓
|
|
Working Protocol / Distribution Protocol / Knowledge Objects
|
|
```
|
|
|
|
Accepted future Meeting Context V2 preparation flow:
|
|
|
|
```text
|
|
Whisper
|
|
↓
|
|
Entity Detection
|
|
↓
|
|
User Confirmation
|
|
↓
|
|
Entity Registry Update
|
|
↓
|
|
Meeting Context Builder
|
|
↓
|
|
meeting_context.yaml
|
|
↓
|
|
Extraction Pipeline
|
|
```
|
|
|
|
Each stage solves one clearly defined problem.
|
|
|
|
No module should perform multiple semantic tasks simultaneously.
|
|
|
|
---
|
|
|
|
# Module Overview
|
|
|
|
The current architecture consists of the following processing stages.
|
|
|
|
## normalization/
|
|
|
|
Deterministic transcript cleanup.
|
|
|
|
Responsibilities:
|
|
|
|
- remove filler words
|
|
- remove immediate repetitions
|
|
- whitespace cleanup
|
|
- generate change log
|
|
|
|
---
|
|
|
|
## chunking/
|
|
|
|
Creates model-sized chunks.
|
|
|
|
Chunking is purely technical.
|
|
|
|
It does **not** recognize discussion topics.
|
|
|
|
---
|
|
|
|
## segmentation/
|
|
|
|
Identifies discussion topics.
|
|
|
|
Responsibilities:
|
|
|
|
- detect topic start
|
|
- detect topic end
|
|
- detect topic switches
|
|
- recognize resumed topics
|
|
|
|
This is the next major development milestone.
|
|
|
|
---
|
|
|
|
## extraction/
|
|
|
|
Contains specialized LLM modules.
|
|
|
|
Planned extractors include:
|
|
|
|
- facts
|
|
- questions
|
|
- positions
|
|
- decisions
|
|
- todos
|
|
- technical information
|
|
|
|
Each extractor has exactly one task and one prompt.
|
|
|
|
---
|
|
|
|
## Meeting Context
|
|
|
|
Meeting Context V1 is a manually maintained YAML scaffold for reliable meeting
|
|
metadata such as title, language, participants, aliases, departments,
|
|
abbreviations and known entities.
|
|
|
|
It is documented in `docs/meeting-context.md` and templated at
|
|
`samples/templates/meeting_context.template.yaml`. It is implemented for
|
|
loading, validation and optional injection into chunk extraction prompts.
|
|
Extraction results record only minimal context provenance. It is not yet
|
|
connected to consolidation, Canonical Meeting Knowledge or output rendering.
|
|
|
|
The context can help prevent non-participants from being interpreted as
|
|
attendees and can normalize known aliases for extraction. It must not infer
|
|
roles, departments, responsibilities or decisions.
|
|
|
|
Accepted future direction:
|
|
|
|
Meeting Context V2 should be generated or assisted from an interactive entity
|
|
confirmation workflow and a persistent Entity Registry. The Entity Registry is
|
|
the persistent cross-meeting knowledge source for confirmed people,
|
|
organizations, departments, products, projects, locations, aliases and
|
|
organizational metadata under stable internal IDs. Display names may change,
|
|
but internal IDs remain stable. Aliases are first-class data.
|
|
|
|
The registry never learns automatically. It may propose matches and aliases,
|
|
including spelling variants, Whisper transcription variants, umlaut variants
|
|
and OCR-like mistakes, but only explicit user confirmation updates registry
|
|
state. Unknown names should be presented to the user as meeting participant,
|
|
mentioned person, external person, transcription error or ignore.
|
|
|
|
In this architecture, `meeting_context.yaml` remains the authoritative
|
|
meeting-specific Point of Truth and reproducible input artifact consumed by the
|
|
extraction pipeline. It is a meeting-specific snapshot built from the Entity
|
|
Registry, user confirmations and meeting metadata. The Registry must not
|
|
override explicit meeting-specific confirmations, and Registry changes after a
|
|
meeting run must not silently change the historical Meeting Context used for
|
|
that run.
|
|
|
|
Status: Accepted Architecture; implementation deferred. See
|
|
`docs/adr-meeting-context-v2-entity-registry.md`.
|
|
|
|
---
|
|
|
|
## consolidation/
|
|
|
|
Planned area for canonicalization and consolidation.
|
|
|
|
The next milestone splits this into two stages.
|
|
|
|
Deterministic Canonicalizer:
|
|
|
|
- implemented in Python
|
|
- uses no LLM
|
|
- validates and normalizes extraction objects
|
|
- assigns stable source references and IDs
|
|
- normalizes category names and basic field structure
|
|
- validates and normalizes action-item responsible fields against Meeting
|
|
Context when available: known participant and mentioned-person aliases are
|
|
normalized to canonical display names, while dates, locations, projects,
|
|
products, technical terms, generic process words and unknown free text are
|
|
cleared with a structured validation record
|
|
- performs only safe deterministic cleanup
|
|
- may group exact duplicates
|
|
- preserves all source evidence
|
|
- must not perform uncertain semantic merging
|
|
|
|
Semantic Consolidator:
|
|
|
|
- uses the local LLM
|
|
- V0 is implemented for facts-only semantic duplicate detection
|
|
- V0 merges semantically equivalent fact items conservatively
|
|
- V0 preserves source references and evidence
|
|
- V0 validates that every source fact appears exactly once
|
|
- V0 sizes its Ollama output budget from the actual fact payload instead of
|
|
using a fixed response cap for every meeting
|
|
- V0 may apply deterministic source-coverage repair after valid model JSON is
|
|
parsed: duplicate source IDs are removed after their first occurrence, empty
|
|
groups are removed and missing source facts are restored as singleton groups
|
|
from canonicalized input before strict validation runs
|
|
- V0 does not process non-fact categories semantically
|
|
- later versions should group content by topic, mark contradictions and
|
|
uncertainty, separate durable information from transient discussion and
|
|
prepare Canonical Meeting Knowledge
|
|
- does not directly write a protocol
|
|
|
|
---
|
|
|
|
## protocol/
|
|
|
|
Generates output views from Canonical Meeting Knowledge.
|
|
|
|
Output generation never invents information.
|
|
|
|
It only reformulates the analysis results for a specific audience and purpose.
|
|
Depending on the output and maturity of the implementation, a renderer may be
|
|
deterministic, template-based or LLM-assisted.
|
|
|
|
LLM-assisted renderers preserve raw model output separately and write the final
|
|
output artifact only after deterministic contract validation succeeds. Renderer
|
|
post-processing may remove non-semantic wrapper text, but must not fabricate
|
|
missing semantic sections or relabel an invalid summary as a valid output view.
|
|
|
|
The planned output products are:
|
|
|
|
- Working Protocol (`working_protocol.md`, Arbeitsprotokoll)
|
|
- Distribution Protocol (`distribution_protocol.md`,
|
|
Verteilerprotokoll)
|
|
- Knowledge Objects, which may be rendered as a Knowledge-base Entry
|
|
(`knowledge_entry.md`) and later stored in a structured format such as
|
|
`knowledge_entry.json` (Wissensdatenbankeintrag)
|
|
|
|
These are parallel renderings of the same canonical semantic model, not
|
|
documents derived from one another.
|
|
|
|
Rendered protocol language should normally match the dominant language of the
|
|
source transcript or consolidated meeting knowledge unless an explicit output
|
|
language is requested.
|
|
|
|
---
|
|
|
|
# Repository Layout
|
|
|
|
```text
|
|
meeting-lab/
|
|
│
|
|
├── src/
|
|
├── prompts/
|
|
├── experiments/
|
|
├── samples/
|
|
├── tests/
|
|
└── docs/
|
|
```
|
|
|
|
Additional documentation is intentionally split into focused documents.
|
|
|
|
Examples:
|
|
|
|
- pipeline.md
|
|
- segmentation.md
|
|
- prompts.md
|
|
- experiments.md
|
|
- output-views.md
|
|
|
|
The architecture document only describes the overall system.
|
|
|
|
---
|
|
|
|
# Version 2 Accepted Architectural Direction
|
|
|
|
The following topics are accepted architectural goals for Version 2. They
|
|
record direction reached through the BUG-011 through BUG-015 investigations;
|
|
they are not descriptions of implemented behavior or authorization to change
|
|
the current pipeline.
|
|
|
|
## Speaker diarization before semantic analysis
|
|
|
|
Version 2 should determine **who is speaking before semantic analysis**.
|
|
Speaker identity contains evidence that cannot reliably be reconstructed from
|
|
text alone. It helps distinguish, for example, who answers a question, accepts
|
|
work, agrees with a proposal, or advances the discussion after another
|
|
speaker. It also preserves conversational flow that anonymous transcript text
|
|
can erase.
|
|
|
|
Diarization is therefore a semantic prerequisite in the intended Version 2
|
|
architecture, not merely a display enhancement. Its output should remain
|
|
traceable to transcript segments so later stages can preserve speaker and
|
|
source provenance.
|
|
|
|
## Persistent speaker identification
|
|
|
|
Version 2 should add a persistent speaker database and an interactive identity
|
|
workflow during import:
|
|
|
|
```text
|
|
Unknown speaker detected
|
|
->
|
|
Representative audio sample (approximately 20 seconds)
|
|
->
|
|
User selects an existing identity or creates a new identity
|
|
->
|
|
Known speaker available for future recognition
|
|
```
|
|
|
|
Automatic recognition may suggest an identity, but user confirmation governs
|
|
the persistent association. Over the long term, speaker embeddings rather
|
|
than raw meeting recordings should be the persistent recognition
|
|
representation. Representative raw audio is an import and confirmation aid,
|
|
not the intended durable identity store. Privacy, deletion and false-match
|
|
handling require separate design before implementation.
|
|
|
|
## Meeting Context as a probabilistic prior
|
|
|
|
Known meeting participants should influence semantic interpretation, but
|
|
Meeting Context is a **probabilistic prior**, not a deterministic semantic
|
|
rule. It can make one interpretation more plausible and help focus review; it
|
|
must never manufacture a commitment, decision or responsibility assignment.
|
|
|
|
For example, if Marleen is confirmed as present, “Marleen müsste sich mal
|
|
äußern” is more likely to be conversation management directed at a current
|
|
participant than future project work. Presence alone does not prove this
|
|
interpretation, and it does not establish an Action Item or responsibility.
|
|
Explicit meeting evidence remains authoritative. This extends, rather than
|
|
weakens, the responsibility attribution invariant.
|
|
|
|
The existing Meeting Context V2 entity direction is documented in
|
|
[`adr-meeting-context-v2-entity-registry.md`](adr-meeting-context-v2-entity-registry.md).
|
|
Speaker identities and meeting-specific participant confirmation should
|
|
eventually feed that context without turning registry metadata into semantic
|
|
facts.
|
|
|
|
## Conversation Management versus Meeting Content
|
|
|
|
BUG-015 reinforced that not every utterance is protocol-worthy content.
|
|
Version 2 should conceptually distinguish:
|
|
|
|
- **Conversation Management**: utterances that coordinate the meeting itself,
|
|
such as asking a present participant to speak, moderation, requesting a
|
|
slide, or asking someone to repeat something.
|
|
- **Meeting Content**: propositions that may contribute to the meeting's
|
|
durable knowledge, including facts, technical findings, decisions, action
|
|
items and open questions.
|
|
|
|
This is an architectural concept, not a currently implemented category or
|
|
filter. The distinction should prevent conversational coordination from being
|
|
promoted into project commitments while retaining sufficient provenance to
|
|
understand dialogue. Context and diarization can inform the distinction, but
|
|
neither should act as a deterministic keyword or participant rule.
|
|
|
|
## Evidence and commitment before protocol eligibility
|
|
|
|
The BUG-015 design study concludes that semantic state should be classified
|
|
before policy determines whether an item is eligible for a protocol. A binary
|
|
keep/reject verifier conflates evidence recognition with publication policy
|
|
and loses valid intermediate states.
|
|
|
|
The proposed decision progression is:
|
|
|
|
```text
|
|
idea -> option -> proposal -> preferred option -> tentative agreement -> decision
|
|
```
|
|
|
|
The proposed action progression is:
|
|
|
|
```text
|
|
possible next step -> recommendation -> requested action
|
|
-> established action -> ongoing work -> completed
|
|
```
|
|
|
|
Questions combine a communicative **kind** with an independent **resolution
|
|
state**, rather than treating every uncertainty or interrogative as an Open
|
|
Question. Responsibility remains an independent dimension and may be recorded
|
|
only when explicitly assigned, accepted or confirmed. Evidence strength is
|
|
also independent: it describes support for a semantic label, not semantic
|
|
maturity or protocol eligibility.
|
|
|
|
After semantic state classification, explicit policy should select Decisions,
|
|
Action Items and Open Questions for a particular output view. Renderers should
|
|
receive policy-selected semantic content and must not promote proposals or
|
|
conversation management into commitments. The complete taxonomy, trade-offs,
|
|
architecture interactions and migration questions are recorded in
|
|
[`design/evidence_commitment_model.md`](design/evidence_commitment_model.md).
|
|
|
|
## Topic-oriented primary protocol
|
|
|
|
One of the highest-level Version 2 requirements is:
|
|
|
|
> **The protocol is primarily a topic-oriented reconstruction of the meeting, not a category-oriented listing of extracted information.**
|
|
|
|
Semantic categories remain metadata and supporting structure within topics;
|
|
they must not dictate the main document structure. The conceptual target flow
|
|
is `Transcript -> Evidence Extraction -> Topic Reconstruction -> Semantic
|
|
Synthesis -> Protocol Rendering`. The primary renderer should eventually
|
|
receive topic-oriented semantic knowledge. Category-oriented Action Item,
|
|
Decision, Open Question and management views remain useful derived outputs.
|
|
|
|
The detailed rationale, example, architectural implications and explicit
|
|
deferral of a final Topic Reconstruction schema are documented in
|
|
[`design/evidence_commitment_model.md`](design/evidence_commitment_model.md#thematic-protocol-as-the-primary-structure).
|
|
|
|
Implementation is intentionally postponed until this architecture has been
|
|
reviewed. No Version 2 goal in this section changes current extraction,
|
|
canonicalization, consolidation or rendering behavior.
|
|
|
|
---
|
|
|
|
# Current State
|
|
|
|
Implemented:
|
|
|
|
- Transcript normalization
|
|
- Technical chunk generation
|
|
- Experimental LLM-based information extraction
|
|
- Meeting Context V1 loading, validation and extraction prompt integration
|
|
- Canonicalizer V1 deterministic extraction canonicalization
|
|
|
|
The current extraction step still performs multiple tasks simultaneously.
|
|
|
|
This was sufficient as a proof of concept but does not reflect the intended long-term architecture.
|
|
|
|
The current protocol builder is also an interim implementation. It concatenates
|
|
extraction results into `meeting_protocol.md` for technical validation. The
|
|
planned architecture separates Canonical Meeting Knowledge from the final Output
|
|
Views documented in `output-views.md`.
|
|
|
|
## Extraction classification contract
|
|
|
|
Decision, Action Item and Open Question extraction shares one evidence-oriented
|
|
classification contract. It is placed after the transcript so it remains the
|
|
final classification instruction in the single multi-category extraction call.
|
|
Decisions require a settled outcome; Action Items require established work;
|
|
Open Questions require a concrete unresolved need. Unsupported candidates must
|
|
not be moved into another category.
|
|
|
|
Action existence and responsibility attribution are separate checks. A valid
|
|
Action Item may have no known owner, while a named owner requires explicit
|
|
assignment, volunteering or acceptance. These are semantic LLM classifications;
|
|
deterministic validation must not guess intent from keywords.
|
|
|
|
### Classification Verifier
|
|
|
|
Before extraction output is normalized for Canonicalizer input, Decision,
|
|
Action Item and Open Question candidates pass through a semantic precision
|
|
gate. Facts and technical details pass through unchanged. Each candidate is
|
|
reviewed independently with its evidence and bounded local chunk context; the
|
|
verifier may only keep or reject the existing candidate. It cannot add or
|
|
rewrite semantic items.
|
|
|
|
Verifier output contains the stable candidate ID, category, `keep|reject`
|
|
verdict, evidence-based reason and `responsibility_supported`. A kept Action
|
|
Item with an unsupported named owner is retained with its responsibility
|
|
cleared. Malformed output fails the extraction verification substage closed,
|
|
after preserving candidate input and raw response. Per-candidate results and an
|
|
aggregate audit remain traceable before Canonicalizer input is written.
|
|
|
|
## Semantic Consolidator failure handling
|
|
|
|
Semantic Consolidator V0 preserves every raw model response before parsing.
|
|
Its normal path accepts parseable grouping JSON and leaves duplicate source-ID
|
|
unknown-source-ID and missing-source-ID correction to the deterministic
|
|
coverage repair. Unknown source IDs are removed without attempting numeric or
|
|
semantic remapping. A group retains its valid IDs and remains present when at
|
|
least one valid ID survives. A group containing only unknown IDs is removed
|
|
after it becomes empty. Existing coverage repair then restores every genuinely
|
|
missing canonical fact as a singleton. Strict validation runs against the
|
|
repaired result and still rejects any unknown ID that survives this process.
|
|
|
|
An invalid or truncated response is not retried by default. One controlled
|
|
retry is allowed only when deterministic inspection finds at least three
|
|
consecutive complete groups with an identical structural signature:
|
|
`canonical_text`, ordered `source_item_ids` and `merge_reason`. The retry keeps
|
|
the same model, temperature, context window and generation limit and adds only
|
|
an instruction not to emit an identical group more than once. Both attempts
|
|
and the detected repetition metadata are preserved. If the retry also fails,
|
|
the stage fails normally; it does not make another LLM call.
|
|
|
|
## Working Protocol V2 renderer contract
|
|
|
|
The renderer deterministically projects consolidated items to the semantic
|
|
fields required for presentation and omits bulky provenance fields from the
|
|
LLM request. Every renderable item remains represented; structurally empty
|
|
items are recorded separately rather than turned into invented prose. The
|
|
compact renderer input is preserved as an artifact.
|
|
|
|
The exact Markdown structure is generated from the same heading constants used
|
|
by the validator and appended after the renderer input so it remains visible
|
|
within the evaluated context. Decisions, action items and open questions carry
|
|
input-derived hidden coverage markers. Strict validation requires every such
|
|
renderable priority item exactly once in its matching section and rejects
|
|
missing, duplicate, wrong-section or invented markers. Facts and technical
|
|
details remain condensable as background.
|
|
|
|
Renderer output budgeting is adaptive to required priority content and prompt
|
|
size, while an explicit `num_predict` override remains authoritative. Raw model
|
|
output is always preserved, and `working_protocol.md` is written only after
|
|
strict structure and coverage validation passes. The renderer does not retry
|
|
automatically.
|
|
|
|
---
|
|
|
|
# Next Milestone
|
|
|
|
The next architecture milestone is the implementation of a deterministic
|
|
canonicalization stage followed by a semantic consolidation stage. These stages
|
|
convert raw chunk extraction JSON into evidence-preserving Canonical Meeting
|
|
Knowledge before any Output View renderer writes a protocol.
|
|
|
|
---
|
|
|
|
# Guiding Principle
|
|
|
|
The Meeting Lab assumes that the greatest improvement in transcript quality will not come from increasingly powerful language models.
|
|
|
|
Instead, quality is expected to emerge from a pipeline that decomposes a complex problem into many small, clearly defined and independently testable processing steps.
|