Author SHA1 Message Date
admin 490d557cb3 Reconcile Meeting Assistant architecture with Meeting Lab 2026-08-23 20:19:47 +02:00
admin e028f1c82c Document development workflow 2026-08-23 19:56:15 +02:00
13 changed files with 466 additions and 658 deletions
+14
View File
@@ -1,5 +1,19 @@
# Changelog
## [Unreleased]
### Changed
- Reconciled the documented MVP with the validated Meeting Lab backend.
- Selected `whisper.cpp` for transcription and made diarization optional.
- Defined the Python API, progress-event and Meeting Context integration
boundaries.
- Set the user-facing Meeting Assistant GUI as the next milestone.
### Added
- ADR 0011 for the Meeting Lab MVP backend architecture.
## [0.1.0] - 2026-07-14
### Added
+103 -153
View File
@@ -2,162 +2,112 @@
## Vision
The project creates a complete pipeline that transforms spoken meetings into structured organizational knowledge.
Meeting Assistant turns recorded meetings into reviewed, user-facing protocols
and, over time, searchable organizational knowledge. It is the application
layer over the reusable Meeting Lab processing backend.
The transcript is the canonical source.
The source audio and canonical transcript are preserved. Generated protocols
are derived, reproducible artifacts and require human review.
AI-generated summaries are reproducible artefacts.
## System Boundary
---
### Meeting Lab
Meeting Lab owns reusable and experimental processing:
- FFmpeg audio preparation
- `whisper.cpp` transcription
- optional `pyannote.audio` diarization
- protocol generation
- pipeline orchestration and progress events
Its reusable integration surface is the Python API:
- `MvpMeetingConfig`
- `MvpRunResult`
- `run_mvp_meeting(...)`
The CLI is a thin adapter, not the Meeting Assistant integration boundary.
### Meeting Assistant
Meeting Assistant owns user interaction and product workflow:
- audio-file selection
- structured Meeting Context editing
- participant management
- optional explicit speaker mapping
- pipeline launch and progress display
- protocol review and editing
- export and presentation of protocol versions
Processing logic must not be duplicated in the application.
## Validated MVP Pipeline
```text
Imported audio
-> prepare as mono, 16 kHz PCM WAV with FFmpeg
-> transcribe with whisper.cpp and large-v3-turbo
-> optionally diarize with pyannote.audio Community-1
-> generate a protocol directly from the full transcript
-> review and edit by a human
```
Mandatory fixed five-minute chunking is not part of the productive MVP path.
Chunking or segmentation used internally by an engine does not shape the
Meeting Assistant storage or domain model.
On the North reference machine, `whisper.cpp` with Vulkan on an AMD RX 9070
transcribed approximately 93.5 minutes in about 100 seconds. Pyannote through
PyTorch/ROCm processed the same duration in approximately 98.6 seconds. These
are validation observations, not performance guarantees or hardware
requirements. CPU execution remains supported and may be substantially slower.
## Meeting Context and Speakers
`MeetingContext` is a structured domain object containing meeting metadata and
participants. It can also contain optional explicit mappings from
`SPEAKER_XX` labels to `participant_id` values.
A speaker mapping is authoritative only when a user explicitly confirms it.
Speakers otherwise remain anonymous. Automatic speaker-name inference is not
allowed. The GUI must create and edit Meeting Context; hand-written YAML is not
a product requirement.
## Progress Contract
Meeting Lab emits stage-based progress events for:
- `preparing`
- `transcription`
- `diarization`
- `protocol_generation`
- `completed`
- `failed`
The GUI consumes this event interface. Percentages are shown only when real
measurable progress is available.
## Protocol Policy
Direct full-transcript protocol generation is the practical MVP direction.
`qwen3.8:27B` has shown strong readability and contextual synthesis;
`qwen3.6:35B-A3B` has shown somewhat more conservative behavior in some areas.
Experimental dual-model and diarization-assisted hard-fact extraction has not
demonstrated reliably better strict attribution accuracy and is not mandatory.
The product should ultimately offer both a detailed contextual protocol and a
shorter participant/distribution version. The short version remains follow-up
work if it is not available for the first GUI milestone.
## Design Principles
Offline-first where practical.
Replaceable AI components.
Vendor independence.
Modular architecture.
Small focused modules.
Simple interfaces.
---
## Core Pipeline
Recorder
↓
Transcription
↓
Speaker Identification
↓
LLM Analysis
↓
Knowledge Extraction
↓
Storage
↓
Export
---
## AI Components
Recorder
Responsible only for recording.
No AI.
---
Transcription
Responsible only for speech-to-text.
No summarization.
---
Speaker Identification
Responsible only for identifying speakers.
Must not modify transcript text.
---
LLM Analysis
Responsible for:
- Summary
- Decisions
- Action Items
- Risks
- Questions
- Knowledge extraction
---
Storage
Stores:
Audio
Transcript
Metadata
AI Results
Knowledge Objects
---
## Architectural Rule
Every module has exactly one responsibility.
---
## Transcript Policy
Never modify the original transcript.
Corrections must create a new derived version.
---
## AI Output Policy
Every AI output should be reproducible.
Prompt version should be stored.
Model should be stored.
Timestamp should be stored.
---
## Long-Term Goal
Every meeting becomes searchable organizational knowledge.
No information should be lost after the meeting.
---
## Engineering Principles
The project follows an architecture-first development approach.
Before implementing a feature:
- define the domain model
- define module boundaries
- document architectural decisions
Implementation is intentionally delayed until the architecture is considered sufficiently stable.
- offline-first where practical
- explicit application/backend boundaries
- replaceable AI components
- CPU-compatible operation with optional acceleration
- immutable source artifacts and traceable derived artifacts
- no automatic identity claims
- human review of generated protocols
- small modules and simple interfaces
+47 -108
View File
@@ -1,126 +1,65 @@
# Meeting Knowledge Assistant
# Meeting Assistant
## Vision
Meeting Assistant is the user-facing application for turning recorded meetings
into reviewed, distributable protocols. It uses Meeting Lab as its reusable
local processing backend.
The Meeting Knowledge Assistant (MKA) is an AI-assisted desktop application that records meetings, creates high-quality transcripts and transforms them into structured knowledge.
## Current MVP Direction
Unlike traditional meeting assistants, MKA does not aim to replace human interaction during meetings.
Its goal is to create an accurate, searchable and long-term knowledge base from spoken communication.
The first product milestone is a desktop GUI that lets a user:
The project follows an "AI-first" architecture:
- select an existing audio recording
- create and edit structured meeting metadata and participants
- optionally map anonymous `SPEAKER_XX` labels to known participants
- start the Meeting Lab processing pipeline
- follow stage-based progress
- review and edit the generated protocol
- export the reviewed result
Audio
→ Transcription
→ Speaker Identification
→ AI Analysis
→ Structured Knowledge
→ Search
The MVP does not include integrated recording. OBS Studio remains a recommended
reference recorder, but imported audio is not tied to OBS-specific behavior.
---
## Core Features
- Import of existing meeting recordings
- OBS Studio as the recommended reference recorder
- High-quality local transcription
- Speaker diarization
- Editable working transcripts
- AI-generated meeting minutes
- Executive summaries
- Action items
- Decision tracking
- Knowledge extraction
- Local file-based storage
- Export to Markdown, PDF and DOCX
---
## Recording Workflow
The MVP does not include its own audio recorder.
OBS Studio is the recommended reference recorder for online meetings. It can
capture system audio and microphone audio without joining the meeting as a bot.
Initial workflow:
## Processing Pipeline
```text
Audio file
-> FFmpeg preparation (mono, 16 kHz PCM WAV)
-> whisper.cpp transcription (large-v3-turbo)
-> optional pyannote.audio Community-1 diarization
-> direct full-transcript protocol generation
-> human review and export
```
Online Meeting
Meeting Assistant calls Meeting Lab's reusable Python API
(`MvpMeetingConfig`, `MvpRunResult` and `run_mvp_meeting(...)`). It does not run
the Meeting Lab CLI as a subprocess. Fixed five-minute audio chunking is not a
required part of this workflow.
↓
GPU acceleration is supported where available, but CPU execution remains a
compatibility path. Diarization is optional and produces anonymous speaker
labels. A label identifies a participant only when the user explicitly
confirms the mapping; automatic speaker-name inference is not allowed.
OBS Studio
## Product Outputs
↓
The product direction includes:
WAV Recording
- a detailed, contextual protocol for review and continued work
- a shorter version suitable for participants and distribution
↓
The detailed direct-protocol flow is the practical MVP direction. The shorter
version may be implemented after the initial GUI. Human review remains part of
the workflow for all generated protocols.
Meeting Assistant Import
## Project Boundary
↓
Meeting Lab owns reusable audio preparation, transcription, diarization,
orchestration and protocol-generation logic. Meeting Assistant owns the GUI,
context and participant editing, explicit speaker mapping, progress display,
protocol editing and export.
Transcription
The code currently contains the initial Meeting Assistant project and domain
foundation. The GUI has not yet been implemented.
↓
Speaker Diarization
↓
Analysis and Export
---
## Design Goals
- High transcription quality
- Modular architecture
- Offline-first whenever practical
- Replaceable AI components
- Vendor independence
- Long-term maintainability
---
## Philosophy
The transcript is the primary asset.
Everything else (summaries, reports, action items, knowledge extraction)
can always be regenerated using better AI models in the future.
Therefore:
Audio
→ Transcript
is considered immutable.
AI output is considered reproducible.
---
## Planned Architecture
Recorder
↓
Transcription
↓
Speaker Identification
↓
LLM Processing
↓
Knowledge Database
↓
Export
---
## Current Status
Project planning.
No implementation has started yet.
See [Architecture](docs/architecture.md), [Project Knowledge](PROJECT_KNOWLEDGE.md),
[Roadmap](ROADMAP.md) and [ADR 0011](docs/adr/0011-use-meeting-lab-mvp-backend.md).
+38 -100
View File
@@ -1,112 +1,50 @@
# Roadmap
## Phase 1: Minimum Viable Processing Pipeline
## Next Milestone: Meeting Assistant MVP GUI
- Import existing WAV recordings
- Validate recording metadata
- Create meeting directory
- Create canonical transcript
- Save transcript as JSON
- Export transcript as Markdown
- Archive confirmed recordings as FLAC
Build the user-facing application over the validated Meeting Lab Python API.
---
- select an existing audio file
- create and edit Meeting Context through structured fields
- add and manage participants
- optionally map `SPEAKER_XX` labels to participants with explicit confirmation
- configure and start `run_mvp_meeting(...)`
- display stage-based status and real progress when available
- handle completed and failed runs clearly
- display and edit the detailed generated protocol
- export the reviewed protocol
## Phase 2
The milestone must keep diarization, GPU acceleration and fixed-duration audio
chunking optional. It must not launch the Meeting Lab CLI as a subprocess or
duplicate Meeting Lab processing logic.
Speaker diarization
## Following Milestone: Distribution Protocol
- Speaker detection
- Speaker naming
- Timeline view
- derive or generate a shorter participant-facing version
- allow review and editing before distribution
- export both detailed and short versions
- retain provenance linking each version to its transcript, configuration,
prompt and model where practical
---
## Later Product Work
## Phase 3
- transcript viewing and correction workflows
- recording and artifact lifecycle management
- search, tags, projects and meeting history
- richer Markdown, PDF and DOCX export
- optional SQLite indexing and full-text search
- integrated recording controls
- calendar and conferencing integrations
- semantic search and organizational knowledge features
Meeting intelligence
- Executive Summary
- Action Items
- Decisions
- Risks
- Open Questions
---
## Phase 4
Knowledge system
- SQLite database
- Full text search
- Semantic search
- Tags
- Projects
- Participants
---
## Phase 5
Desktop application
- Recording UI
- Transcript viewer
- Search
- Export
---
## Phase 6
Integrations
- Outlook
- Teams
- Google Calendar
- Local LLM
- OpenAI
- Ollama
---
## Future Phase: Integrated Recording
- Capture microphone audio
- Capture system audio
- Support separate audio channels
- Provide recording controls in the desktop application
- Replace or complement the external OBS workflow
- Live transcription
- Real-time summaries
- Company glossary
- Custom vocabulary
- Automatic project detection
- Meeting templates
- Voice identification
- Audio cleanup
## Research and Optional Extensions
- live transcription and real-time summaries
- company glossary and custom vocabulary
- voice identification with explicit consent and confirmation
- audio cleanup
- OCR for shared screens
- Timeline with bookmarks
- RAG knowledge integration
- Multi-language meetings
- Translation
- Automatic follow-up generation
- RAG and knowledge-graph integration
- multi-language meetings and translation
- dual-model or evidence pipelines only if later validation demonstrates a
material reliability benefit
+6 -1
View File
@@ -2,7 +2,7 @@
## Status
Accepted
Accepted, amended by [ADR 0011](0011-use-meeting-lab-mvp-backend.md)
## Context
@@ -30,3 +30,8 @@ The transcript is treated as the canonical source. AI-generated summaries and kn
- AI outputs must be reproducible where practical.
- Modules should remain replaceable.
- The project should avoid unnecessary vendor lock-in.
ADR 0011 narrows the practical MVP to a user-facing application over the
reusable Meeting Lab backend. Direct protocol generation is the validated MVP
path; a separate knowledge-extraction stage remains a possible future
extension rather than an MVP requirement.
+3 -47
View File
@@ -2,7 +2,7 @@
## Status
Accepted
Superseded by [ADR 0011](0011-use-meeting-lab-mvp-backend.md)
## Context
@@ -10,60 +10,16 @@ The project requires high-quality speech-to-text transcription for German and En
The transcription component should support local execution where practical and should remain replaceable in the future. The project should not depend on a meeting bot or on a proprietary meeting platform.
Practical testing showed that long continuous recordings may suffer from error propagation when the decoder continuously conditions on previously generated text.
Independent transcription of smaller audio chunks produced significantly more stable results.
## Decision
The project will use `faster-whisper` as the initial transcription engine.
The preferred first model is `large-v3-turbo`, with `large-v3` as a quality-oriented fallback if required.
The transcription pipeline shall process recordings as independent audio chunks rather than as a single continuous recording.
The default processing strategy is:
```text
Recording (WAV)
│
▼
Split into 5-minute chunks
│
▼
Independent Whisper transcription
(condition_on_previous_text = False)
│
▼
JSON transcript per chunk
│
▼
Merge chunk transcripts
│
▼
Canonical transcript
```
Chunk duration shall be configurable, with **5 minutes** as the default value.
Each chunk shall be processed independently.
The decoder shall not use text generated from previous chunks as context.
Each chunk shall produce its own intermediate JSON transcript before merging.
The merge step shall preserve timestamps and chunk ordering.
The preferred first model is `large-v3-turbo`, with `large-v3` as quality-oriented fallback if required.
## Consequences
- Long recordings become more robust.
- Error propagation across chunk boundaries is prevented.
- Failed chunks can be retranscribed independently.
- Parallel processing of multiple chunks becomes possible.
- Intermediate JSON files simplify debugging and future reprocessing.
- Chunk size can be tuned later without changing the overall architecture.
- Transcription can run locally on suitable hardware.
- NVIDIA CUDA acceleration can be used later on an office AI PC.
- The transcription module must hide the concrete engine behind an internal interface.
- Model name, language setting, timestamp mode, engine version, chunk duration and relevant decoding parameters shall be stored with every transcript.
- Model name, language setting, timestamp and engine version should be stored with every transcript.
- The decision can be revisited if another engine provides clearly better quality, speed or deployment characteristics.
+9 -1
View File
@@ -2,7 +2,7 @@
## Status
Accepted
Accepted, amended by [ADR 0011](0011-use-meeting-lab-mvp-backend.md)
## Context
@@ -23,3 +23,11 @@ Speaker diarization will be treated as a separate pipeline step after transcript
- Human correction of speaker names should be supported later.
- The diarization module must hide the concrete engine behind an internal interface.
- Model name, engine version and timestamp should be stored with every diarization result.
## Amendment
ADR 0011 makes diarization optional for the MVP and selects `pyannote.audio`
Community-1 as the current preferred backend. CPU execution remains supported,
with GPU acceleration used when available. Speaker labels remain anonymous
unless a user explicitly confirms a `SPEAKER_XX` to participant mapping in the
Meeting Context. Automatic speaker-name inference is not allowed.
+3 -3
View File
@@ -2,7 +2,7 @@
## Status
Accepted
Accepted, amended by [ADR 0011](0011-use-meeting-lab-mvp-backend.md)
## Context
@@ -18,8 +18,8 @@ removing temporary source files without risking data loss.
Recordings shall initially be captured in a processing-friendly lossless
format, normally WAV.
After transcription and diarization have completed, the recording may be
converted to an archive format.
After transcription and, when enabled, diarization have completed, the
recording may be converted to an archive format.
The preferred archive formats are:
+1 -6
View File
@@ -54,12 +54,7 @@ meeting/
├── transcript/
│ ├── canonical.json
│ ├── working_v1.json
│
├── chunks/
│ ├── chunk_000.json
│ ├── chunk_001.json
│ ├── chunk_002.json
│ ├── ...
│ └── working_v2.json
│
├── artifacts/
│ ├── summary.md
@@ -2,7 +2,7 @@
## Status
Accepted
Accepted, amended by [ADR 0011](0011-use-meeting-lab-mvp-backend.md)
## Context
@@ -13,8 +13,9 @@ OBS Studio already provides stable, configurable and cross-platform audio
capture. Reimplementing audio capture during the initial project phase would
add substantial complexity without directly improving transcription quality.
The primary value of the Meeting Assistant lies in transcription, speaker
diarization, analysis, structured knowledge extraction and export.
The primary value of the Meeting Assistant lies in transcription, optional
speaker diarization, protocol generation, review and export. ADR 0011 defines
the validated processing boundary and current MVP protocol path.
## Decision
@@ -0,0 +1,94 @@
# ADR 0011: Use Meeting Lab as the MVP Processing Backend
## Status
Accepted
Supersedes ADR 0002 and amends ADRs 0001 and 0003.
## Context
Meeting Lab experiments have validated a practical local pipeline for turning
a recorded meeting into an editable protocol. Earlier Meeting Assistant plans
selected `faster-whisper`, treated diarization as required and described a
separate knowledge-extraction stage as part of the primary pipeline. Work on an
old feature branch also proposed mandatory fixed five-minute audio chunks and
chunk-oriented storage.
The validated backend now has a reusable Python API, optional diarization and a
direct full-transcript protocol path. The product boundary must keep backend
processing in Meeting Lab and user interaction in Meeting Assistant.
## Decision
Meeting Lab is the reusable processing backend and experimental engine.
Meeting Assistant is the user-facing application and calls Meeting Lab through
its Python API rather than launching its CLI as a subprocess.
The MVP processing flow is:
```text
Source audio
-> FFmpeg preparation (mono, 16 kHz PCM WAV)
-> whisper.cpp transcription (large-v3-turbo)
-> optional pyannote.audio Community-1 diarization
-> direct full-transcript protocol generation
-> human review and editing
```
The productive path does not require fixed five-minute audio chunks. Internal
streaming or segmentation remains an implementation detail of Meeting Lab and
must not define Meeting Assistant's domain or storage model.
Meeting Assistant integrates with these public Meeting Lab interfaces:
- `MvpMeetingConfig`
- `MvpRunResult`
- `run_mvp_meeting(...)`
- stage-based progress events: `preparing`, `transcription`, `diarization`,
`protocol_generation`, `completed` and `failed`
Progress percentages are displayed only when the backend reports real
measurable progress.
`MeetingContext` is the structured input containing meeting metadata and
participants. It may contain explicitly confirmed `SPEAKER_XX -> participant_id`
mappings. Such mappings are authoritative only when confirmed by the user.
The system must not infer speaker names automatically.
The preferred local engines are:
- `whisper.cpp` with `large-v3-turbo` for transcription
- `pyannote.audio` Community-1 for optional diarization
GPU acceleration is optional. Vulkan transcription and PyTorch/ROCm
diarization are validated acceleration paths on North's AMD RX 9070. CPU
execution remains the compatibility fallback.
The direct full-transcript protocol path is the practical MVP direction.
`qwen3.8:27B` has produced strong readability and contextual synthesis, while
`qwen3.6:35B-A3B` has sometimes behaved more conservatively. Dual-model and
diarization-assisted hard-fact experiments have not established sufficiently
reliable attribution gains to justify mandatory MVP complexity. Model choice
remains backend configuration, and human review is required.
Meeting Assistant owns the GUI, structured context editor, participant and
speaker-mapping controls, progress presentation, protocol editor and export.
Meeting Lab owns audio preparation, transcription, diarization, orchestration
and protocol-generation logic.
The product should ultimately provide a detailed contextual protocol and a
shorter participant/distribution version. The shorter version may follow the
first GUI milestone.
## Consequences
- ADR 0002's `faster-whisper` selection is no longer current.
- Fixed-duration chunking and chunk-centric storage are not MVP requirements.
- Diarization and GPU acceleration remain optional.
- The Meeting Lab CLI remains useful for command-line operation but is not the
application integration boundary.
- Meeting Context is created through the future GUI; users are not expected to
hand-write YAML.
- Backend experiments can evolve without moving processing logic into the GUI.
- Protocols remain AI-assisted outputs that require human review.
+125 -233
View File
@@ -2,254 +2,146 @@
## Purpose
The Meeting Knowledge Assistant transforms recorded meetings into structured organizational knowledge through a modular processing pipeline.
Meeting Assistant is the user-facing application for preparing meeting
context, running the Meeting Lab pipeline and reviewing its results. Meeting
Lab is the reusable processing backend and experimental engine.
---
# Design Principles
- Architecture first
- Offline first where practical
- Immutable source data
- Replaceable AI components
- Small, focused modules
- Explicit interfaces
- Reproducible AI outputs
- External recording and internal processing are separate responsibilities.
- The processing pipeline must not depend on OBS-specific metadata or behavior.
- Imported recordings must be treated like recordings from any other supported source.
---
## High-Level Pipeline
External Recorder
↓
Audio Import
↓
Recording Validation
↓
Meeting Storage
↓
Transcription
↓
Speaker Diarization
↓
Working Transcript
↓
LLM Analysis
↓
Knowledge Extraction
↓
Export
---
## Core Pipeline
External Recording
↓
Audio Import and Validation
↓
Transcription
↓
Speaker Diarization
↓
Working Transcript
↓
LLM Analysis
↓
Knowledge Extraction
↓
Storage and Export
# Domain Model
Meeting
│
├── Recording
├── Transcript
│ ├── Canonical
│ └── Working
├── Speakers
├── Participants
├── AI Artifacts
├── Knowledge Objects
├── Attachments
└── Exports
---
# Module Responsibilities
## Input and Ingest
Responsible for importing existing recordings into the application.
Initial reference source:
- OBS Studio
Initial supported formats:
- WAV
- FLAC
Responsibilities:
- validate the input file
- collect technical metadata
- calculate checksums
- copy or move the recording into the meeting directory
- create the initial Recording domain object
The ingest component must not perform transcription or modify the audio content.
## Recorder
The recorder module is reserved for a future integrated recording implementation.
It is not required for the MVP.
The initial application workflow uses externally created recordings, with OBS
Studio as the recommended reference recorder.
---
## Transcription
Responsible only for speech-to-text conversion.
Input:
Recording
Output:
Canonical Transcript
---
## Diarization
Responsible only for speaker identification.
Input:
Recording + Canonical Transcript
Output:
Working Transcript
---
## LLM
Responsible for semantic analysis.
Input:
Transcript
Output:
AI Artifacts
---
## Knowledge Extraction
Responsible for creating structured knowledge.
Input:
AI Artifacts
Output:
Knowledge Objects
---
## Export
Responsible for creating user-facing documents.
Input:
Knowledge Objects
Output:
Markdown
PDF
DOCX
---
## Artifact Lifecycle
## System Boundary
```text
Imported Recording
↓
Validated Source Recording
↓
Canonical Transcript
↓
Working Transcript
↓
AI Artifacts
↓
Knowledge Objects
↓
Exports
Meeting Assistant Meeting Lab
----------------- -----------
Audio selection ------> Audio preparation
Meeting Context editor ------> Transcription
Participant management ------> Optional diarization
Explicit speaker mapping ------> Protocol generation
Progress presentation <------ Progress events
Protocol editor and export <------ Run result and artifacts
```
---
Meeting Assistant calls the Meeting Lab Python API directly. The Meeting Lab
CLI is a thin adapter over that API and must not be launched as an application
subprocess.
## Initial Recording Strategy
## MVP Processing Flow
The MVP does not implement platform-specific audio capture.
```text
Source audio
-> FFmpeg normalization/preparation
-> mono, 16 kHz PCM WAV
-> whisper.cpp transcription with large-v3-turbo
-> optional pyannote.audio Community-1 diarization
-> direct full-transcript protocol generation
-> human review and editing
-> export
```
OBS Studio is the recommended reference recorder for online meetings.
The MVP does not require fixed five-minute audio chunks, semantic chunking, a
separate evidence-extraction pipeline or multiple LLMs. Any internal
segmentation remains a Meeting Lab implementation detail.
The application initially processes existing WAV or FLAC recordings. Integrated
## Backend Integration
recording remains a future extension and must not be required by transcription,
The reusable Meeting Lab interface consists of:
diarization or analysis modules.
- `MvpMeetingConfig`, the run configuration
- `run_mvp_meeting(...)`, the orchestration entry point
- `MvpRunResult`, the completed run result
- stage-based progress events
---
The known stages are:
# Storage Strategy
```text
preparing
transcription
diarization
protocol_generation
completed
failed
```
The domain model is independent of the storage backend.
The GUI shows the current stage. It shows a percentage only when the event
contains real measurable progress; stage changes must not be presented as
invented percentages.
Possible implementations:
## Meeting Context
- File System
- SQLite
- PostgreSQL
- Cloud Storage
`MeetingContext` is structured domain input rather than an informal prompt or a
file users must author manually. It contains meeting metadata and participants
and can include optional mappings:
---
```text
SPEAKER_XX -> participant_id
```
# Future Extensions
Mappings are authoritative only after explicit user confirmation. Diarization
labels otherwise remain anonymous, and the application must not infer speaker
names automatically. The Meeting Assistant GUI owns creation and editing of
this context and its mappings.
- Live transcription
- Video processing
- OCR
- Semantic search
- Knowledge graph
- Company glossary
- Multi-language meetings
- Local LLM support
## Processing Components
### Audio Preparation
Meeting Lab uses FFmpeg to prepare a consistent local-processing input. The
current practical target is mono, 16 kHz PCM WAV. The imported source remains a
separate source artifact.
### Transcription
`whisper.cpp` with `large-v3-turbo` is the preferred local backend. Vulkan is a
validated acceleration path on North's AMD RX 9070, while CPU execution remains
a compatibility fallback. Engine and model details should be retained with the
result for traceability.
### Diarization
`pyannote.audio` Community-1 is preferred when diarization is enabled. It can
run on CPU and may use GPU acceleration such as PyTorch/ROCm where available.
Diarization is optional and its anonymous labels do not modify the canonical
transcript or assert participant identity.
### Protocol Generation
The direct full-transcript path is the practical MVP. Current model experiments
favor `qwen3.8:27B` for readability and contextual synthesis, while
`qwen3.6:35B-A3B` has shown more conservative behavior in some areas. Neither a
specific dual-model arrangement nor an evidence pipeline is an application
requirement. Generated protocols require human review.
## Application Responsibilities
The first GUI milestone provides:
- audio-file selection
- structured Meeting Context and participant editing
- optional explicit speaker mapping
- pipeline start and configuration
- progress and failure presentation
- detailed protocol display and editing
- export of the reviewed result
The product direction also includes a shorter participant/distribution
protocol. It may be delivered after the first GUI milestone.
## Artifact and Storage Principles
The existing Meeting domain remains the application-level container for source
recordings, canonical and working transcripts, context, generated artifacts and
exports. Storage remains independent of processing segmentation. In particular,
chunks are not primary Meeting Assistant domain objects or the required unit of
MVP persistence.
Source audio and the canonical transcription output are preserved. Corrections
and reviewed protocols are derived versions. Generated artifacts should retain
their input version, backend/model configuration, prompt version and timestamp
where practical.
## Future Extensions
- shorter distribution protocols
- integrated recording
- searchable meeting history and optional database indexes
- live transcription
- OCR and video processing
- semantic search and organizational knowledge features
+17 -1
View File
@@ -2,7 +2,23 @@ Transcript
Original speech-to-text output.
Diarization
Assignment of transcript segments to speakers.
Assignment of transcript segments to anonymous speaker labels. It does not by
itself identify participants.
Meeting Context
Structured meeting metadata, participants and optional explicitly confirmed
speaker-to-participant mappings supplied to the processing pipeline.
Speaker Mapping
An explicitly confirmed association from a diarization label such as
`SPEAKER_00` to a participant ID. It must not be inferred automatically.
Detailed Protocol
Contextual protocol generated from the full transcript for human review and
continued work.
Distribution Protocol
Shorter, participant-facing version of a reviewed meeting protocol.
Knowledge Object
Structured information extracted from meetings.