Author SHA1 Message Date
admin 31801729ed Refine chunked transcription strategy 2026-07-20 12:21:04 +02:00
29 changed files with 642 additions and 2060 deletions
-17
View File
@@ -1,17 +0,0 @@
MKA_WHISPER_MODEL=/path/to/ggml-large-v3-turbo.bin
MKA_WHISPER_EXECUTABLE=whisper-cli
MKA_FFMPEG_EXECUTABLE=ffmpeg
MKA_PROTOCOL_MODEL=qwen3.8:27B
MKA_OLLAMA_ENDPOINT=http://127.0.0.1:11434
MKA_DATA_ROOT=data/meetings
MKA_WHISPER_THREADS=auto
MKA_DIARIZATION_MODE=auto
MKA_DIARIZATION_RUNTIME=native
# MKA_DIARIZATION_CONTAINER_IMAGE=
# MKA_DIARIZATION_CONTAINER_ARGS=[]
# Example AMD ROCm container configuration for a compatible workstation:
# MKA_DIARIZATION_MODE=gpu
# MKA_DIARIZATION_RUNTIME=container
# MKA_DIARIZATION_CONTAINER_IMAGE=rocm/pytorch:rocm7.2.1_ubuntu24.04_py3.12_pytorch_release_2.9.1
# MKA_DIARIZATION_CONTAINER_ARGS='["--device=/dev/kfd","--device=/dev/dri","--group-add","video"]'
-2
View File
@@ -34,7 +34,6 @@ dist/
*.sqlite3
# Meeting data
data/meetings/*
data/recordings/*
data/transcripts/*
data/database/*
@@ -43,4 +42,3 @@ data/database/*
!data/recordings/.gitkeep
!data/transcripts/.gitkeep
!data/database/.gitkeep
!data/meetings/.gitkeep
-21
View File
@@ -1,26 +1,5 @@
# Changelog
## [Unreleased]
### Added
- First Streamlit MVP for audio upload, Meeting Context entry, participant
management, processing progress and protocol editing.
- Application service and Meeting Lab adapter with environment-based runtime
configuration.
- Per-meeting upload and edited-protocol persistence.
- Unit tests for context construction, configuration translation, progress,
failure and result handling.
- ADR 0011 for the Meeting Lab MVP backend architecture.
### Changed
- Reconciled the documented MVP with the validated Meeting Lab backend.
- Selected `whisper.cpp` for transcription and made diarization optional.
- Defined the Python API, progress-event and Meeting Context integration
boundaries.
- Set the user-facing Meeting Assistant GUI as the next milestone.
## [0.1.0] - 2026-07-14
### Added
+153 -123
View File
@@ -2,132 +2,162 @@
## Vision
Meeting Assistant turns recorded meetings into reviewed, user-facing protocols
and, over time, searchable organizational knowledge. It is the application
layer over the reusable Meeting Lab processing backend.
The project creates a complete pipeline that transforms spoken meetings into structured organizational knowledge.
The source audio and canonical transcript are preserved. Generated protocols
are derived, reproducible artifacts and require human review.
The transcript is the canonical source.
## System Boundary
AI-generated summaries are reproducible artefacts.
### Meeting Lab
Meeting Lab owns reusable and experimental processing:
- FFmpeg audio preparation, with optional normalization enabled by default
- `whisper.cpp` transcription
- optional `pyannote.audio` diarization
- protocol generation
- pipeline orchestration and progress events
Its reusable integration surface is the Python API:
- `MvpMeetingConfig`
- `MvpRunResult`
- `run_mvp_meeting(...)`
The CLI is a thin adapter, not the Meeting Assistant integration boundary.
### Meeting Assistant
Meeting Assistant owns user interaction and product workflow:
- audio-file selection
- structured Meeting Context editing
- participant management
- versioned YAML import/export for reusable People lists, using replace semantics
- optional explicit speaker mapping
- pipeline launch and progress display
- protocol review and editing
- export and presentation of protocol versions
Processing logic must not be duplicated in the application.
Diarization runtime, container image and ordered container arguments are
machine configuration. `MKA_DIARIZATION_CONTAINER_ARGS` is a JSON array of
strings forwarded unchanged to Meeting Lab; these details are not normal UI
controls.
People-list YAML is a Meeting Assistant application concern and contains only
stable IDs, display names, roles, organizations and attendance states. It does
not contain meeting metadata or processing settings. Named team or meeting
templates may build on this later, but are not part of the current mechanism.
## Validated MVP Pipeline
```text
Imported audio
-> always prepare as mono, 16 kHz PCM WAV with FFmpeg
(optionally normalize loudness; default on)
-> transcribe with whisper.cpp and large-v3-turbo
-> optionally diarize with pyannote.audio Community-1
-> generate a protocol directly from the full transcript
-> review and edit by a human
```
Mandatory fixed five-minute chunking is not part of the productive MVP path.
Chunking or segmentation used internally by an engine does not shape the
Meeting Assistant storage or domain model.
On the North reference machine, `whisper.cpp` with Vulkan on an AMD RX 9070
transcribed approximately 93.5 minutes in about 100 seconds. Pyannote through
PyTorch/ROCm processed the same duration in approximately 98.6 seconds. These
are validation observations, not performance guarantees or hardware
requirements. CPU execution remains supported and may be substantially slower.
## Meeting Context and Speakers
`MeetingContext` is a structured domain object containing meeting metadata,
present participants, and people who were mentioned without attending. The GUI
records exactly `present` or `mentioned_only`; missing legacy participant status
defaults to `present` in Meeting Lab. It can also contain optional explicit mappings from
`SPEAKER_XX` labels to `participant_id` values.
A speaker mapping is authoritative only when a user explicitly confirms it and
may reference only a present participant.
Speakers otherwise remain anonymous. Automatic speaker-name inference is not
allowed. The GUI must create and edit Meeting Context; hand-written YAML is not
a product requirement.
## Progress Contract
Meeting Lab emits stage-based progress events for:
- `preparing`
- `transcription`
- `diarization`
- `protocol_generation`
- `completed`
- `failed`
The GUI consumes this event interface. Percentages are shown only when real
measurable progress is available.
## Protocol Policy
Direct full-transcript protocol generation is the practical MVP direction.
`qwen3.8:27B` has shown strong readability and contextual synthesis;
`qwen3.6:35B-A3B` has shown somewhat more conservative behavior in some areas.
Experimental dual-model and diarization-assisted hard-fact extraction has not
demonstrated reliably better strict attribution accuracy and is not mandatory.
The product should ultimately offer both a detailed contextual protocol and a
shorter participant/distribution version. The short version remains follow-up
work if it is not available for the first GUI milestone.
Confirmed meeting-specific corrections should ultimately be reusable by both
protocol views. The detailed correction and selective-regeneration decision is
recorded in [ADR 0012](docs/adr/0012-post-run-corrections.md); it is later
product work, not a current MVP requirement.
---
## Design Principles
- offline-first where practical
- explicit application/backend boundaries
- replaceable AI components
- CPU-compatible operation with optional acceleration
- immutable source artifacts and traceable derived artifacts
- no automatic identity claims
- human review of generated protocols
- small modules and simple interfaces
Offline-first where practical.
Replaceable AI components.
Vendor independence.
Modular architecture.
Small focused modules.
Simple interfaces.
---
## Core Pipeline
Recorder
↓
Transcription
↓
Speaker Identification
↓
LLM Analysis
↓
Knowledge Extraction
↓
Storage
↓
Export
---
## AI Components
Recorder
Responsible only for recording.
No AI.
---
Transcription
Responsible only for speech-to-text.
No summarization.
---
Speaker Identification
Responsible only for identifying speakers.
Must not modify transcript text.
---
LLM Analysis
Responsible for:
- Summary
- Decisions
- Action Items
- Risks
- Questions
- Knowledge extraction
---
Storage
Stores:
Audio
Transcript
Metadata
AI Results
Knowledge Objects
---
## Architectural Rule
Every module has exactly one responsibility.
---
## Transcript Policy
Never modify the original transcript.
Corrections must create a new derived version.
---
## AI Output Policy
Every AI output should be reproducible.
Prompt version should be stored.
Model should be stored.
Timestamp should be stored.
---
## Long-Term Goal
Every meeting becomes searchable organizational knowledge.
No information should be lost after the meeting.
---
## Engineering Principles
The project follows an architecture-first development approach.
Before implementing a feature:
- define the domain model
- define module boundaries
- document architectural decisions
Implementation is intentionally delayed until the architecture is considered sufficiently stable.
+91 -114
View File
@@ -1,149 +1,126 @@
# Meeting Assistant
# Meeting Knowledge Assistant
Meeting Assistant is the user-facing application for turning recorded meetings
into reviewed, distributable protocols. It uses Meeting Lab as its reusable
local processing backend.
## Vision
## Current MVP Direction
The Meeting Knowledge Assistant (MKA) is an AI-assisted desktop application that records meetings, creates high-quality transcripts and transforms them into structured knowledge.
The first product milestone is a desktop GUI that lets a user:
Unlike traditional meeting assistants, MKA does not aim to replace human interaction during meetings.
Its goal is to create an accurate, searchable and long-term knowledge base from spoken communication.
- select an existing WAV, FLAC, or M4A recording
- create and edit structured meeting metadata and relevant people, including whether
they were present or only mentioned
- import and export the reusable People list as versioned YAML
- optionally map anonymous `SPEAKER_XX` labels to known participants
- start the Meeting Lab processing pipeline
- follow stage-based progress
- review and edit the generated protocol
- export the reviewed result
The project follows an "AI-first" architecture:
The MVP does not include integrated recording. OBS Studio remains a recommended
reference recorder, but imported audio is not tied to OBS-specific behavior.
Audio
→ Transcription
→ Speaker Identification
→ AI Analysis
→ Structured Knowledge
→ Search
## Processing Pipeline
---
## Core Features
- Import of existing meeting recordings
- OBS Studio as the recommended reference recorder
- High-quality local transcription
- Speaker diarization
- Editable working transcripts
- AI-generated meeting minutes
- Executive summaries
- Action items
- Decision tracking
- Knowledge extraction
- Local file-based storage
- Export to Markdown, PDF and DOCX
---
## Recording Workflow
The MVP does not include its own audio recorder.
OBS Studio is the recommended reference recorder for online meetings. It can
capture system audio and microphone audio without joining the meeting as a bot.
Initial workflow:
```text
Audio file
-> FFmpeg preparation (mono, 16 kHz PCM WAV; normalization optional)
-> whisper.cpp transcription (large-v3-turbo)
-> optional pyannote.audio Community-1 diarization
-> direct full-transcript protocol generation
-> human review and export
```
Meeting Assistant calls Meeting Lab's reusable Python API
(`MvpMeetingConfig`, `MvpRunResult` and `run_mvp_meeting(...)`). It does not run
the Meeting Lab CLI as a subprocess. Fixed five-minute audio chunking is not a
required part of this workflow.
Online Meeting
GPU acceleration is supported where available, but CPU execution remains a
compatibility path. Diarization is optional and produces anonymous speaker
labels. A label identifies a participant only when the user explicitly
confirms the mapping; automatic speaker-name inference is not allowed.
↓
## Product Outputs
OBS Studio
The product direction includes:
↓
- a detailed, contextual protocol for review and continued work
- a shorter version suitable for participants and distribution
WAV Recording
The detailed direct-protocol flow is the practical MVP direction. The shorter
version may be implemented after the initial GUI. Human review remains part of
the workflow for all generated protocols.
↓
## Project Boundary
Meeting Assistant Import
Meeting Lab owns reusable audio preparation, transcription, diarization,
orchestration and protocol-generation logic. Meeting Assistant owns the GUI,
context and participant editing, explicit speaker mapping, progress display,
protocol editing and export.
↓
Audio normalization is enabled by default and can be disabled in the processing
options. This controls loudness normalization only: Meeting Lab still prepares
every WAV, FLAC or M4A source as canonical audio before transcription.
Transcription
The People section can export its current entries to a UTF-8 `people.yaml` file
and replace them from a previous `.yaml` or `.yml` export. Stable person IDs,
names, roles, organizations and attendance states are retained. This is a small
reuse mechanism, not a server-side participant library or named meeting-template
system.
↓
## Run the Streamlit MVP
Speaker Diarization
The development setup expects `meeting-assistant` and `meeting-lab` to be
sibling repositories. Meeting Lab currently imports its API through the
`src.meeting_lab` package path, so its repository root must be supplied on
`PYTHONPATH`. Meeting Assistant does not modify `sys.path` at runtime.
↓
Create a virtual environment and install Meeting Assistant, including
Streamlit:
Analysis and Export
```bash
python3 -m venv .venv
.venv/bin/pip install -e '.[dev]'
.venv/bin/pip install -e ../meeting-lab
```
---
Configure at least the local whisper.cpp model. All supported values are shown
in `.env.example`; export them in the shell because the application does not
load `.env` files implicitly:
## Design Goals
```bash
export MKA_WHISPER_MODEL=/path/to/ggml-large-v3-turbo.bin
export MKA_WHISPER_EXECUTABLE=whisper-cli
export MKA_FFMPEG_EXECUTABLE=ffmpeg
export MKA_PROTOCOL_MODEL=qwen3.8:27B
```
- High transcription quality
- Modular architecture
- Offline-first whenever practical
- Replaceable AI components
- Vendor independence
- Long-term maintainability
Optional machine-specific settings include:
---
- `MKA_DATA_ROOT` (default: `data/meetings`)
- `MKA_OLLAMA_ENDPOINT` (default: `http://127.0.0.1:11434`)
- `MKA_WHISPER_THREADS` (default: `auto`)
- `MKA_FFMPEG_EXECUTABLE` (default: `ffmpeg` found on `PATH`)
- `MKA_DIARIZATION_MODE` (`auto`, `cpu`, or `gpu`; default: `auto`)
- `MKA_DIARIZATION_RUNTIME` (`native` or `container`; default: `native`)
- `MKA_DIARIZATION_CONTAINER_IMAGE` (required for container diarization)
- `MKA_DIARIZATION_CONTAINER_ARGS` (default: no extra arguments), encoded as a
JSON array of strings so ordering and leading dashes are preserved exactly
## Philosophy
For example, a compatible AMD ROCm workstation can configure validated
container access without adding controls to the Streamlit UI:
The transcript is the primary asset.
```bash
export MKA_DIARIZATION_MODE=gpu
export MKA_DIARIZATION_RUNTIME=container
export MKA_DIARIZATION_CONTAINER_IMAGE=rocm/pytorch:rocm7.2.1_ubuntu24.04_py3.12_pytorch_release_2.9.1
export MKA_DIARIZATION_CONTAINER_ARGS='["--device=/dev/kfd","--device=/dev/dri","--group-add","video"]'
```
Everything else (summaries, reports, action items, knowledge extraction)
can always be regenerated using better AI models in the future.
The JSON elements are forwarded as four separate, ordered Meeting Lab
container arguments. Runtime and hardware details remain machine-specific
environment configuration; the UI continues to expose only the diarization
on/off choice.
Therefore:
From the Meeting Assistant repository, start the UI with:
Audio
→ Transcript
is considered immutable.
```bash
PYTHONPATH=src:../meeting-lab .venv/bin/streamlit run src/mka/ui/streamlit_app.py
```
AI output is considered reproducible.
Streamlit opens `http://localhost:8501` by default. Installing the sibling
Meeting Lab project supplies its runtime requirements such as PyYAML and
Requests. Optional diarization dependencies are needed only when diarization
is enabled.
---
Uploaded source files are stored under
`data/meetings/<meeting-id>/uploads/`. Meeting Lab run artifacts are stored
under `data/meetings/<meeting-id>/runs/`. The reviewed protocol is saved as
`protocol_edited.md` inside its run directory; the generated `protocol.md`
remains unchanged.
## Planned Architecture
The upload remains in its original format and is passed unchanged to Meeting Lab.
Meeting Lab creates the canonical mono 16 kHz signed PCM16 WAV run artifact used by
transcription; Meeting Assistant does not duplicate audio conversion.
Recorder
↓
Transcription
↓
Speaker Identification
↓
LLM Processing
↓
Knowledge Database
↓
Export
See [Architecture](docs/architecture.md), [Project Knowledge](PROJECT_KNOWLEDGE.md),
[Roadmap](ROADMAP.md) and [ADR 0011](docs/adr/0011-use-meeting-lab-mvp-backend.md).
---
## Current Status
Project planning.
No implementation has started yet.
+99 -49
View File
@@ -1,62 +1,112 @@
# Roadmap
## Next Milestone: Meeting Assistant MVP GUI
## Phase 1: Minimum Viable Processing Pipeline
Build the user-facing application over the validated Meeting Lab Python API.
- Import existing WAV recordings
- Validate recording metadata
- Create meeting directory
- Create canonical transcript
- Save transcript as JSON
- Export transcript as Markdown
- Archive confirmed recordings as FLAC
- select an existing audio file
- reliably import WAV, FLAC and M4A through canonical FFmpeg preparation
- offer optional audio normalization, enabled by default
- create and edit Meeting Context through structured fields
- add and manage present and `mentioned_only` people
- optionally map `SPEAKER_XX` labels to participants with explicit confirmation
- configure and start `run_mvp_meeting(...)`
- display stage-based status and real progress when available
- handle completed and failed runs clearly
- display and edit the detailed generated protocol
- export the reviewed protocol
---
The milestone must keep diarization, GPU acceleration and fixed-duration audio
chunking optional. It must not launch the Meeting Lab CLI as a subprocess or
duplicate Meeting Lab processing logic.
## Phase 2
Complete real-meeting validation of this vertical slice before expanding the
post-run workflow.
Speaker diarization
## Following Milestone: Distribution Protocol
- Speaker detection
- Speaker naming
- Timeline view
- derive or generate a shorter participant-facing version
- allow review and editing before distribution
- export both detailed and short versions
- retain provenance linking each version to its transcript, configuration,
prompt and model where practical
---
## Later Product Work
## Phase 3
- speaker/name review with explicit human confirmation
- deterministic correction of suitable derived artifacts, beginning with a
simple search-and-replace path
- protocol-only regeneration from existing transcription and diarization plus
corrected mappings and Meeting Context
- transcript viewing and broader correction workflows
- recording and artifact lifecycle management
- search, tags, projects and meeting history
- richer Markdown, PDF and DOCX export
- optional SQLite indexing and full-text search
- integrated recording controls
- calendar and conferencing integrations
- semantic search and organizational knowledge features
Meeting intelligence
## Research and Optional Extensions
- Executive Summary
- Action Items
- Decisions
- Risks
- Open Questions
---
## Phase 4
Knowledge system
- SQLite database
- Full text search
- Semantic search
- Tags
- Projects
- Participants
---
## Phase 5
Desktop application
- Recording UI
- Transcript viewer
- Search
- Export
---
## Phase 6
Integrations
- Outlook
- Teams
- Google Calendar
- Local LLM
- OpenAI
- Ollama
---
## Future Phase: Integrated Recording
- Capture microphone audio
- Capture system audio
- Support separate audio channels
- Provide recording controls in the desktop application
- Replace or complement the external OBS workflow
- Live transcription
- Real-time summaries
- Company glossary
- Custom vocabulary
- Automatic project detection
- Meeting templates
- Voice identification
- Audio cleanup
- live transcription and real-time summaries
- company glossary and custom vocabulary
- voice identification with explicit consent and confirmation
- automatic name and speaker suggestions only as non-authoritative candidates
for later human review
- audio cleanup
- OCR for shared screens
- RAG and knowledge-graph integration
- multi-language meetings and translation
- dual-model or evidence pipelines only if later validation demonstrates a
material reliability benefit
- Timeline with bookmarks
- RAG knowledge integration
- Multi-language meetings
- Translation
- Automatic follow-up generation
View File
+2 -7
View File
@@ -2,7 +2,7 @@
## Status
Accepted, amended by [ADR 0011](0011-use-meeting-lab-mvp-backend.md)
Accepted
## Context
@@ -29,9 +29,4 @@ The transcript is treated as the canonical source. AI-generated summaries and kn
- The system must preserve original recordings and original transcripts.
- AI outputs must be reproducible where practical.
- Modules should remain replaceable.
- The project should avoid unnecessary vendor lock-in.
ADR 0011 narrows the practical MVP to a user-facing application over the
reusable Meeting Lab backend. Direct protocol generation is the validated MVP
path; a separate knowledge-extraction stage remains a possible future
extension rather than an MVP requirement.
- The project should avoid unnecessary vendor lock-in.
+48 -4
View File
@@ -2,7 +2,7 @@
## Status
Superseded by [ADR 0011](0011-use-meeting-lab-mvp-backend.md)
Accepted
## Context
@@ -10,16 +10,60 @@ The project requires high-quality speech-to-text transcription for German and En
The transcription component should support local execution where practical and should remain replaceable in the future. The project should not depend on a meeting bot or on a proprietary meeting platform.
Practical testing showed that long continuous recordings may suffer from error propagation when the decoder continuously conditions on previously generated text.
Independent transcription of smaller audio chunks produced significantly more stable results.
## Decision
The project will use `faster-whisper` as the initial transcription engine.
The preferred first model is `large-v3-turbo`, with `large-v3` as quality-oriented fallback if required.
The preferred first model is `large-v3-turbo`, with `large-v3` as a quality-oriented fallback if required.
The transcription pipeline shall process recordings as independent audio chunks rather than as a single continuous recording.
The default processing strategy is:
```text
Recording (WAV)
│
▼
Split into 5-minute chunks
│
▼
Independent Whisper transcription
(condition_on_previous_text = False)
│
▼
JSON transcript per chunk
│
▼
Merge chunk transcripts
│
▼
Canonical transcript
```
Chunk duration shall be configurable, with **5 minutes** as the default value.
Each chunk shall be processed independently.
The decoder shall not use text generated from previous chunks as context.
Each chunk shall produce its own intermediate JSON transcript before merging.
The merge step shall preserve timestamps and chunk ordering.
## Consequences
- Long recordings become more robust.
- Error propagation across chunk boundaries is prevented.
- Failed chunks can be retranscribed independently.
- Parallel processing of multiple chunks becomes possible.
- Intermediate JSON files simplify debugging and future reprocessing.
- Chunk size can be tuned later without changing the overall architecture.
- Transcription can run locally on suitable hardware.
- NVIDIA CUDA acceleration can be used later on an office AI PC.
- The transcription module must hide the concrete engine behind an internal interface.
- Model name, language setting, timestamp and engine version should be stored with every transcript.
- The decision can be revisited if another engine provides clearly better quality, speed or deployment characteristics.
- Model name, language setting, timestamp mode, engine version, chunk duration and relevant decoding parameters shall be stored with every transcript.
- The decision can be revisited if another engine provides clearly better quality, speed or deployment characteristics.
+1 -9
View File
@@ -2,7 +2,7 @@
## Status
Accepted, amended by [ADR 0011](0011-use-meeting-lab-mvp-backend.md)
Accepted
## Context
@@ -23,11 +23,3 @@ Speaker diarization will be treated as a separate pipeline step after transcript
- Human correction of speaker names should be supported later.
- The diarization module must hide the concrete engine behind an internal interface.
- Model name, engine version and timestamp should be stored with every diarization result.
## Amendment
ADR 0011 makes diarization optional for the MVP and selects `pyannote.audio`
Community-1 as the current preferred backend. CPU execution remains supported,
with GPU acceleration used when available. Speaker labels remain anonymous
unless a user explicitly confirms a `SPEAKER_XX` to participant mapping in the
Meeting Context. Automatic speaker-name inference is not allowed.
+3 -3
View File
@@ -2,7 +2,7 @@
## Status
Accepted, amended by [ADR 0011](0011-use-meeting-lab-mvp-backend.md)
Accepted
## Context
@@ -18,8 +18,8 @@ removing temporary source files without risking data loss.
Recordings shall initially be captured in a processing-friendly lossless
format, normally WAV.
After transcription and, when enabled, diarization have completed, the
recording may be converted to an archive format.
After transcription and diarization have completed, the recording may be
converted to an archive format.
The preferred archive formats are:
+6 -1
View File
@@ -54,7 +54,12 @@ meeting/
├── transcript/
│ ├── canonical.json
│ ├── working_v1.json
│ └── working_v2.json
│
├── chunks/
│ ├── chunk_000.json
│ ├── chunk_001.json
│ ├── chunk_002.json
│ ├── ...
│
├── artifacts/
│ ├── summary.md
@@ -2,7 +2,7 @@
## Status
Accepted, amended by [ADR 0011](0011-use-meeting-lab-mvp-backend.md)
Accepted
## Context
@@ -13,9 +13,8 @@ OBS Studio already provides stable, configurable and cross-platform audio
capture. Reimplementing audio capture during the initial project phase would
add substantial complexity without directly improving transcription quality.
The primary value of the Meeting Assistant lies in transcription, optional
speaker diarization, protocol generation, review and export. ADR 0011 defines
the validated processing boundary and current MVP protocol path.
The primary value of the Meeting Assistant lies in transcription, speaker
diarization, analysis, structured knowledge extraction and export.
## Decision
@@ -1,94 +0,0 @@
# ADR 0011: Use Meeting Lab as the MVP Processing Backend
## Status
Accepted
Supersedes ADR 0002 and amends ADRs 0001 and 0003.
## Context
Meeting Lab experiments have validated a practical local pipeline for turning
a recorded meeting into an editable protocol. Earlier Meeting Assistant plans
selected `faster-whisper`, treated diarization as required and described a
separate knowledge-extraction stage as part of the primary pipeline. Work on an
old feature branch also proposed mandatory fixed five-minute audio chunks and
chunk-oriented storage.
The validated backend now has a reusable Python API, optional diarization and a
direct full-transcript protocol path. The product boundary must keep backend
processing in Meeting Lab and user interaction in Meeting Assistant.
## Decision
Meeting Lab is the reusable processing backend and experimental engine.
Meeting Assistant is the user-facing application and calls Meeting Lab through
its Python API rather than launching its CLI as a subprocess.
The MVP processing flow is:
```text
Source audio
-> FFmpeg preparation (mono, 16 kHz PCM WAV)
-> whisper.cpp transcription (large-v3-turbo)
-> optional pyannote.audio Community-1 diarization
-> direct full-transcript protocol generation
-> human review and editing
```
The productive path does not require fixed five-minute audio chunks. Internal
streaming or segmentation remains an implementation detail of Meeting Lab and
must not define Meeting Assistant's domain or storage model.
Meeting Assistant integrates with these public Meeting Lab interfaces:
- `MvpMeetingConfig`
- `MvpRunResult`
- `run_mvp_meeting(...)`
- stage-based progress events: `preparing`, `transcription`, `diarization`,
`protocol_generation`, `completed` and `failed`
Progress percentages are displayed only when the backend reports real
measurable progress.
`MeetingContext` is the structured input containing meeting metadata and
participants. It may contain explicitly confirmed `SPEAKER_XX -> participant_id`
mappings. Such mappings are authoritative only when confirmed by the user.
The system must not infer speaker names automatically.
The preferred local engines are:
- `whisper.cpp` with `large-v3-turbo` for transcription
- `pyannote.audio` Community-1 for optional diarization
GPU acceleration is optional. Vulkan transcription and PyTorch/ROCm
diarization are validated acceleration paths on North's AMD RX 9070. CPU
execution remains the compatibility fallback.
The direct full-transcript protocol path is the practical MVP direction.
`qwen3.8:27B` has produced strong readability and contextual synthesis, while
`qwen3.6:35B-A3B` has sometimes behaved more conservatively. Dual-model and
diarization-assisted hard-fact experiments have not established sufficiently
reliable attribution gains to justify mandatory MVP complexity. Model choice
remains backend configuration, and human review is required.
Meeting Assistant owns the GUI, structured context editor, participant and
speaker-mapping controls, progress presentation, protocol editor and export.
Meeting Lab owns audio preparation, transcription, diarization, orchestration
and protocol-generation logic.
The product should ultimately provide a detailed contextual protocol and a
shorter participant/distribution version. The shorter version may follow the
first GUI milestone.
## Consequences
- ADR 0002's `faster-whisper` selection is no longer current.
- Fixed-duration chunking and chunk-centric storage are not MVP requirements.
- Diarization and GPU acceleration remain optional.
- The Meeting Lab CLI remains useful for command-line operation but is not the
application integration boundary.
- Meeting Context is created through the future GUI; users are not expected to
hand-write YAML.
- Backend experiments can evolve without moving processing logic into the GUI.
- Protocols remain AI-assisted outputs that require human review.
-67
View File
@@ -1,67 +0,0 @@
# ADR 0012: Preserve Corrections as Post-Run Knowledge
## Status
Accepted as a future product direction; not required for the current MVP.
## Context
Real-meeting review exposes corrections that should not require repeating
expensive processing. These include misspelled or recurring name variants,
people mentioned but omitted from the initial Meeting Context, and confirmed
`SPEAKER_XX -> participant_id` mappings. A destructive edit would lose useful
provenance, while rerunning Whisper or diarization would add cost without
improving a deterministic correction.
The product is also expected to provide both a detailed contextual protocol and
a short participant/distribution protocol. Both should eventually use the same
confirmed context and corrections.
## Decision
Original machine-generated artifacts remain immutable. Corrected or reviewed
artifacts are separate derivatives. Human corrections should evolve into
structured, meeting-specific knowledge rather than opaque destructive edits;
the correction schema is deliberately not defined by this ADR.
An early correction tool may be simple deterministic search and replace, for
example `Grossman` to `Herr Grossmann`. Applying such a confirmed correction to
a suitable editable artifact requires neither Whisper, diarization nor an LLM
call.
A later review stage may run after diarization or the complete initial run. It
may propose likely person/name matches, allow addition of previously omitted
mentioned people, and present anonymous speaker mappings for review. Every
suggestion is non-authoritative: anonymous labels remain anonymous until the
user explicitly confirms or corrects them, and `mentioned_only` people cannot
be mapped as speakers.
After confirmation, the product should offer two paths:
1. A fast path applies confirmed speaker, name or text corrections
deterministically where that is semantically safe.
2. A quality path regenerates protocol output using the existing transcription,
existing diarization, corrected Meeting Context and confirmed mappings or
corrections. It reruns protocol generation only. Whisper and diarization run
again only when separately requested or technically necessary.
The likely later workflow is therefore:
```text
Audio -> transcription -> optional diarization -> initial protocol
-> name/person/speaker review -> human confirmation
-> deterministic correction or protocol-only regeneration
-> final reviewed detailed and/or distribution protocol
```
The exact ordering may evolve, and this review stage is not mandatory for the
current Streamlit MVP.
## Consequences
- Human knowledge can improve outputs without unnecessary upstream work.
- Provenance is retained because originals and corrected derivatives coexist.
- Correction data can later be reused consistently across detailed and short
protocol views.
- Search/replace, correction storage, review UI, identity suggestions and
protocol-only rerun controls remain future implementation work.
+233 -138
View File
@@ -2,159 +2,254 @@
## Purpose
Meeting Assistant is the user-facing application for preparing meeting
context, running the Meeting Lab pipeline and reviewing its results. Meeting
Lab is the reusable processing backend and experimental engine.
The Meeting Knowledge Assistant transforms recorded meetings into structured organizational knowledge through a modular processing pipeline.
## System Boundary
---
# Design Principles
- Architecture first
- Offline first where practical
- Immutable source data
- Replaceable AI components
- Small, focused modules
- Explicit interfaces
- Reproducible AI outputs
- External recording and internal processing are separate responsibilities.
- The processing pipeline must not depend on OBS-specific metadata or behavior.
- Imported recordings must be treated like recordings from any other supported source.
---
## High-Level Pipeline
External Recorder
↓
Audio Import
↓
Recording Validation
↓
Meeting Storage
↓
Transcription
↓
Speaker Diarization
↓
Working Transcript
↓
LLM Analysis
↓
Knowledge Extraction
↓
Export
---
## Core Pipeline
External Recording
↓
Audio Import and Validation
↓
Transcription
↓
Speaker Diarization
↓
Working Transcript
↓
LLM Analysis
↓
Knowledge Extraction
↓
Storage and Export
# Domain Model
Meeting
│
├── Recording
├── Transcript
│ ├── Canonical
│ └── Working
├── Speakers
├── Participants
├── AI Artifacts
├── Knowledge Objects
├── Attachments
└── Exports
---
# Module Responsibilities
## Input and Ingest
Responsible for importing existing recordings into the application.
Initial reference source:
- OBS Studio
Initial supported formats:
- WAV
- FLAC
Responsibilities:
- validate the input file
- collect technical metadata
- calculate checksums
- copy or move the recording into the meeting directory
- create the initial Recording domain object
The ingest component must not perform transcription or modify the audio content.
## Recorder
The recorder module is reserved for a future integrated recording implementation.
It is not required for the MVP.
The initial application workflow uses externally created recordings, with OBS
Studio as the recommended reference recorder.
---
## Transcription
Responsible only for speech-to-text conversion.
Input:
Recording
Output:
Canonical Transcript
---
## Diarization
Responsible only for speaker identification.
Input:
Recording + Canonical Transcript
Output:
Working Transcript
---
## LLM
Responsible for semantic analysis.
Input:
Transcript
Output:
AI Artifacts
---
## Knowledge Extraction
Responsible for creating structured knowledge.
Input:
AI Artifacts
Output:
Knowledge Objects
---
## Export
Responsible for creating user-facing documents.
Input:
Knowledge Objects
Output:
Markdown
PDF
DOCX
---
## Artifact Lifecycle
```text
Meeting Assistant Meeting Lab
----------------- -----------
Audio selection ------> Audio preparation
Meeting Context editor ------> Transcription
Participant management ------> Optional diarization
Explicit speaker mapping ------> Protocol generation
Progress presentation <------ Progress events
Protocol editor and export <------ Run result and artifacts
```
Imported Recording
↓
Validated Source Recording
↓
Canonical Transcript
↓
Working Transcript
↓
AI Artifacts
↓
Knowledge Objects
↓
Exports
Meeting Assistant calls the Meeting Lab Python API directly. The Meeting Lab
CLI is a thin adapter over that API and must not be launched as an application
subprocess.
---
## MVP Processing Flow
## Initial Recording Strategy
```text
Source audio
-> FFmpeg preparation (normalization optional, default on)
-> mono, 16 kHz PCM WAV
-> whisper.cpp transcription with large-v3-turbo
-> optional pyannote.audio Community-1 diarization
-> direct full-transcript protocol generation
-> human review and editing
-> export
```
The MVP does not implement platform-specific audio capture.
The MVP does not require fixed five-minute audio chunks, semantic chunking, a
separate evidence-extraction pipeline or multiple LLMs. Any internal
segmentation remains a Meeting Lab implementation detail.
OBS Studio is the recommended reference recorder for online meetings.
## Backend Integration
The application initially processes existing WAV or FLAC recordings. Integrated
The reusable Meeting Lab interface consists of:
recording remains a future extension and must not be required by transcription,
- `MvpMeetingConfig`, the run configuration
- `run_mvp_meeting(...)`, the orchestration entry point
- `MvpRunResult`, the completed run result
- stage-based progress events
diarization or analysis modules.
The known stages are:
---
```text
preparing
transcription
diarization
protocol_generation
completed
failed
```
# Storage Strategy
The GUI shows the current stage. It shows a percentage only when the event
contains real measurable progress; stage changes must not be presented as
invented percentages.
The domain model is independent of the storage backend.
## Meeting Context
Possible implementations:
`MeetingContext` is structured domain input rather than an informal prompt or a
file users must author manually. It contains meeting metadata and participants
and can include optional mappings:
- File System
- SQLite
- PostgreSQL
- Cloud Storage
```text
SPEAKER_XX -> participant_id
```
---
Mappings are authoritative only after explicit user confirmation. Diarization
labels otherwise remain anonymous, and the application must not infer speaker
names automatically. The Meeting Assistant GUI owns creation and editing of
this context and its mappings.
# Future Extensions
## Processing Components
### Audio Preparation
Meeting Lab uses FFmpeg to prepare a consistent local-processing input. The
current practical target is mono, 16 kHz PCM WAV. The imported source remains a
separate source artifact. Preparation always runs for WAV, FLAC and M4A,
regardless of the normalization switch. When enabled, Meeting Lab currently
uses `loudnorm=I=-16:LRA=11:TP=-1.5`, an isolated conservative default for
speech recordings that may be revisited after empirical comparison. Meeting
Assistant passes only an on/off choice and does not own filter parameters.
### Transcription
`whisper.cpp` with `large-v3-turbo` is the preferred local backend. Vulkan is a
validated acceleration path on North's AMD RX 9070, while CPU execution remains
a compatibility fallback. Engine and model details should be retained with the
result for traceability.
### Diarization
`pyannote.audio` Community-1 is preferred when diarization is enabled. It can
run on CPU and may use GPU acceleration such as PyTorch/ROCm where available.
Diarization is optional and its anonymous labels do not modify the canonical
transcript or assert participant identity.
### Protocol Generation
The direct full-transcript path is the practical MVP. Current model experiments
favor `qwen3.8:27B` for readability and contextual synthesis, while
`qwen3.6:35B-A3B` has shown more conservative behavior in some areas. Neither a
specific dual-model arrangement nor an evidence pipeline is an application
requirement. Generated protocols require human review.
## Application Responsibilities
The first GUI milestone provides:
- audio-file selection
- structured Meeting Context and participant editing
- optional explicit speaker mapping
- pipeline start and configuration
- progress and failure presentation
- detailed protocol display and editing
- export of the reviewed result
The product direction also includes a shorter participant/distribution
protocol. It may be delivered after the first GUI milestone.
## Artifact and Storage Principles
The existing Meeting domain remains the application-level container for source
recordings, canonical and working transcripts, context, generated artifacts and
exports. Storage remains independent of processing segmentation. In particular,
chunks are not primary Meeting Assistant domain objects or the required unit of
MVP persistence.
Source audio and the canonical transcription output are preserved. Corrections
and reviewed protocols are derived versions. Generated artifacts should retain
their input version, backend/model configuration, prompt version and timestamp
where practical.
## Later Post-Run Correction Flow
Post-run corrections are planned as a separate, non-mandatory workflow after
initial protocol generation. As detailed in [ADR 0012](adr/0012-post-run-corrections.md),
confirmed name, person and anonymous-speaker corrections should become
meeting-specific knowledge. They may then be applied deterministically to safe
derived artifacts or used for protocol-only regeneration without needlessly
rerunning transcription or diarization.
## Future Extensions
- shorter distribution protocols
- integrated recording
- searchable meeting history and optional database indexes
- live transcription
- OCR and video processing
- semantic search and organizational knowledge features
- Live transcription
- Video processing
- OCR
- Semantic search
- Knowledge graph
- Company glossary
- Multi-language meetings
- Local LLM support
+1 -17
View File
@@ -2,23 +2,7 @@ Transcript
Original speech-to-text output.
Diarization
Assignment of transcript segments to anonymous speaker labels. It does not by
itself identify participants.
Meeting Context
Structured meeting metadata, participants and optional explicitly confirmed
speaker-to-participant mappings supplied to the processing pipeline.
Speaker Mapping
An explicitly confirmed association from a diarization label such as
`SPEAKER_00` to a participant ID. It must not be inferred automatically.
Detailed Protocol
Contextual protocol generated from the full transcript for human review and
continued work.
Distribution Protocol
Shorter, participant-facing version of a reviewed meeting protocol.
Assignment of transcript segments to speakers.
Knowledge Object
Structured information extracted from meetings.
+1 -3
View File
@@ -12,9 +12,7 @@ authors = [
{ name = "Martin Tazl" }
]
dependencies = [
"pydantic>=2,<3",
"PyYAML>=6,<7",
"streamlit>=1.40,<2",
"pydantic>=2,<3"
]
[project.optional-dependencies]
-17
View File
@@ -1,17 +0,0 @@
"""Application services for Meeting Assistant workflows."""
from mka.application.meeting_service import (
MeetingDetails,
MeetingProcessingService,
ParticipantInput,
ProcessingOptions,
ProcessingOutcome,
)
__all__ = [
"MeetingDetails",
"MeetingProcessingService",
"ParticipantInput",
"ProcessingOptions",
"ProcessingOutcome",
]
-100
View File
@@ -1,100 +0,0 @@
"""Environment-backed configuration for the Meeting Assistant application."""
from __future__ import annotations
import json
import os
from dataclasses import dataclass
from pathlib import Path
class ConfigurationError(ValueError):
"""Raised when required processing configuration is unavailable."""
@dataclass(frozen=True)
class AppSettings:
"""Machine-specific Meeting Lab settings kept outside the UI."""
data_root: Path
whisper_model: Path | None
whisper_executable: str = "whisper-cli"
ffmpeg_executable: str = "ffmpeg"
protocol_model: str = "qwen3.8:27B"
ollama_endpoint: str = "http://127.0.0.1:11434"
language: str = "de"
threads: str | int = "auto"
diarization_mode: str = "auto"
diarization_runtime: str = "native"
diarization_container_image: str | None = None
diarization_container_args: tuple[str, ...] = ()
@classmethod
def from_environment(cls) -> AppSettings:
"""Load settings from MKA_* environment variables."""
whisper_model = os.getenv("MKA_WHISPER_MODEL")
threads = os.getenv("MKA_WHISPER_THREADS", "auto")
parsed_threads: str | int = int(threads) if threads.isdigit() else threads
return cls(
data_root=Path(os.getenv("MKA_DATA_ROOT", "data/meetings")),
whisper_model=Path(whisper_model) if whisper_model else None,
whisper_executable=os.getenv("MKA_WHISPER_EXECUTABLE", "whisper-cli"),
ffmpeg_executable=os.getenv("MKA_FFMPEG_EXECUTABLE", "ffmpeg"),
protocol_model=os.getenv("MKA_PROTOCOL_MODEL", "qwen3.8:27B"),
ollama_endpoint=os.getenv("MKA_OLLAMA_ENDPOINT", "http://127.0.0.1:11434"),
language=os.getenv("MKA_LANGUAGE", "de"),
threads=parsed_threads,
diarization_mode=os.getenv("MKA_DIARIZATION_MODE", "auto"),
diarization_runtime=os.getenv("MKA_DIARIZATION_RUNTIME", "native"),
diarization_container_image=os.getenv("MKA_DIARIZATION_CONTAINER_IMAGE"),
diarization_container_args=_parse_container_args(
os.getenv("MKA_DIARIZATION_CONTAINER_ARGS")
),
)
def validate_for_processing(self) -> None:
"""Validate settings needed before a Meeting Lab run starts."""
if self.whisper_model is None:
raise ConfigurationError("MKA_WHISPER_MODEL is not configured.")
if not self.whisper_model.is_file():
raise ConfigurationError(
f"Configured Whisper model does not exist: {self.whisper_model}"
)
if not self.whisper_executable.strip():
raise ConfigurationError("MKA_WHISPER_EXECUTABLE must not be empty.")
if not self.ffmpeg_executable.strip():
raise ConfigurationError("MKA_FFMPEG_EXECUTABLE must not be empty.")
if self.diarization_mode not in {"auto", "cpu", "gpu"}:
raise ConfigurationError("MKA_DIARIZATION_MODE must be one of: auto, cpu, gpu.")
if self.diarization_runtime not in {"native", "container"}:
raise ConfigurationError("MKA_DIARIZATION_RUNTIME must be native or container.")
if self.diarization_runtime == "container" and not self.diarization_container_image:
raise ConfigurationError(
"MKA_DIARIZATION_CONTAINER_IMAGE is required for container runtime."
)
if any(
not isinstance(argument, str) or not argument
for argument in self.diarization_container_args
):
raise ConfigurationError(
"MKA_DIARIZATION_CONTAINER_ARGS must contain only non-empty strings."
)
def _parse_container_args(value: str | None) -> tuple[str, ...]:
"""Parse an ordered JSON array of opaque container command arguments."""
if value is None or not value.strip():
return ()
try:
parsed = json.loads(value)
except json.JSONDecodeError as exc:
raise ConfigurationError(
"MKA_DIARIZATION_CONTAINER_ARGS must be a JSON array of strings."
) from exc
if not isinstance(parsed, list) or any(
not isinstance(argument, str) or not argument for argument in parsed
):
raise ConfigurationError(
"MKA_DIARIZATION_CONTAINER_ARGS must be a JSON array of non-empty strings."
)
return tuple(parsed)
-272
View File
@@ -1,272 +0,0 @@
"""Use-case service for processing one uploaded meeting recording."""
from __future__ import annotations
import json
import re
import shutil
from collections.abc import Callable
from dataclasses import dataclass
from datetime import date
from pathlib import Path
from typing import Any, Protocol
from uuid import uuid4
from mka.application.config import AppSettings
STAGES = ("preparing", "transcription", "diarization", "protocol_generation")
class MeetingLabPort(Protocol):
"""Operations the application requires from Meeting Lab."""
def create_context(self, data: dict[str, Any]) -> Any: ...
def create_config(self, values: dict[str, Any]) -> Any: ...
def run(
self,
config: Any,
meeting_context: Any,
progress_sink: Callable[[Any], None],
) -> Any: ...
@dataclass(frozen=True)
class MeetingDetails:
title: str
language: str = "de"
meeting_date: date | None = None
description: str = ""
meeting_id: str | None = None
@dataclass(frozen=True)
class ParticipantInput:
participant_id: str
display_name: str
role: str = ""
organization: str = ""
attendance_status: str = "present"
@dataclass(frozen=True)
class ProcessingOptions:
diarization_enabled: bool = False
audio_normalization: bool = True
@dataclass(frozen=True)
class AppProgressEvent:
stage: str
status: str
elapsed_seconds: float
progress: float | None = None
message: str | None = None
@dataclass(frozen=True)
class ProcessingOutcome:
succeeded: bool
run_dir: Path | None
original_protocol: str | None
protocol_path: Path | None
failed_stage: str | None = None
error_message: str | None = None
def stable_id(value: str) -> str:
"""Return a schema-safe ID based on user text, with a random fallback."""
normalized = re.sub(r"[^a-z0-9]+", "-", value.lower()).strip("-")
return normalized or f"participant-{uuid4().hex[:8]}"
class MeetingProcessingService:
"""Translate UI input into Meeting Lab calls and persisted artifacts."""
def __init__(self, settings: AppSettings, meeting_lab: MeetingLabPort) -> None:
self.settings = settings
self.meeting_lab = meeting_lab
def build_context(
self,
meeting: MeetingDetails,
participants: list[ParticipantInput],
speaker_mappings: dict[str, str] | None = None,
) -> Any:
"""Build and validate Meeting Context V1 through Meeting Lab."""
meeting_id = meeting.meeting_id or stable_id(meeting.title)
organizations = sorted(
{item.organization.strip() for item in participants if item.organization.strip()}
)
departments = [
{"id": stable_id(name), "name": name, "aliases": []} for name in organizations
]
department_ids = {item["name"]: item["id"] for item in departments}
data = {
"schema_version": "1",
"meeting": {
"meeting_id": meeting_id,
"title": meeting.title.strip(),
"language": meeting.language,
"date": meeting.meeting_date.isoformat() if meeting.meeting_date else None,
"objective": "",
"notes": meeting.description.strip(),
},
"participants": [
{
"participant_id": item.participant_id.strip(),
"display_name": item.display_name.strip(),
"aliases": [],
"role": item.role.strip() or None,
"department": department_ids.get(item.organization.strip()),
"attendance_status": "present",
"notes": None,
}
for item in participants
if item.attendance_status == "present"
],
"speaker_mappings": dict(speaker_mappings or {}),
"mentioned_people": [
{
"person_id": item.participant_id.strip(),
"display_name": item.display_name.strip(),
"aliases": [],
"role": item.role.strip() or None,
"department": department_ids.get(item.organization.strip()),
"attendance_status": "mentioned_only",
"notes": None,
}
for item in participants
if item.attendance_status == "mentioned_only"
],
"organization": {
"name": None,
"departments": departments,
"abbreviations": {},
},
"known_entities": {},
"context_rules": {
"participant_list_is_authoritative": True,
"do_not_infer_roles": True,
"do_not_infer_departments": True,
"do_not_infer_responsibilities": True,
"mentioned_people_are_not_participants": True,
},
}
return self.meeting_lab.create_context(data)
def preserve_upload(self, meeting_id: str, filename: str, source: Any) -> Path:
"""Persist an uploaded source before processing starts."""
safe_filename = Path(filename).name
upload_dir = self.settings.data_root / stable_id(meeting_id) / "uploads"
upload_dir.mkdir(parents=True, exist_ok=True)
destination = upload_dir / f"{uuid4().hex[:8]}_{safe_filename}"
with destination.open("wb") as target:
if hasattr(source, "getbuffer"):
target.write(source.getbuffer())
else:
shutil.copyfileobj(source, target)
return destination
def process(
self,
audio_file: Path,
meeting: MeetingDetails,
participants: list[ParticipantInput],
options: ProcessingOptions,
progress_sink: Callable[[AppProgressEvent], None] | None = None,
speaker_mappings: dict[str, str] | None = None,
) -> ProcessingOutcome:
"""Validate input, run Meeting Lab and expose a UI-oriented result."""
self.settings.validate_for_processing()
context = self.build_context(meeting, participants, speaker_mappings)
meeting_id = context.meeting_id
output_root = self.settings.data_root / stable_id(meeting_id) / "runs"
config = self.meeting_lab.create_config(
{
"audio_file": audio_file,
"whisper_model": self.settings.whisper_model,
"whisper_executable": self.settings.whisper_executable,
"ffmpeg_executable": self.settings.ffmpeg_executable,
"audio_normalization": options.audio_normalization,
"output_root": output_root,
"language": meeting.language or self.settings.language,
"threads": self.settings.threads,
"model": self.settings.protocol_model,
"ollama_endpoint": self.settings.ollama_endpoint,
"diarization": (
self.settings.diarization_mode if options.diarization_enabled else "off"
),
"diarization_runtime": self.settings.diarization_runtime,
"diarization_container_image": (self.settings.diarization_container_image),
"diarization_container_args": self.settings.diarization_container_args,
}
)
current_stage: str | None = None
def relay(event: Any) -> None:
nonlocal current_stage
if event.stage in STAGES and event.status == "started":
current_stage = event.stage
app_event = AppProgressEvent(
stage=event.stage,
status=event.status,
elapsed_seconds=event.elapsed_seconds,
progress=event.progress,
message=event.message,
)
if progress_sink is not None:
progress_sink(app_event)
result = self.meeting_lab.run(config, context, relay)
if result.exit_code != 0:
failure = self._read_failure(result.run_dir)
backend_stage = failure.get("stage")
failed_stage = {
"validation": "preparing",
"setup": "preparing",
"audio_preparation": "preparing",
"whisper": "transcription",
"transcript_validation": "transcription",
"protocol": "protocol_generation",
}.get(backend_stage, backend_stage)
failure_message = failure.get("message") or "Meeting Lab processing failed."
failure_type = failure.get("type")
if failure_type and not failure_message.startswith(f"{failure_type}:"):
failure_message = f"{failure_type}: {failure_message}"
return ProcessingOutcome(
succeeded=False,
run_dir=result.run_dir,
original_protocol=None,
protocol_path=None,
failed_stage=failed_stage or current_stage or "preparing",
error_message=failure_message,
)
protocol_path = Path(result.protocol_path)
return ProcessingOutcome(
succeeded=True,
run_dir=result.run_dir,
original_protocol=protocol_path.read_text(encoding="utf-8"),
protocol_path=protocol_path,
)
@staticmethod
def _read_failure(run_dir: Path | None) -> dict[str, str]:
if run_dir is None:
return {}
metadata_path = Path(run_dir) / "run_metadata.json"
if not metadata_path.is_file():
return {}
try:
failure = json.loads(metadata_path.read_text(encoding="utf-8")).get("failure")
except (OSError, json.JSONDecodeError):
return {}
return failure if isinstance(failure, dict) else {}
@staticmethod
def save_edited_protocol(run_dir: Path, text: str) -> Path:
"""Persist user edits beside, never over, the generated protocol."""
destination = Path(run_dir) / "protocol_edited.md"
destination.write_text(text, encoding="utf-8")
return destination
-108
View File
@@ -1,108 +0,0 @@
"""Versioned YAML import and export for reusable People lists."""
from __future__ import annotations
import re
from collections.abc import Sequence
import yaml
from mka.application.meeting_service import ParticipantInput
PEOPLE_YAML_VERSION = 1
ATTENDANCE_VALUES = frozenset({"present", "mentioned_only"})
PARTICIPANT_ID_PATTERN = re.compile(r"^[A-Za-z0-9][A-Za-z0-9_.:-]*$")
class PeopleYamlError(ValueError):
"""Raised when a reusable People-list document is invalid."""
def export_people_yaml(people: Sequence[ParticipantInput]) -> str:
"""Serialize people deterministically without meeting-specific data."""
document = {
"version": PEOPLE_YAML_VERSION,
"people": [
{
"participant_id": person.participant_id,
"display_name": person.display_name,
"role": person.role,
"organization": person.organization,
"attendance_status": person.attendance_status,
}
for person in people
],
}
return yaml.safe_dump(
document,
allow_unicode=True,
sort_keys=False,
default_flow_style=False,
)
def import_people_yaml(content: str | bytes) -> list[ParticipantInput]:
"""Parse and validate a complete replacement People list."""
try:
if isinstance(content, bytes):
content = content.decode("utf-8")
document = yaml.safe_load(content)
except (yaml.YAMLError, UnicodeDecodeError) as exc:
raise PeopleYamlError(f"Malformed People YAML: {exc}") from exc
if not isinstance(document, dict):
raise PeopleYamlError("People YAML must contain a top-level mapping.")
version = document.get("version")
if type(version) is not int or version != PEOPLE_YAML_VERSION:
raise PeopleYamlError(
f"Unsupported People YAML version {version!r}; expected {PEOPLE_YAML_VERSION}."
)
entries = document.get("people")
if not isinstance(entries, list):
raise PeopleYamlError("People YAML must contain a top-level 'people' list.")
people: list[ParticipantInput] = []
seen_ids: set[str] = set()
for index, entry in enumerate(entries, start=1):
if not isinstance(entry, dict):
raise PeopleYamlError(f"Person {index} must be a mapping.")
participant_id = _required_text(entry, "participant_id", index)
if PARTICIPANT_ID_PATTERN.fullmatch(participant_id) is None:
raise PeopleYamlError(f"Person {index} has invalid participant_id {participant_id!r}.")
if participant_id in seen_ids:
raise PeopleYamlError(f"Duplicate participant_id: {participant_id!r}.")
seen_ids.add(participant_id)
display_name = _required_text(entry, "display_name", index)
role = _optional_text(entry, "role", index)
organization = _optional_text(entry, "organization", index)
attendance_status = entry.get("attendance_status")
if attendance_status not in ATTENDANCE_VALUES:
allowed = ", ".join(sorted(ATTENDANCE_VALUES))
raise PeopleYamlError(
f"Person {index} has invalid attendance_status; expected one of: {allowed}."
)
people.append(
ParticipantInput(
participant_id=participant_id,
display_name=display_name,
role=role,
organization=organization,
attendance_status=attendance_status,
)
)
return people
def _required_text(entry: dict[object, object], field: str, index: int) -> str:
value = entry.get(field)
if not isinstance(value, str) or not value.strip():
raise PeopleYamlError(f"Person {index} requires a non-empty {field}.")
return value
def _optional_text(entry: dict[object, object], field: str, index: int) -> str:
value = entry.get(field, "")
if not isinstance(value, str):
raise PeopleYamlError(f"Person {index} field {field} must be text.")
return value
-1
View File
@@ -1 +0,0 @@
"""Adapters for external processing systems."""
-51
View File
@@ -1,51 +0,0 @@
"""Narrow adapter around the reusable Meeting Lab Python API."""
from __future__ import annotations
from collections.abc import Callable, Mapping
from typing import Any
class MeetingLabUnavailableError(RuntimeError):
"""Raised when the Meeting Lab package cannot be imported."""
class MeetingLabGateway:
"""Load and delegate to Meeting Lab without coupling Streamlit to it."""
def __init__(self) -> None:
try:
from src.meeting_lab.models.meeting_context import create_meeting_context
from src.meeting_lab.orchestration.mvp import (
MvpMeetingConfig,
run_mvp_meeting,
)
except ImportError as exc:
raise MeetingLabUnavailableError(
"Meeting Lab is unavailable. Add the Meeting Lab repository root "
"to PYTHONPATH as described in README.md."
) from exc
self._config_type = MvpMeetingConfig
self._create_context = create_meeting_context
self._run_mvp_meeting = run_mvp_meeting
def create_context(self, data: dict[str, Any]) -> Any:
"""Validate structured context through Meeting Lab's domain boundary."""
return self._create_context(data)
def create_config(self, values: Mapping[str, Any]) -> Any:
"""Create the concrete Meeting Lab run configuration."""
return self._config_type(**values)
def run(
self,
config: Any,
meeting_context: Any,
progress_sink: Callable[[Any], None],
) -> Any:
"""Run Meeting Lab synchronously and relay progress callbacks."""
return self._run_mvp_meeting(
config,
meeting_context=meeting_context,
progress_sink=progress_sink,
)
+1 -1
View File
@@ -22,4 +22,4 @@ class DomainModel(BaseModel):
id: UUID = Field(default_factory=uuid4)
created_at: datetime = Field(default_factory=utc_now)
updated_at: datetime = Field(default_factory=utc_now)
updated_at: datetime = Field(default_factory=utc_now)
-302
View File
@@ -1,302 +0,0 @@
"""Streamlit presentation layer for the first Meeting Assistant MVP."""
from __future__ import annotations
from datetime import date
from typing import Any
from uuid import uuid4
import streamlit as st
from mka.application.config import AppSettings, ConfigurationError
from mka.application.meeting_service import (
STAGES,
AppProgressEvent,
MeetingDetails,
MeetingProcessingService,
ParticipantInput,
ProcessingOptions,
stable_id,
)
from mka.application.people_yaml import (
PeopleYamlError,
export_people_yaml,
import_people_yaml,
)
from mka.integrations.meeting_lab import (
MeetingLabGateway,
MeetingLabUnavailableError,
)
STAGE_LABELS = {
"preparing": "Preparation",
"transcription": "Transcription",
"diarization": "Diarization",
"protocol_generation": "Protocol generation",
}
def _new_participant() -> dict[str, str]:
return {
"row_id": uuid4().hex,
"participant_id": "",
"display_name": "",
"role": "",
"organization": "",
"attendance_status": "present",
}
def _initialize_state() -> None:
st.session_state.setdefault("participants", [_new_participant()])
st.session_state.setdefault("outcome", None)
st.session_state.setdefault("edited_protocol", "")
def _people_to_rows(people: list[ParticipantInput]) -> list[dict[str, str]]:
"""Create fresh widget rows while preserving reusable person IDs."""
return [
{
"row_id": uuid4().hex,
"participant_id": person.participant_id,
"display_name": person.display_name,
"role": person.role,
"organization": person.organization,
"attendance_status": person.attendance_status,
}
for person in people
]
def _render_participants() -> list[ParticipantInput]:
st.subheader("People")
st.caption("Record whether each relevant person attended or was only mentioned.")
imported_file = st.file_uploader(
"Import people",
type=["yaml", "yml"],
help="Replace the current People list from a versioned YAML export.",
)
if st.button("Import people list", disabled=imported_file is None):
try:
imported_people = import_people_yaml(imported_file.getvalue())
except PeopleYamlError as exc:
st.error(str(exc))
else:
st.session_state.participants = _people_to_rows(imported_people)
st.session_state.people_import_message = f"Imported {len(imported_people)} people."
st.rerun()
if message := st.session_state.pop("people_import_message", None):
st.success(message)
rows = st.session_state.participants
remove_index: int | None = None
for index, row in enumerate(rows):
row_id = row["row_id"]
columns = st.columns([2, 2, 2, 2, 2, 0.6])
row["display_name"] = columns[0].text_input(
"Name", value=row["display_name"], key=f"name_{row_id}"
)
suggested_id = row["participant_id"] or stable_id(row["display_name"])
row["participant_id"] = columns[1].text_input(
"Participant ID", value=suggested_id, key=f"id_{row_id}"
)
row["role"] = columns[2].text_input("Role", value=row["role"], key=f"role_{row_id}")
row["organization"] = columns[3].text_input(
"Organization / department",
value=row["organization"],
key=f"organization_{row_id}",
)
row["attendance_status"] = columns[4].selectbox(
"Attendance",
options=["present", "mentioned_only"],
format_func=lambda value: {
"present": "Present / participated",
"mentioned_only": "Mentioned, but not present",
}[value],
index=0 if row["attendance_status"] == "present" else 1,
key=f"attendance_{row_id}",
)
if columns[5].button("Remove", key=f"remove_{row_id}"):
remove_index = index
if remove_index is not None:
rows.pop(remove_index)
st.rerun()
action_columns = st.columns(2)
if action_columns[0].button("Add person"):
rows.append(_new_participant())
st.rerun()
people = [
ParticipantInput(
participant_id=row["participant_id"],
display_name=row["display_name"],
role=row["role"],
organization=row["organization"],
attendance_status=row["attendance_status"],
)
for row in rows
]
exported_yaml = export_people_yaml(people).encode("utf-8")
action_columns[1].download_button(
"Export people",
data=exported_yaml,
file_name="people.yaml",
mime="application/yaml",
)
return people
def _progress_callback(
status_box: Any,
stage_table: Any,
progress_slot: Any,
states: dict[str, str],
) -> Any:
progress_bar = None
def update(event: AppProgressEvent) -> None:
nonlocal progress_bar
if event.stage in states:
states[event.stage] = "running" if event.status == "started" else event.status
if event.stage == "failed":
running = next(
(stage for stage, status in states.items() if status == "running"),
None,
)
if running:
states[running] = "failed"
elapsed = f"{event.elapsed_seconds:.1f} s"
message = event.message or STAGE_LABELS.get(event.stage, event.stage)
status_box.info(f"{message} — elapsed {elapsed}")
stage_table.table(
[{"Stage": STAGE_LABELS[stage], "Status": states[stage]} for stage in STAGES]
)
if event.progress is not None:
if progress_bar is None:
progress_bar = progress_slot.progress(0.0)
progress_bar.progress(
min(max(event.progress, 0.0), 1.0),
text=f"{STAGE_LABELS.get(event.stage, event.stage)}: {event.progress:.0%}",
)
return update
def _render_result() -> None:
outcome = st.session_state.outcome
if outcome is None:
return
st.divider()
st.header("Protocol result")
if not outcome.succeeded:
st.error(f"Processing failed during {outcome.failed_stage}: {outcome.error_message}")
if outcome.run_dir:
st.code(str(outcome.run_dir))
st.caption("Intermediate artifacts and run metadata were preserved here.")
return
st.success("Processing completed. Review the generated protocol before use.")
st.caption(f"Run artifacts: {outcome.run_dir}")
with st.expander("Original generated protocol", expanded=False):
st.markdown(outcome.original_protocol or "")
edited = st.text_area(
"Editable protocol",
key="edited_protocol",
height=500,
help="The original protocol.md remains unchanged.",
)
if st.button("Save edited protocol", type="primary"):
path = MeetingProcessingService.save_edited_protocol(outcome.run_dir, edited)
st.success(f"Saved edited protocol to {path}")
def main() -> None:
st.set_page_config(page_title="Meeting Assistant", layout="wide")
_initialize_state()
st.title("Meeting Assistant")
st.caption("Create meeting context, run Meeting Lab, and review the protocol.")
st.header("Meeting and audio")
audio = st.file_uploader("Audio recording", type=["wav", "flac", "m4a"])
title = st.text_input("Meeting title")
description = st.text_area("Description / context", height=100)
metadata_columns = st.columns(3)
language = metadata_columns[0].selectbox("Meeting language", options=["de", "en"])
has_date = metadata_columns[1].checkbox("Meeting date is known", value=True)
selected_date: date | None = (
metadata_columns[2].date_input("Meeting date") if has_date else None
)
participants = _render_participants()
st.header("Processing options")
audio_normalization = st.checkbox(
"Audio normalization",
value=True,
help=(
"Normalize speech loudness during preparation. WAV, FLAC, and M4A "
"are always converted to the canonical Meeting Lab audio format, "
"regardless of this setting."
),
)
diarization_enabled = st.checkbox(
"Enable speaker diarization",
help="Speaker labels remain anonymous; identities are never inferred.",
)
if st.button("Start processing", type="primary", disabled=audio is None):
if not title.strip():
st.error("Meeting title is required.")
elif any(
not item.display_name.strip() or not item.participant_id.strip()
for item in participants
):
st.error("Every person row needs a name and participant ID.")
else:
settings = AppSettings.from_environment()
try:
service = MeetingProcessingService(settings, MeetingLabGateway())
meeting = MeetingDetails(
title=title,
language=language,
meeting_date=selected_date,
description=description,
)
meeting_id = stable_id(title)
audio_path = service.preserve_upload(meeting_id, audio.name, audio)
st.header("Processing status")
status_box = st.empty()
stage_table = st.empty()
progress_slot = st.empty()
states = {stage: "pending" for stage in STAGES}
if not diarization_enabled:
states["diarization"] = "skipped"
callback = _progress_callback(status_box, stage_table, progress_slot, states)
outcome = service.process(
audio_path,
meeting,
participants,
ProcessingOptions(
diarization_enabled=diarization_enabled,
audio_normalization=audio_normalization,
),
progress_sink=callback,
)
st.session_state.outcome = outcome
st.session_state.edited_protocol = outcome.original_protocol or ""
if outcome.succeeded:
status_box.success("Processing completed.")
else:
status_box.error(
f"Processing failed during {outcome.failed_stage}: "
f"{outcome.error_message} Artifacts were preserved."
)
except (ConfigurationError, MeetingLabUnavailableError, ValueError) as exc:
st.error(str(exc))
except Exception as exc:
st.error(f"Unable to process meeting: {exc}")
_render_result()
if __name__ == "__main__":
main()
-68
View File
@@ -1,68 +0,0 @@
from pathlib import Path
import pytest
from mka.application.config import AppSettings, ConfigurationError
def test_settings_require_whisper_model() -> None:
settings = AppSettings(data_root=Path("runs"), whisper_model=None)
with pytest.raises(ConfigurationError, match="MKA_WHISPER_MODEL"):
settings.validate_for_processing()
def test_settings_reject_missing_whisper_model(tmp_path: Path) -> None:
settings = AppSettings(
data_root=tmp_path / "runs",
whisper_model=tmp_path / "missing.bin",
)
with pytest.raises(ConfigurationError, match="does not exist"):
settings.validate_for_processing()
def test_settings_accept_valid_native_configuration(tmp_path: Path) -> None:
model = tmp_path / "model.bin"
model.write_bytes(b"model")
settings = AppSettings(data_root=tmp_path / "runs", whisper_model=model)
settings.validate_for_processing()
assert settings.diarization_runtime == "native"
assert settings.diarization_container_args == ()
def test_environment_defaults_to_no_diarization_container_args(monkeypatch) -> None:
monkeypatch.delenv("MKA_DIARIZATION_CONTAINER_ARGS", raising=False)
settings = AppSettings.from_environment()
assert settings.diarization_container_args == ()
def test_environment_parses_multiple_ordered_container_args(monkeypatch) -> None:
monkeypatch.setenv(
"MKA_DIARIZATION_CONTAINER_ARGS",
'["--device=/dev/kfd", "--device=/dev/dri", "--group-add", "video"]',
)
settings = AppSettings.from_environment()
assert settings.diarization_container_args == (
"--device=/dev/kfd",
"--device=/dev/dri",
"--group-add",
"video",
)
@pytest.mark.parametrize(
"value",
["not-json", '"--flag"', '["valid", ""]', '["valid", 1]'],
)
def test_environment_rejects_invalid_container_args(monkeypatch, value: str) -> None:
monkeypatch.setenv("MKA_DIARIZATION_CONTAINER_ARGS", value)
with pytest.raises(ConfigurationError, match="JSON array"):
AppSettings.from_environment()
-340
View File
@@ -1,340 +0,0 @@
import json
from dataclasses import dataclass, replace
from datetime import date
from pathlib import Path
from types import SimpleNamespace
from typing import Any
from mka.application.config import AppSettings
from mka.application.meeting_service import (
MeetingDetails,
MeetingProcessingService,
ParticipantInput,
ProcessingOptions,
)
@dataclass
class FakeContext:
data: dict[str, Any]
@property
def meeting_id(self) -> str:
return self.data["meeting"]["meeting_id"]
class FakeMeetingLab:
def __init__(self, run_dir: Path) -> None:
self.run_dir = run_dir
self.context_data: dict[str, Any] | None = None
self.config_values: dict[str, Any] | None = None
self.fail = False
def create_context(self, data: dict[str, Any]) -> FakeContext:
self.context_data = data
if not data["meeting"]["title"]:
raise ValueError("meeting.title must be present and non-empty")
participant_ids = [item["participant_id"] for item in data["participants"]]
if len(participant_ids) != len(set(participant_ids)):
raise ValueError("Duplicate participant_id")
return FakeContext(data)
def create_config(self, values: dict[str, Any]) -> dict[str, Any]:
self.config_values = values
return values
def run(self, config: Any, meeting_context: Any, progress_sink: Any) -> Any:
self.run_dir.mkdir(parents=True, exist_ok=True)
progress_sink(
SimpleNamespace(
stage="preparing",
status="started",
elapsed_seconds=0.1,
progress=None,
message=None,
)
)
progress_sink(
SimpleNamespace(
stage="preparing",
status="completed",
elapsed_seconds=0.2,
progress=None,
message=None,
)
)
progress_sink(
SimpleNamespace(
stage="transcription",
status="started",
elapsed_seconds=0.2,
progress=0.25,
message="transcribing",
)
)
if self.fail:
(self.run_dir / "run_metadata.json").write_text(
json.dumps(
{
"failure": {
"stage": "whisper",
"type": "TranscriptionError",
"message": "model failed",
}
}
),
encoding="utf-8",
)
progress_sink(
SimpleNamespace(
stage="failed",
status="failed",
elapsed_seconds=0.3,
progress=None,
message="transcription: model failed",
)
)
return SimpleNamespace(exit_code=2, run_dir=self.run_dir, protocol_path=None)
protocol = self.run_dir / "protocol.md"
protocol.write_text("# Generated protocol\n", encoding="utf-8")
return SimpleNamespace(exit_code=0, run_dir=self.run_dir, protocol_path=protocol)
def make_service(tmp_path: Path) -> tuple[MeetingProcessingService, FakeMeetingLab]:
model = tmp_path / "model.bin"
model.write_bytes(b"model")
gateway = FakeMeetingLab(tmp_path / "backend-run")
settings = AppSettings(
data_root=tmp_path / "meetings",
whisper_model=model,
whisper_executable="/opt/whisper-cli",
protocol_model="test:model",
diarization_mode="gpu",
)
return MeetingProcessingService(settings, gateway), gateway
def meeting() -> MeetingDetails:
return MeetingDetails(
title="Architecture Review",
language="de",
meeting_date=date(2026, 8, 23),
description="Review the MVP.",
)
def participants() -> list[ParticipantInput]:
return [
ParticipantInput(
participant_id="martin",
display_name="Martin",
role="Project lead",
organization="Engineering",
),
ParticipantInput(
participant_id="alex",
display_name="Alex",
organization="Engineering",
),
]
def test_build_context_uses_actual_v1_shape(tmp_path: Path) -> None:
service, gateway = make_service(tmp_path)
context = service.build_context(meeting(), participants())
assert context.meeting_id == "architecture-review"
assert gateway.context_data is not None
assert gateway.context_data["meeting"]["date"] == "2026-08-23"
assert gateway.context_data["meeting"]["notes"] == "Review the MVP."
assert gateway.context_data["participants"][0] == {
"participant_id": "martin",
"display_name": "Martin",
"aliases": [],
"role": "Project lead",
"department": "engineering",
"attendance_status": "present",
"notes": None,
}
assert gateway.context_data["organization"]["departments"] == [
{"id": "engineering", "name": "Engineering", "aliases": []}
]
def test_build_context_preserves_explicit_speaker_mapping(tmp_path: Path) -> None:
service, gateway = make_service(tmp_path)
service.build_context(meeting(), participants(), {"SPEAKER_00": "martin"})
assert gateway.context_data is not None
assert gateway.context_data["speaker_mappings"] == {"SPEAKER_00": "martin"}
def test_build_context_translates_mentioned_only_person(tmp_path: Path) -> None:
service, gateway = make_service(tmp_path)
people = participants() + [
ParticipantInput(
participant_id="sam",
display_name="Sam",
attendance_status="mentioned_only",
)
]
service.build_context(meeting(), people)
assert gateway.context_data is not None
assert [item["participant_id"] for item in gateway.context_data["participants"]] == [
"martin",
"alex",
]
assert gateway.context_data["mentioned_people"] == [
{
"person_id": "sam",
"display_name": "Sam",
"aliases": [],
"role": None,
"department": None,
"attendance_status": "mentioned_only",
"notes": None,
}
]
def test_process_translates_configuration_and_disables_diarization(
tmp_path: Path,
) -> None:
service, gateway = make_service(tmp_path)
audio = tmp_path / "meeting.wav"
audio.write_bytes(b"audio")
outcome = service.process(
audio, meeting(), participants(), ProcessingOptions(diarization_enabled=False)
)
assert outcome.succeeded
assert gateway.config_values is not None
assert gateway.config_values["diarization"] == "off"
assert gateway.config_values["model"] == "test:model"
assert gateway.config_values["whisper_executable"] == "/opt/whisper-cli"
assert gateway.config_values["ffmpeg_executable"] == "ffmpeg"
assert gateway.config_values["audio_normalization"] is True
assert gateway.config_values["diarization_runtime"] == "native"
assert gateway.config_values["diarization_container_args"] == ()
assert gateway.config_values["output_root"] == (
tmp_path / "meetings" / "architecture-review" / "runs"
)
def test_process_propagates_disabled_audio_normalization(tmp_path: Path) -> None:
service, gateway = make_service(tmp_path)
audio = tmp_path / "meeting.m4a"
audio.write_bytes(b"audio")
outcome = service.process(
audio,
meeting(),
participants(),
ProcessingOptions(audio_normalization=False),
)
assert outcome.succeeded
assert gateway.config_values is not None
assert gateway.config_values["audio_normalization"] is False
def test_processing_options_default_to_audio_normalization_on() -> None:
assert ProcessingOptions().audio_normalization is True
def test_process_propagates_enabled_diarization_and_progress(tmp_path: Path) -> None:
service, gateway = make_service(tmp_path)
audio = tmp_path / "meeting.flac"
audio.write_bytes(b"audio")
events = []
service.process(
audio,
meeting(),
participants(),
ProcessingOptions(diarization_enabled=True),
progress_sink=events.append,
)
assert gateway.config_values is not None
assert gateway.config_values["diarization"] == "gpu"
assert [(event.stage, event.status) for event in events[:2]] == [
("preparing", "started"),
("preparing", "completed"),
]
assert events[2].progress == 0.25
def test_process_propagates_ordered_diarization_container_args(tmp_path: Path) -> None:
service, gateway = make_service(tmp_path)
service.settings = replace(
service.settings,
diarization_runtime="container",
diarization_container_image="runtime/image:tag",
diarization_container_args=(
"--network=host",
"--label",
"meeting-test",
),
)
audio = tmp_path / "meeting.wav"
audio.write_bytes(b"audio")
outcome = service.process(
audio,
meeting(),
participants(),
ProcessingOptions(diarization_enabled=True),
)
assert outcome.succeeded
assert gateway.config_values is not None
assert gateway.config_values["diarization_container_args"] == (
"--network=host",
"--label",
"meeting-test",
)
def test_result_and_user_edit_are_preserved_separately(tmp_path: Path) -> None:
service, _ = make_service(tmp_path)
audio = tmp_path / "meeting.wav"
audio.write_bytes(b"audio")
outcome = service.process(audio, meeting(), participants(), ProcessingOptions())
edited_path = service.save_edited_protocol(outcome.run_dir, "# Reviewed\n")
assert outcome.original_protocol == "# Generated protocol\n"
assert outcome.protocol_path.read_text(encoding="utf-8") == "# Generated protocol\n"
assert edited_path.read_text(encoding="utf-8") == "# Reviewed\n"
def test_failure_reports_stage_and_preserves_run_dir(tmp_path: Path) -> None:
service, gateway = make_service(tmp_path)
gateway.fail = True
audio = tmp_path / "meeting.wav"
audio.write_bytes(b"audio")
outcome = service.process(audio, meeting(), participants(), ProcessingOptions())
assert not outcome.succeeded
assert outcome.failed_stage == "transcription"
assert outcome.error_message == "TranscriptionError: model failed"
assert outcome.run_dir == gateway.run_dir
assert (gateway.run_dir / "run_metadata.json").is_file()
def test_uploaded_source_is_preserved_in_meeting_directory(tmp_path: Path) -> None:
service, _ = make_service(tmp_path)
source = SimpleNamespace(getbuffer=lambda: b"source audio")
destination = service.preserve_upload("meeting-1", "../unsafe.wav", source)
assert destination.parent == tmp_path / "meetings" / "meeting-1" / "uploads"
assert destination.name.endswith("_unsafe.wav")
assert destination.read_bytes() == b"source audio"
-127
View File
@@ -1,127 +0,0 @@
import pytest
from mka.application.meeting_service import ParticipantInput
from mka.application.people_yaml import (
PeopleYamlError,
export_people_yaml,
import_people_yaml,
)
def sample_people() -> list[ParticipantInput]:
return [
ParticipantInput(
participant_id="participant-35c1528c",
display_name="Martin Tazl",
role="Beirat",
organization="Verwaltungsbeirat",
attendance_status="present",
),
ParticipantInput(
participant_id="participant-56aae811",
display_name="Norbert Hebbelmann",
role="Gast",
organization="Eigentümergemeinschaft",
attendance_status="mentioned_only",
),
]
def test_export_import_round_trip_preserves_all_people_fields() -> None:
people = sample_people()
exported = export_people_yaml(people)
assert exported.startswith("version: 1\npeople:\n")
assert import_people_yaml(exported) == people
assert export_people_yaml(import_people_yaml(exported)) == exported
def test_import_preserves_stable_ids_roles_organizations_and_attendance() -> None:
imported = import_people_yaml(export_people_yaml(sample_people()))
assert [person.participant_id for person in imported] == [
"participant-35c1528c",
"participant-56aae811",
]
assert imported[0].role == "Beirat"
assert imported[1].organization == "Eigentümergemeinschaft"
assert [person.attendance_status for person in imported] == [
"present",
"mentioned_only",
]
@pytest.mark.parametrize("content", ["people: [", b"\xff\xfe"])
def test_malformed_yaml_is_rejected(content: str | bytes) -> None:
with pytest.raises(PeopleYamlError, match="Malformed People YAML"):
import_people_yaml(content)
def test_unsupported_version_is_rejected() -> None:
with pytest.raises(PeopleYamlError, match="Unsupported People YAML version"):
import_people_yaml("version: 2\npeople: []\n")
def test_missing_id_is_rejected() -> None:
content = """\
version: 1
people:
- display_name: Martin Tazl
attendance_status: present
"""
with pytest.raises(PeopleYamlError, match="participant_id"):
import_people_yaml(content)
def test_invalid_id_is_rejected_without_generating_a_replacement() -> None:
content = """\
version: 1
people:
- participant_id: invalid id
display_name: Martin
attendance_status: present
"""
with pytest.raises(PeopleYamlError, match="invalid participant_id"):
import_people_yaml(content)
def test_duplicate_id_is_rejected() -> None:
content = """\
version: 1
people:
- participant_id: martin
display_name: Martin
attendance_status: present
- participant_id: martin
display_name: Martin Duplicate
attendance_status: mentioned_only
"""
with pytest.raises(PeopleYamlError, match="Duplicate participant_id"):
import_people_yaml(content)
def test_invalid_attendance_is_rejected() -> None:
content = """\
version: 1
people:
- participant_id: martin
display_name: Martin
attendance_status: absent
"""
with pytest.raises(PeopleYamlError, match="invalid attendance_status"):
import_people_yaml(content)
def test_failed_import_does_not_mutate_existing_people() -> None:
existing = sample_people()
snapshot = list(existing)
with pytest.raises(PeopleYamlError):
import_people_yaml("version: 1\npeople: not-a-list\n")
assert existing == snapshot