28 changed files with 3656 additions and 72 deletions
+2 -1
View File
@@ -1,9 +1,10 @@
MKA_WHISPER_MODEL=/path/to/ggml-large-v3-turbo.bin MKA_WHISPER_MODEL=/path/to/ggml-large-v3-turbo.bin
MKA_WHISPER_EXECUTABLE=whisper-cli MKA_WHISPER_EXECUTABLE=whisper-cli
MKA_FFMPEG_EXECUTABLE=ffmpeg MKA_FFMPEG_EXECUTABLE=ffmpeg
MKA_PROTOCOL_MODEL=qwen3.8:27B MKA_PROTOCOL_MODEL=qwen3.8:27b
MKA_OLLAMA_ENDPOINT=http://127.0.0.1:11434 MKA_OLLAMA_ENDPOINT=http://127.0.0.1:11434
MKA_DATA_ROOT=data/meetings MKA_DATA_ROOT=data/meetings
MKA_GLOSSARY_DATABASE=data/database/glossary.sqlite3
MKA_WHISPER_THREADS=auto MKA_WHISPER_THREADS=auto
MKA_DIARIZATION_MODE=auto MKA_DIARIZATION_MODE=auto
MKA_DIARIZATION_RUNTIME=native MKA_DIARIZATION_RUNTIME=native
+10
View File
@@ -4,6 +4,14 @@
### Added ### Added
- Retained live stage and total timings for processing and protocol regeneration.
- Immutable diagnostic generation history with atomic latest-result publication.
- Auto/Fast/Efficient/Powersave protocol profiles with backend-selected threads
by default, also applied during mapped-speaker regeneration.
- Versioned UTF-8 YAML import/export for the SQLite terminology glossary, with
stable entry IDs, complete validation, and atomic replacement semantics.
- First Streamlit MVP for audio upload, Meeting Context entry, participant - First Streamlit MVP for audio upload, Meeting Context entry, participant
management, processing progress and protocol editing. management, processing progress and protocol editing.
- Application service and Meeting Lab adapter with environment-based runtime - Application service and Meeting Lab adapter with environment-based runtime
@@ -15,6 +23,8 @@
### Changed ### Changed
- Record active glossary configuration without rewriting protocol transcript input.
- Reconciled the documented MVP with the validated Meeting Lab backend. - Reconciled the documented MVP with the validated Meeting Lab backend.
- Selected `whisper.cpp` for transcription and made diarization optional. - Selected `whisper.cpp` for transcription and made diarization optional.
- Defined the Python API, progress-event and Meeting Context integration - Defined the Python API, progress-event and Meeting Context integration
+55 -2
View File
@@ -1,5 +1,14 @@
# Project Knowledge # Project Knowledge
## Meeting language
The GUI's `meeting_language` is passed through `MeetingDetails.language` to both
Meeting Context `meeting.language` and `MvpMeetingConfig.language` (Whisper).
Meeting Lab derives explicit protocol output language from the persisted context,
also during regeneration, and records derived `output_language` provenance in
protocol runtime metadata. Missing context language defaults to German. There is
no separate protocol-language setting or automatic context/transcript translation.
## Vision ## Vision
Meeting Assistant turns recorded meetings into reviewed, user-facing protocols Meeting Assistant turns recorded meetings into reviewed, user-facing protocols
@@ -37,9 +46,19 @@ Meeting Assistant owns user interaction and product workflow:
- structured Meeting Context editing - structured Meeting Context editing
- participant management - participant management
- versioned YAML import/export for reusable People lists, using replace semantics - versioned YAML import/export for reusable People lists, using replace semantics
- versioned JSON import/export for the complete user-configurable run-input
form; source media is represented only by optional filename metadata and must
be selected again after import
- optional explicit speaker mapping - optional explicit speaker mapping
- post-diarization speaker review and protocol-only regeneration from existing
run artifacts
- pipeline launch and progress display - pipeline launch and progress display
- protocol review and editing - protocol review and editing
- a global SQLite terminology glossary whose active canonical core terms and
recognition aliases are rendered through Meeting Context into direct
protocol prompts
- versioned UTF-8 YAML glossary interchange (`version: 1`) that preserves entry
IDs and atomically replaces SQLite state only after complete validation
- export and presentation of protocol versions - export and presentation of protocol versions
Processing logic must not be duplicated in the application. Processing logic must not be duplicated in the application.
@@ -62,7 +81,8 @@ Imported audio
(optionally normalize loudness; default on) (optionally normalize loudness; default on)
-> transcribe with whisper.cpp and large-v3-turbo -> transcribe with whisper.cpp and large-v3-turbo
-> optionally diarize with pyannote.audio Community-1 -> optionally diarize with pyannote.audio Community-1
-> generate a protocol directly from the full transcript -> when diarized, wait for optional speaker review before explicit protocol generation
-> otherwise generate a protocol directly from the full transcript
-> review and edit by a human -> review and edit by a human
``` ```
@@ -76,6 +96,14 @@ PyTorch/ROCm processed the same duration in approximately 98.6 seconds. These
are validation observations, not performance guarantees or hardware are validation observations, not performance guarantees or hardware
requirements. CPU execution remains supported and may be substantially slower. requirements. CPU execution remains supported and may be substantially slower.
Protocol-generation performance is selected through intentionally abstract
profiles rather than hardware controls in the GUI. `auto` is the default and
delegates thread selection to the inference backend; with Ollama this means
omitting `num_thread` entirely. The explicit resource profiles currently map
`fast` to 16 Ollama CPU threads, `efficient` to 10, and `powersave` to 4. These
concrete mappings belong to the backend/configuration boundary and may evolve
independently of the user-facing profile semantics.
## Meeting Context and Speakers ## Meeting Context and Speakers
`MeetingContext` is a structured domain object containing meeting metadata, `MeetingContext` is a structured domain object containing meeting metadata,
@@ -90,6 +118,15 @@ Speakers otherwise remain anonymous. Automatic speaker-name inference is not
allowed. The GUI must create and edit Meeting Context; hand-written YAML is not allowed. The GUI must create and edit Meeting Context; hand-written YAML is not
a product requirement. a product requirement.
Speaker selectors filter participants using a snapshot of all current widget
selections, falling back to saved mappings before first interaction. Filtering
retains existing assignments; clearing a selection releases the participant.
Completeness counts and warnings cover detected speakers only and do not gate
protocol generation or change mapping persistence. A diarized run persists the
`awaiting_speaker_review` checkpoint after diarization; an empty, partial or
complete mapping can then explicitly generate a protocol from the existing
artifacts. Failed generation preserves that checkpoint and saved mappings for retry.
## Progress Contract ## Progress Contract
Meeting Lab emits stage-based progress events for: Meeting Lab emits stage-based progress events for:
@@ -107,7 +144,7 @@ measurable progress is available.
## Protocol Policy ## Protocol Policy
Direct full-transcript protocol generation is the practical MVP direction. Direct full-transcript protocol generation is the practical MVP direction.
`qwen3.8:27B` has shown strong readability and contextual synthesis; `qwen3.8:27b` has shown strong readability and contextual synthesis;
`qwen3.6:35B-A3B` has shown somewhat more conservative behavior in some areas. `qwen3.6:35B-A3B` has shown somewhat more conservative behavior in some areas.
Experimental dual-model and diarization-assisted hard-fact extraction has not Experimental dual-model and diarization-assisted hard-fact extraction has not
demonstrated reliably better strict attribution accuracy and is not mandatory. demonstrated reliably better strict attribution accuracy and is not mandatory.
@@ -131,3 +168,19 @@ product work, not a current MVP requirement.
- no automatic identity claims - no automatic identity claims
- human review of generated protocols - human review of generated protocols
- small modules and simple interfaces - small modules and simple interfaces
Active glossary aliases are forwarded to Meeting Lab as provenance metadata.
Canonical terminology remains Meeting Context guidance; the exact protocol
transcript input and the raw Whisper/diarization artifacts are not rewritten.
`glossary_replacements` remains empty.
Processing and protocol regeneration show measured monotonic stage/total timings,
including frozen failure durations. The latest timing display survives ordinary
Streamlit reruns. Initial-run durations are saved in `run_metadata.json`;
regeneration display timings remain session-local. Worker callbacks capture UI
configuration before dispatch and never read Streamlit session state.
Meeting Lab retains immutable generation records including prompt, transcript
input, model response/metadata, glossary configuration, mappings and Meeting
Context. Latest protocol/diagnostic/context paths resolve through an atomic
`protocol/current` link. Copy whole runs with relative symlinks preserved.
+117 -2
View File
@@ -23,12 +23,20 @@ reference recorder, but imported audio is not tied to OBS-specific behavior.
## Processing Pipeline ## Processing Pipeline
The existing **Meeting language** selector controls both transcription and protocol
output: `de` means German for both, and `en` means English for both. The selection
is saved in Meeting Context (`meeting.language`) and reused for protocol-only
regeneration without retranscription. Older contexts without a language retain
German protocol output. Names, speaker mappings and authored Meeting Context are
preserved; no transcript translation stage is added.
```text ```text
Audio file Audio file
-> FFmpeg preparation (mono, 16 kHz PCM WAV; normalization optional) -> FFmpeg preparation (mono, 16 kHz PCM WAV; normalization optional)
-> whisper.cpp transcription (large-v3-turbo) -> whisper.cpp transcription (large-v3-turbo)
-> optional pyannote.audio Community-1 diarization -> optional pyannote.audio Community-1 diarization
-> direct full-transcript protocol generation -> optional speaker review and explicit protocol generation (when diarized)
or direct full-transcript protocol generation (without diarization)
-> human review and export -> human review and export
``` ```
@@ -42,6 +50,17 @@ compatibility path. Diarization is optional and produces anonymous speaker
labels. A label identifies a participant only when the user explicitly labels. A label identifies a participant only when the user explicitly
confirms the mapping; automatic speaker-name inference is not allowed. confirms the mapping; automatic speaker-name inference is not allowed.
Speaker selectors hide participants already assigned to other speakers while keeping
the current assignment available. The UI shows detected, assigned and unassigned
counts, marks unassigned speakers, and warns when mappings remain incomplete.
Anonymous-speaker protocol generation remains available.
After a diarized run, the result view lists detected `SPEAKER_XX` labels with
short transcript excerpts. Confirmed mappings regenerate only the protocol
from the existing diarized transcript; audio preparation, Whisper and Pyannote
are not rerun. `Unmapped / Unknown` remains valid, and the original anonymous
diarized transcript is preserved.
## Product Outputs ## Product Outputs
The product direction includes: The product direction includes:
@@ -64,6 +83,14 @@ Audio normalization is enabled by default and can be disabled in the processing
options. This controls loudness normalization only: Meeting Lab still prepares options. This controls loudness normalization only: Meeting Lab still prepares
every WAV, FLAC or M4A source as canonical audio before transcription. every WAV, FLAC or M4A source as canonical audio before transcription.
The **Performance profile** selector controls protocol-generation runtime using
the abstract Auto (default), Fast, Efficient, and Powersave profiles. Auto lets
the inference backend select its own thread configuration; for Ollama, Meeting
Assistant intentionally sends no `num_thread` option. Fast, Efficient, and
Powersave are explicit resource profiles currently mapped to 16, 10, and 4
Ollama CPU threads. These concrete mappings may evolve independently of the UI
semantics.
The People section can export its current entries to a UTF-8 `people.yaml` file The People section can export its current entries to a UTF-8 `people.yaml` file
and replace them from a previous `.yaml` or `.yml` export. Stable person IDs, and replace them from a previous `.yaml` or `.yml` export. Stable person IDs,
names, roles, organizations and attendance states are retained. This is a small names, roles, organizations and attendance states are retained. This is a small
@@ -94,7 +121,7 @@ load `.env` files implicitly:
export MKA_WHISPER_MODEL=/path/to/ggml-large-v3-turbo.bin export MKA_WHISPER_MODEL=/path/to/ggml-large-v3-turbo.bin
export MKA_WHISPER_EXECUTABLE=whisper-cli export MKA_WHISPER_EXECUTABLE=whisper-cli
export MKA_FFMPEG_EXECUTABLE=ffmpeg export MKA_FFMPEG_EXECUTABLE=ffmpeg
export MKA_PROTOCOL_MODEL=qwen3.8:27B export MKA_PROTOCOL_MODEL=qwen3.8:27b
``` ```
Optional machine-specific settings include: Optional machine-specific settings include:
@@ -108,6 +135,62 @@ Optional machine-specific settings include:
- `MKA_DIARIZATION_CONTAINER_IMAGE` (required for container diarization) - `MKA_DIARIZATION_CONTAINER_IMAGE` (required for container diarization)
- `MKA_DIARIZATION_CONTAINER_ARGS` (default: no extra arguments), encoded as a - `MKA_DIARIZATION_CONTAINER_ARGS` (default: no extra arguments), encoded as a
JSON array of strings so ordering and leading dashes are preserved exactly JSON array of strings so ordering and leading dashes are preserved exactly
- `MKA_GLOSSARY_DATABASE` (default: `data/database/glossary.sqlite3`), the local
SQLite file used by the global terminology glossary
## Terminology glossary
The Streamlit **Terminology glossary** section manages recurring product,
material, organization, acronym, and technical names. Store canonical core
terms such as `Secugrid HS`, not every compound such as `Secugrid HS Düse`.
Aliases help the protocol model recognize transcript variants while retaining
the surrounding wording. Only active entries are added to Meeting Context as
authoritative terminology for direct protocol generation; inactive entries
remain stored but are omitted. The production database starts empty.
The database defaults to `data/database/glossary.sqlite3`. It is created and
bootstrapped automatically and can be moved with `MKA_GLOSSARY_DATABASE`.
Glossary integration does not rewrite raw or diarized transcript artifacts.
Use **Export glossary** to download the complete SQLite-backed glossary as
UTF-8 YAML, and **Import glossary** to replace it from a previously exported
file. Imports are fully validated before a single SQLite transaction replaces
the current glossary, so malformed or conflicting files leave existing data
unchanged. Schema version 1 is:
```yaml
version: 1
glossary:
- id: 1
canonical_term: Secugrid HS
aliases:
- Sikirgut
category: product
description: Canonical product spelling
active: true
```
`id` is the stable SQLite glossary-entry identifier. Categories are `product`,
`material`, `organization`, `technical_term`, `acronym`, or `other`.
## Input configuration import and export
Use **Export inputs** to save the current Meeting Assistant run-input form as a
small, versioned JSON file and **Import inputs** to restore it later. The file
contains meeting metadata, participant records, language, date, audio
normalization, and diarization choices. It contains form configuration only:
generated prompts, protocols, run artifacts, and source-media contents are not
included.
The original source filename may be retained as a reminder, but importing does
not restore an upload. Select the audio file explicitly before starting the new
run.
Protocol generation with `qwen3.8:27b` explicitly requests a 32,768-token
Ollama context with thinking disabled; Ollama's machine default may otherwise
be only 4,096. Keep normal prompt input at approximately 29,000 tokens or less.
Although a 31,038-token synthetic prompt passed, larger input is not assumed
safe merely because the model advertises a 262,144-token native context.
For example, a compatible AMD ROCm workstation can configure validated For example, a compatible AMD ROCm workstation can configure validated
container access without adding controls to the Streamlit UI: container access without adding controls to the Streamlit UI:
@@ -130,6 +213,22 @@ From the Meeting Assistant repository, start the UI with:
PYTHONPATH=src:../meeting-lab .venv/bin/streamlit run src/mka/ui/streamlit_app.py PYTHONPATH=src:../meeting-lab .venv/bin/streamlit run src/mka/ui/streamlit_app.py
``` ```
### North launcher
In the post-diarization Assistant worktree on North, use the checked-in
launcher instead of creating another virtual environment:
```bash
./run_north.sh
```
It reuses North's validated Assistant virtual environment and external
Whisper/Ollama/Docker-ROCm runtime while loading Meeting Assistant from this
worktree and Meeting Lab from
`/opt/git-projekts/meeting-lab`. `HF_TOKEN`, when present, is
passed through to the disposable diarization container without being stored by
the launcher.
Streamlit opens `http://localhost:8501` by default. Installing the sibling Streamlit opens `http://localhost:8501` by default. Installing the sibling
Meeting Lab project supplies its runtime requirements such as PyYAML and Meeting Lab project supplies its runtime requirements such as PyYAML and
Requests. Optional diarization dependencies are needed only when diarization Requests. Optional diarization dependencies are needed only when diarization
@@ -147,3 +246,19 @@ transcription; Meeting Assistant does not duplicate audio conversion.
See [Architecture](docs/architecture.md), [Project Knowledge](PROJECT_KNOWLEDGE.md), See [Architecture](docs/architecture.md), [Project Knowledge](PROJECT_KNOWLEDGE.md),
[Roadmap](ROADMAP.md) and [ADR 0011](docs/adr/0011-use-meeting-lab-mvp-backend.md). [Roadmap](ROADMAP.md) and [ADR 0011](docs/adr/0011-use-meeting-lab-mvp-backend.md).
Active glossary aliases are forwarded to Meeting Lab as provenance metadata.
Canonical terminology remains Meeting Context guidance; the exact protocol
transcript input and the raw Whisper/diarization artifacts are not rewritten.
`glossary_replacements` remains empty.
Processing and protocol regeneration show measured monotonic stage/total timings,
including frozen failure durations. The latest timing display survives ordinary
Streamlit reruns. Initial-run durations are saved in `run_metadata.json`;
regeneration display timings remain session-local. Worker callbacks capture UI
configuration before dispatch and never read Streamlit session state.
Meeting Lab retains immutable generation records including prompt, transcript
input, model response/metadata, glossary configuration, mappings and Meeting
Context. Latest protocol/diagnostic/context paths resolve through an atomic
`protocol/current` link. Copy whole runs with relative symlinks preserved.
+40
View File
@@ -0,0 +1,40 @@
# Alpha paired validation
Run the complete Assistant suite with `PYTHONPATH=src:<Lab candidate root>`.
`tests/test_alpha_pair.py` exercises the real adapter, context, orchestration,
transcript selection and generation persistence with mocked audio preparation,
Whisper, diarization and model responses. It covers de/en, all performance
profiles, anonymous generation, mapped regeneration, glossary provenance and
immutable history. The test skips when Lab is unavailable; a release validation
must run it with Lab available and report no skips.
Run Lab's full `python -m unittest discover -s tests` at its candidate revision.
History tests inject write/publication failures and a process interruption.
UI tests use the real executor and reject worker access to session state.
These tests do not measure recognition quality or model language compliance.
Before tagging, perform the final local GUI/model smoke with the exact paired
revisions and configured model/runtime. Reuse copied historical transcripts
where possible; do not rewrite original regression evidence. Preserve the known
GTM-Hub human-reference whitespace. Record both candidate SHAs in release notes.
## Historical Lab fixture dependency
Four older experiment tests require six ignored files, absent from a fresh Git
worktree. Do not add them to the alpha source boundary or silently skip the tests:
- `artifacts/experiments/evidence_observations_v3/20260819_v3_single_run/h_resulting_action/parsed_observations.json`
- `artifacts/experiments/negative_act_form_v0/20260820_qwen35_9b_single_run/na-01/v3_style_input_observations.json`
- The same negative-act filename under `na-03`, `na-04`, `na-05`, and `na-06`.
For the alpha audit, these files were copied from the original Lab worktree to a
separate temporary test directory. That directory exposed the candidate's
`tests`, `prompts`, `samples`, `src`, and `scripts` as resource links. The full
suite ran there with the candidate Lab root on `PYTHONPATH` and the candidate's
absolute tests directory supplied to unittest discovery. All 413 tests passed.
The bare candidate worktree run instead reports four missing-fixture errors.
This is a historical test portability limitation, not a runtime dependency.
The paired Assistant suite passed all 107 tests with the candidate Lab backend;
no Whisper, diarization or Ollama inference was executed. These counts describe
the preparation validation and do not replace the final smoke-test record.
+39
View File
@@ -0,0 +1,39 @@
# Meeting Assistant — First functional alpha
Version: `0.1.0a1`
Tag: `v0.1.0-alpha.1`
This is the first functional end-to-end alpha of the local Meeting Assistant.
It provides local meeting configuration and audio workflow, FFmpeg preparation
and normalization, Whisper transcription, optional diarization, explicit
speaker mapping, anonymous-speaker protocol generation, mapped-speaker
regeneration, German and English meeting-language handling, glossary handling
with YAML interchange, glossary provenance without transcript mutation,
configurable protocol performance profiles, processing timings, generation
history and provenance, and atomic publication of complete successful protocol
generations. The GTM-Hub qualitative regression case is included.
The validated candidate revisions for this release were:
- Assistant: `b50b87c32b3941310c719f459fd45faa87658b68`
- Meeting Lab: `d2e3636b949280402728308022c99bdee6ea1066`
The immutable release revisions are the commits targeted by the paired
`v0.1.0-alpha.1` tags in the two repositories.
## Accepted alpha limitations
- A local runtime, model, and container setup is required.
- Generated protocols require human review, especially responsibility and
action attribution; speaker identity requires explicit confirmation.
- Diarization can be imperfect. The final native smoke environment did not
contain `pyannote.audio==4.0.7` and PyTorch, so real diarization and mapped
speaker regeneration were not rerun there; automated tests cover these paths.
- Real German and English `qwen3.8:27b` protocol generation was validated.
- Context limits can cause attribution fallback or oversized-input rejection.
- Historical ignored Lab fixtures are required by some old experiment tests and
are not automatically present in a completely fresh Lab worktree.
- Generation publication assumes a local POSIX filesystem; power-loss
durability is not guaranteed.
- Visual redesign is post-alpha, and concise distribution-protocol rendering
remains future work.
+1 -1
View File
@@ -4,7 +4,7 @@ build-backend = "hatchling.build"
[project] [project]
name = "meeting-assistant" name = "meeting-assistant"
version = "0.1.0" version = "0.1.0a1"
description = "Local-first meeting transcription and knowledge capture assistant" description = "Local-first meeting transcription and knowledge capture assistant"
readme = "README.md" readme = "README.md"
requires-python = ">=3.11" requires-python = ">=3.11"
Executable
+57
View File
@@ -0,0 +1,57 @@
#!/usr/bin/env bash
set -euo pipefail
assistant_root="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd -P)"
assistant_src="$assistant_root/src"
lab_root=/opt/git-projekts/meeting-lab
north_venv=/opt/git-projekts/meeting-assistant/.venv
whisper_executable=/home/martin/whisper.cpp/whisper.cpp/build/bin/whisper-cli
whisper_model=/home/martin/whisper.cpp/whisper.cpp/models/ggml-large-v3-turbo.bin
require_directory() {
if [[ ! -d "$1" ]]; then
printf 'Required directory is missing: %s\n' "$1" >&2
exit 1
fi
}
require_file() {
if [[ ! -f "$1" ]]; then
printf 'Required file is missing: %s\n' "$1" >&2
exit 1
fi
}
require_executable() {
if [[ ! -x "$1" ]]; then
printf 'Required executable is missing or not executable: %s\n' "$1" >&2
exit 1
fi
}
require_directory "$assistant_src"
require_file "$assistant_src/mka/ui/streamlit_app.py"
require_directory "$lab_root"
require_file "$lab_root/src/meeting_lab/__init__.py"
require_directory "$north_venv"
require_executable "$north_venv/bin/streamlit"
require_executable "$whisper_executable"
require_file "$whisper_model"
if ! command -v docker >/dev/null 2>&1; then
printf 'Required executable is not on PATH: docker\n' >&2
exit 1
fi
export MKA_WHISPER_EXECUTABLE="$whisper_executable"
export MKA_WHISPER_MODEL="$whisper_model"
export MKA_PROTOCOL_MODEL=qwen3.8:27b
export MKA_OLLAMA_ENDPOINT=http://127.0.0.1:11434
export MKA_DIARIZATION_MODE=auto
export MKA_DIARIZATION_RUNTIME=container
export MKA_DIARIZATION_CONTAINER_IMAGE=rocm/pytorch:rocm7.2.1_ubuntu24.04_py3.12_pytorch_release_2.9.1
export MKA_DIARIZATION_CONTAINER_ARGS='["--device=/dev/kfd","--device=/dev/dri","--group-add","video"]'
# Keep feature sources ahead of the reused venv's editable installations.
export PYTHONPATH="$assistant_src:$lab_root"
exec "$north_venv/bin/streamlit" run "$assistant_src/mka/ui/streamlit_app.py"
+8 -2
View File
@@ -18,10 +18,13 @@ class AppSettings:
data_root: Path data_root: Path
whisper_model: Path | None whisper_model: Path | None
glossary_database: Path = Path("data/database/glossary.sqlite3")
whisper_executable: str = "whisper-cli" whisper_executable: str = "whisper-cli"
ffmpeg_executable: str = "ffmpeg" ffmpeg_executable: str = "ffmpeg"
protocol_model: str = "qwen3.8:27B" protocol_model: str = "qwen3.8:27b"
ollama_endpoint: str = "http://127.0.0.1:11434" ollama_endpoint: str = "http://127.0.0.1:11434"
protocol_num_ctx: int = 32_768
protocol_safe_input_token_budget: int = 29_000
language: str = "de" language: str = "de"
threads: str | int = "auto" threads: str | int = "auto"
diarization_mode: str = "auto" diarization_mode: str = "auto"
@@ -38,9 +41,12 @@ class AppSettings:
return cls( return cls(
data_root=Path(os.getenv("MKA_DATA_ROOT", "data/meetings")), data_root=Path(os.getenv("MKA_DATA_ROOT", "data/meetings")),
whisper_model=Path(whisper_model) if whisper_model else None, whisper_model=Path(whisper_model) if whisper_model else None,
glossary_database=Path(
os.getenv("MKA_GLOSSARY_DATABASE", "data/database/glossary.sqlite3")
),
whisper_executable=os.getenv("MKA_WHISPER_EXECUTABLE", "whisper-cli"), whisper_executable=os.getenv("MKA_WHISPER_EXECUTABLE", "whisper-cli"),
ffmpeg_executable=os.getenv("MKA_FFMPEG_EXECUTABLE", "ffmpeg"), ffmpeg_executable=os.getenv("MKA_FFMPEG_EXECUTABLE", "ffmpeg"),
protocol_model=os.getenv("MKA_PROTOCOL_MODEL", "qwen3.8:27B"), protocol_model=os.getenv("MKA_PROTOCOL_MODEL", "qwen3.8:27b"),
ollama_endpoint=os.getenv("MKA_OLLAMA_ENDPOINT", "http://127.0.0.1:11434"), ollama_endpoint=os.getenv("MKA_OLLAMA_ENDPOINT", "http://127.0.0.1:11434"),
language=os.getenv("MKA_LANGUAGE", "de"), language=os.getenv("MKA_LANGUAGE", "de"),
threads=parsed_threads, threads=parsed_threads,
+369
View File
@@ -0,0 +1,369 @@
"""SQLite-backed global terminology glossary."""
from __future__ import annotations
import sqlite3
from collections.abc import Iterable, Iterator
from contextlib import contextmanager
from dataclasses import dataclass
from datetime import UTC, datetime
from pathlib import Path
GLOSSARY_CATEGORIES = (
"product",
"material",
"organization",
"technical_term",
"acronym",
"other",
)
class GlossaryConflictError(ValueError):
"""Raised when a canonical term or alias conflicts with existing terminology."""
@dataclass(frozen=True)
class GlossaryEntry:
id: int
canonical_term: str
category: str
description: str | None
is_active: bool
aliases: tuple[str, ...]
created_at: str
updated_at: str
@dataclass(frozen=True)
class GlossaryReplacementEntry:
"""Validated values used to atomically replace the persisted glossary."""
id: int
canonical_term: str
category: str
description: str | None
is_active: bool
aliases: tuple[str, ...]
class GlossaryRepository:
"""Small data-access boundary for the local glossary database."""
def __init__(self, database_path: Path) -> None:
self.database_path = Path(database_path)
def initialize(self) -> None:
"""Create the database and current schema when absent."""
self.database_path.parent.mkdir(parents=True, exist_ok=True)
with self._connect() as connection:
connection.executescript(
"""
CREATE TABLE IF NOT EXISTS glossary_entries (
id INTEGER PRIMARY KEY,
canonical_term TEXT NOT NULL COLLATE NOCASE UNIQUE,
category TEXT NOT NULL CHECK (category IN (
'product', 'material', 'organization',
'technical_term', 'acronym', 'other'
)),
description TEXT,
is_active INTEGER NOT NULL DEFAULT 1 CHECK (is_active IN (0, 1)),
created_at TEXT NOT NULL,
updated_at TEXT NOT NULL
);
CREATE TABLE IF NOT EXISTS glossary_aliases (
id INTEGER PRIMARY KEY,
entry_id INTEGER NOT NULL REFERENCES glossary_entries(id)
ON DELETE CASCADE,
alias TEXT NOT NULL COLLATE NOCASE UNIQUE,
created_at TEXT NOT NULL
);
CREATE INDEX IF NOT EXISTS idx_glossary_aliases_entry_id
ON glossary_aliases(entry_id);
PRAGMA user_version = 1;
"""
)
def create(
self,
canonical_term: str,
category: str,
*,
aliases: Iterable[str] = (),
description: str | None = None,
is_active: bool = True,
) -> GlossaryEntry:
canonical, normalized_aliases = self._validate_values(canonical_term, category, aliases)
now = _timestamp()
try:
with self._connect() as connection:
self._ensure_terms_available(connection, canonical, normalized_aliases)
cursor = connection.execute(
"""INSERT INTO glossary_entries
(canonical_term, category, description, is_active, created_at, updated_at)
VALUES (?, ?, ?, ?, ?, ?)""",
(canonical, category, _optional_text(description), is_active, now, now),
)
entry_id = int(cursor.lastrowid)
connection.executemany(
"INSERT INTO glossary_aliases (entry_id, alias, created_at) VALUES (?, ?, ?)",
((entry_id, alias, now) for alias in normalized_aliases),
)
except sqlite3.IntegrityError as exc:
raise GlossaryConflictError("Canonical term or alias already exists.") from exc
return self.get(entry_id)
def get(self, entry_id: int) -> GlossaryEntry:
with self._connect() as connection:
row = connection.execute(
"SELECT * FROM glossary_entries WHERE id = ?", (entry_id,)
).fetchone()
if row is None:
raise KeyError(f"Unknown glossary entry: {entry_id}")
return self._to_entry(connection, row)
def list(self, search: str = "", *, active_only: bool = False) -> list[GlossaryEntry]:
clauses: list[str] = []
parameters: list[object] = []
if active_only:
clauses.append("entry.is_active = 1")
if search.strip():
clauses.append(
"(entry.canonical_term LIKE ? COLLATE NOCASE "
"OR entry.category LIKE ? COLLATE NOCASE "
"OR entry.description LIKE ? COLLATE NOCASE "
"OR alias.alias LIKE ? COLLATE NOCASE)"
)
pattern = f"%{search.strip()}%"
parameters.extend([pattern] * 4)
where = f"WHERE {' AND '.join(clauses)}" if clauses else ""
with self._connect() as connection:
rows = connection.execute(
f"""SELECT DISTINCT entry.* FROM glossary_entries AS entry
LEFT JOIN glossary_aliases AS alias ON alias.entry_id = entry.id
{where} ORDER BY entry.canonical_term COLLATE NOCASE""", # noqa: S608
parameters,
).fetchall()
return [self._to_entry(connection, row) for row in rows]
def update(
self,
entry_id: int,
canonical_term: str,
category: str,
*,
aliases: Iterable[str] = (),
description: str | None = None,
is_active: bool = True,
) -> GlossaryEntry:
canonical, normalized_aliases = self._validate_values(canonical_term, category, aliases)
try:
with self._connect() as connection:
exists = connection.execute(
"SELECT 1 FROM glossary_entries WHERE id = ?", (entry_id,)
).fetchone()
if exists is None:
raise KeyError(f"Unknown glossary entry: {entry_id}")
self._ensure_terms_available(
connection, canonical, normalized_aliases, excluding_entry_id=entry_id
)
now = _timestamp()
connection.execute(
"""UPDATE glossary_entries SET canonical_term = ?, category = ?,
description = ?, is_active = ?, updated_at = ? WHERE id = ?""",
(
canonical,
category,
_optional_text(description),
is_active,
now,
entry_id,
),
)
connection.execute("DELETE FROM glossary_aliases WHERE entry_id = ?", (entry_id,))
connection.executemany(
"INSERT INTO glossary_aliases (entry_id, alias, created_at) VALUES (?, ?, ?)",
((entry_id, alias, now) for alias in normalized_aliases),
)
except sqlite3.IntegrityError as exc:
raise GlossaryConflictError("Canonical term or alias already exists.") from exc
return self.get(entry_id)
def set_active(self, entry_id: int, is_active: bool) -> GlossaryEntry:
entry = self.get(entry_id)
return self.update(
entry.id,
entry.canonical_term,
entry.category,
aliases=entry.aliases,
description=entry.description,
is_active=is_active,
)
def delete(self, entry_id: int) -> None:
with self._connect() as connection:
cursor = connection.execute("DELETE FROM glossary_entries WHERE id = ?", (entry_id,))
if cursor.rowcount == 0:
raise KeyError(f"Unknown glossary entry: {entry_id}")
def replace_all(self, entries: Iterable[GlossaryReplacementEntry]) -> None:
"""Atomically replace every glossary entry, preserving supplied stable IDs."""
replacements = tuple(entries)
validated: list[GlossaryReplacementEntry] = []
seen_ids: set[int] = set()
seen_terms: set[str] = set()
for entry in replacements:
if type(entry.id) is not int or entry.id <= 0:
raise ValueError("Glossary entry IDs must be positive integers.")
if entry.id in seen_ids:
raise GlossaryConflictError(f"Duplicate glossary entry ID: {entry.id}.")
canonical, aliases = self._validate_values(
entry.canonical_term, entry.category, entry.aliases
)
folded_terms = {canonical.casefold(), *(alias.casefold() for alias in aliases)}
if seen_terms.intersection(folded_terms):
raise GlossaryConflictError("Canonical term or alias already exists.")
seen_ids.add(entry.id)
seen_terms.update(folded_terms)
validated.append(
GlossaryReplacementEntry(
id=entry.id,
canonical_term=canonical,
category=entry.category,
description=_optional_text(entry.description),
is_active=entry.is_active,
aliases=aliases,
)
)
now = _timestamp()
try:
with self._connect() as connection:
connection.execute("DELETE FROM glossary_entries")
connection.executemany(
"""INSERT INTO glossary_entries
(id, canonical_term, category, description, is_active, created_at, updated_at)
VALUES (?, ?, ?, ?, ?, ?, ?)""",
(
(
entry.id,
entry.canonical_term,
entry.category,
entry.description,
entry.is_active,
now,
now,
)
for entry in validated
),
)
connection.executemany(
"INSERT INTO glossary_aliases (entry_id, alias, created_at) VALUES (?, ?, ?)",
((entry.id, alias, now) for entry in validated for alias in entry.aliases),
)
except sqlite3.IntegrityError as exc:
raise GlossaryConflictError("Canonical term or alias already exists.") from exc
@contextmanager
def _connect(self) -> Iterator[sqlite3.Connection]:
connection = sqlite3.connect(self.database_path, timeout=5)
try:
connection.row_factory = sqlite3.Row
connection.execute("PRAGMA foreign_keys = ON")
connection.execute("PRAGMA journal_mode = WAL")
connection.execute("PRAGMA busy_timeout = 5000")
with connection:
yield connection
finally:
connection.close()
@staticmethod
def _validate_values(
canonical_term: str, category: str, aliases: Iterable[str]
) -> tuple[str, tuple[str, ...]]:
canonical = canonical_term.strip()
if not canonical:
raise ValueError("Canonical term is required.")
if category not in GLOSSARY_CATEGORIES:
raise ValueError(f"Unsupported glossary category: {category}")
normalized_aliases = tuple(
dict.fromkeys(alias.strip() for alias in aliases if alias.strip())
)
folded = [alias.casefold() for alias in normalized_aliases]
if len(folded) != len(set(folded)) or canonical.casefold() in folded:
raise GlossaryConflictError(
"Aliases must be unique and differ from the canonical term."
)
return canonical, normalized_aliases
@staticmethod
def _ensure_terms_available(
connection: sqlite3.Connection,
canonical: str,
aliases: tuple[str, ...],
*,
excluding_entry_id: int | None = None,
) -> None:
terms = (canonical, *aliases)
placeholders = ", ".join("?" for _ in terms)
exclusion = "AND entry_id != ?" if excluding_entry_id is not None else ""
alias_parameters: list[object] = [*terms]
if excluding_entry_id is not None:
alias_parameters.append(excluding_entry_id)
alias_conflict = connection.execute(
f"SELECT 1 FROM glossary_aliases WHERE alias IN ({placeholders}) {exclusion} LIMIT 1", # noqa: S608
alias_parameters,
).fetchone()
entry_exclusion = "AND id != ?" if excluding_entry_id is not None else ""
entry_parameters: list[object] = [*terms]
if excluding_entry_id is not None:
entry_parameters.append(excluding_entry_id)
canonical_conflict = connection.execute(
f"SELECT 1 FROM glossary_entries WHERE canonical_term IN ({placeholders}) " # noqa: S608
f"{entry_exclusion} LIMIT 1",
entry_parameters,
).fetchone()
if alias_conflict or canonical_conflict:
raise GlossaryConflictError("Canonical term or alias already exists.")
@staticmethod
def _to_entry(connection: sqlite3.Connection, row: sqlite3.Row) -> GlossaryEntry:
aliases = connection.execute(
"SELECT alias FROM glossary_aliases WHERE entry_id = ? ORDER BY alias COLLATE NOCASE",
(row["id"],),
).fetchall()
return GlossaryEntry(
id=row["id"],
canonical_term=row["canonical_term"],
category=row["category"],
description=row["description"],
is_active=bool(row["is_active"]),
aliases=tuple(alias["alias"] for alias in aliases),
created_at=row["created_at"],
updated_at=row["updated_at"],
)
def render_glossary_terms(entries: Iterable[GlossaryEntry]) -> list[str]:
"""Render concise canonical terms with recognition aliases for Meeting Context."""
rendered = []
for entry in entries:
item = entry.canonical_term
if entry.aliases:
item += f" (aliases: {', '.join(entry.aliases)})"
rendered.append(item)
return rendered
def glossary_alias_mapping(entries: Iterable[GlossaryEntry]) -> dict[str, str]:
"""Return explicit alias configuration for provenance, not substitution."""
return {alias: entry.canonical_term for entry in entries for alias in entry.aliases}
def _timestamp() -> str:
return datetime.now(UTC).isoformat(timespec="seconds")
def _optional_text(value: str | None) -> str | None:
stripped = value.strip() if value else ""
return stripped or None
+122
View File
@@ -0,0 +1,122 @@
"""Versioned UTF-8 YAML import and export for the terminology glossary."""
from __future__ import annotations
from collections.abc import Sequence
import yaml
from mka.application.glossary import (
GLOSSARY_CATEGORIES,
GlossaryEntry,
GlossaryReplacementEntry,
)
GLOSSARY_YAML_VERSION = 1
class GlossaryYamlError(ValueError):
"""Raised when a glossary YAML document is malformed or invalid."""
def export_glossary_yaml(entries: Sequence[GlossaryEntry]) -> str:
"""Serialize all glossary state in a deterministic, versioned document."""
document = {
"version": GLOSSARY_YAML_VERSION,
"glossary": [
{
"id": entry.id,
"canonical_term": entry.canonical_term,
"aliases": list(entry.aliases),
"category": entry.category,
"description": entry.description,
"active": entry.is_active,
}
for entry in entries
],
}
return yaml.safe_dump(document, allow_unicode=True, sort_keys=False, default_flow_style=False)
def import_glossary_yaml(content: str | bytes) -> list[GlossaryReplacementEntry]:
"""Parse and completely validate a glossary replacement document."""
try:
if isinstance(content, bytes):
content = content.decode("utf-8")
document = yaml.safe_load(content)
except (yaml.YAMLError, UnicodeDecodeError) as exc:
raise GlossaryYamlError(f"Malformed glossary YAML: {exc}") from exc
if not isinstance(document, dict):
raise GlossaryYamlError("Glossary YAML must contain a top-level mapping.")
version = document.get("version")
if type(version) is not int or version != GLOSSARY_YAML_VERSION:
raise GlossaryYamlError(
f"Unsupported glossary YAML version {version!r}; expected {GLOSSARY_YAML_VERSION}."
)
raw_entries = document.get("glossary")
if not isinstance(raw_entries, list):
raise GlossaryYamlError("Glossary YAML must contain a top-level 'glossary' list.")
result: list[GlossaryReplacementEntry] = []
seen_ids: set[int] = set()
seen_terms: set[str] = set()
for index, raw in enumerate(raw_entries, start=1):
if not isinstance(raw, dict):
raise GlossaryYamlError(f"Glossary entry {index} must be a mapping.")
entry_id = raw.get("id")
if type(entry_id) is not int or entry_id <= 0:
raise GlossaryYamlError(f"Glossary entry {index} requires a positive integer id.")
if entry_id in seen_ids:
raise GlossaryYamlError(f"Duplicate glossary entry id: {entry_id}.")
seen_ids.add(entry_id)
canonical = _required_text(raw, "canonical_term", index).strip()
category = raw.get("category")
if category not in GLOSSARY_CATEGORIES:
allowed = ", ".join(GLOSSARY_CATEGORIES)
raise GlossaryYamlError(
f"Glossary entry {index} has invalid category; expected one of: {allowed}."
)
aliases_value = raw.get("aliases")
if not isinstance(aliases_value, list) or any(
not isinstance(alias, str) or not alias.strip() for alias in aliases_value
):
raise GlossaryYamlError(f"Glossary entry {index} aliases must be a list of text.")
aliases = tuple(alias.strip() for alias in aliases_value)
folded = [canonical.casefold(), *(alias.casefold() for alias in aliases)]
if len(folded) != len(set(folded)):
raise GlossaryYamlError(
f"Glossary entry {index} aliases must be unique and differ from its canonical term."
)
duplicate = next((term for term in folded if term in seen_terms), None)
if duplicate is not None:
raise GlossaryYamlError(
f"Glossary entry {index} contains a duplicate canonical term or alias."
)
seen_terms.update(folded)
description = raw.get("description")
if description is not None and not isinstance(description, str):
raise GlossaryYamlError(f"Glossary entry {index} description must be text or null.")
active = raw.get("active")
if type(active) is not bool:
raise GlossaryYamlError(f"Glossary entry {index} active must be true or false.")
result.append(
GlossaryReplacementEntry(
id=entry_id,
canonical_term=canonical,
aliases=aliases,
category=category,
description=description,
is_active=active,
)
)
return result
def _required_text(entry: dict[object, object], field: str, index: int) -> str:
value = entry.get(field)
if not isinstance(value, str) or not value.strip():
raise GlossaryYamlError(f"Glossary entry {index} requires a non-empty {field}.")
return value
+482 -2
View File
@@ -12,9 +12,49 @@ from pathlib import Path
from typing import Any, Protocol from typing import Any, Protocol
from uuid import uuid4 from uuid import uuid4
import yaml
from mka.application.config import AppSettings from mka.application.config import AppSettings
from mka.application.glossary import (
GlossaryRepository,
glossary_alias_mapping,
render_glossary_terms,
)
from mka.application.performance import DEFAULT_PERFORMANCE_PROFILE, resolve_ollama_num_thread
from mka.application.progress_timing import ProcessingTimer
STAGES = ("preparing", "transcription", "diarization", "protocol_generation") STAGES = ("preparing", "transcription", "diarization", "protocol_generation")
_EXCERPT_TOKEN_PATTERN = re.compile(r"[^\W\d_]+|\d+", re.UNICODE)
_EXCERPT_STOP_WORDS = frozenset(
{
"a", "an", "and", "are", "das", "der", "die", "ein", "eine", "for", "i",
"ich", "im", "in", "is", "it", "ja", "mit", "of", "or", "the", "to", "und",
"we", "wir", "with",
}
)
_ACKNOWLEDGEMENT_WORDS = frozenset(
{
"absolutely", "alles", "danke", "genau", "gut", "ja", "klar", "okay", "ok",
"perfekt", "right", "sure", "thanks", "verstanden", "yes",
}
)
_FIRST_PERSON_WORDS = frozenset(
{"i", "ich", "me", "mein", "meine", "my", "unser", "unsere", "we", "wir"}
)
_ACTION_WORDS = frozenset(
{
"agree", "approve", "believe", "can", "decide", "habe", "kann", "möchte", "plan",
"recommend", "think", "übernehme", "werde", "will",
}
)
_DATE_WORDS = frozenset(
{
"april", "august", "december", "dienstag", "donnerstag", "february", "freitag",
"friday", "january", "juli", "june", "mai", "march", "monday", "mittwoch", "montag",
"november", "october", "saturday", "september", "sonntag", "sunday", "thursday",
"tuesday", "wednesday",
}
)
class MeetingLabPort(Protocol): class MeetingLabPort(Protocol):
@@ -31,6 +71,14 @@ class MeetingLabPort(Protocol):
progress_sink: Callable[[Any], None], progress_sink: Callable[[Any], None],
) -> Any: ... ) -> Any: ...
def regenerate_protocol(
self,
run_dir: Path,
meeting_context: Any,
progress_sink: Callable[[Any], None],
**options: Any,
) -> Any: ...
@dataclass(frozen=True) @dataclass(frozen=True)
class MeetingDetails: class MeetingDetails:
@@ -54,6 +102,7 @@ class ParticipantInput:
class ProcessingOptions: class ProcessingOptions:
diarization_enabled: bool = False diarization_enabled: bool = False
audio_normalization: bool = True audio_normalization: bool = True
performance_profile: str = DEFAULT_PERFORMANCE_PROFILE
@dataclass(frozen=True) @dataclass(frozen=True)
@@ -73,6 +122,105 @@ class ProcessingOutcome:
protocol_path: Path | None protocol_path: Path | None
failed_stage: str | None = None failed_stage: str | None = None
error_message: str | None = None error_message: str | None = None
speaker_attribution_available: bool | None = None
awaiting_speaker_review: bool = False
detected_speaker_count: int | None = None
@dataclass(frozen=True)
class SpeakerReview:
speaker_label: str
excerpts: tuple[str, ...]
@dataclass(frozen=True)
class SpeakerMappingReview:
speakers: tuple[SpeakerReview, ...]
participants: tuple[tuple[str, str], ...]
current_mappings: dict[str, str]
@dataclass(frozen=True)
class _ExcerptCandidate:
speaker_label: str
text: str
segment_index: int
tokens: frozenset[str]
lexical_tokens: frozenset[str]
previous_other_speaker_tokens: frozenset[str]
def _excerpt_tokens(text: str) -> tuple[str, ...]:
return tuple(token.lower() for token in _EXCERPT_TOKEN_PATTERN.findall(text))
def _excerpt_score(candidate: _ExcerptCandidate, distinctive_token_count: int) -> float:
"""Score an excerpt using deterministic lexical signals only."""
word_count = len(candidate.tokens)
lexical_count = len(candidate.lexical_tokens)
score = min(word_count, 24) * 0.3 + min(lexical_count, 16) * 0.6
if word_count < 5:
score -= 8.0
elif word_count < 9:
score -= 3.0
if lexical_count <= 2:
score -= 4.0
if candidate.tokens and candidate.tokens <= _ACKNOWLEDGEMENT_WORDS:
score -= 14.0
if candidate.text.rstrip().endswith("?"):
score -= 5.0
if any(token.isdigit() for token in candidate.tokens):
score += 2.5
if candidate.tokens & _DATE_WORDS:
score += 2.0
if candidate.tokens & _FIRST_PERSON_WORDS:
score += 1.5
if candidate.tokens & _ACTION_WORDS:
score += 2.0
proper_name_count = len(re.findall(r"\b[A-ZÄÖÜ][a-zäöüß]{2,}\b", candidate.text))
score += min(proper_name_count, 3) * 0.75
score += min(sum(len(token) >= 10 for token in candidate.lexical_tokens), 4) * 0.4
score += min(distinctive_token_count, 6) * 0.65
if candidate.previous_other_speaker_tokens and candidate.lexical_tokens:
overlap = len(candidate.lexical_tokens & candidate.previous_other_speaker_tokens)
if overlap / len(candidate.lexical_tokens) >= 0.7:
score -= 5.0
return score
def _segment_bucket(segment_index: int, segment_count: int) -> int:
return min(5, segment_index * 6 // max(segment_count, 1))
def _selection_score(
candidate: _ExcerptCandidate,
base_score: float,
selected: list[_ExcerptCandidate],
selected_buckets: set[int],
segment_count: int,
) -> float:
"""Favor excerpts from new meeting positions and avoid near duplicates."""
score = base_score
bucket = _segment_bucket(candidate.segment_index, segment_count)
if bucket not in selected_buckets:
score += 2.5
if not selected or not candidate.lexical_tokens:
return score
similarities = [
len(candidate.lexical_tokens & item.lexical_tokens)
/ len(candidate.lexical_tokens | item.lexical_tokens)
for item in selected
if item.lexical_tokens
]
maximum_similarity = max(similarities, default=0.0)
if maximum_similarity >= 0.72:
score -= 18.0
elif maximum_similarity >= 0.55:
score -= 6.0
nearest_position = min(abs(candidate.segment_index - item.segment_index) for item in selected)
if nearest_position / max(segment_count - 1, 1) < 0.12:
score -= 1.5
return score
def stable_id(value: str) -> str: def stable_id(value: str) -> str:
@@ -84,9 +232,18 @@ def stable_id(value: str) -> str:
class MeetingProcessingService: class MeetingProcessingService:
"""Translate UI input into Meeting Lab calls and persisted artifacts.""" """Translate UI input into Meeting Lab calls and persisted artifacts."""
def __init__(self, settings: AppSettings, meeting_lab: MeetingLabPort) -> None: def __init__(
self,
settings: AppSettings,
meeting_lab: MeetingLabPort,
glossary: GlossaryRepository | None = None,
timer: ProcessingTimer | None = None,
) -> None:
self.settings = settings self.settings = settings
self.meeting_lab = meeting_lab self.meeting_lab = meeting_lab
self.glossary = glossary or GlossaryRepository(settings.glossary_database)
self.glossary.initialize()
self.timer = timer or ProcessingTimer()
def build_context( def build_context(
self, self,
@@ -154,6 +311,7 @@ class MeetingProcessingService:
"mentioned_people_are_not_participants": True, "mentioned_people_are_not_participants": True,
}, },
} }
self._apply_glossary(data)
return self.meeting_lab.create_context(data) return self.meeting_lab.create_context(data)
def preserve_upload(self, meeting_id: str, filename: str, source: Any) -> Path: def preserve_upload(self, meeting_id: str, filename: str, source: Any) -> Path:
@@ -195,20 +353,29 @@ class MeetingProcessingService:
"threads": self.settings.threads, "threads": self.settings.threads,
"model": self.settings.protocol_model, "model": self.settings.protocol_model,
"ollama_endpoint": self.settings.ollama_endpoint, "ollama_endpoint": self.settings.ollama_endpoint,
"protocol_num_ctx": self.settings.protocol_num_ctx,
"protocol_safe_input_token_budget": (
self.settings.protocol_safe_input_token_budget
),
"protocol_num_thread": resolve_ollama_num_thread(options.performance_profile),
"glossary_aliases": glossary_alias_mapping(self.glossary.list(active_only=True)),
"diarization": ( "diarization": (
self.settings.diarization_mode if options.diarization_enabled else "off" self.settings.diarization_mode if options.diarization_enabled else "off"
), ),
"diarization_runtime": self.settings.diarization_runtime, "diarization_runtime": self.settings.diarization_runtime,
"diarization_container_image": (self.settings.diarization_container_image), "diarization_container_image": (self.settings.diarization_container_image),
"diarization_container_args": self.settings.diarization_container_args, "diarization_container_args": self.settings.diarization_container_args,
"stop_after_diarization": options.diarization_enabled,
} }
) )
current_stage: str | None = None current_stage: str | None = None
self.timer.start_run()
def relay(event: Any) -> None: def relay(event: Any) -> None:
nonlocal current_stage nonlocal current_stage
if event.stage in STAGES and event.status == "started": if event.stage in STAGES and event.status == "started":
current_stage = event.stage current_stage = event.stage
self._update_timer(event)
app_event = AppProgressEvent( app_event = AppProgressEvent(
stage=event.stage, stage=event.stage,
status=event.status, status=event.status,
@@ -219,7 +386,13 @@ class MeetingProcessingService:
if progress_sink is not None: if progress_sink is not None:
progress_sink(app_event) progress_sink(app_event)
result = self.meeting_lab.run(config, context, relay) try:
result = self.meeting_lab.run(config, context, relay)
except Exception:
self.timer.finish_run()
raise
self.timer.finish_run()
self._persist_timing(result.run_dir)
if result.exit_code != 0: if result.exit_code != 0:
failure = self._read_failure(result.run_dir) failure = self._read_failure(result.run_dir)
backend_stage = failure.get("stage") backend_stage = failure.get("stage")
@@ -243,14 +416,286 @@ class MeetingProcessingService:
failed_stage=failed_stage or current_stage or "preparing", failed_stage=failed_stage or current_stage or "preparing",
error_message=failure_message, error_message=failure_message,
) )
if result.protocol_path is None:
return ProcessingOutcome(
succeeded=True,
run_dir=result.run_dir,
original_protocol=None,
protocol_path=None,
awaiting_speaker_review=True,
detected_speaker_count=self._detected_speaker_count(result.run_dir),
)
protocol_path = Path(result.protocol_path) protocol_path = Path(result.protocol_path)
return ProcessingOutcome( return ProcessingOutcome(
succeeded=True, succeeded=True,
run_dir=result.run_dir, run_dir=result.run_dir,
original_protocol=protocol_path.read_text(encoding="utf-8"), original_protocol=protocol_path.read_text(encoding="utf-8"),
protocol_path=protocol_path, protocol_path=protocol_path,
speaker_attribution_available=self._speaker_attribution_available(result.run_dir),
) )
def load_speaker_mapping_review(
self,
run_dir: Path,
*,
excerpts_per_speaker: int = 6,
minimum_excerpt_characters: int = 20,
) -> SpeakerMappingReview | None:
"""Load detected labels and small deterministic identification excerpts."""
run_dir = Path(run_dir)
transcript_path = run_dir / "diarization" / "transcript_diarized.json"
context_path = run_dir / "context" / "meeting_context.yaml"
if not transcript_path.is_file() or not context_path.is_file():
return None
transcript = json.loads(transcript_path.read_text(encoding="utf-8-sig"))
context_data = yaml.safe_load(context_path.read_text(encoding="utf-8"))
if not isinstance(transcript, dict) or not isinstance(context_data, dict):
raise ValueError("Existing run contains malformed speaker review artifacts.")
segments = transcript.get("segments")
if not isinstance(segments, list):
raise ValueError("Diarized transcript has no segments list.")
candidates_by_speaker: dict[str, list[_ExcerptCandidate]] = {}
short_candidates_by_speaker: dict[str, list[_ExcerptCandidate]] = {}
previous_label: str | None = None
previous_tokens: frozenset[str] = frozenset()
for segment_index, segment in enumerate(segments):
if not isinstance(segment, dict):
continue
label = segment.get("speaker_id")
text = segment.get("text")
normalized = " ".join(text.split()) if isinstance(text, str) else ""
tokens = frozenset(_excerpt_tokens(normalized))
if (
not isinstance(label, str)
or not label.startswith("SPEAKER_")
or label == "SPEAKER_UNASSIGNED"
or not normalized
):
if normalized:
previous_label = label if isinstance(label, str) else None
previous_tokens = tokens
continue
candidate = _ExcerptCandidate(
speaker_label=label,
text=normalized,
segment_index=segment_index,
tokens=tokens,
lexical_tokens=tokens - _EXCERPT_STOP_WORDS,
previous_other_speaker_tokens=(
previous_tokens
if previous_label is not None and previous_label != label
else frozenset()
),
)
target = (
candidates_by_speaker
if len(normalized) >= minimum_excerpt_characters
else short_candidates_by_speaker
)
target.setdefault(label, []).append(candidate)
previous_label = label
previous_tokens = tokens
labels = sorted(set(candidates_by_speaker) | set(short_candidates_by_speaker))
token_speakers: dict[str, set[str]] = {}
for candidate_group in (*candidates_by_speaker.values(), *short_candidates_by_speaker.values()):
for candidate in candidate_group:
for token in candidate.lexical_tokens:
token_speakers.setdefault(token, set()).add(candidate.speaker_label)
speakers = []
for label in labels:
candidates = candidates_by_speaker.get(label) or short_candidates_by_speaker.get(
label, []
)
excerpts = self._select_identifying_excerpts(
candidates,
excerpts_per_speaker=excerpts_per_speaker,
segment_count=len(segments),
token_speakers=token_speakers,
)
speakers.append(SpeakerReview(label, excerpts))
participants = tuple(
(participant["participant_id"], participant["display_name"])
for participant in context_data.get("participants", [])
if isinstance(participant, dict)
and isinstance(participant.get("participant_id"), str)
and isinstance(participant.get("display_name"), str)
)
mappings = context_data.get("speaker_mappings")
return SpeakerMappingReview(
speakers=tuple(speakers),
participants=participants,
current_mappings=dict(mappings) if isinstance(mappings, dict) else {},
)
@staticmethod
def _select_identifying_excerpts(
candidates: list[_ExcerptCandidate],
*,
excerpts_per_speaker: int,
segment_count: int,
token_speakers: dict[str, set[str]],
) -> tuple[str, ...]:
"""Choose informative, distinct excerpts while retaining transcript order."""
if len(candidates) <= excerpts_per_speaker:
return tuple(candidate.text for candidate in candidates)
base_scores = {
candidate: _excerpt_score(
candidate,
sum(
len(token_speakers.get(token, set())) == 1
for token in candidate.lexical_tokens
),
)
for candidate in candidates
}
selected: list[_ExcerptCandidate] = []
selected_buckets: set[int] = set()
while len(selected) < excerpts_per_speaker:
choice = max(
(candidate for candidate in candidates if candidate not in selected),
key=lambda candidate: (
_selection_score(
candidate,
base_scores[candidate],
selected,
selected_buckets,
segment_count,
),
-candidate.segment_index,
),
)
selected.append(choice)
selected_buckets.add(_segment_bucket(choice.segment_index, segment_count))
return tuple(
candidate.text for candidate in sorted(selected, key=lambda item: item.segment_index)
)
def regenerate_protocol(
self,
run_dir: Path,
speaker_mappings: dict[str, str],
progress_sink: Callable[[AppProgressEvent], None] | None = None,
performance_profile: str = DEFAULT_PERFORMANCE_PROFILE,
) -> ProcessingOutcome:
"""Persist confirmed mappings and regenerate only the direct protocol."""
review = self.load_speaker_mapping_review(run_dir)
if review is None:
raise ValueError("This run has no diarization artifacts to map.")
detected = {speaker.speaker_label for speaker in review.speakers}
participant_ids = {participant_id for participant_id, _ in review.participants}
unknown_labels = sorted(set(speaker_mappings) - detected)
if unknown_labels:
raise ValueError(f"Unknown diarization speaker label: {unknown_labels[0]}")
unknown_participants = sorted(set(speaker_mappings.values()) - participant_ids)
if unknown_participants:
raise ValueError(f"Unknown participant ID: {unknown_participants[0]}")
assigned_participants = list(speaker_mappings.values())
if len(assigned_participants) != len(set(assigned_participants)):
raise ValueError("A participant may be assigned to only one speaker label.")
context_path = Path(run_dir) / "context" / "meeting_context.yaml"
context_data = yaml.safe_load(context_path.read_text(encoding="utf-8"))
if not isinstance(context_data, dict):
raise ValueError("Existing run contains malformed Meeting Context.")
context_data["speaker_mappings"] = dict(sorted(speaker_mappings.items()))
self._apply_glossary(context_data)
context = self.meeting_lab.create_context(context_data)
# A checkpoint mapping is user input, not a derived protocol artifact. Save it
# before inference so a failed generation remains retryable with the same review.
if context_path.is_symlink():
context_path.unlink()
context_path.write_text(
yaml.safe_dump(context.data, allow_unicode=True, sort_keys=False), encoding="utf-8"
)
def relay(event: Any) -> None:
self._update_timer(event)
if progress_sink is not None:
progress_sink(
AppProgressEvent(
stage=event.stage,
status=event.status,
elapsed_seconds=event.elapsed_seconds,
progress=event.progress,
message=event.message,
)
)
self.timer.start_run()
try:
result = self.meeting_lab.regenerate_protocol(
Path(run_dir),
context,
relay,
model=self.settings.protocol_model,
ollama_endpoint=self.settings.ollama_endpoint,
protocol_num_ctx=self.settings.protocol_num_ctx,
protocol_safe_input_token_budget=(self.settings.protocol_safe_input_token_budget),
protocol_num_thread=resolve_ollama_num_thread(performance_profile),
glossary_aliases=glossary_alias_mapping(self.glossary.list(active_only=True)),
)
except Exception:
self.timer.finish_run()
raise
self.timer.finish_run()
self._persist_timing(Path(result.run_dir))
protocol_path = Path(result.protocol_path)
return ProcessingOutcome(
succeeded=True,
run_dir=Path(result.run_dir),
original_protocol=protocol_path.read_text(encoding="utf-8"),
protocol_path=protocol_path,
speaker_attribution_available=self._speaker_attribution_available(result.run_dir),
)
def _apply_glossary(self, context_data: dict[str, Any]) -> None:
"""Merge active terminology into context without discarding other entities."""
glossary_terms = render_glossary_terms(self.glossary.list(active_only=True))
known_entities = context_data.setdefault("known_entities", {})
if glossary_terms:
known_entities["Authoritative terminology"] = glossary_terms
else:
known_entities.pop("Authoritative terminology", None)
rules = context_data.setdefault("context_rules", {})
rules["glossary_canonical_spelling"] = (
"Use canonical glossary spellings only when the meeting clearly refers to "
"those terms; do not invent matches or replace unrelated words."
)
rules["glossary_core_terms"] = (
"Preserve surrounding context and use canonical core terms inside compounds "
"where appropriate."
)
@staticmethod
def _speaker_attribution_available(run_dir: Path | None) -> bool | None:
if run_dir is None:
return None
metadata_path = Path(run_dir) / "protocol" / "runtime_metadata.json"
if not metadata_path.is_file():
return None
try:
metadata = json.loads(metadata_path.read_text(encoding="utf-8"))
except (OSError, json.JSONDecodeError):
return None
value = metadata.get("speaker_attribution_available")
return value if isinstance(value, bool) else None
@staticmethod
def _detected_speaker_count(run_dir: Path | None) -> int | None:
if run_dir is None:
return None
metadata_path = Path(run_dir) / "run_metadata.json"
try:
metadata = json.loads(metadata_path.read_text(encoding="utf-8"))
except (OSError, json.JSONDecodeError):
return None
value = metadata.get("diarization", {}).get("speaker_count")
return value if isinstance(value, int) else None
@staticmethod @staticmethod
def _read_failure(run_dir: Path | None) -> dict[str, str]: def _read_failure(run_dir: Path | None) -> dict[str, str]:
if run_dir is None: if run_dir is None:
@@ -270,3 +715,38 @@ class MeetingProcessingService:
destination = Path(run_dir) / "protocol_edited.md" destination = Path(run_dir) / "protocol_edited.md"
destination.write_text(text, encoding="utf-8") destination.write_text(text, encoding="utf-8")
return destination return destination
def _update_timer(self, event: Any) -> None:
"""Apply a backend progress event to the shared per-run timer."""
if event.stage in STAGES and event.status == "started":
self.timer.start_stage(event.stage)
elif event.stage in STAGES and event.status in {"completed", "skipped"}:
self.timer.finish_stage(event.stage)
if event.stage in {"completed", "failed"}:
self.timer.finish_run()
def _persist_timing(self, run_dir: Path | None) -> None:
"""Merge final monotonic durations into Meeting Lab run metadata."""
if run_dir is None:
return
metadata_path = Path(run_dir) / "run_metadata.json"
metadata: dict[str, Any] = {}
if metadata_path.is_file():
try:
existing = json.loads(metadata_path.read_text(encoding="utf-8"))
if isinstance(existing, dict):
metadata = existing
except (OSError, json.JSONDecodeError):
return
snapshot = self.timer.snapshot()
metadata["timing"] = {
"stages_seconds": snapshot.stage_durations,
"total_seconds": snapshot.total_duration,
}
metadata_path.parent.mkdir(parents=True, exist_ok=True)
try:
metadata_path.write_text(
json.dumps(metadata, indent=2, sort_keys=True) + "\n", encoding="utf-8"
)
except OSError:
return
+23
View File
@@ -0,0 +1,23 @@
"""Abstract performance profiles resolved to current backend runtime options."""
from __future__ import annotations
DEFAULT_PERFORMANCE_PROFILE = "auto"
PERFORMANCE_PROFILES = ("auto", "fast", "efficient", "powersave")
_OLLAMA_THREADS_BY_PROFILE = {
"fast": 16,
"efficient": 10,
"powersave": 4,
}
def resolve_ollama_num_thread(profile: str) -> int | None:
"""Resolve a profile, leaving automatic thread selection to Ollama for Auto."""
if profile == "auto":
return None
try:
return _OLLAMA_THREADS_BY_PROFILE[profile]
except KeyError as exc:
allowed = ", ".join(PERFORMANCE_PROFILES)
raise ValueError(f"Unknown performance profile {profile!r}; expected: {allowed}.") from exc
+80
View File
@@ -0,0 +1,80 @@
"""Thread-safe monotonic timing state for meeting-processing progress."""
from __future__ import annotations
import time
from collections.abc import Callable
from dataclasses import dataclass
from threading import Lock
@dataclass(frozen=True)
class TimingSnapshot:
"""Presentation-neutral runtime values measured in seconds."""
stage_durations: dict[str, float]
active_stage: str | None
total_duration: float
running: bool
class ProcessingTimer:
"""Track pipeline stages independently from pipeline business logic."""
def __init__(self, clock: Callable[[], float] = time.perf_counter) -> None:
self._clock = clock
self._lock = Lock()
self._run_started: float | None = None
self._run_finished: float | None = None
self._stage_started: dict[str, float] = {}
self._stage_finished: dict[str, float] = {}
self._active_stage: str | None = None
def start_run(self) -> None:
with self._lock:
self._run_started = self._clock()
self._run_finished = None
self._stage_started.clear()
self._stage_finished.clear()
self._active_stage = None
def start_stage(self, stage: str) -> None:
with self._lock:
now = self._clock()
if self._active_stage is not None and self._active_stage != stage:
self._stage_finished.setdefault(self._active_stage, now)
self._stage_started.setdefault(stage, now)
self._active_stage = stage
def finish_stage(self, stage: str) -> None:
with self._lock:
if stage in self._stage_started:
self._stage_finished.setdefault(stage, self._clock())
if self._active_stage == stage:
self._active_stage = None
def finish_run(self) -> None:
with self._lock:
if self._run_finished is not None:
return
now = self._clock()
if self._active_stage is not None:
self._stage_finished.setdefault(self._active_stage, now)
self._active_stage = None
if self._run_started is not None:
self._run_finished = now
def snapshot(self) -> TimingSnapshot:
with self._lock:
now = self._run_finished if self._run_finished is not None else self._clock()
durations = {
stage: max(0.0, self._stage_finished.get(stage, now) - started)
for stage, started in self._stage_started.items()
}
total = max(0.0, now - self._run_started) if self._run_started is not None else 0.0
return TimingSnapshot(
stage_durations=durations,
active_stage=self._active_stage,
total_duration=total,
running=self._run_started is not None and self._run_finished is None,
)
+218
View File
@@ -0,0 +1,218 @@
"""Versioned JSON import/export for the user-facing run input form."""
from __future__ import annotations
import json
from collections.abc import Sequence
from dataclasses import dataclass
from datetime import date
from pathlib import Path
from typing import Any
from mka.application.meeting_service import ParticipantInput, stable_id
from mka.application.people_yaml import ATTENDANCE_VALUES, PARTICIPANT_ID_PATTERN
RUN_INPUT_SCHEMA_VERSION = 1
SUPPORTED_LANGUAGES = frozenset({"de", "en"})
class RunInputJsonError(ValueError):
"""Raised when a run-input JSON document is malformed or unsupported."""
@dataclass(frozen=True)
class RunInputState:
"""All user-configurable values needed before starting a processing run."""
title: str
description: str
language: str
meeting_date: date | None
participants: tuple[ParticipantInput, ...]
audio_normalization: bool
diarization_enabled: bool
source_file_name: str | None = None
@classmethod
def defaults(cls) -> RunInputState:
return cls(
title="",
description="",
language="de",
meeting_date=date.today(),
participants=(ParticipantInput(participant_id="", display_name=""),),
audio_normalization=True,
diarization_enabled=False,
)
def export_run_inputs(state: RunInputState) -> str:
"""Serialize form state as deterministic, human-readable JSON."""
source_file_name = _source_file_name(state.source_file_name)
document = {
"schema_version": RUN_INPUT_SCHEMA_VERSION,
"meeting": {
"title": state.title,
"description": state.description,
"language": state.language,
"date": state.meeting_date.isoformat() if state.meeting_date else None,
"participants": [
{
"participant_id": person.participant_id,
"display_name": person.display_name,
"role": person.role,
"organization": person.organization,
"attendance_status": person.attendance_status,
}
for person in state.participants
],
},
"processing": {
"audio_normalization": state.audio_normalization,
"diarization_enabled": state.diarization_enabled,
},
"source_file_name": source_file_name,
}
return json.dumps(document, ensure_ascii=False, indent=2) + "\n"
def import_run_inputs(content: str | bytes) -> RunInputState:
"""Parse supported fields, applying current defaults to omitted optional fields."""
try:
if isinstance(content, bytes):
content = content.decode("utf-8")
document = json.loads(content)
except (UnicodeDecodeError, json.JSONDecodeError) as exc:
raise RunInputJsonError(f"Malformed input JSON: {exc}") from exc
if not isinstance(document, dict):
raise RunInputJsonError("Input JSON must contain a top-level object.")
version = document.get("schema_version")
if type(version) is not int or version != RUN_INPUT_SCHEMA_VERSION:
raise RunInputJsonError(
f"Unsupported input schema_version {version!r}; expected {RUN_INPUT_SCHEMA_VERSION}."
)
defaults = RunInputState.defaults()
meeting = _optional_mapping(document, "meeting")
processing = _optional_mapping(document, "processing")
title = _optional_text(meeting, "title", defaults.title)
description = _optional_text(meeting, "description", defaults.description)
language = _optional_text(meeting, "language", defaults.language)
if language not in SUPPORTED_LANGUAGES:
raise RunInputJsonError(f"Unsupported meeting language: {language!r}.")
meeting_date = _meeting_date(meeting, defaults.meeting_date)
participants = _participants(meeting, defaults.participants)
audio_normalization = _optional_bool(
processing, "audio_normalization", defaults.audio_normalization
)
diarization_enabled = _optional_bool(
processing, "diarization_enabled", defaults.diarization_enabled
)
source_file_name = _source_file_name(document.get("source_file_name"))
return RunInputState(
title=title,
description=description,
language=language,
meeting_date=meeting_date,
participants=participants,
audio_normalization=audio_normalization,
diarization_enabled=diarization_enabled,
source_file_name=source_file_name,
)
def run_input_filename(title: str) -> str:
"""Build a stable download filename without filesystem-specific characters."""
suffix = stable_id(title) if title.strip() else "untitled"
return f"meeting-inputs-{suffix}.json"
def _optional_mapping(document: dict[str, Any], field: str) -> dict[str, Any]:
value = document.get(field, {})
if not isinstance(value, dict):
raise RunInputJsonError(f"Field {field!r} must be an object.")
return value
def _optional_text(document: dict[str, Any], field: str, default: str) -> str:
value = document.get(field, default)
if not isinstance(value, str):
raise RunInputJsonError(f"Field {field!r} must be text.")
return value
def _optional_bool(document: dict[str, Any], field: str, default: bool) -> bool:
value = document.get(field, default)
if type(value) is not bool:
raise RunInputJsonError(f"Field {field!r} must be true or false.")
return value
def _meeting_date(document: dict[str, Any], default: date | None) -> date | None:
if "date" not in document:
return default
value = document["date"]
if value is None:
return None
if not isinstance(value, str):
raise RunInputJsonError("Field 'date' must be an ISO date or null.")
try:
return date.fromisoformat(value)
except ValueError as exc:
raise RunInputJsonError("Field 'date' must be a valid ISO date or null.") from exc
def _participants(
document: dict[str, Any], default: Sequence[ParticipantInput]
) -> tuple[ParticipantInput, ...]:
if "participants" not in document:
return tuple(default)
entries = document["participants"]
if not isinstance(entries, list):
raise RunInputJsonError("Field 'participants' must be a list.")
people: list[ParticipantInput] = []
seen_ids: set[str] = set()
for index, entry in enumerate(entries, start=1):
if not isinstance(entry, dict):
raise RunInputJsonError(f"Participant {index} must be an object.")
participant_id = _optional_text(entry, "participant_id", "")
display_name = _optional_text(entry, "display_name", "")
role = _optional_text(entry, "role", "")
organization = _optional_text(entry, "organization", "")
attendance = _optional_text(entry, "attendance_status", "present")
if bool(participant_id) != bool(display_name):
raise RunInputJsonError(
f"Participant {index} must provide both participant_id and display_name."
)
if participant_id and PARTICIPANT_ID_PATTERN.fullmatch(participant_id) is None:
raise RunInputJsonError(
f"Participant {index} has invalid participant_id {participant_id!r}."
)
if participant_id in seen_ids:
raise RunInputJsonError(f"Duplicate participant_id: {participant_id!r}.")
if participant_id:
seen_ids.add(participant_id)
if attendance not in ATTENDANCE_VALUES:
raise RunInputJsonError(
f"Participant {index} has invalid attendance_status {attendance!r}."
)
people.append(
ParticipantInput(
participant_id=participant_id,
display_name=display_name,
role=role,
organization=organization,
attendance_status=attendance,
)
)
return tuple(people)
def _source_file_name(value: Any) -> str | None:
if value is None:
return None
if not isinstance(value, str) or not value.strip():
raise RunInputJsonError("Field 'source_file_name' must be non-empty text or null.")
if Path(value).name != value:
raise RunInputJsonError("Field 'source_file_name' must be a filename, not a path.")
return value
+17
View File
@@ -18,6 +18,7 @@ class MeetingLabGateway:
from src.meeting_lab.models.meeting_context import create_meeting_context from src.meeting_lab.models.meeting_context import create_meeting_context
from src.meeting_lab.orchestration.mvp import ( from src.meeting_lab.orchestration.mvp import (
MvpMeetingConfig, MvpMeetingConfig,
regenerate_mvp_protocol,
run_mvp_meeting, run_mvp_meeting,
) )
except ImportError as exc: except ImportError as exc:
@@ -28,6 +29,7 @@ class MeetingLabGateway:
self._config_type = MvpMeetingConfig self._config_type = MvpMeetingConfig
self._create_context = create_meeting_context self._create_context = create_meeting_context
self._run_mvp_meeting = run_mvp_meeting self._run_mvp_meeting = run_mvp_meeting
self._regenerate_mvp_protocol = regenerate_mvp_protocol
def create_context(self, data: dict[str, Any]) -> Any: def create_context(self, data: dict[str, Any]) -> Any:
"""Validate structured context through Meeting Lab's domain boundary.""" """Validate structured context through Meeting Lab's domain boundary."""
@@ -49,3 +51,18 @@ class MeetingLabGateway:
meeting_context=meeting_context, meeting_context=meeting_context,
progress_sink=progress_sink, progress_sink=progress_sink,
) )
def regenerate_protocol(
self,
run_dir: Any,
meeting_context: Any,
progress_sink: Callable[[Any], None],
**options: Any,
) -> Any:
"""Regenerate protocol artifacts without rerunning media processing."""
return self._regenerate_mvp_protocol(
run_dir,
meeting_context=meeting_context,
progress_sink=progress_sink,
**options,
)
+543 -60
View File
@@ -2,13 +2,27 @@
from __future__ import annotations from __future__ import annotations
import time
from collections.abc import Callable
from concurrent.futures import ThreadPoolExecutor
from datetime import date from datetime import date
from queue import Empty, Queue
from typing import Any from typing import Any
from uuid import uuid4 from uuid import uuid4
import streamlit as st import streamlit as st
from mka.application.config import AppSettings, ConfigurationError from mka.application.config import AppSettings, ConfigurationError
from mka.application.glossary import (
GLOSSARY_CATEGORIES,
GlossaryConflictError,
GlossaryRepository,
)
from mka.application.glossary_yaml import (
GlossaryYamlError,
export_glossary_yaml,
import_glossary_yaml,
)
from mka.application.meeting_service import ( from mka.application.meeting_service import (
STAGES, STAGES,
AppProgressEvent, AppProgressEvent,
@@ -16,6 +30,7 @@ from mka.application.meeting_service import (
MeetingProcessingService, MeetingProcessingService,
ParticipantInput, ParticipantInput,
ProcessingOptions, ProcessingOptions,
SpeakerMappingReview,
stable_id, stable_id,
) )
from mka.application.people_yaml import ( from mka.application.people_yaml import (
@@ -23,6 +38,15 @@ from mka.application.people_yaml import (
export_people_yaml, export_people_yaml,
import_people_yaml, import_people_yaml,
) )
from mka.application.performance import DEFAULT_PERFORMANCE_PROFILE, PERFORMANCE_PROFILES
from mka.application.progress_timing import TimingSnapshot
from mka.application.run_inputs import (
RunInputJsonError,
RunInputState,
export_run_inputs,
import_run_inputs,
run_input_filename,
)
from mka.integrations.meeting_lab import ( from mka.integrations.meeting_lab import (
MeetingLabGateway, MeetingLabGateway,
MeetingLabUnavailableError, MeetingLabUnavailableError,
@@ -36,6 +60,114 @@ STAGE_LABELS = {
} }
def _parse_aliases(value: str) -> tuple[str, ...]:
"""Parse one alias per line while tolerating comma-separated input."""
return tuple(
alias.strip() for line in value.splitlines() for alias in line.split(",") if alias.strip()
)
def _render_glossary(repository: GlossaryRepository) -> None:
"""Render simple global glossary CRUD controls."""
with st.expander("Terminology glossary", expanded=False):
st.caption(
"Store canonical core terms and recognition aliases. Compound phrases are "
"composed from meeting context during protocol generation."
)
import_file = st.file_uploader(
"Import glossary",
type=["yaml", "yml"],
key="glossary_import_file",
help="Validate and atomically replace the SQLite glossary from versioned YAML.",
)
action_columns = st.columns(2)
if action_columns[0].button("Import glossary", disabled=import_file is None):
try:
imported = import_glossary_yaml(import_file.getvalue())
repository.replace_all(imported)
except (GlossaryYamlError, GlossaryConflictError, ValueError) as exc:
st.error(str(exc))
else:
st.session_state.glossary_import_message = (
f"Imported {len(imported)} glossary entries."
)
st.rerun()
action_columns[1].download_button(
"Export glossary",
data=export_glossary_yaml(repository.list()).encode("utf-8"),
file_name="terminology-glossary.yaml",
mime="application/yaml",
)
if message := st.session_state.pop("glossary_import_message", None):
st.success(message)
with st.form("glossary_add"):
columns = st.columns(2)
canonical = columns[0].text_input("Canonical term")
category = columns[1].selectbox("Category", GLOSSARY_CATEGORIES)
aliases = st.text_area("Aliases", help="One per line or comma-separated.")
description = st.text_area("Optional description")
if st.form_submit_button("Add glossary entry"):
try:
repository.create(
canonical,
category,
aliases=_parse_aliases(aliases),
description=description,
)
except (GlossaryConflictError, ValueError) as exc:
st.error(str(exc))
else:
st.success("Glossary entry added.")
st.rerun()
search = st.text_input("Search glossary")
active_only = st.checkbox("Show active entries only")
entries = repository.list(search, active_only=active_only)
st.caption(f"{len(entries)} glossary entries")
for entry in entries:
status = "active" if entry.is_active else "inactive"
with (
st.expander(f"{entry.canonical_term} · {entry.category} · {status}"),
st.form(f"glossary_edit_{entry.id}"),
):
columns = st.columns(2)
edited_canonical = columns[0].text_input("Canonical term", entry.canonical_term)
edited_category = columns[1].selectbox(
"Category",
GLOSSARY_CATEGORIES,
index=GLOSSARY_CATEGORIES.index(entry.category),
)
edited_aliases = st.text_area("Aliases", "\n".join(entry.aliases))
edited_description = st.text_area("Optional description", entry.description or "")
edited_active = st.checkbox("Active", value=entry.is_active)
delete_confirmed = st.checkbox(
"Permanently delete this entry",
help="Prefer clearing Active for normal use.",
)
action_columns = st.columns(2)
save = action_columns[0].form_submit_button("Save changes")
delete = action_columns[1].form_submit_button(
"Delete", disabled=not delete_confirmed
)
try:
if save:
repository.update(
entry.id,
edited_canonical,
edited_category,
aliases=_parse_aliases(edited_aliases),
description=edited_description,
is_active=edited_active,
)
st.rerun()
if delete:
repository.delete(entry.id)
st.rerun()
except (GlossaryConflictError, ValueError) as exc:
st.error(str(exc))
def _new_participant() -> dict[str, str]: def _new_participant() -> dict[str, str]:
return { return {
"row_id": uuid4().hex, "row_id": uuid4().hex,
@@ -51,6 +183,51 @@ def _initialize_state() -> None:
st.session_state.setdefault("participants", [_new_participant()]) st.session_state.setdefault("participants", [_new_participant()])
st.session_state.setdefault("outcome", None) st.session_state.setdefault("outcome", None)
st.session_state.setdefault("edited_protocol", "") st.session_state.setdefault("edited_protocol", "")
st.session_state.setdefault("source_media_widget_generation", 0)
_apply_pending_run_inputs()
_apply_pending_edited_protocol()
def _queue_run_inputs(value: RunInputState) -> None:
"""Defer form restoration until the beginning of the next Streamlit run."""
st.session_state["pending_run_inputs"] = value
def _apply_pending_run_inputs() -> None:
"""Restore imported values before any corresponding widget is instantiated."""
value = st.session_state.pop("pending_run_inputs", None)
if value is None:
return
st.session_state.update(
{
"meeting_title": value.title,
"meeting_description": value.description,
"meeting_language": value.language,
"meeting_has_date": value.meeting_date is not None,
"meeting_date": value.meeting_date or date.today(),
"participants": _people_to_rows(list(value.participants)),
"audio_normalization": value.audio_normalization,
"diarization_enabled": value.diarization_enabled,
"imported_source_file_name": value.source_file_name,
"run_input_import_message": (
"Meeting inputs imported. Select the source media before processing."
),
"source_media_widget_generation": (
st.session_state.get("source_media_widget_generation", 0) + 1
),
}
)
def _apply_pending_edited_protocol() -> None:
"""Apply a deferred widget value before the widget is instantiated."""
if "pending_edited_protocol" in st.session_state:
st.session_state["edited_protocol"] = st.session_state.pop("pending_edited_protocol")
def _queue_edited_protocol(value: str) -> None:
"""Defer an edited-protocol widget update until the next Streamlit run."""
st.session_state["pending_edited_protocol"] = value
def _people_to_rows(people: list[ParticipantInput]) -> list[dict[str, str]]: def _people_to_rows(people: list[ParticipantInput]) -> list[dict[str, str]]:
@@ -68,6 +245,27 @@ def _people_to_rows(people: list[ParticipantInput]) -> list[dict[str, str]]:
] ]
def _render_run_input_import() -> None:
"""Render import controls before the widgets whose state they restore."""
st.subheader("Input configuration")
imported_file = st.file_uploader(
"Import inputs",
type=["json"],
key="run_input_import_file",
help="Restore form values only; source media is never included.",
)
if st.button("Import input configuration", disabled=imported_file is None):
try:
imported = import_run_inputs(imported_file.getvalue())
except RunInputJsonError as exc:
st.error(str(exc))
else:
_queue_run_inputs(imported)
st.rerun()
if message := st.session_state.pop("run_input_import_message", None):
st.success(message)
def _render_participants() -> list[ParticipantInput]: def _render_participants() -> list[ParticipantInput]:
st.subheader("People") st.subheader("People")
st.caption("Record whether each relevant person attended or was only mentioned.") st.caption("Record whether each relevant person attended or was only mentioned.")
@@ -145,40 +343,199 @@ def _render_participants() -> list[ParticipantInput]:
return people return people
def _progress_callback( def _format_duration(seconds: float) -> str:
"""Format numeric seconds as MM:SS or H:MM:SS."""
whole_seconds = max(0, int(seconds))
hours, remainder = divmod(whole_seconds, 3600)
minutes, seconds = divmod(remainder, 60)
return f"{hours}:{minutes:02d}:{seconds:02d}" if hours else f"{minutes:02d}:{seconds:02d}"
def _render_progress_status(
status_box: Any, status_box: Any,
stage_table: Any, stage_table: Any,
progress_slot: Any,
states: dict[str, str], states: dict[str, str],
) -> Any: timing: TimingSnapshot,
progress_bar = None message: str,
) -> None:
suffix = " …" if timing.running else ""
status_box.info(f"{message} — total runtime {_format_duration(timing.total_duration)}{suffix}")
stage_table.table(
[
{
"Stage": STAGE_LABELS[stage],
"Status": states[stage],
"Runtime": (
_format_duration(timing.stage_durations[stage])
+ (" …" if timing.active_stage == stage else "")
if stage in timing.stage_durations
else ""
),
}
for stage in STAGES
]
+ [
{
"Stage": "Total runtime",
"Status": "",
"Runtime": _format_duration(timing.total_duration) + suffix,
}
]
)
def update(event: AppProgressEvent) -> None:
nonlocal progress_bar def _apply_progress_event(event: AppProgressEvent, states: dict[str, str]) -> None:
if event.stage in states: if event.stage in states:
states[event.stage] = "running" if event.status == "started" else event.status states[event.stage] = "running" if event.status == "started" else event.status
if event.stage == "failed": if event.stage == "failed":
running = next( running = next((stage for stage, status in states.items() if status == "running"), None)
if running:
states[running] = "failed"
def _run_with_live_progress(
service: MeetingProcessingService,
work: Callable[[Callable[[AppProgressEvent], None]], Any],
states: dict[str, str],
initial_message: str,
) -> tuple[Any, dict[str, str], str, TimingSnapshot]:
"""Run pipeline work while rendering the shared timer and progress events."""
status_box = st.empty()
stage_table = st.empty()
progress_slot = st.empty()
event_queue: Queue[AppProgressEvent] = Queue()
latest_message = initial_message
progress_bar = None
with ThreadPoolExecutor(max_workers=1) as executor:
future = executor.submit(work, event_queue.put)
while not future.done():
try:
while True:
event = event_queue.get_nowait()
_apply_progress_event(event, states)
latest_message = event.message or STAGE_LABELS.get(event.stage, event.stage)
if event.progress is not None:
if progress_bar is None:
progress_bar = progress_slot.progress(0.0)
progress_bar.progress(
min(max(event.progress, 0.0), 1.0),
text=(
f"{STAGE_LABELS.get(event.stage, event.stage)}: "
f"{event.progress:.0%}"
),
)
except Empty:
pass
_render_progress_status(
status_box,
stage_table,
states,
service.timer.snapshot(),
latest_message,
)
time.sleep(0.2)
try:
result = future.result()
except Exception:
while not event_queue.empty():
event = event_queue.get_nowait()
_apply_progress_event(event, states)
latest_message = event.message or STAGE_LABELS.get(event.stage, event.stage)
service.timer.finish_run()
running_stage = next(
(stage for stage, status in states.items() if status == "running"), (stage for stage, status in states.items() if status == "running"),
None, None,
) )
if running: if running_stage is not None:
states[running] = "failed" states[running_stage] = "failed"
elapsed = f"{event.elapsed_seconds:.1f} s" snapshot = service.timer.snapshot()
message = event.message or STAGE_LABELS.get(event.stage, event.stage) _render_progress_status(status_box, stage_table, states, snapshot, latest_message)
status_box.info(f"{message} — elapsed {elapsed}") st.session_state["processing_progress"] = (states.copy(), latest_message, snapshot)
stage_table.table( raise
[{"Stage": STAGE_LABELS[stage], "Status": states[stage]} for stage in STAGES] while not event_queue.empty():
) event = event_queue.get_nowait()
if event.progress is not None: _apply_progress_event(event, states)
if progress_bar is None: latest_message = event.message or STAGE_LABELS.get(event.stage, event.stage)
progress_bar = progress_slot.progress(0.0) running_stage = next((stage for stage, status in states.items() if status == "running"), None)
progress_bar.progress( if running_stage is not None:
min(max(event.progress, 0.0), 1.0), states[running_stage] = "completed" if getattr(result, "succeeded", True) else "failed"
text=f"{STAGE_LABELS.get(event.stage, event.stage)}: {event.progress:.0%}", snapshot = service.timer.snapshot()
) _render_progress_status(status_box, stage_table, states, snapshot, latest_message)
st.session_state["processing_progress"] = (states.copy(), latest_message, snapshot)
return result, states, latest_message, snapshot
return update
def _remember_regeneration_timing(
states: dict[str, str], message: str, timing: TimingSnapshot
) -> None:
st.session_state["regeneration_progress"] = (states.copy(), message, timing)
def _speaker_options(
participants: tuple[str, ...], mappings: dict[str, str | None], speaker_label: str
) -> list[str | None]:
"""Reserve other speakers' participants while retaining this speaker's mapping."""
current = mappings.get(speaker_label)
reserved = {value for label, value in mappings.items() if label != speaker_label}
options: list[str | None] = [None]
options.extend(person for person in participants if person == current or person not in reserved)
if current is not None and current not in options:
options.append(current)
return options
def _mapping_counts(
speaker_labels: tuple[str, ...], mappings: dict[str, str | None]
) -> tuple[int, int, int]:
"""Count only detected speakers, including explicit cleared selections."""
detected = len(speaker_labels)
assigned = sum(mappings.get(label) is not None for label in speaker_labels)
return detected, assigned, detected - assigned
def _render_speaker_mapping(review: SpeakerMappingReview, run_name: str) -> dict[str, str]:
st.subheader("Identify diarized speakers")
st.caption(
"Confirm identities explicitly. Unmapped speakers remain anonymous; "
"the diarized source transcript is not modified."
)
participant_names = dict(review.participants)
# Read every widget before rendering so later speakers also reserve their person.
keys = {
speaker.speaker_label: f"speaker_mapping_{run_name}_{speaker.speaker_label}"
for speaker in review.speakers
}
mappings = {
label: st.session_state.get(key, review.current_mappings.get(label))
for label, key in keys.items()
}
detected, assigned, unassigned = _mapping_counts(tuple(keys), mappings)
st.markdown(f"**{detected} speakers detected · {assigned} assigned · {unassigned} unassigned**")
if unassigned:
st.warning(
"Some detected speakers have no confirmed participant mapping. "
"Check whether a participant is missing or speaker assignment is incomplete. "
"You can still generate a protocol with anonymous speakers."
)
selections: dict[str, str] = {}
for speaker in review.speakers:
label = speaker.speaker_label
current = mappings[label]
options = _speaker_options(tuple(participant_names), mappings, label)
st.session_state[keys[label]] = current
selected = st.selectbox(
f"{label} — Unassigned" if current is None else label,
options=options,
format_func=lambda value, names=participant_names: (
"Unmapped / Unknown" if value is None else names.get(value, value)
),
key=keys[label],
)
if selected is not None:
selections[label] = selected
for excerpt in speaker.excerpts:
st.caption(f"“{excerpt}”")
return selections
def _render_result() -> None: def _render_result() -> None:
@@ -187,6 +544,9 @@ def _render_result() -> None:
return return
st.divider() st.divider()
st.header("Protocol result") st.header("Protocol result")
if retained_progress := st.session_state.get("processing_progress"):
states, message, timing = retained_progress
_render_progress_status(st.empty(), st.empty(), states, timing, message)
if not outcome.succeeded: if not outcome.succeeded:
st.error(f"Processing failed during {outcome.failed_stage}: {outcome.error_message}") st.error(f"Processing failed during {outcome.failed_stage}: {outcome.error_message}")
if outcome.run_dir: if outcome.run_dir:
@@ -194,19 +554,88 @@ def _render_result() -> None:
st.caption("Intermediate artifacts and run metadata were preserved here.") st.caption("Intermediate artifacts and run metadata were preserved here.")
return return
st.success("Processing completed. Review the generated protocol before use.") awaiting_review = getattr(outcome, "awaiting_speaker_review", False)
if awaiting_review:
count = getattr(outcome, "detected_speaker_count", None)
detected = f"{count} speakers detected" if count is not None else "speakers detected"
st.info(
f"Diarization complete — {detected}. Assign speakers if desired, then generate the protocol."
)
else:
st.success("Processing completed. Review the generated protocol before use.")
st.caption(f"Run artifacts: {outcome.run_dir}") st.caption(f"Run artifacts: {outcome.run_dir}")
with st.expander("Original generated protocol", expanded=False): if outcome.speaker_attribution_available is False:
st.markdown(outcome.original_protocol or "") st.warning(
edited = st.text_area( "Speaker attribution was unavailable for this protocol because the "
"Editable protocol", "safe-budget fallback removed diarization labels from the prompt."
key="edited_protocol", )
height=500, if message := st.session_state.pop("speaker_mapping_message", None):
help="The original protocol.md remains unchanged.", st.success(message)
) if not awaiting_review:
if st.button("Save edited protocol", type="primary"): with st.expander("Original generated protocol", expanded=False):
path = MeetingProcessingService.save_edited_protocol(outcome.run_dir, edited) st.markdown(outcome.original_protocol or "")
st.success(f"Saved edited protocol to {path}") edited = st.text_area(
"Editable protocol",
key="edited_protocol",
height=500,
help="The original protocol.md remains unchanged.",
)
if st.button("Save edited protocol", type="primary"):
path = MeetingProcessingService.save_edited_protocol(outcome.run_dir, edited)
st.success(f"Saved edited protocol to {path}")
try:
service = MeetingProcessingService(AppSettings.from_environment(), MeetingLabGateway())
review = service.load_speaker_mapping_review(outcome.run_dir)
except (MeetingLabUnavailableError, OSError, ValueError) as exc:
st.warning(f"Speaker mapping is unavailable: {exc}")
return
if review is None or not review.speakers:
return
selections = _render_speaker_mapping(review, outcome.run_dir.name)
duplicate_assignments = len(selections.values()) != len(set(selections.values()))
if duplicate_assignments:
st.error("Each participant can be assigned to only one detected speaker.")
if st.button(
"Generate protocol" if awaiting_review else "Regenerate protocol with confirmed speakers",
disabled=duplicate_assignments,
type="primary",
):
st.session_state.pop("regeneration_progress", None)
run_dir = outcome.run_dir
mapping_items = tuple(selections.items())
selected_profile = st.session_state.get("performance_profile", DEFAULT_PERFORMANCE_PROFILE)
states = {stage: "skipped" for stage in STAGES}
states["protocol_generation"] = "pending"
try:
regenerated, states, message, timing = _run_with_live_progress(
service,
lambda progress_sink: service.regenerate_protocol(
run_dir,
dict(mapping_items),
progress_sink=progress_sink,
performance_profile=selected_profile,
),
states,
"Starting protocol generation" if awaiting_review else "Starting protocol regeneration",
)
except (OSError, RuntimeError, ValueError) as exc:
_remember_regeneration_timing(
states,
"Protocol regeneration failed",
service.timer.snapshot(),
)
st.error(f"Protocol generation failed: {exc}")
else:
_remember_regeneration_timing(states, message, timing)
st.session_state.outcome = regenerated
_queue_edited_protocol(regenerated.original_protocol or "")
st.session_state.speaker_mapping_message = (
"Speaker mappings saved and protocol generated."
)
st.rerun()
def main() -> None: def main() -> None:
@@ -215,23 +644,55 @@ def main() -> None:
st.title("Meeting Assistant") st.title("Meeting Assistant")
st.caption("Create meeting context, run Meeting Lab, and review the protocol.") st.caption("Create meeting context, run Meeting Lab, and review the protocol.")
settings = AppSettings.from_environment()
glossary = GlossaryRepository(settings.glossary_database)
glossary.initialize()
_render_glossary(glossary)
_render_run_input_import()
st.header("Meeting and audio") st.header("Meeting and audio")
audio = st.file_uploader("Audio recording", type=["wav", "flac", "m4a"]) audio = st.file_uploader(
title = st.text_input("Meeting title") "Audio recording",
description = st.text_area("Description / context", height=100) type=["wav", "flac", "m4a"],
key=f"source_media_{st.session_state.source_media_widget_generation}",
)
imported_source_name = st.session_state.get("imported_source_file_name")
if audio is None and imported_source_name:
st.caption(
f"Previous source filename: {imported_source_name}. "
"Select the media file again before processing."
)
title = st.text_input("Meeting title", key="meeting_title")
description = st.text_area("Description / context", height=100, key="meeting_description")
metadata_columns = st.columns(3) metadata_columns = st.columns(3)
language = metadata_columns[0].selectbox("Meeting language", options=["de", "en"]) language = metadata_columns[0].selectbox(
has_date = metadata_columns[1].checkbox("Meeting date is known", value=True) "Meeting language",
options=["de", "en"],
key="meeting_language",
help="Language for both transcription and the generated protocol.",
)
has_date = metadata_columns[1].checkbox(
"Meeting date is known", value=True, key="meeting_has_date"
)
selected_date: date | None = ( selected_date: date | None = (
metadata_columns[2].date_input("Meeting date") if has_date else None metadata_columns[2].date_input("Meeting date", key="meeting_date") if has_date else None
) )
participants = _render_participants() participants = _render_participants()
st.header("Processing options") st.header("Processing options")
performance_profile = st.selectbox(
"Performance profile",
options=PERFORMANCE_PROFILES,
format_func=str.title,
key="performance_profile",
help="Controls protocol-generation performance using an abstract runtime profile.",
)
audio_normalization = st.checkbox( audio_normalization = st.checkbox(
"Audio normalization", "Audio normalization",
value=True, value=True,
key="audio_normalization",
help=( help=(
"Normalize speech loudness during preparation. WAV, FLAC, and M4A " "Normalize speech loudness during preparation. WAV, FLAC, and M4A "
"are always converted to the canonical Meeting Lab audio format, " "are always converted to the canonical Meeting Lab audio format, "
@@ -240,9 +701,28 @@ def main() -> None:
) )
diarization_enabled = st.checkbox( diarization_enabled = st.checkbox(
"Enable speaker diarization", "Enable speaker diarization",
key="diarization_enabled",
help="Speaker labels remain anonymous; identities are never inferred.", help="Speaker labels remain anonymous; identities are never inferred.",
) )
current_inputs = RunInputState(
title=title,
description=description,
language=language,
meeting_date=selected_date,
participants=tuple(participants),
audio_normalization=audio_normalization,
diarization_enabled=diarization_enabled,
source_file_name=(audio.name if audio is not None else imported_source_name),
)
st.download_button(
"Export inputs",
data=export_run_inputs(current_inputs).encode("utf-8"),
file_name=run_input_filename(title),
mime="application/json",
help="Download the current form configuration without source media or run results.",
)
if st.button("Start processing", type="primary", disabled=audio is None): if st.button("Start processing", type="primary", disabled=audio is None):
if not title.strip(): if not title.strip():
st.error("Meeting title is required.") st.error("Meeting title is required.")
@@ -252,9 +732,8 @@ def main() -> None:
): ):
st.error("Every person row needs a name and participant ID.") st.error("Every person row needs a name and participant ID.")
else: else:
settings = AppSettings.from_environment()
try: try:
service = MeetingProcessingService(settings, MeetingLabGateway()) service = MeetingProcessingService(settings, MeetingLabGateway(), glossary)
meeting = MeetingDetails( meeting = MeetingDetails(
title=title, title=title,
language=language, language=language,
@@ -263,30 +742,34 @@ def main() -> None:
) )
meeting_id = stable_id(title) meeting_id = stable_id(title)
audio_path = service.preserve_upload(meeting_id, audio.name, audio) audio_path = service.preserve_upload(meeting_id, audio.name, audio)
st.session_state.pop("regeneration_progress", None)
st.session_state.pop("processing_progress", None)
st.header("Processing status") st.header("Processing status")
status_box = st.empty()
stage_table = st.empty()
progress_slot = st.empty()
states = {stage: "pending" for stage in STAGES} states = {stage: "pending" for stage in STAGES}
if not diarization_enabled: if not diarization_enabled:
states["diarization"] = "skipped" states["diarization"] = "skipped"
callback = _progress_callback(status_box, stage_table, progress_slot, states) outcome, _, _, _ = _run_with_live_progress(
outcome = service.process( service,
audio_path, lambda progress_sink: service.process(
meeting, audio_path,
participants, meeting,
ProcessingOptions( participants,
diarization_enabled=diarization_enabled, ProcessingOptions(
audio_normalization=audio_normalization, diarization_enabled=diarization_enabled,
audio_normalization=audio_normalization,
performance_profile=performance_profile,
),
progress_sink=progress_sink,
), ),
progress_sink=callback, states,
"Starting processing",
) )
st.session_state.outcome = outcome st.session_state.outcome = outcome
st.session_state.edited_protocol = outcome.original_protocol or "" st.session_state.edited_protocol = outcome.original_protocol or ""
if outcome.succeeded: if outcome.succeeded:
status_box.success("Processing completed.") st.success("Processing completed.")
else: else:
status_box.error( st.error(
f"Processing failed during {outcome.failed_stage}: " f"Processing failed during {outcome.failed_stage}: "
f"{outcome.error_message} Artifacts were preserved." f"{outcome.error_message} Artifacts were preserved."
) )
+138
View File
@@ -0,0 +1,138 @@
"""Paired Assistant/Lab smoke with only expensive media/model boundaries mocked."""
import json
import shutil
from dataclasses import replace
from unittest.mock import Mock
import pytest
from mka.application.meeting_service import MeetingProcessingService, ProcessingOptions
from mka.integrations.meeting_lab import MeetingLabGateway
from test_meeting_service import make_service, meeting, participants
mvp = pytest.importorskip("src.meeting_lab.orchestration.mvp")
from src.meeting_lab.audio import PreparedAudio # noqa: E402
from src.meeting_lab.diarization.backend import DiarizationResult # noqa: E402
from src.meeting_lab.llm.ollama import OllamaGeneration # noqa: E402
from src.meeting_lab.protocol.generate_direct_protocol import generate_direct_protocol # noqa: E402
from src.meeting_lab.transcription.whisper import TranscriptionResult # noqa: E402
def fake_prepare(source, destination, **kwargs):
destination.parent.mkdir(parents=True, exist_ok=True)
shutil.copyfile(source, destination)
return PreparedAudio(source, source.suffix[1:], destination, "ffmpeg", "ffmpeg")
def fake_transcribe(audio, model, output, language, **kwargs):
output.mkdir(parents=True, exist_ok=True)
raw, transcript, text, metadata = [
output / n
for n in ("whisper_raw.json", "transcript.json", "transcript.txt", "runtime_metadata.json")
]
raw.write_text("{}")
transcript.write_text(
json.dumps(
{
"text": "Lumini project discussion.",
"segments": [{"id": 0, "start": 0, "end": 2, "text": "Lumini project discussion."}],
}
)
)
text.write_text("Lumini project discussion.")
metadata.write_text("{}")
return TranscriptionResult(output, raw, transcript, text, metadata, 0.1)
def fake_diarize(audio, output, mode, **kwargs):
output.mkdir(parents=True, exist_ok=True)
paths = [
output / name
for name in (
"metadata.json",
"diarization.rttm",
"exclusive_diarization.rttm",
"turns.json",
"exclusive_turns.json",
)
]
metadata = {"speaker_count": 1, "runtime_seconds": 0.1}
paths[0].write_text(json.dumps(metadata))
for p in paths[1:]:
p.write_text("[]")
paths[-1].write_text(json.dumps([{"start": 0, "end": 2, "speaker_id": "SPEAKER_00"}]))
return DiarizationResult(output, *paths, metadata)
@pytest.mark.parametrize("language,expected", [("de", "German"), ("en", "English")])
@pytest.mark.parametrize(
"profile,threads", [("auto", None), ("fast", 16), ("efficient", 10), ("powersave", 4)]
)
def test_paired_anonymous_generation_and_mapped_regeneration(
tmp_path, monkeypatch, language, expected, profile, threads
):
template, _ = make_service(tmp_path)
service = MeetingProcessingService(template.settings, MeetingLabGateway())
service.glossary.create("Luminy", "product", aliases=("Lumini",))
calls = []
def model_call(endpoint, model, prompt, **kwargs):
calls.append((prompt, kwargs))
return OllamaGeneration(
{"response": "# Mock protocol", "done": True}, "# Mock protocol", 0.1
)
def generate(transcript, context, **kwargs):
return generate_direct_protocol(
transcript, context, **kwargs, model_check=lambda *_: {}, generation_call=model_call
)
prepare = Mock(side_effect=fake_prepare)
transcribe = Mock(side_effect=fake_transcribe)
diarize = Mock(side_effect=fake_diarize)
monkeypatch.setattr(mvp, "prepare_audio", prepare)
monkeypatch.setattr(mvp, "transcribe_audio", transcribe)
monkeypatch.setattr(mvp, "diarize_audio", diarize)
monkeypatch.setattr(mvp, "generate_direct_protocol", generate)
audio = tmp_path / "sample.wav"
audio.write_bytes(b"mock audio")
outcome = service.process(
audio,
replace(meeting(), language=language),
participants(),
ProcessingOptions(diarization_enabled=True, performance_profile=profile),
)
assert outcome.succeeded
assert outcome.awaiting_speaker_review
assert prepare.call_args.kwargs["normalization_enabled"] is True
assert transcribe.call_args.args[3] == language
root = outcome.run_dir
assert calls == []
service.regenerate_protocol(root, {}, performance_profile=profile)
first = root / "protocol/generations/001"
before = {p.name: p.read_bytes() for p in first.iterdir()}
original = (root / "diarization/transcript_diarized.json").read_bytes()
first_meta = json.loads((first / "runtime_metadata.json").read_text())
assert first_meta["speaker_mapping"] == {}
assert first_meta["output_language"] == language
assert "SPEAKER_00" in (first / "transcript_input.txt").read_text()
assert "timing" in json.loads((root / "run_metadata.json").read_text())
service.regenerate_protocol(root, {"SPEAKER_00": "martin"}, performance_profile=profile)
assert prepare.call_count == transcribe.call_count == diarize.call_count == 1
assert (root / "diarization/transcript_diarized.json").read_bytes() == original
assert {p.name: p.read_bytes() for p in first.iterdir()} == before
second = root / "protocol/generations/002"
metadata = json.loads((second / "runtime_metadata.json").read_text())
assert metadata["speaker_mapping"] == {"SPEAKER_00": "martin"}
assert metadata["speaker_mapping_names"] == {"SPEAKER_00": "Martin"}
assert metadata["output_language"] == language
assert metadata["num_thread"] == threads
assert metadata["glossary_aliases_configured"]["Lumini"] == "Luminy"
assert metadata["glossary_replacements"] == []
assert "Lumini" in (second / "transcript_input.txt").read_text()
assert (root / "protocol.md").resolve() == second / "protocol.md"
for prompt, options in calls:
assert f"Write the meeting protocol in {expected}." in prompt
assert "Luminy" in prompt
assert options["num_thread"] == threads
+8
View File
@@ -39,6 +39,14 @@ def test_environment_defaults_to_no_diarization_container_args(monkeypatch) -> N
settings = AppSettings.from_environment() settings = AppSettings.from_environment()
assert settings.diarization_container_args == () assert settings.diarization_container_args == ()
assert settings.glossary_database == Path("data/database/glossary.sqlite3")
def test_environment_configures_glossary_database(monkeypatch, tmp_path: Path) -> None:
database = tmp_path / "terms.sqlite3"
monkeypatch.setenv("MKA_GLOSSARY_DATABASE", str(database))
assert AppSettings.from_environment().glossary_database == database
def test_environment_parses_multiple_ordered_container_args(monkeypatch) -> None: def test_environment_parses_multiple_ordered_container_args(monkeypatch) -> None:
+78
View File
@@ -0,0 +1,78 @@
import sqlite3
from pathlib import Path
import pytest
from mka.application.glossary import GlossaryConflictError, GlossaryRepository
def repository(tmp_path: Path) -> GlossaryRepository:
result = GlossaryRepository(tmp_path / "database" / "glossary.sqlite3")
result.initialize()
return result
def test_initialization_creates_empty_versioned_database(tmp_path: Path) -> None:
glossary = repository(tmp_path)
assert glossary.database_path.is_file()
assert glossary.list() == []
with sqlite3.connect(glossary.database_path) as connection:
assert connection.execute("PRAGMA user_version").fetchone()[0] == 1
tables = {
row[0]
for row in connection.execute("SELECT name FROM sqlite_master WHERE type = 'table'")
}
assert {"glossary_entries", "glossary_aliases"} <= tables
def test_create_read_update_and_deactivate_with_multiple_aliases(tmp_path: Path) -> None:
glossary = repository(tmp_path)
created = glossary.create(
"Secugrid HS",
"product",
aliases=("Sikirgut", "Secugrid H S"),
description="Canonical core product name",
)
assert glossary.get(created.id).aliases == ("Secugrid H S", "Sikirgut")
assert glossary.list("sikir")[0].canonical_term == "Secugrid HS"
updated = glossary.update(
created.id,
"Secugrid HS",
"technical_term",
aliases=("Sekugrid HS",),
description="Updated",
is_active=False,
)
assert updated.category == "technical_term"
assert updated.aliases == ("Sekugrid HS",)
assert not updated.is_active
assert glossary.list(active_only=True) == []
def test_terms_and_aliases_are_unique_case_insensitively_across_entries(
tmp_path: Path,
) -> None:
glossary = repository(tmp_path)
glossary.create("PBAT", "acronym", aliases=("Polybutylene adipate terephthalate",))
with pytest.raises(GlossaryConflictError):
glossary.create("pbat", "material")
with pytest.raises(GlossaryConflictError):
glossary.create("Other", "other", aliases=("PBAT",))
with pytest.raises(GlossaryConflictError):
glossary.create("Polybutylene adipate terephthalate", "material")
def test_delete_removes_entry_and_aliases(tmp_path: Path) -> None:
glossary = repository(tmp_path)
entry = glossary.create("Luminy", "product", aliases=("Lumini",))
glossary.delete(entry.id)
assert glossary.list() == []
with sqlite3.connect(glossary.database_path) as connection:
assert connection.execute("SELECT COUNT(*) FROM glossary_aliases").fetchone()[0] == 0
+147
View File
@@ -0,0 +1,147 @@
from pathlib import Path
import pytest
from mka.application.glossary import GlossaryRepository
from mka.application.glossary_yaml import (
GlossaryYamlError,
export_glossary_yaml,
import_glossary_yaml,
)
def repository(tmp_path: Path) -> GlossaryRepository:
result = GlossaryRepository(tmp_path / "glossary.sqlite3")
result.initialize()
return result
def test_round_trip_preserves_ids_aliases_and_active_state(tmp_path: Path) -> None:
source = repository(tmp_path / "source")
active = source.create(
"Größenmaß",
"technical_term",
aliases=("Groessenmass", "Größen-Maß"),
description="Canonical UTF-8 spelling",
)
inactive = source.create("Legacy Product", "product", is_active=False)
encoded = export_glossary_yaml(source.list()).encode("utf-8")
target = repository(tmp_path / "target")
target.create("Will be replaced", "other")
target.replace_all(import_glossary_yaml(encoded))
entries = target.list()
assert [entry.id for entry in entries] == [active.id, inactive.id]
assert entries[0].canonical_term == "Größenmaß"
assert entries[0].aliases == ("Groessenmass", "Größen-Maß")
assert entries[0].description == "Canonical UTF-8 spelling"
assert entries[0].is_active is True
assert entries[1].is_active is False
def test_malformed_yaml_does_not_replace_existing_glossary(tmp_path: Path) -> None:
glossary = repository(tmp_path)
existing = glossary.create("PBAT", "acronym")
with pytest.raises(GlossaryYamlError, match="Malformed glossary YAML"):
import_glossary_yaml(b"version: 1\nglossary: [")
assert glossary.list() == [existing]
@pytest.mark.parametrize(
"entries, message",
[
(
[
{
"id": 1,
"canonical_term": "PBAT",
"aliases": [],
"category": "acronym",
"description": None,
"active": True,
},
{
"id": 2,
"canonical_term": "pbat",
"aliases": [],
"category": "material",
"description": None,
"active": True,
},
],
"duplicate canonical term or alias",
),
(
[
{
"id": 1,
"canonical_term": "Secugrid",
"aliases": ["Sikirgut", "SIKIRGUT"],
"category": "product",
"description": None,
"active": True,
},
],
"aliases must be unique",
),
(
[
{
"id": 1,
"canonical_term": "Secugrid",
"aliases": ["PBAT"],
"category": "product",
"description": None,
"active": True,
},
{
"id": 2,
"canonical_term": "PBAT",
"aliases": [],
"category": "acronym",
"description": None,
"active": False,
},
],
"duplicate canonical term or alias",
),
],
)
def test_duplicate_terms_and_aliases_are_rejected(
entries: list[dict[str, object]], message: str
) -> None:
import yaml
content = yaml.safe_dump({"version": 1, "glossary": entries})
with pytest.raises(GlossaryYamlError, match=message):
import_glossary_yaml(content)
def test_validation_failure_leaves_database_unchanged(tmp_path: Path) -> None:
glossary = repository(tmp_path)
existing = glossary.create("Existing", "other", aliases=("Existing alias",))
invalid = b"""version: 1
glossary:
- id: 10
canonical_term: Duplicate
aliases: []
category: other
description: null
active: true
- id: 11
canonical_term: duplicate
aliases: []
category: other
description: null
active: false
"""
with pytest.raises(GlossaryYamlError):
imported = import_glossary_yaml(invalid)
glossary.replace_all(imported)
assert glossary.list() == [existing]
+88
View File
@@ -0,0 +1,88 @@
"""Exercise the real executor and Streamlit reruns without model calls."""
from pathlib import Path
from threading import current_thread
from types import SimpleNamespace
import pytest
import streamlit
from streamlit.testing.v1 import AppTest
from mka.application.meeting_service import AppProgressEvent, SpeakerMappingReview, SpeakerReview
from mka.application.progress_timing import ProcessingTimer
from mka.ui import streamlit_app as ui
class GuardedStreamlit:
"""Fail if a worker touches the UI session proxy, including callback reads."""
def __getattr__(self, name):
if name == "session_state" and current_thread().name.startswith("ThreadPoolExecutor"):
raise AssertionError("Worker accessed Streamlit session state")
return getattr(streamlit, name)
@pytest.mark.parametrize("fails", [False, True])
@pytest.mark.parametrize("mapped", [False, True])
def test_regeneration_captures_ui_values_before_real_worker_and_retains_timing(
monkeypatch, fails, mapped
):
outcome = SimpleNamespace(
succeeded=True,
run_dir=Path("/tmp/alpha-mapping-test"),
speaker_attribution_available=True,
original_protocol="Existing protocol",
)
review = SpeakerMappingReview((SpeakerReview("SPEAKER_00", ()),), (("a", "A"),), {})
calls = []
class Service:
def __init__(self, *args):
self.timer = ProcessingTimer()
def load_speaker_mapping_review(self, run_dir):
return review
def regenerate_protocol(self, run_dir, selections, *, progress_sink, performance_profile):
assert current_thread().name.startswith("ThreadPoolExecutor")
calls.append((run_dir, selections, performance_profile))
self.timer.start_run()
self.timer.start_stage("protocol_generation")
progress_sink(AppProgressEvent("protocol_generation", "started", 0.0))
if fails:
raise RuntimeError("model unavailable")
self.timer.finish_stage("protocol_generation")
self.timer.finish_run()
progress_sink(AppProgressEvent("protocol_generation", "completed", 0.0))
return outcome
monkeypatch.setattr(ui, "st", GuardedStreamlit())
monkeypatch.setattr(ui, "MeetingProcessingService", Service)
monkeypatch.setattr(ui.AppSettings, "from_environment", lambda: None)
monkeypatch.setattr(ui, "MeetingLabGateway", lambda: None)
def result_app(outcome):
import streamlit as st
from mka.ui.streamlit_app import _render_result
st.session_state.outcome = outcome
st.session_state.performance_profile = "fast"
_render_result()
app = AppTest.from_function(result_app, args=(outcome,)).run()
if mapped:
app.selectbox[0].select("a").run()
button = next(button for button in app.button if "Regenerate" in button.label)
assert not button.disabled
button.click().run()
assert not app.exception
assert calls == [(outcome.run_dir, {"SPEAKER_00": "a"} if mapped else {}, "fast")]
states, message, timing = app.session_state["processing_progress"]
assert states["protocol_generation"] == ("failed" if fails else "completed")
assert not timing.running
assert timing.total_duration >= timing.stage_durations["protocol_generation"] >= 0
app.run()
assert not app.exception
assert app.session_state["processing_progress"][2] == timing
assert any("total runtime" in item.value for item in app.info)
+583 -2
View File
@@ -5,6 +5,9 @@ from pathlib import Path
from types import SimpleNamespace from types import SimpleNamespace
from typing import Any from typing import Any
import pytest
import yaml
from mka.application.config import AppSettings from mka.application.config import AppSettings
from mka.application.meeting_service import ( from mka.application.meeting_service import (
MeetingDetails, MeetingDetails,
@@ -12,6 +15,7 @@ from mka.application.meeting_service import (
ParticipantInput, ParticipantInput,
ProcessingOptions, ProcessingOptions,
) )
from mka.application.progress_timing import ProcessingTimer
@dataclass @dataclass
@@ -29,6 +33,7 @@ class FakeMeetingLab:
self.context_data: dict[str, Any] | None = None self.context_data: dict[str, Any] | None = None
self.config_values: dict[str, Any] | None = None self.config_values: dict[str, Any] | None = None
self.fail = False self.fail = False
self.regeneration: dict[str, Any] | None = None
def create_context(self, data: dict[str, Any]) -> FakeContext: def create_context(self, data: dict[str, Any]) -> FakeContext:
self.context_data = data self.context_data = data
@@ -95,10 +100,58 @@ class FakeMeetingLab:
) )
) )
return SimpleNamespace(exit_code=2, run_dir=self.run_dir, protocol_path=None) return SimpleNamespace(exit_code=2, run_dir=self.run_dir, protocol_path=None)
if config["stop_after_diarization"]:
write_speaker_review_artifacts(self.run_dir)
(self.run_dir / "run_metadata.json").write_text(
json.dumps(
{
"status": "awaiting_speaker_review",
"diarization": {"speaker_count": 2},
}
),
encoding="utf-8",
)
progress_sink(
SimpleNamespace(
stage="diarization", status="completed", elapsed_seconds=0.3,
progress=None, message=None,
)
)
return SimpleNamespace(exit_code=0, run_dir=self.run_dir, protocol_path=None)
protocol = self.run_dir / "protocol.md" protocol = self.run_dir / "protocol.md"
protocol.write_text("# Generated protocol\n", encoding="utf-8") protocol.write_text("# Generated protocol\n", encoding="utf-8")
return SimpleNamespace(exit_code=0, run_dir=self.run_dir, protocol_path=protocol) return SimpleNamespace(exit_code=0, run_dir=self.run_dir, protocol_path=protocol)
def regenerate_protocol(
self,
run_dir: Path,
meeting_context: Any,
progress_sink: Any,
**options: Any,
) -> Any:
self.regeneration = {
"run_dir": run_dir,
"meeting_context": meeting_context,
"options": options,
}
progress_sink(
SimpleNamespace(
stage="protocol_generation",
status="started",
elapsed_seconds=0.0,
progress=None,
message=None,
)
)
protocol = Path(run_dir) / "protocol.md"
protocol.write_text("# Regenerated protocol\n", encoding="utf-8")
protocol_dir = Path(run_dir) / "protocol"
protocol_dir.mkdir(exist_ok=True)
(protocol_dir / "runtime_metadata.json").write_text(
json.dumps({"speaker_attribution_available": True}), encoding="utf-8"
)
return SimpleNamespace(exit_code=0, run_dir=run_dir, protocol_path=protocol)
def make_service(tmp_path: Path) -> tuple[MeetingProcessingService, FakeMeetingLab]: def make_service(tmp_path: Path) -> tuple[MeetingProcessingService, FakeMeetingLab]:
model = tmp_path / "model.bin" model = tmp_path / "model.bin"
@@ -107,6 +160,7 @@ def make_service(tmp_path: Path) -> tuple[MeetingProcessingService, FakeMeetingL
settings = AppSettings( settings = AppSettings(
data_root=tmp_path / "meetings", data_root=tmp_path / "meetings",
whisper_model=model, whisper_model=model,
glossary_database=tmp_path / "glossary.sqlite3",
whisper_executable="/opt/whisper-cli", whisper_executable="/opt/whisper-cli",
protocol_model="test:model", protocol_model="test:model",
diarization_mode="gpu", diarization_mode="gpu",
@@ -139,6 +193,72 @@ def participants() -> list[ParticipantInput]:
] ]
def write_speaker_review_artifacts(run_dir: Path) -> Path:
diarization_dir = run_dir / "diarization"
context_dir = run_dir / "context"
diarization_dir.mkdir(parents=True)
context_dir.mkdir()
transcript_path = diarization_dir / "transcript_diarized.json"
transcript_path.write_text(
json.dumps(
{
"speaker_labels_anonymous": True,
"segments": [
{
"speaker_id": "SPEAKER_01",
"text": "I will prepare all raw materials before Wednesday.",
},
{"speaker_id": "SPEAKER_00", "text": "Yes."},
{
"speaker_id": "SPEAKER_00",
"text": "We will run the production trial on Wednesday.",
},
{
"speaker_id": "SPEAKER_00",
"text": "The trial requires the complete production team.",
},
{"speaker_id": None, "text": "Unassigned text."},
],
}
),
encoding="utf-8",
)
(context_dir / "meeting_context.yaml").write_text(
json.dumps(
{
"schema_version": "1",
"meeting": {
"meeting_id": "speaker-review",
"title": "Speaker review",
"language": "en",
},
"participants": [
{"participant_id": "martin", "display_name": "Martin"},
{"participant_id": "anna", "display_name": "Anna"},
],
"speaker_mappings": {},
"mentioned_people": [],
"organization": {"departments": []},
"known_entities": {},
}
),
encoding="utf-8",
)
return transcript_path
def write_custom_speaker_review_artifacts(
run_dir: Path, segments: list[dict[str, str | None]]
) -> Path:
"""Write custom diarized segments using the standard review context fixture."""
transcript_path = write_speaker_review_artifacts(run_dir)
transcript_path.write_text(
json.dumps({"speaker_labels_anonymous": True, "segments": segments}),
encoding="utf-8",
)
return transcript_path
def test_build_context_uses_actual_v1_shape(tmp_path: Path) -> None: def test_build_context_uses_actual_v1_shape(tmp_path: Path) -> None:
service, gateway = make_service(tmp_path) service, gateway = make_service(tmp_path)
@@ -171,6 +291,40 @@ def test_build_context_preserves_explicit_speaker_mapping(tmp_path: Path) -> Non
assert gateway.context_data["speaker_mappings"] == {"SPEAKER_00": "martin"} assert gateway.context_data["speaker_mappings"] == {"SPEAKER_00": "martin"}
def test_build_context_includes_only_active_authoritative_glossary_terms(
tmp_path: Path,
) -> None:
service, gateway = make_service(tmp_path)
service.glossary.create("Secugrid HS", "product", aliases=("Sikirgut", "Secugrid H S"))
inactive = service.glossary.create("Old Name", "other")
service.glossary.set_active(inactive.id, False)
service.build_context(meeting(), participants())
assert gateway.context_data is not None
assert gateway.context_data["known_entities"] == {
"Authoritative terminology": ["Secugrid HS (aliases: Secugrid H S, Sikirgut)"]
}
rules = gateway.context_data["context_rules"]
assert "do not invent matches" in rules["glossary_canonical_spelling"]
assert "inside compounds" in rules["glossary_core_terms"]
assert "Old Name" not in str(gateway.context_data)
def test_glossary_is_rendered_into_meeting_lab_protocol_context(tmp_path: Path) -> None:
meeting_context = pytest.importorskip("src.meeting_lab.models.meeting_context")
service, gateway = make_service(tmp_path)
service.glossary.create("PBAT", "acronym", aliases=("P B A T",))
context = service.build_context(meeting(), participants())
real_context = meeting_context.create_meeting_context(context.data)
prompt_context = meeting_context.render_meeting_context_for_prompt(real_context)
assert "Authoritative terminology: PBAT (aliases: P B A T)" in prompt_context
assert "Use canonical glossary spellings" in prompt_context
assert "use canonical core terms inside compounds" in prompt_context
def test_build_context_translates_mentioned_only_person(tmp_path: Path) -> None: def test_build_context_translates_mentioned_only_person(tmp_path: Path) -> None:
service, gateway = make_service(tmp_path) service, gateway = make_service(tmp_path)
people = participants() + [ people = participants() + [
@@ -201,21 +355,32 @@ def test_build_context_translates_mentioned_only_person(tmp_path: Path) -> None:
] ]
@pytest.mark.parametrize("language", ["de", "en"])
def test_process_translates_configuration_and_disables_diarization( def test_process_translates_configuration_and_disables_diarization(
tmp_path: Path, tmp_path: Path,
language: str,
) -> None: ) -> None:
service, gateway = make_service(tmp_path) service, gateway = make_service(tmp_path)
audio = tmp_path / "meeting.wav" audio = tmp_path / "meeting.wav"
audio.write_bytes(b"audio") audio.write_bytes(b"audio")
outcome = service.process( outcome = service.process(
audio, meeting(), participants(), ProcessingOptions(diarization_enabled=False) audio,
replace(meeting(), language=language),
participants(),
ProcessingOptions(diarization_enabled=False),
) )
assert outcome.succeeded assert outcome.succeeded
assert gateway.config_values is not None assert gateway.config_values is not None
assert gateway.config_values["language"] == language
assert service.build_context(replace(meeting(), language=language), participants()).data[
"meeting"
]["language"] == language
assert gateway.config_values["diarization"] == "off" assert gateway.config_values["diarization"] == "off"
assert gateway.config_values["model"] == "test:model" assert gateway.config_values["model"] == "test:model"
assert gateway.config_values["protocol_num_ctx"] == 32_768
assert gateway.config_values["protocol_safe_input_token_budget"] == 29_000
assert gateway.config_values["whisper_executable"] == "/opt/whisper-cli" assert gateway.config_values["whisper_executable"] == "/opt/whisper-cli"
assert gateway.config_values["ffmpeg_executable"] == "ffmpeg" assert gateway.config_values["ffmpeg_executable"] == "ffmpeg"
assert gateway.config_values["audio_normalization"] is True assert gateway.config_values["audio_normalization"] is True
@@ -253,7 +418,7 @@ def test_process_propagates_enabled_diarization_and_progress(tmp_path: Path) ->
audio.write_bytes(b"audio") audio.write_bytes(b"audio")
events = [] events = []
service.process( outcome = service.process(
audio, audio,
meeting(), meeting(),
participants(), participants(),
@@ -263,6 +428,9 @@ def test_process_propagates_enabled_diarization_and_progress(tmp_path: Path) ->
assert gateway.config_values is not None assert gateway.config_values is not None
assert gateway.config_values["diarization"] == "gpu" assert gateway.config_values["diarization"] == "gpu"
assert gateway.config_values["stop_after_diarization"] is True
assert outcome.awaiting_speaker_review is True
assert outcome.detected_speaker_count == 2
assert [(event.stage, event.status) for event in events[:2]] == [ assert [(event.stage, event.status) for event in events[:2]] == [
("preparing", "started"), ("preparing", "started"),
("preparing", "completed"), ("preparing", "completed"),
@@ -270,6 +438,33 @@ def test_process_propagates_enabled_diarization_and_progress(tmp_path: Path) ->
assert events[2].progress == 0.25 assert events[2].progress == 0.25
def test_diarization_checkpoint_defers_protocol_but_disabled_run_does_not(tmp_path: Path) -> None:
service, gateway = make_service(tmp_path)
audio = tmp_path / "meeting.wav"
audio.write_bytes(b"audio")
checkpoint = service.process(
audio, meeting(), participants(), ProcessingOptions(diarization_enabled=True)
)
assert checkpoint.succeeded
assert checkpoint.awaiting_speaker_review
assert checkpoint.original_protocol is None
assert checkpoint.protocol_path is None
normal_root = tmp_path / "without-diarization"
normal_root.mkdir()
normal_service, normal_gateway = make_service(normal_root)
normal_audio = normal_root / "meeting.wav"
normal_audio.write_bytes(b"audio")
normal = normal_service.process(
normal_audio, meeting(), participants(), ProcessingOptions(diarization_enabled=False)
)
assert normal.succeeded
assert not normal.awaiting_speaker_review
assert normal.protocol_path is not None
assert normal_gateway.config_values["stop_after_diarization"] is False
def test_process_propagates_ordered_diarization_container_args(tmp_path: Path) -> None: def test_process_propagates_ordered_diarization_container_args(tmp_path: Path) -> None:
service, gateway = make_service(tmp_path) service, gateway = make_service(tmp_path)
service.settings = replace( service.settings = replace(
@@ -329,6 +524,223 @@ def test_failure_reports_stage_and_preserves_run_dir(tmp_path: Path) -> None:
assert (gateway.run_dir / "run_metadata.json").is_file() assert (gateway.run_dir / "run_metadata.json").is_file()
def test_speaker_review_lists_detected_labels_participants_and_excerpts(
tmp_path: Path,
) -> None:
service, gateway = make_service(tmp_path)
source = write_speaker_review_artifacts(gateway.run_dir)
review = service.load_speaker_mapping_review(gateway.run_dir, excerpts_per_speaker=2)
assert review is not None
assert [speaker.speaker_label for speaker in review.speakers] == [
"SPEAKER_00",
"SPEAKER_01",
]
assert review.speakers[0].excerpts == (
"We will run the production trial on Wednesday.",
"The trial requires the complete production team.",
)
assert review.participants == (("martin", "Martin"), ("anna", "Anna"))
assert review.current_mappings == {}
assert "SPEAKER_00" in source.read_text(encoding="utf-8")
def test_speaker_review_prefers_content_rich_excerpts_over_acknowledgements_and_questions(
tmp_path: Path,
) -> None:
service, gateway = make_service(tmp_path)
acknowledgement = "Okay, that sounds good."
question = "Could we continue now?"
rich_excerpts = [
"I will complete the Atlas deployment in Berlin on 14 September with the database team.",
"I recommend moving the Orion product launch because the supplier test is incomplete.",
"We decided that I will own the Kubernetes migration for the payments service.",
"My experience with the Frankfurt factory rollout suggests a two-week validation period.",
"I will present the security audit results to the Apollo project steering group.",
"We should reserve twelve hours for the API integration and production monitoring.",
]
write_custom_speaker_review_artifacts(
gateway.run_dir,
[{"speaker_id": "SPEAKER_00", "text": text} for text in [acknowledgement, question, *rich_excerpts]],
)
review = service.load_speaker_mapping_review(gateway.run_dir)
assert review is not None
excerpts = review.speakers[0].excerpts
assert len(excerpts) == 6
assert acknowledgement not in excerpts
assert question not in excerpts
assert set(excerpts) == set(rich_excerpts)
def test_speaker_review_avoids_near_duplicate_excerpts(tmp_path: Path) -> None:
service, gateway = make_service(tmp_path)
primary = "I will lead the Atlas API migration for the Berlin platform team next Monday."
duplicate = "I will lead the Atlas API migration for the Berlin platform group next Monday."
alternatives = [
"I recommend documenting the supplier quality decision in the project register.",
"My team will complete the production calibration before the factory trial.",
"I will present the customer research findings at the Munich planning workshop.",
"Our database monitoring threshold needs approval from the operations group.",
"I will coordinate the Orion release checklist with the support organization.",
]
write_custom_speaker_review_artifacts(
gateway.run_dir,
[{"speaker_id": "SPEAKER_00", "text": text} for text in [primary, duplicate, *alternatives]],
)
review = service.load_speaker_mapping_review(gateway.run_dir)
assert review is not None
excerpts = review.speakers[0].excerpts
assert primary in excerpts
assert duplicate not in excerpts
assert set(excerpts) == {primary, *alternatives}
def test_speaker_review_spreads_selected_excerpts_across_meeting_positions(
tmp_path: Path,
) -> None:
service, gateway = make_service(tmp_path)
positions = {
0: "I will lead the Atlas deployment review with the engineering team marker alpha.",
1: "I will lead the Atlas deployment review with the engineering team marker bravo.",
6: "I will lead the Atlas deployment review with the engineering team marker charlie.",
7: "I will lead the Atlas deployment review with the engineering team marker delta.",
12: "I will lead the Atlas deployment review with the engineering team marker echo.",
13: "I will lead the Atlas deployment review with the engineering team marker foxtrot.",
18: "I will lead the Atlas deployment review with the engineering team marker golf.",
19: "I will lead the Atlas deployment review with the engineering team marker hotel.",
}
segments = [
{
"speaker_id": "SPEAKER_00" if index in positions else "SPEAKER_01",
"text": positions.get(index, "Acknowledged."),
}
for index in range(20)
]
write_custom_speaker_review_artifacts(gateway.run_dir, segments)
review = service.load_speaker_mapping_review(gateway.run_dir)
assert review is not None
assert review.speakers[0].excerpts == tuple(positions[index] for index in (0, 1, 6, 7, 12, 18))
def test_protocol_only_regeneration_persists_mappings_without_rewriting_transcript(
tmp_path: Path,
) -> None:
service, gateway = make_service(tmp_path)
service.glossary.create("ENLYZE", "organization", aliases=("Enlyse",))
source = write_speaker_review_artifacts(gateway.run_dir)
original_source = source.read_bytes()
events = []
outcome = service.regenerate_protocol(
gateway.run_dir,
{"SPEAKER_00": "martin"},
progress_sink=events.append,
)
assert outcome.succeeded
assert outcome.original_protocol == "# Regenerated protocol\n"
assert outcome.speaker_attribution_available is True
assert gateway.context_data is not None
assert gateway.context_data["speaker_mappings"] == {"SPEAKER_00": "martin"}
assert gateway.context_data["known_entities"] == {
"Authoritative terminology": ["ENLYZE (aliases: Enlyse)"]
}
assert gateway.regeneration is not None
assert gateway.regeneration["meeting_context"].data["speaker_mappings"] == {
"SPEAKER_00": "martin"
}
assert gateway.regeneration["options"]["protocol_num_ctx"] == 32_768
assert gateway.regeneration["options"]["glossary_aliases"] == {"Enlyse": "ENLYZE"}
assert events[0].stage == "protocol_generation"
assert source.read_bytes() == original_source
def test_protocol_regeneration_allows_unmapped_and_rejects_duplicate_participant(
tmp_path: Path,
) -> None:
service, gateway = make_service(tmp_path)
write_speaker_review_artifacts(gateway.run_dir)
outcome = service.regenerate_protocol(gateway.run_dir, {})
assert outcome.succeeded
assert gateway.context_data is not None
assert gateway.context_data["speaker_mappings"] == {}
with pytest.raises(ValueError, match="only one speaker"):
service.regenerate_protocol(
gateway.run_dir,
{"SPEAKER_00": "martin", "SPEAKER_01": "martin"},
)
@pytest.mark.parametrize(
"mappings",
[
{"SPEAKER_00": "martin"},
{"SPEAKER_00": "martin", "SPEAKER_01": "anna"},
],
)
def test_protocol_generation_uses_partial_and_complete_confirmed_mappings(
tmp_path: Path, mappings: dict[str, str]
) -> None:
service, gateway = make_service(tmp_path)
write_speaker_review_artifacts(gateway.run_dir)
outcome = service.regenerate_protocol(gateway.run_dir, mappings)
assert outcome.succeeded
assert gateway.regeneration is not None
assert gateway.regeneration["meeting_context"].data["speaker_mappings"] == mappings
def test_failed_protocol_generation_keeps_checkpoint_mapping_for_retry(tmp_path: Path) -> None:
service, gateway = make_service(tmp_path)
source = write_speaker_review_artifacts(gateway.run_dir)
original_source = source.read_bytes()
original_regeneration = gateway.regenerate_protocol
def fail_regeneration(*args, **kwargs):
raise RuntimeError("model unavailable")
gateway.regenerate_protocol = fail_regeneration
with pytest.raises(RuntimeError, match="model unavailable"):
service.regenerate_protocol(gateway.run_dir, {"SPEAKER_00": "martin"})
persisted = yaml.safe_load(
(gateway.run_dir / "context" / "meeting_context.yaml").read_text(encoding="utf-8")
)
assert persisted["speaker_mappings"] == {"SPEAKER_00": "martin"}
gateway.regenerate_protocol = original_regeneration
retry = service.regenerate_protocol(gateway.run_dir, {"SPEAKER_00": "martin"})
assert retry.succeeded
assert source.read_bytes() == original_source
def test_fallback_attribution_loss_is_read_from_runtime_metadata(tmp_path: Path) -> None:
run_dir = tmp_path / "run"
protocol_dir = run_dir / "protocol"
protocol_dir.mkdir(parents=True)
(protocol_dir / "runtime_metadata.json").write_text(
json.dumps(
{
"speaker_attribution_available": False,
"speaker_attribution_loss_reason": "plain_transcript_fallback",
}
),
encoding="utf-8",
)
assert MeetingProcessingService._speaker_attribution_available(run_dir) is False
def test_uploaded_source_is_preserved_in_meeting_directory(tmp_path: Path) -> None: def test_uploaded_source_is_preserved_in_meeting_directory(tmp_path: Path) -> None:
service, _ = make_service(tmp_path) service, _ = make_service(tmp_path)
source = SimpleNamespace(getbuffer=lambda: b"source audio") source = SimpleNamespace(getbuffer=lambda: b"source audio")
@@ -338,3 +750,172 @@ def test_uploaded_source_is_preserved_in_meeting_directory(tmp_path: Path) -> No
assert destination.parent == tmp_path / "meetings" / "meeting-1" / "uploads" assert destination.parent == tmp_path / "meetings" / "meeting-1" / "uploads"
assert destination.name.endswith("_unsafe.wav") assert destination.name.endswith("_unsafe.wav")
assert destination.read_bytes() == b"source audio" assert destination.read_bytes() == b"source audio"
@pytest.mark.parametrize(
("profile", "threads"),
[("fast", 16), ("efficient", 10), ("powersave", 4)],
)
def test_process_resolves_profile_for_protocol_generation(
tmp_path: Path, profile: str, threads: int
) -> None:
service, gateway = make_service(tmp_path)
audio = tmp_path / "meeting.wav"
audio.write_bytes(b"audio")
service.process(
audio,
meeting(),
participants(),
ProcessingOptions(performance_profile=profile),
)
assert gateway.config_values is not None
assert gateway.config_values["protocol_num_thread"] == threads
def test_process_defaults_to_backend_thread_selection(tmp_path: Path) -> None:
service, gateway = make_service(tmp_path)
audio = tmp_path / "meeting.wav"
audio.write_bytes(b"audio")
service.process(audio, meeting(), participants(), ProcessingOptions())
assert gateway.config_values is not None
assert gateway.config_values["protocol_num_thread"] is None
def test_protocol_regeneration_propagates_explicit_profile(tmp_path: Path) -> None:
service, gateway = make_service(tmp_path)
write_speaker_review_artifacts(gateway.run_dir)
service.regenerate_protocol(
gateway.run_dir,
{},
performance_profile="fast",
)
assert gateway.regeneration is not None
assert gateway.regeneration["options"]["protocol_num_thread"] == 16
def test_process_passes_only_active_explicit_glossary_aliases(tmp_path: Path) -> None:
service, gateway = make_service(tmp_path)
service.glossary.create("Luminy", "product", aliases=("Lumini",))
inactive = service.glossary.create("PBAT", "acronym", aliases=("PBRT",))
service.glossary.set_active(inactive.id, False)
audio = tmp_path / "meeting.wav"
audio.write_bytes(b"audio")
service.process(audio, meeting(), participants(), ProcessingOptions())
assert gateway.config_values is not None
assert gateway.config_values["glossary_aliases"] == {"Lumini": "Luminy"}
@pytest.mark.parametrize("fails", [False, True])
def test_process_persists_completed_and_failed_timings(tmp_path: Path, fails: bool) -> None:
service, gateway = make_service(tmp_path)
gateway.fail = fails
values = iter([0.0, 1.0, 9.0, 10.0, 20.0])
service.timer = ProcessingTimer(lambda: next(values))
audio = tmp_path / "meeting.wav"
audio.write_bytes(b"audio")
outcome = service.process(audio, meeting(), participants(), ProcessingOptions())
metadata = json.loads((gateway.run_dir / "run_metadata.json").read_text(encoding="utf-8"))
assert metadata["timing"] == {
"stages_seconds": {"preparing": 8.0, "transcription": 10.0},
"total_seconds": 20.0,
}
assert outcome.succeeded is not fails
if fails:
assert metadata["failure"]["type"] == "TranscriptionError"
def test_regeneration_starts_fresh_timer_and_retains_final_duration(
tmp_path: Path, monkeypatch: pytest.MonkeyPatch
) -> None:
service, gateway = make_service(tmp_path)
write_speaker_review_artifacts(gateway.run_dir)
clock = SimpleNamespace(value=10.0)
service.timer = ProcessingTimer(lambda: clock.value)
audio = tmp_path / "meeting.wav"
audio.write_bytes(b"audio")
service.process(audio, meeting(), participants(), ProcessingOptions())
clock.value = 100.0
observed = []
def regenerate(run_dir: Path, context: Any, progress_sink: Any, **options: Any) -> Any:
progress_sink(
SimpleNamespace(
stage="protocol_generation",
status="started",
elapsed_seconds=0.0,
progress=None,
message=None,
)
)
clock.value = 105.0
observed.append(service.timer.snapshot())
progress_sink(
SimpleNamespace(
stage="protocol_generation",
status="completed",
elapsed_seconds=5.0,
progress=None,
message=None,
)
)
clock.value = 109.0
protocol = run_dir / "protocol.md"
protocol.write_text("# Regenerated protocol\n", encoding="utf-8")
return SimpleNamespace(run_dir=run_dir, protocol_path=protocol)
monkeypatch.setattr(gateway, "regenerate_protocol", regenerate)
service.regenerate_protocol(gateway.run_dir, {"SPEAKER_00": "martin"})
final = service.timer.snapshot()
assert observed[0].active_stage == "protocol_generation"
assert observed[0].stage_durations == {"protocol_generation": 5.0}
assert observed[0].total_duration == 5.0
assert observed[0].running is True
assert final.stage_durations == {"protocol_generation": 5.0}
assert final.total_duration == 9.0
assert final.running is False
def test_failed_regeneration_freezes_active_and_total_durations(
tmp_path: Path, monkeypatch: pytest.MonkeyPatch
) -> None:
service, gateway = make_service(tmp_path)
write_speaker_review_artifacts(gateway.run_dir)
clock = SimpleNamespace(value=20.0)
service.timer = ProcessingTimer(lambda: clock.value)
def fail_regeneration(run_dir: Path, context: Any, progress_sink: Any, **options: Any) -> Any:
progress_sink(
SimpleNamespace(
stage="protocol_generation",
status="started",
elapsed_seconds=0.0,
progress=None,
message=None,
)
)
clock.value = 27.0
raise RuntimeError("generation failed")
monkeypatch.setattr(gateway, "regenerate_protocol", fail_regeneration)
with pytest.raises(RuntimeError, match="generation failed"):
service.regenerate_protocol(gateway.run_dir, {})
clock.value = 99.0
final = service.timer.snapshot()
assert final.stage_durations == {"protocol_generation": 7.0}
assert final.total_duration == 7.0
assert final.running is False
+21
View File
@@ -0,0 +1,21 @@
import pytest
from mka.application.meeting_service import ProcessingOptions
from mka.application.performance import PERFORMANCE_PROFILES, resolve_ollama_num_thread
def test_default_profile_is_auto() -> None:
assert ProcessingOptions().performance_profile == "auto"
assert PERFORMANCE_PROFILES[0] == "auto"
def test_auto_profile_has_no_ollama_thread_override() -> None:
assert resolve_ollama_num_thread("auto") is None
@pytest.mark.parametrize(
("profile", "threads"),
[("fast", 16), ("efficient", 10), ("powersave", 4)],
)
def test_profiles_resolve_to_current_ollama_thread_counts(profile: str, threads: int) -> None:
assert resolve_ollama_num_thread(profile) == threads
+45
View File
@@ -0,0 +1,45 @@
from mka.application.progress_timing import ProcessingTimer
class FakeClock:
def __init__(self, value: float = 0.0) -> None:
self.value = value
def __call__(self) -> float:
return self.value
def test_completed_duration_is_retained_while_active_stage_increases() -> None:
clock = FakeClock(10.0)
timer = ProcessingTimer(clock)
timer.start_run()
timer.start_stage("preparing")
clock.value = 18.0
timer.finish_stage("preparing")
timer.start_stage("transcription")
clock.value = 20.0
first = timer.snapshot()
clock.value = 25.5
second = timer.snapshot()
assert first.stage_durations == {"preparing": 8.0, "transcription": 2.0}
assert second.stage_durations == {"preparing": 8.0, "transcription": 7.5}
assert second.total_duration == 15.5
def test_failed_active_stage_and_total_duration_are_frozen() -> None:
clock = FakeClock(2.0)
timer = ProcessingTimer(clock)
timer.start_run()
clock.value = 5.0
timer.start_stage("transcription")
clock.value = 14.0
timer.finish_run()
clock.value = 99.0
snapshot = timer.snapshot()
assert snapshot.stage_durations["transcription"] == 9.0
assert snapshot.total_duration == 12.0
assert snapshot.running is False
+136
View File
@@ -0,0 +1,136 @@
import json
from datetime import date
import pytest
from mka.application.meeting_service import ParticipantInput
from mka.application.run_inputs import (
RunInputJsonError,
RunInputState,
export_run_inputs,
import_run_inputs,
run_input_filename,
)
def populated_state() -> RunInputState:
return RunInputState(
title="F&E Technik Abteilungs Jour Fixe",
description="Review the production trial.",
language="de",
meeting_date=date(2026, 8, 25),
participants=(
ParticipantInput(
participant_id="martin",
display_name="Martin Tazl",
role="Project lead",
organization="Engineering",
),
ParticipantInput(
participant_id="alex",
display_name="Alexander Funk",
attendance_status="mentioned_only",
),
),
audio_normalization=False,
diarization_enabled=True,
source_file_name="2026-08-25_jour_fixe.wav",
)
def test_current_form_state_serializes_with_schema_and_all_supported_values() -> None:
document = json.loads(export_run_inputs(populated_state()))
assert document["schema_version"] == 1
assert document["meeting"]["title"] == "F&E Technik Abteilungs Jour Fixe"
assert document["meeting"]["description"] == "Review the production trial."
assert document["meeting"]["language"] == "de"
assert document["meeting"]["date"] == "2026-08-25"
assert document["meeting"]["participants"][1] == {
"participant_id": "alex",
"display_name": "Alexander Funk",
"role": "",
"organization": "",
"attendance_status": "mentioned_only",
}
assert document["processing"] == {
"audio_normalization": False,
"diarization_enabled": True,
}
def test_export_import_round_trip_restores_form_and_participants() -> None:
original = populated_state()
restored = import_run_inputs(export_run_inputs(original))
assert restored == original
assert restored.participants[0].participant_id == "martin"
assert restored.participants[0].display_name == "Martin Tazl"
def test_missing_optional_fields_use_current_defaults() -> None:
restored = import_run_inputs('{"schema_version": 1}')
defaults = RunInputState.defaults()
assert restored == defaults
def test_unknown_safe_fields_are_ignored() -> None:
restored = import_run_inputs(
json.dumps(
{
"schema_version": 1,
"meeting": {"title": "Known", "future_field": {"value": 1}},
"processing": {"future_toggle": True},
"future_section": [1, 2, 3],
}
)
)
assert restored.title == "Known"
@pytest.mark.parametrize("content", ["not-json", "[]", b"\xff"])
def test_malformed_json_is_rejected(content: str | bytes) -> None:
with pytest.raises(RunInputJsonError, match="Malformed|top-level"):
import_run_inputs(content)
def test_unsupported_future_schema_is_rejected() -> None:
with pytest.raises(RunInputJsonError, match="Unsupported.*schema_version"):
import_run_inputs('{"schema_version": 2}')
def test_media_contents_are_never_serialized() -> None:
exported = export_run_inputs(populated_state())
assert "2026-08-25_jour_fixe.wav" in exported
assert "audio_bytes" not in exported
assert "base64" not in exported
def test_source_filename_must_not_be_a_machine_specific_path() -> None:
with pytest.raises(RunInputJsonError, match="filename, not a path"):
import_run_inputs('{"schema_version": 1, "source_file_name": "/tmp/meeting.wav"}')
def test_invalid_or_duplicate_participants_are_not_silently_reinterpreted() -> None:
document = {
"schema_version": 1,
"meeting": {
"participants": [
{"participant_id": "martin", "display_name": "Martin"},
{"participant_id": "martin", "display_name": "Someone else"},
]
},
}
with pytest.raises(RunInputJsonError, match="Duplicate participant_id"):
import_run_inputs(json.dumps(document))
def test_export_filename_is_readable_and_machine_independent() -> None:
assert run_input_filename("F&E Technik Jour Fixe") == (
"meeting-inputs-f-e-technik-jour-fixe.json"
)
+134
View File
@@ -0,0 +1,134 @@
from pathlib import Path
from types import SimpleNamespace
import pytest
from streamlit.testing.v1 import AppTest
from mka.ui.streamlit_app import _mapping_counts, _speaker_options
@pytest.mark.parametrize("label", ["SPEAKER_00", "SPEAKER_01", "SPEAKER_02"])
def test_unmapped_options_include_all_participants(label):
assert _speaker_options(("a", "b"), {}, label) == [None, "a", "b"]
def test_options_reserve_other_assignments_and_preserve_current():
mappings = {"SPEAKER_00": "a"}
assert _speaker_options(("a", "b"), mappings, "SPEAKER_00") == [None, "a", "b"]
for label in ("SPEAKER_01", "SPEAKER_02"):
assert _speaker_options(("a", "b"), mappings, label) == [None, "b"]
mappings["SPEAKER_00"] = "b"
assert _speaker_options(("a", "b"), mappings, "SPEAKER_01") == [None, "a"]
mappings["SPEAKER_00"] = None
assert _speaker_options(("a", "b"), mappings, "SPEAKER_01") == [None, "a", "b"]
def test_existing_duplicate_or_missing_participant_is_not_removed():
assert _speaker_options(("b",), {"s0": "a", "s1": "a"}, "s0") == [None, "b", "a"]
@pytest.mark.parametrize(
("mappings", "expected"),
[
({}, (2, 0, 2)),
({"s0": "a", "other": "b"}, (2, 1, 1)),
({"s0": "a", "s1": "b"}, (2, 2, 0)),
({"s0": None}, (2, 0, 2)),
],
)
def test_counts_only_include_detected_speakers(mappings, expected):
assert _mapping_counts(("s0", "s1"), mappings) == expected
def mapping_app():
from mka.application.meeting_service import SpeakerMappingReview, SpeakerReview
from mka.ui.streamlit_app import _render_speaker_mapping
_render_speaker_mapping(
SpeakerMappingReview(
tuple(SpeakerReview(f"SPEAKER_0{i}", ()) for i in range(3)),
(("a", "Participant A"), ("b", "Participant B"), ("c", "Participant C")),
{"SPEAKER_00": "a"},
),
"test",
)
def test_widget_reruns_filter_release_and_update_status():
app = AppTest.from_function(mapping_app).run()
assert not app.exception
assert app.selectbox[0].value == "a"
assert "Participant A" not in app.selectbox[1].options
assert len(app.warning) == 1
assert "3 speakers detected · 1 assigned · 2 unassigned" in app.markdown[0].value
assert "Unassigned" in app.selectbox[1].label
app.selectbox[2].select("c").run()
assert app.selectbox[0].value == "a"
assert "Participant C" not in app.selectbox[0].options
app.selectbox[1].select("b").run()
assert not app.warning
assert "3 assigned · 0 unassigned" in app.markdown[0].value
app.selectbox[0].select(None).run()
assert app.warning
assert "Participant A" in app.selectbox[1].options
assert app.selectbox[1].value == "b"
app.selectbox[1].select("a").run()
assert "Participant B" in app.selectbox[0].options
assert "Participant A" not in app.selectbox[0].options
assert app.selectbox[0].value is None
assert not app.exception
def test_checkpoint_allows_anonymous_protocol_generation(monkeypatch):
from mka.application.meeting_service import SpeakerMappingReview, SpeakerReview
from mka.ui import streamlit_app as ui
outcome = SimpleNamespace(
succeeded=True,
run_dir=Path("/tmp/mapping-test"),
speaker_attribution_available=True,
original_protocol="Anonymous protocol",
awaiting_speaker_review=True,
detected_speaker_count=1,
)
review = SpeakerMappingReview((SpeakerReview("SPEAKER_00", ()),), (("a", "A"),), {})
calls = []
class Service:
def __init__(self, *args):
pass
def load_speaker_mapping_review(self, run_dir):
return review
def regenerate_protocol(self, run_dir, selections, **kwargs):
calls.append(selections)
return outcome
monkeypatch.setattr(ui, "MeetingProcessingService", Service)
monkeypatch.setattr(ui.AppSettings, "from_environment", lambda: None)
monkeypatch.setattr(ui, "MeetingLabGateway", lambda: None)
monkeypatch.setattr(ui, "_remember_regeneration_timing", lambda *args: None, raising=False)
monkeypatch.setattr(
ui,
"_run_with_live_progress",
lambda service, action, states, message: (action(None), states, message, None),
raising=False,
)
def result_app(outcome):
import streamlit as st
from mka.ui.streamlit_app import _render_result
st.session_state.outcome = outcome
_render_result()
app = AppTest.from_function(result_app, args=(outcome,)).run()
generate = next(button for button in app.button if button.label == "Generate protocol")
assert not generate.disabled
assert app.warning
generate.click().run()
assert calls == [{}]
assert not app.exception
+95
View File
@@ -0,0 +1,95 @@
from datetime import date
from mka.application.meeting_service import ParticipantInput
from mka.application.progress_timing import TimingSnapshot
from mka.application.run_inputs import RunInputState
from mka.ui import streamlit_app
def test_alias_input_accepts_lines_and_commas() -> None:
assert streamlit_app._parse_aliases("Sikirgut\nSekugrid HS, Secugrid H S") == (
"Sikirgut",
"Sekugrid HS",
"Secugrid H S",
)
def test_regenerated_protocol_widget_value_is_deferred_until_next_run(
monkeypatch,
) -> None:
state = {"edited_protocol": "old protocol"}
monkeypatch.setattr(streamlit_app.st, "session_state", state)
streamlit_app._queue_edited_protocol("regenerated protocol")
assert state["edited_protocol"] == "old protocol"
assert state["pending_edited_protocol"] == "regenerated protocol"
streamlit_app._apply_pending_edited_protocol()
assert state["edited_protocol"] == "regenerated protocol"
assert "pending_edited_protocol" not in state
def test_imported_inputs_are_applied_via_pending_state_before_widgets(
monkeypatch,
) -> None:
existing_upload = object()
state = {"source_media_0": existing_upload, "source_media_widget_generation": 0}
monkeypatch.setattr(streamlit_app.st, "session_state", state)
imported = RunInputState(
title="Imported meeting",
description="Imported context",
language="en",
meeting_date=date(2026, 8, 25),
participants=(ParticipantInput("martin", "Martin"),),
audio_normalization=False,
diarization_enabled=True,
source_file_name="meeting.wav",
)
streamlit_app._queue_run_inputs(imported)
assert "meeting_title" not in state
assert state["pending_run_inputs"] == imported
streamlit_app._apply_pending_run_inputs()
assert state["meeting_title"] == "Imported meeting"
assert state["meeting_description"] == "Imported context"
assert state["meeting_language"] == "en"
assert state["meeting_has_date"] is True
assert state["participants"][0]["participant_id"] == "martin"
assert state["audio_normalization"] is False
assert state["diarization_enabled"] is True
assert state["imported_source_file_name"] == "meeting.wav"
assert state["source_media_widget_generation"] == 1
assert "source_media_1" not in state
assert "pending_run_inputs" not in state
assert state["source_media_0"] is existing_upload
def test_duration_formatting_covers_seconds_minutes_and_hours() -> None:
assert streamlit_app._format_duration(8.9) == "00:08"
assert streamlit_app._format_duration(7 * 60 + 18) == "07:18"
assert streamlit_app._format_duration(3600 + 3 * 60 + 42) == "1:03:42"
def test_completed_regeneration_timing_survives_frontend_rerun(monkeypatch) -> None:
state = {}
monkeypatch.setattr(streamlit_app.st, "session_state", state)
timing = TimingSnapshot(
stage_durations={"protocol_generation": 4.5},
active_stage=None,
total_duration=4.5,
running=False,
)
states = {stage: "skipped" for stage in streamlit_app.STAGES}
states["protocol_generation"] = "completed"
streamlit_app._remember_regeneration_timing(states, "Protocol regenerated", timing)
saved_states, saved_message, saved_timing = state["regeneration_progress"]
assert saved_states["protocol_generation"] == "completed"
assert saved_message == "Protocol regenerated"
assert saved_timing == timing