Compare commits
5
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
6fc07690d9 | ||
|
|
8a0f38fce4 | ||
|
|
df89a38829 | ||
|
|
d77bfedb6e | ||
|
|
d94436af43 |
@@ -1,5 +1,18 @@
|
|||||||
# Project Knowledge
|
# Project Knowledge
|
||||||
|
|
||||||
|
## Meeting language in direct protocols
|
||||||
|
|
||||||
|
Direct protocol prompts derive their explicit output language from persisted
|
||||||
|
Meeting Context `meeting.language`, including regeneration and diarized fallback.
|
||||||
|
Missing context/language explicitly defaults to `de`; omitted language is accepted
|
||||||
|
without modifying context data. Protocol runtime `output_language` is derived
|
||||||
|
provenance, not a separate setting. No transcript/context translation is performed.
|
||||||
|
|
||||||
|
Prompt evolution (2026-09-11): replaced unconditional German output with the
|
||||||
|
meeting-language instruction in the shared prompt builder. Names remain verbatim.
|
||||||
|
Validation uses mocked generation for German/English, plain/diarized inputs,
|
||||||
|
regeneration and legacy contexts; no live model or extraction gold run is involved.
|
||||||
|
|
||||||
This is a compact operational summary of the current Meeting Lab state.
|
This is a compact operational summary of the current Meeting Lab state.
|
||||||
|
|
||||||
## Objective
|
## Objective
|
||||||
@@ -25,7 +38,28 @@ Implemented:
|
|||||||
- Prompt loading from `src/meeting_lab/llm/prompts.py`.
|
- Prompt loading from `src/meeting_lab/llm/prompts.py`.
|
||||||
- Meeting Context V1 loading, validation and optional extraction prompt
|
- Meeting Context V1 loading, validation and optional extraction prompt
|
||||||
injection with minimal extraction JSON provenance.
|
injection with minimal extraction JSON provenance.
|
||||||
|
- FFmpeg-backed WAV, FLAC and M4A preparation into a per-run canonical mono
|
||||||
|
16 kHz signed PCM16 WAV artifact before transcription or diarization. Audio
|
||||||
|
preparation always runs. Optional loudness normalization defaults to on and
|
||||||
|
currently uses the isolated FFmpeg filter
|
||||||
|
`loudnorm=I=-16:LRA=11:TP=-1.5`. This is a conservative speech-recording
|
||||||
|
default and may be revisited after empirical comparison without changing the
|
||||||
|
orchestration API.
|
||||||
- Interim Markdown protocol generation in `src/meeting_lab/protocol/`.
|
- Interim Markdown protocol generation in `src/meeting_lab/protocol/`.
|
||||||
|
- Direct protocol prompt input protection: diarized transcripts are rendered as
|
||||||
|
compact adjacent-speaker blocks without per-segment timestamps. Every source
|
||||||
|
segment remains represented in order. A deterministic heuristic enforces a
|
||||||
|
configurable safe input budget, falls back to complete plain transcript text
|
||||||
|
when necessary, and fails before any Ollama request if even that input is too
|
||||||
|
large. Silent head/tail truncation is prohibited.
|
||||||
|
- The `qwen3.8:27b` direct-protocol stage explicitly requests `num_ctx=32768`
|
||||||
|
and `think=false`; the practical prompt target is approximately 29,000 tokens.
|
||||||
|
A 31,038-token synthetic prompt passed, but larger prompts are not assumed safe
|
||||||
|
from the model's advertised 262,144-token native context alone.
|
||||||
|
- `regenerate_mvp_protocol` updates the run's validated Meeting Context and
|
||||||
|
regenerates protocol artifacts from the existing diarized transcript when
|
||||||
|
available. It never reruns audio preparation, Whisper or Pyannote, and it
|
||||||
|
preserves anonymous speaker labels in the source transcript.
|
||||||
- Non-LLM unit tests for chunking, extraction helpers, protocol rendering and
|
- Non-LLM unit tests for chunking, extraction helpers, protocol rendering and
|
||||||
gold-test runner validation.
|
gold-test runner validation.
|
||||||
- Meeting Context V1 scaffold and documentation for manually maintained
|
- Meeting Context V1 scaffold and documentation for manually maintained
|
||||||
@@ -139,6 +173,9 @@ departments only when they are explicitly supplied as metadata. It must not be
|
|||||||
used to infer responsibilities. In the current implementation this context can
|
used to infer responsibilities. In the current implementation this context can
|
||||||
be injected into chunk extraction prompts as authoritative metadata, and only
|
be injected into chunk extraction prompts as authoritative metadata, and only
|
||||||
minimal provenance is written to extraction JSON.
|
minimal provenance is written to extraction JSON.
|
||||||
|
The implemented MVP statuses are exactly `present` and `mentioned_only`.
|
||||||
|
Legacy entries without a status receive collection-appropriate defaults. Only
|
||||||
|
present participants may be targets of explicit `SPEAKER_XX` mappings.
|
||||||
|
|
||||||
A `responsible` or future `owner` / `assignee` value may be recorded only when
|
A `responsible` or future `owner` / `assignee` value may be recorded only when
|
||||||
source evidence explicitly assigns, accepts or confirms responsibility. If the
|
source evidence explicitly assigns, accepts or confirms responsibility. If the
|
||||||
|
|||||||
@@ -1,5 +1,12 @@
|
|||||||
# Meeting Lab
|
# Meeting Lab
|
||||||
|
|
||||||
|
The direct protocol uses saved Meeting Context `meeting.language`: `de` requests
|
||||||
|
German output and `en` requests English output, for plain and diarized transcripts
|
||||||
|
and protocol-only regeneration. Meeting Assistant supplies the same selection to
|
||||||
|
Whisper. Missing context/language retains German output for older runs. Effective
|
||||||
|
output language is recorded as `output_language` in protocol runtime metadata.
|
||||||
|
Transcript text, authored context, names and speaker mappings are not translated.
|
||||||
|
|
||||||
Experimentierumgebung zur Entwicklung eines lokalen Diskussionsanalyzers für Meetingtranskripte.
|
Experimentierumgebung zur Entwicklung eines lokalen Diskussionsanalyzers für Meetingtranskripte.
|
||||||
|
|
||||||
## Ziel
|
## Ziel
|
||||||
|
|||||||
+25
-1
@@ -35,4 +35,28 @@ torchcodec file decoder.
|
|||||||
Anonymous `SPEAKER_XX` labels are aligned to Whisper segments by maximum
|
Anonymous `SPEAKER_XX` labels are aligned to Whisper segments by maximum
|
||||||
temporal overlap with Community-1 exclusive diarization. The original Whisper
|
temporal overlap with Community-1 exclusive diarization. The original Whisper
|
||||||
transcript is preserved; the derived transcript under `diarization/` is used as
|
transcript is preserved; the derived transcript under `diarization/` is used as
|
||||||
the unchanged direct-protocol generator's input.
|
the direct-protocol generator's source input.
|
||||||
|
|
||||||
|
The full diarized JSON and timestamped text remain immutable audit artifacts,
|
||||||
|
but their per-segment formatting is too verbose for a full-meeting LLM prompt:
|
||||||
|
timestamps and repeated speaker labels can more than double input size. For
|
||||||
|
protocol generation, Meeting Lab deterministically groups only adjacent
|
||||||
|
segments assigned to the same anonymous speaker and omits timestamps. A later
|
||||||
|
return by the same speaker starts a new block, and unassigned segments remain
|
||||||
|
under `SPEAKER_UNASSIGNED`. `protocol/transcript_input.txt` preserves the exact
|
||||||
|
derived representation sent to prompt construction.
|
||||||
|
|
||||||
|
Before contacting Ollama, Meeting Lab conservatively estimates prompt tokens
|
||||||
|
from UTF-8 byte count without adding a model tokenizer dependency. The default safe
|
||||||
|
budget is 29,000 estimated tokens within the explicitly configured 32,768-token
|
||||||
|
Ollama context. The estimate is calibrated against the currently validated
|
||||||
|
German BPD input and is configurable through
|
||||||
|
`MvpMeetingConfig.protocol_safe_input_token_budget` or
|
||||||
|
`--protocol-safe-input-token-budget`.
|
||||||
|
|
||||||
|
If compact diarized input exceeds the budget, the generator deterministically
|
||||||
|
uses the complete plain segment transcript and records the fallback. If that
|
||||||
|
also exceeds the budget, generation fails before model lookup or generation;
|
||||||
|
it never truncates, chunks, summarizes, retries, or makes multiple protocol
|
||||||
|
calls implicitly. Full diarization artifacts are never overwritten by this
|
||||||
|
selection.
|
||||||
|
|||||||
@@ -107,8 +107,10 @@ or aliases are corrected.
|
|||||||
|
|
||||||
`department`: Organizational unit. Optional and nullable.
|
`department`: Organizational unit. Optional and nullable.
|
||||||
|
|
||||||
`attendance_status`: `present` for participants. This distinguishes attendees
|
`attendance_status`: exactly `present` for participants or `mentioned_only` for
|
||||||
from mentioned people.
|
people who are relevant but did not attend. For backward compatibility, a
|
||||||
|
missing status defaults to `present` in `participants` and `mentioned_only` in
|
||||||
|
`mentioned_people`.
|
||||||
|
|
||||||
`mentioned_people`: People discussed or referenced but not present. They are
|
`mentioned_people`: People discussed or referenced but not present. They are
|
||||||
not participants and must not be treated as speakers.
|
not participants and must not be treated as speakers.
|
||||||
@@ -153,7 +155,7 @@ mentioned_people:
|
|||||||
aliases: []
|
aliases: []
|
||||||
role: null
|
role: null
|
||||||
department: null
|
department: null
|
||||||
attendance_status: "not_present"
|
attendance_status: "mentioned_only"
|
||||||
notes: "Wurde erwaehnt, war aber nicht anwesend."
|
notes: "Wurde erwaehnt, war aber nicht anwesend."
|
||||||
|
|
||||||
organization:
|
organization:
|
||||||
@@ -282,6 +284,8 @@ The current validator checks that:
|
|||||||
- participant ids and mentioned-person ids do not collide
|
- participant ids and mentioned-person ids do not collide
|
||||||
- referenced departments exist in `organization.departments`
|
- referenced departments exist in `organization.departments`
|
||||||
- `attendance_status` values are valid
|
- `attendance_status` values are valid
|
||||||
|
- speaker mappings reference present participants only; mentioned-only people
|
||||||
|
cannot be diarized speakers
|
||||||
- participants are marked `present`
|
- participants are marked `present`
|
||||||
- mentioned people are not marked `present`
|
- mentioned people are not marked `present`
|
||||||
|
|
||||||
|
|||||||
@@ -75,7 +75,7 @@ mentioned_people:
|
|||||||
aliases: []
|
aliases: []
|
||||||
role: null
|
role: null
|
||||||
department: null
|
department: null
|
||||||
attendance_status: "not_present"
|
attendance_status: "mentioned_only"
|
||||||
notes: null
|
notes: null
|
||||||
|
|
||||||
organization:
|
organization:
|
||||||
@@ -127,4 +127,4 @@ context_rules:
|
|||||||
do_not_infer_departments: true
|
do_not_infer_departments: true
|
||||||
do_not_infer_responsibilities: true
|
do_not_infer_responsibilities: true
|
||||||
do_not_infer_attendance: true
|
do_not_infer_attendance: true
|
||||||
mentioned_people_are_not_participants: true
|
mentioned_people_are_not_participants: true
|
||||||
|
|||||||
@@ -64,7 +64,7 @@ mentioned_people:
|
|||||||
- "Giovana"
|
- "Giovana"
|
||||||
role: "Leiterin Business Development"
|
role: "Leiterin Business Development"
|
||||||
department_id: "bd"
|
department_id: "bd"
|
||||||
attendance_status: "not_present"
|
attendance_status: "mentioned_only"
|
||||||
notes: null
|
notes: null
|
||||||
|
|
||||||
organization:
|
organization:
|
||||||
|
|||||||
@@ -45,7 +45,7 @@ mentioned_people:
|
|||||||
aliases: []
|
aliases: []
|
||||||
role: null
|
role: null
|
||||||
department: null
|
department: null
|
||||||
attendance_status: "not_present"
|
attendance_status: "mentioned_only"
|
||||||
notes: null
|
notes: null
|
||||||
|
|
||||||
organization:
|
organization:
|
||||||
|
|||||||
@@ -20,6 +20,7 @@ if str(REPO_ROOT) not in sys.path:
|
|||||||
from src.meeting_lab.llm.ollama import DEFAULT_ENDPOINT # noqa: E402
|
from src.meeting_lab.llm.ollama import DEFAULT_ENDPOINT # noqa: E402
|
||||||
from src.meeting_lab.protocol.generate_direct_protocol import ( # noqa: E402
|
from src.meeting_lab.protocol.generate_direct_protocol import ( # noqa: E402
|
||||||
DEFAULT_MODEL,
|
DEFAULT_MODEL,
|
||||||
|
DEFAULT_SAFE_INPUT_TOKEN_BUDGET,
|
||||||
DirectProtocolResult,
|
DirectProtocolResult,
|
||||||
generate_direct_protocol,
|
generate_direct_protocol,
|
||||||
)
|
)
|
||||||
@@ -35,6 +36,11 @@ def parse_args(argv: list[str] | None = None) -> argparse.Namespace:
|
|||||||
parser.add_argument("--output-root", type=Path, default=DEFAULT_OUTPUT_ROOT)
|
parser.add_argument("--output-root", type=Path, default=DEFAULT_OUTPUT_ROOT)
|
||||||
parser.add_argument("--model", default=DEFAULT_MODEL)
|
parser.add_argument("--model", default=DEFAULT_MODEL)
|
||||||
parser.add_argument("--ollama-endpoint", default=DEFAULT_ENDPOINT)
|
parser.add_argument("--ollama-endpoint", default=DEFAULT_ENDPOINT)
|
||||||
|
parser.add_argument(
|
||||||
|
"--safe-input-token-budget",
|
||||||
|
type=int,
|
||||||
|
default=DEFAULT_SAFE_INPUT_TOKEN_BUDGET,
|
||||||
|
)
|
||||||
return parser.parse_args(argv)
|
return parser.parse_args(argv)
|
||||||
|
|
||||||
|
|
||||||
@@ -64,6 +70,11 @@ def persist_result(run_dir: Path, result: DirectProtocolResult) -> Path:
|
|||||||
(protocol_dir / "exact_prompt.txt").write_text(result.exact_prompt, encoding="utf-8")
|
(protocol_dir / "exact_prompt.txt").write_text(result.exact_prompt, encoding="utf-8")
|
||||||
write_json(protocol_dir / "raw_response.json", result.raw_response)
|
write_json(protocol_dir / "raw_response.json", result.raw_response)
|
||||||
write_json(protocol_dir / "runtime_metadata.json", result.runtime_metadata)
|
write_json(protocol_dir / "runtime_metadata.json", result.runtime_metadata)
|
||||||
|
transcript_input = getattr(result, "transcript_input", None)
|
||||||
|
if transcript_input is not None:
|
||||||
|
(protocol_dir / "transcript_input.txt").write_text(
|
||||||
|
transcript_input, encoding="utf-8"
|
||||||
|
)
|
||||||
protocol_path = run_dir / "protocol.md"
|
protocol_path = run_dir / "protocol.md"
|
||||||
protocol_path.write_text(result.protocol_text, encoding="utf-8")
|
protocol_path.write_text(result.protocol_text, encoding="utf-8")
|
||||||
return protocol_path
|
return protocol_path
|
||||||
@@ -116,6 +127,7 @@ def run(args: argparse.Namespace) -> tuple[int, Path, Path | None]:
|
|||||||
preserved_context,
|
preserved_context,
|
||||||
model=args.model,
|
model=args.model,
|
||||||
endpoint=args.ollama_endpoint,
|
endpoint=args.ollama_endpoint,
|
||||||
|
safe_input_token_budget=args.safe_input_token_budget,
|
||||||
)
|
)
|
||||||
protocol_path = persist_result(run_dir, result)
|
protocol_path = persist_result(run_dir, result)
|
||||||
metadata["status"] = "completed"
|
metadata["status"] = "completed"
|
||||||
|
|||||||
@@ -19,6 +19,7 @@ from src.meeting_lab.orchestration.mvp import ( # noqa: E402
|
|||||||
DEFAULT_DIARIZATION_MODEL,
|
DEFAULT_DIARIZATION_MODEL,
|
||||||
DEFAULT_MODEL,
|
DEFAULT_MODEL,
|
||||||
DEFAULT_OUTPUT_ROOT,
|
DEFAULT_OUTPUT_ROOT,
|
||||||
|
DEFAULT_SAFE_INPUT_TOKEN_BUDGET,
|
||||||
MvpMeetingConfig,
|
MvpMeetingConfig,
|
||||||
create_unique_run_dir,
|
create_unique_run_dir,
|
||||||
run_mvp_meeting,
|
run_mvp_meeting,
|
||||||
@@ -33,6 +34,16 @@ def parse_args(argv: list[str] | None = None) -> argparse.Namespace:
|
|||||||
parser.add_argument("audio_file", type=Path)
|
parser.add_argument("audio_file", type=Path)
|
||||||
parser.add_argument("--whisper-model", type=Path, required=True)
|
parser.add_argument("--whisper-model", type=Path, required=True)
|
||||||
parser.add_argument("--whisper-executable", default="whisper-cli")
|
parser.add_argument("--whisper-executable", default="whisper-cli")
|
||||||
|
parser.add_argument("--ffmpeg-executable", default="ffmpeg")
|
||||||
|
parser.add_argument(
|
||||||
|
"--audio-normalization",
|
||||||
|
action=argparse.BooleanOptionalAction,
|
||||||
|
default=True,
|
||||||
|
help=(
|
||||||
|
"Enable FFmpeg loudness normalization during canonical audio preparation "
|
||||||
|
"(default: enabled)."
|
||||||
|
),
|
||||||
|
)
|
||||||
parser.add_argument("--context", type=Path)
|
parser.add_argument("--context", type=Path)
|
||||||
parser.add_argument("--output-root", type=Path, default=DEFAULT_OUTPUT_ROOT)
|
parser.add_argument("--output-root", type=Path, default=DEFAULT_OUTPUT_ROOT)
|
||||||
parser.add_argument("--language", default="de")
|
parser.add_argument("--language", default="de")
|
||||||
@@ -43,6 +54,12 @@ def parse_args(argv: list[str] | None = None) -> argparse.Namespace:
|
|||||||
)
|
)
|
||||||
parser.add_argument("--model", default=DEFAULT_MODEL)
|
parser.add_argument("--model", default=DEFAULT_MODEL)
|
||||||
parser.add_argument("--ollama-endpoint", default=DEFAULT_ENDPOINT)
|
parser.add_argument("--ollama-endpoint", default=DEFAULT_ENDPOINT)
|
||||||
|
parser.add_argument(
|
||||||
|
"--protocol-safe-input-token-budget",
|
||||||
|
type=int,
|
||||||
|
default=DEFAULT_SAFE_INPUT_TOKEN_BUDGET,
|
||||||
|
help="Conservative estimated prompt-token limit before any Ollama request.",
|
||||||
|
)
|
||||||
parser.add_argument(
|
parser.add_argument(
|
||||||
"--diarization",
|
"--diarization",
|
||||||
choices=("auto", "gpu", "cpu", "off"),
|
choices=("auto", "gpu", "cpu", "off"),
|
||||||
@@ -73,12 +90,15 @@ def config_from_args(args: argparse.Namespace) -> MvpMeetingConfig:
|
|||||||
audio_file=args.audio_file,
|
audio_file=args.audio_file,
|
||||||
whisper_model=args.whisper_model,
|
whisper_model=args.whisper_model,
|
||||||
whisper_executable=args.whisper_executable,
|
whisper_executable=args.whisper_executable,
|
||||||
|
ffmpeg_executable=args.ffmpeg_executable,
|
||||||
|
audio_normalization=args.audio_normalization,
|
||||||
context_file=args.context,
|
context_file=args.context,
|
||||||
output_root=args.output_root,
|
output_root=args.output_root,
|
||||||
language=args.language,
|
language=args.language,
|
||||||
threads=args.threads,
|
threads=args.threads,
|
||||||
model=args.model,
|
model=args.model,
|
||||||
ollama_endpoint=args.ollama_endpoint,
|
ollama_endpoint=args.ollama_endpoint,
|
||||||
|
protocol_safe_input_token_budget=args.protocol_safe_input_token_budget,
|
||||||
diarization=args.diarization,
|
diarization=args.diarization,
|
||||||
diarization_runtime=args.diarization_runtime,
|
diarization_runtime=args.diarization_runtime,
|
||||||
diarization_container_image=args.diarization_container_image,
|
diarization_container_image=args.diarization_container_image,
|
||||||
|
|||||||
@@ -0,0 +1,5 @@
|
|||||||
|
"""Canonical audio preparation boundary."""
|
||||||
|
|
||||||
|
from .preparation import AudioPreparationError, PreparedAudio, prepare_audio
|
||||||
|
|
||||||
|
__all__ = ["AudioPreparationError", "PreparedAudio", "prepare_audio"]
|
||||||
@@ -0,0 +1,190 @@
|
|||||||
|
"""Prepare supported recordings for deterministic downstream processing."""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import os
|
||||||
|
import shutil
|
||||||
|
import subprocess
|
||||||
|
import wave
|
||||||
|
from collections.abc import Callable, Sequence
|
||||||
|
from dataclasses import dataclass
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
SUPPORTED_EXTENSIONS = {".wav", ".flac", ".m4a"}
|
||||||
|
CANONICAL_SAMPLE_RATE = 16_000
|
||||||
|
CANONICAL_CHANNELS = 1
|
||||||
|
CANONICAL_SAMPLE_WIDTH_BYTES = 2
|
||||||
|
CANONICAL_CODEC = "pcm_s16le"
|
||||||
|
DEFAULT_NORMALIZATION_FILTER = "loudnorm=I=-16:LRA=11:TP=-1.5"
|
||||||
|
DEFAULT_NORMALIZATION_METHOD = "ffmpeg_loudnorm"
|
||||||
|
|
||||||
|
|
||||||
|
class AudioPreparationError(RuntimeError):
|
||||||
|
"""Raised when source audio cannot be prepared as canonical WAV."""
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass(frozen=True)
|
||||||
|
class PreparedAudio:
|
||||||
|
source_path: Path
|
||||||
|
source_format: str
|
||||||
|
prepared_path: Path
|
||||||
|
method: str
|
||||||
|
ffmpeg_executable: str
|
||||||
|
normalization_enabled: bool = True
|
||||||
|
normalization_method: str | None = DEFAULT_NORMALIZATION_METHOD
|
||||||
|
normalization_filter: str | None = DEFAULT_NORMALIZATION_FILTER
|
||||||
|
|
||||||
|
def metadata(self) -> dict[str, object]:
|
||||||
|
return {
|
||||||
|
"original_source_path": str(self.source_path.resolve()),
|
||||||
|
"original_source_name": self.source_path.name,
|
||||||
|
"original_format": self.source_format,
|
||||||
|
"prepared_audio_path": str(self.prepared_path.resolve()),
|
||||||
|
"preparation_method": self.method,
|
||||||
|
"ffmpeg_executable": self.ffmpeg_executable,
|
||||||
|
"normalization_enabled": self.normalization_enabled,
|
||||||
|
"normalization_method": self.normalization_method,
|
||||||
|
"normalization_filter": self.normalization_filter,
|
||||||
|
"canonical_output": {
|
||||||
|
"container": "wav",
|
||||||
|
"codec": CANONICAL_CODEC,
|
||||||
|
"channels": CANONICAL_CHANNELS,
|
||||||
|
"sample_rate_hz": CANONICAL_SAMPLE_RATE,
|
||||||
|
"bits_per_sample": CANONICAL_SAMPLE_WIDTH_BYTES * 8,
|
||||||
|
},
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
Runner = Callable[..., subprocess.CompletedProcess[str]]
|
||||||
|
|
||||||
|
|
||||||
|
def prepare_audio(
|
||||||
|
source_path: Path,
|
||||||
|
prepared_path: Path,
|
||||||
|
*,
|
||||||
|
ffmpeg_executable: str = "ffmpeg",
|
||||||
|
normalization_enabled: bool = True,
|
||||||
|
runner: Runner = subprocess.run,
|
||||||
|
) -> PreparedAudio:
|
||||||
|
"""Create and validate a canonical mono 16 kHz signed PCM16 WAV artifact."""
|
||||||
|
source_path = Path(source_path)
|
||||||
|
prepared_path = Path(prepared_path)
|
||||||
|
source_format = source_path.suffix.lower()
|
||||||
|
if not source_path.is_file():
|
||||||
|
raise AudioPreparationError(f"Source audio does not exist: {source_path}")
|
||||||
|
if source_format not in SUPPORTED_EXTENSIONS:
|
||||||
|
supported = ", ".join(sorted(SUPPORTED_EXTENSIONS))
|
||||||
|
raise AudioPreparationError(
|
||||||
|
f"Unsupported audio format {source_format or '<none>'!r}; supported: {supported}."
|
||||||
|
)
|
||||||
|
|
||||||
|
resolved_executable = _resolve_executable(ffmpeg_executable)
|
||||||
|
prepared_path.parent.mkdir(parents=True, exist_ok=True)
|
||||||
|
temporary_path = prepared_path.with_name(f".{prepared_path.name}.tmp.wav")
|
||||||
|
command_parts = [
|
||||||
|
resolved_executable,
|
||||||
|
"-nostdin",
|
||||||
|
"-hide_banner",
|
||||||
|
"-loglevel",
|
||||||
|
"error",
|
||||||
|
"-y",
|
||||||
|
"-i",
|
||||||
|
str(source_path),
|
||||||
|
"-map_metadata",
|
||||||
|
"-1",
|
||||||
|
"-vn",
|
||||||
|
]
|
||||||
|
if normalization_enabled:
|
||||||
|
command_parts.extend(("-af", DEFAULT_NORMALIZATION_FILTER))
|
||||||
|
command_parts.extend(
|
||||||
|
(
|
||||||
|
"-ac",
|
||||||
|
str(CANONICAL_CHANNELS),
|
||||||
|
"-ar",
|
||||||
|
str(CANONICAL_SAMPLE_RATE),
|
||||||
|
"-c:a",
|
||||||
|
CANONICAL_CODEC,
|
||||||
|
"-fflags",
|
||||||
|
"+bitexact",
|
||||||
|
str(temporary_path),
|
||||||
|
)
|
||||||
|
)
|
||||||
|
command: Sequence[str] = tuple(command_parts)
|
||||||
|
try:
|
||||||
|
completed = runner(command, capture_output=True, text=True, check=False)
|
||||||
|
except OSError as exc:
|
||||||
|
raise AudioPreparationError(f"Could not run FFmpeg: {exc}") from exc
|
||||||
|
if completed.returncode != 0:
|
||||||
|
detail = (
|
||||||
|
completed.stderr or completed.stdout or "no diagnostic output"
|
||||||
|
).strip()
|
||||||
|
raise AudioPreparationError(
|
||||||
|
f"FFmpeg failed to prepare {source_path.name} (exit {completed.returncode}): "
|
||||||
|
f"{detail}"
|
||||||
|
)
|
||||||
|
try:
|
||||||
|
_validate_canonical_wav(temporary_path)
|
||||||
|
os.replace(temporary_path, prepared_path)
|
||||||
|
except Exception:
|
||||||
|
temporary_path.unlink(missing_ok=True)
|
||||||
|
raise
|
||||||
|
|
||||||
|
return PreparedAudio(
|
||||||
|
source_path=source_path,
|
||||||
|
source_format=source_format.removeprefix("."),
|
||||||
|
prepared_path=prepared_path,
|
||||||
|
method="ffmpeg",
|
||||||
|
ffmpeg_executable=resolved_executable,
|
||||||
|
normalization_enabled=normalization_enabled,
|
||||||
|
normalization_method=(
|
||||||
|
DEFAULT_NORMALIZATION_METHOD if normalization_enabled else None
|
||||||
|
),
|
||||||
|
normalization_filter=(
|
||||||
|
DEFAULT_NORMALIZATION_FILTER if normalization_enabled else None
|
||||||
|
),
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _resolve_executable(executable: str) -> str:
|
||||||
|
value = executable.strip()
|
||||||
|
if not value:
|
||||||
|
raise AudioPreparationError("FFmpeg executable must not be empty.")
|
||||||
|
if Path(value).parent != Path("."):
|
||||||
|
path = Path(value)
|
||||||
|
if path.is_file() and os.access(path, os.X_OK):
|
||||||
|
return str(path)
|
||||||
|
raise AudioPreparationError(f"FFmpeg executable is not available: {value}")
|
||||||
|
resolved = shutil.which(value)
|
||||||
|
if resolved is None:
|
||||||
|
raise AudioPreparationError(
|
||||||
|
f"FFmpeg executable {value!r} was not found on PATH. Install FFmpeg or "
|
||||||
|
"configure its executable path."
|
||||||
|
)
|
||||||
|
return resolved
|
||||||
|
|
||||||
|
|
||||||
|
def _validate_canonical_wav(path: Path) -> None:
|
||||||
|
try:
|
||||||
|
with wave.open(str(path), "rb") as recording:
|
||||||
|
properties = (
|
||||||
|
recording.getnchannels(),
|
||||||
|
recording.getframerate(),
|
||||||
|
recording.getsampwidth(),
|
||||||
|
recording.getcomptype(),
|
||||||
|
)
|
||||||
|
except (OSError, EOFError, wave.Error) as exc:
|
||||||
|
raise AudioPreparationError(
|
||||||
|
f"FFmpeg did not produce a readable WAV file: {path}: {exc}"
|
||||||
|
) from exc
|
||||||
|
expected = (
|
||||||
|
CANONICAL_CHANNELS,
|
||||||
|
CANONICAL_SAMPLE_RATE,
|
||||||
|
CANONICAL_SAMPLE_WIDTH_BYTES,
|
||||||
|
"NONE",
|
||||||
|
)
|
||||||
|
if properties != expected:
|
||||||
|
raise AudioPreparationError(
|
||||||
|
"Prepared audio is not canonical mono 16 kHz PCM16 WAV: "
|
||||||
|
f"channels={properties[0]}, sample_rate={properties[1]}, "
|
||||||
|
f"sample_width={properties[2]}, compression={properties[3]}."
|
||||||
|
)
|
||||||
@@ -11,7 +11,7 @@ from typing import Any
|
|||||||
|
|
||||||
|
|
||||||
SUPPORTED_SCHEMA_VERSIONS = {"1"}
|
SUPPORTED_SCHEMA_VERSIONS = {"1"}
|
||||||
VALID_ATTENDANCE_STATUSES = {"present", "not_present", "absent"}
|
VALID_ATTENDANCE_STATUSES = {"present", "mentioned_only"}
|
||||||
|
|
||||||
|
|
||||||
class MeetingContextValidationError(ValueError):
|
class MeetingContextValidationError(ValueError):
|
||||||
@@ -59,15 +59,16 @@ def load_meeting_context(path: Path) -> MeetingContext:
|
|||||||
if not isinstance(loaded, dict):
|
if not isinstance(loaded, dict):
|
||||||
raise MeetingContextValidationError("Meeting Context must be a YAML object.")
|
raise MeetingContextValidationError("Meeting Context must be a YAML object.")
|
||||||
|
|
||||||
validate_meeting_context(loaded)
|
normalized = _with_attendance_defaults(loaded)
|
||||||
return MeetingContext(data=loaded, source_file=path)
|
validate_meeting_context(normalized)
|
||||||
|
return MeetingContext(data=normalized, source_file=path)
|
||||||
|
|
||||||
|
|
||||||
def create_meeting_context(
|
def create_meeting_context(
|
||||||
data: dict[str, Any], *, source_file: Path = Path("<generated>")
|
data: dict[str, Any], *, source_file: Path = Path("<generated>")
|
||||||
) -> MeetingContext:
|
) -> MeetingContext:
|
||||||
"""Validate structured data and return an immutable context boundary."""
|
"""Validate structured data and return an immutable context boundary."""
|
||||||
validated = copy.deepcopy(data)
|
validated = _with_attendance_defaults(data)
|
||||||
validate_meeting_context(validated)
|
validate_meeting_context(validated)
|
||||||
return MeetingContext(data=validated, source_file=source_file)
|
return MeetingContext(data=validated, source_file=source_file)
|
||||||
|
|
||||||
@@ -107,7 +108,9 @@ def validate_meeting_context(data: dict[str, Any]) -> None:
|
|||||||
meeting = _require_mapping(data, "meeting")
|
meeting = _require_mapping(data, "meeting")
|
||||||
_require_non_empty_string(meeting, "meeting.meeting_id")
|
_require_non_empty_string(meeting, "meeting.meeting_id")
|
||||||
_require_non_empty_string(meeting, "meeting.title")
|
_require_non_empty_string(meeting, "meeting.title")
|
||||||
_require_non_empty_string(meeting, "meeting.language")
|
# Legacy contexts may omit language; protocol generation defaults to German.
|
||||||
|
if "language" in meeting:
|
||||||
|
_require_non_empty_string(meeting, "meeting.language")
|
||||||
|
|
||||||
organization = _optional_mapping(data.get("organization"), "organization")
|
organization = _optional_mapping(data.get("organization"), "organization")
|
||||||
departments = _optional_list(organization.get("departments"), "organization.departments")
|
departments = _optional_list(organization.get("departments"), "organization.departments")
|
||||||
@@ -143,8 +146,8 @@ def validate_meeting_context(data: dict[str, Any]) -> None:
|
|||||||
|
|
||||||
for index, participant in enumerate(participants):
|
for index, participant in enumerate(participants):
|
||||||
item_path = f"participants[{index}]"
|
item_path = f"participants[{index}]"
|
||||||
_validate_attendance(participant, item_path)
|
status = _validate_attendance(participant, item_path, default="present")
|
||||||
if participant.get("attendance_status") != "present":
|
if status != "present":
|
||||||
raise MeetingContextValidationError(
|
raise MeetingContextValidationError(
|
||||||
f"{item_path}.attendance_status must be 'present'."
|
f"{item_path}.attendance_status must be 'present'."
|
||||||
)
|
)
|
||||||
@@ -152,10 +155,10 @@ def validate_meeting_context(data: dict[str, Any]) -> None:
|
|||||||
|
|
||||||
for index, person in enumerate(mentioned_people):
|
for index, person in enumerate(mentioned_people):
|
||||||
item_path = f"mentioned_people[{index}]"
|
item_path = f"mentioned_people[{index}]"
|
||||||
_validate_attendance(person, item_path)
|
status = _validate_attendance(person, item_path, default="mentioned_only")
|
||||||
if person.get("attendance_status") == "present":
|
if status != "mentioned_only":
|
||||||
raise MeetingContextValidationError(
|
raise MeetingContextValidationError(
|
||||||
f"{item_path}.attendance_status must not be 'present'."
|
f"{item_path}.attendance_status must be 'mentioned_only'."
|
||||||
)
|
)
|
||||||
_validate_department_reference(person, item_path, department_ids)
|
_validate_department_reference(person, item_path, department_ids)
|
||||||
|
|
||||||
@@ -334,12 +337,28 @@ def _collect_unique_ids(items: list[Any], key: str, path: str) -> set[str]:
|
|||||||
return ids
|
return ids
|
||||||
|
|
||||||
|
|
||||||
def _validate_attendance(item: dict[str, Any], path: str) -> None:
|
def _validate_attendance(item: dict[str, Any], path: str, *, default: str) -> str:
|
||||||
status = item.get("attendance_status")
|
status = item.get("attendance_status", default)
|
||||||
if status not in VALID_ATTENDANCE_STATUSES:
|
if status not in VALID_ATTENDANCE_STATUSES:
|
||||||
raise MeetingContextValidationError(
|
raise MeetingContextValidationError(
|
||||||
f"{path}.attendance_status has invalid value: {status!r}."
|
f"{path}.attendance_status has invalid value: {status!r}."
|
||||||
)
|
)
|
||||||
|
return status
|
||||||
|
|
||||||
|
|
||||||
|
def _with_attendance_defaults(data: dict[str, Any]) -> dict[str, Any]:
|
||||||
|
normalized = copy.deepcopy(data)
|
||||||
|
participants = normalized.get("participants")
|
||||||
|
if isinstance(participants, list):
|
||||||
|
for participant in participants:
|
||||||
|
if isinstance(participant, dict):
|
||||||
|
participant.setdefault("attendance_status", "present")
|
||||||
|
mentioned_people = normalized.get("mentioned_people")
|
||||||
|
if isinstance(mentioned_people, list):
|
||||||
|
for person in mentioned_people:
|
||||||
|
if isinstance(person, dict):
|
||||||
|
person.setdefault("attendance_status", "mentioned_only")
|
||||||
|
return normalized
|
||||||
|
|
||||||
|
|
||||||
def _validate_department_reference(
|
def _validate_department_reference(
|
||||||
|
|||||||
@@ -7,11 +7,13 @@ import re
|
|||||||
import shutil
|
import shutil
|
||||||
import sys
|
import sys
|
||||||
import time
|
import time
|
||||||
|
from collections.abc import Callable, Mapping, Sequence
|
||||||
from dataclasses import dataclass
|
from dataclasses import dataclass
|
||||||
from datetime import datetime
|
from datetime import datetime
|
||||||
from pathlib import Path
|
from pathlib import Path
|
||||||
from typing import Any, Callable, Mapping, Sequence
|
from typing import Any
|
||||||
|
|
||||||
|
from src.meeting_lab.audio import prepare_audio
|
||||||
from src.meeting_lab.diarization import (
|
from src.meeting_lab.diarization import (
|
||||||
DEFAULT_MODEL as DEFAULT_DIARIZATION_MODEL,
|
DEFAULT_MODEL as DEFAULT_DIARIZATION_MODEL,
|
||||||
diarize_audio,
|
diarize_audio,
|
||||||
@@ -28,6 +30,8 @@ from src.meeting_lab.models.meeting_context import (
|
|||||||
from src.meeting_lab.progress import ProgressEvent, ProgressSink, ProgressStatus
|
from src.meeting_lab.progress import ProgressEvent, ProgressSink, ProgressStatus
|
||||||
from src.meeting_lab.protocol.generate_direct_protocol import (
|
from src.meeting_lab.protocol.generate_direct_protocol import (
|
||||||
DEFAULT_MODEL,
|
DEFAULT_MODEL,
|
||||||
|
DEFAULT_NUM_CTX,
|
||||||
|
DEFAULT_SAFE_INPUT_TOKEN_BUDGET,
|
||||||
DirectProtocolResult,
|
DirectProtocolResult,
|
||||||
generate_direct_protocol,
|
generate_direct_protocol,
|
||||||
load_compact_transcript,
|
load_compact_transcript,
|
||||||
@@ -44,12 +48,16 @@ class MvpMeetingConfig:
|
|||||||
audio_file: Path
|
audio_file: Path
|
||||||
whisper_model: Path
|
whisper_model: Path
|
||||||
whisper_executable: str = "whisper-cli"
|
whisper_executable: str = "whisper-cli"
|
||||||
|
ffmpeg_executable: str = "ffmpeg"
|
||||||
|
audio_normalization: bool = True
|
||||||
context_file: Path | None = None
|
context_file: Path | None = None
|
||||||
output_root: Path = DEFAULT_OUTPUT_ROOT
|
output_root: Path = DEFAULT_OUTPUT_ROOT
|
||||||
language: str = "de"
|
language: str = "de"
|
||||||
threads: str | int = "auto"
|
threads: str | int = "auto"
|
||||||
model: str = DEFAULT_MODEL
|
model: str = DEFAULT_MODEL
|
||||||
ollama_endpoint: str = DEFAULT_ENDPOINT
|
ollama_endpoint: str = DEFAULT_ENDPOINT
|
||||||
|
protocol_num_ctx: int = DEFAULT_NUM_CTX
|
||||||
|
protocol_safe_input_token_budget: int = DEFAULT_SAFE_INPUT_TOKEN_BUDGET
|
||||||
diarization: str = "off"
|
diarization: str = "off"
|
||||||
diarization_runtime: str = "native"
|
diarization_runtime: str = "native"
|
||||||
diarization_container_image: str | None = None
|
diarization_container_image: str | None = None
|
||||||
@@ -63,6 +71,64 @@ class MvpRunResult:
|
|||||||
protocol_path: Path | None
|
protocol_path: Path | None
|
||||||
|
|
||||||
|
|
||||||
|
def regenerate_mvp_protocol(
|
||||||
|
run_dir: Path,
|
||||||
|
*,
|
||||||
|
meeting_context: ContextInput,
|
||||||
|
model: str = DEFAULT_MODEL,
|
||||||
|
ollama_endpoint: str = DEFAULT_ENDPOINT,
|
||||||
|
protocol_num_ctx: int = DEFAULT_NUM_CTX,
|
||||||
|
protocol_safe_input_token_budget: int = DEFAULT_SAFE_INPUT_TOKEN_BUDGET,
|
||||||
|
progress_sink: ProgressSink | None = None,
|
||||||
|
) -> MvpRunResult:
|
||||||
|
"""Regenerate only protocol artifacts from an existing completed run."""
|
||||||
|
started = time.perf_counter()
|
||||||
|
run_dir = Path(run_dir)
|
||||||
|
context = _effective_context(meeting_context)
|
||||||
|
if context is None:
|
||||||
|
raise ValueError("Meeting Context is required for protocol regeneration.")
|
||||||
|
if protocol_num_ctx <= 0:
|
||||||
|
raise ValueError("Protocol Ollama context size must be positive.")
|
||||||
|
if protocol_safe_input_token_budget <= 0:
|
||||||
|
raise ValueError("Protocol safe input token budget must be positive.")
|
||||||
|
|
||||||
|
diarized_transcript = run_dir / "diarization" / "transcript_diarized.json"
|
||||||
|
plain_transcript = run_dir / "transcript" / "transcript.json"
|
||||||
|
transcript_path = (
|
||||||
|
diarized_transcript if diarized_transcript.is_file() else plain_transcript
|
||||||
|
)
|
||||||
|
if not transcript_path.is_file():
|
||||||
|
raise FileNotFoundError(
|
||||||
|
f"Existing run has no protocol transcript artifact: {run_dir}"
|
||||||
|
)
|
||||||
|
|
||||||
|
context_path = run_dir / "context" / "meeting_context.yaml"
|
||||||
|
write_meeting_context(context, context_path)
|
||||||
|
_emit(progress_sink, "protocol_generation", "started", started)
|
||||||
|
try:
|
||||||
|
result = generate_direct_protocol(
|
||||||
|
transcript_path,
|
||||||
|
context_path,
|
||||||
|
model=model,
|
||||||
|
endpoint=ollama_endpoint,
|
||||||
|
num_ctx=protocol_num_ctx,
|
||||||
|
safe_input_token_budget=protocol_safe_input_token_budget,
|
||||||
|
)
|
||||||
|
protocol_path = _persist_protocol(run_dir, result)
|
||||||
|
except Exception as exc:
|
||||||
|
_emit(
|
||||||
|
progress_sink,
|
||||||
|
"failed",
|
||||||
|
"failed",
|
||||||
|
started,
|
||||||
|
message=f"protocol_generation: {type(exc).__name__}: {exc}",
|
||||||
|
)
|
||||||
|
raise
|
||||||
|
_emit(progress_sink, "protocol_generation", "completed", started)
|
||||||
|
_emit(progress_sink, "completed", "completed", started)
|
||||||
|
return MvpRunResult(0, run_dir, protocol_path)
|
||||||
|
|
||||||
|
|
||||||
def create_unique_run_dir(
|
def create_unique_run_dir(
|
||||||
output_root: Path,
|
output_root: Path,
|
||||||
meeting_name: str,
|
meeting_name: str,
|
||||||
@@ -123,6 +189,10 @@ def _validate_inputs(
|
|||||||
and not config.diarization_container_image
|
and not config.diarization_container_image
|
||||||
):
|
):
|
||||||
raise ValueError("A diarization container image is required.")
|
raise ValueError("A diarization container image is required.")
|
||||||
|
if config.protocol_safe_input_token_budget <= 0:
|
||||||
|
raise ValueError("Protocol safe input token budget must be positive.")
|
||||||
|
if config.protocol_num_ctx <= 0:
|
||||||
|
raise ValueError("Protocol Ollama context size must be positive.")
|
||||||
|
|
||||||
|
|
||||||
def _emit(
|
def _emit(
|
||||||
@@ -150,6 +220,11 @@ def _persist_protocol(run_dir: Path, result: DirectProtocolResult) -> Path:
|
|||||||
(protocol_dir / "exact_prompt.txt").write_text(result.exact_prompt, encoding="utf-8")
|
(protocol_dir / "exact_prompt.txt").write_text(result.exact_prompt, encoding="utf-8")
|
||||||
_write_json(protocol_dir / "raw_response.json", result.raw_response)
|
_write_json(protocol_dir / "raw_response.json", result.raw_response)
|
||||||
_write_json(protocol_dir / "runtime_metadata.json", result.runtime_metadata)
|
_write_json(protocol_dir / "runtime_metadata.json", result.runtime_metadata)
|
||||||
|
transcript_input = getattr(result, "transcript_input", None)
|
||||||
|
if transcript_input is not None:
|
||||||
|
(protocol_dir / "transcript_input.txt").write_text(
|
||||||
|
transcript_input, encoding="utf-8"
|
||||||
|
)
|
||||||
protocol_path = run_dir / "protocol.md"
|
protocol_path = run_dir / "protocol.md"
|
||||||
protocol_path.write_text(result.protocol_text, encoding="utf-8")
|
protocol_path.write_text(result.protocol_text, encoding="utf-8")
|
||||||
return protocol_path
|
return protocol_path
|
||||||
@@ -187,6 +262,7 @@ def run_mvp_meeting(
|
|||||||
stage_runtimes: dict[str, float | None] = {
|
stage_runtimes: dict[str, float | None] = {
|
||||||
"validation": round(validation_runtime, 3),
|
"validation": round(validation_runtime, 3),
|
||||||
"setup": None,
|
"setup": None,
|
||||||
|
"audio_preparation": None,
|
||||||
"whisper": None,
|
"whisper": None,
|
||||||
"transcript_validation": None,
|
"transcript_validation": None,
|
||||||
"protocol": None,
|
"protocol": None,
|
||||||
@@ -198,6 +274,7 @@ def run_mvp_meeting(
|
|||||||
"run_id": run_dir.name,
|
"run_id": run_dir.name,
|
||||||
"timestamp": timestamp,
|
"timestamp": timestamp,
|
||||||
"input_audio": str(config.audio_file.resolve()),
|
"input_audio": str(config.audio_file.resolve()),
|
||||||
|
"audio_preparation": None,
|
||||||
"transcript_output": str(transcript_path.resolve()),
|
"transcript_output": str(transcript_path.resolve()),
|
||||||
"protocol_output": str(protocol_path.resolve()),
|
"protocol_output": str(protocol_path.resolve()),
|
||||||
"whisper_model": str(config.whisper_model.resolve()),
|
"whisper_model": str(config.whisper_model.resolve()),
|
||||||
@@ -238,6 +315,33 @@ def run_mvp_meeting(
|
|||||||
},
|
},
|
||||||
)
|
)
|
||||||
|
|
||||||
|
preparation_started = time.perf_counter()
|
||||||
|
current_stage = "audio_preparation"
|
||||||
|
stage_started = preparation_started
|
||||||
|
prepared_audio = prepare_audio(
|
||||||
|
config.audio_file,
|
||||||
|
audio_dir / "prepared.wav",
|
||||||
|
ffmpeg_executable=config.ffmpeg_executable,
|
||||||
|
normalization_enabled=config.audio_normalization,
|
||||||
|
)
|
||||||
|
stage_runtimes["audio_preparation"] = round(
|
||||||
|
time.perf_counter() - preparation_started, 3
|
||||||
|
)
|
||||||
|
current_stage = "preparing"
|
||||||
|
preparation_metadata = prepared_audio.metadata()
|
||||||
|
metadata["audio_preparation"] = preparation_metadata
|
||||||
|
_write_json(audio_dir / "preparation_metadata.json", preparation_metadata)
|
||||||
|
_write_json(
|
||||||
|
audio_dir / "input_manifest.json",
|
||||||
|
{
|
||||||
|
"source_file": str(config.audio_file.resolve()),
|
||||||
|
"filename": config.audio_file.name,
|
||||||
|
"size_bytes": config.audio_file.stat().st_size,
|
||||||
|
"format": config.audio_file.suffix.lower().removeprefix("."),
|
||||||
|
"prepared_audio": preparation_metadata,
|
||||||
|
},
|
||||||
|
)
|
||||||
|
|
||||||
preserved_context: Path | None = None
|
preserved_context: Path | None = None
|
||||||
if effective_context is not None:
|
if effective_context is not None:
|
||||||
preserved_context = context_dir / "meeting_context.yaml"
|
preserved_context = context_dir / "meeting_context.yaml"
|
||||||
@@ -252,7 +356,7 @@ def run_mvp_meeting(
|
|||||||
stage_started = time.perf_counter()
|
stage_started = time.perf_counter()
|
||||||
_emit(progress_sink, "transcription", "started", overall_started)
|
_emit(progress_sink, "transcription", "started", overall_started)
|
||||||
transcription = transcribe_audio(
|
transcription = transcribe_audio(
|
||||||
config.audio_file,
|
prepared_audio.prepared_path,
|
||||||
config.whisper_model,
|
config.whisper_model,
|
||||||
transcript_dir,
|
transcript_dir,
|
||||||
config.language,
|
config.language,
|
||||||
@@ -275,7 +379,7 @@ def run_mvp_meeting(
|
|||||||
_emit(progress_sink, "diarization", "started", overall_started)
|
_emit(progress_sink, "diarization", "started", overall_started)
|
||||||
diarization_dir = run_dir / "diarization"
|
diarization_dir = run_dir / "diarization"
|
||||||
diarization = diarize_audio(
|
diarization = diarize_audio(
|
||||||
config.audio_file,
|
prepared_audio.prepared_path,
|
||||||
diarization_dir,
|
diarization_dir,
|
||||||
config.diarization,
|
config.diarization,
|
||||||
runtime=config.diarization_runtime,
|
runtime=config.diarization_runtime,
|
||||||
@@ -321,6 +425,8 @@ def run_mvp_meeting(
|
|||||||
preserved_context,
|
preserved_context,
|
||||||
model=config.model,
|
model=config.model,
|
||||||
endpoint=config.ollama_endpoint,
|
endpoint=config.ollama_endpoint,
|
||||||
|
num_ctx=config.protocol_num_ctx,
|
||||||
|
safe_input_token_budget=config.protocol_safe_input_token_budget,
|
||||||
)
|
)
|
||||||
stage_runtimes["protocol"] = round(time.perf_counter() - stage_started, 3)
|
stage_runtimes["protocol"] = round(time.perf_counter() - stage_started, 3)
|
||||||
protocol_path = _persist_protocol(run_dir, result)
|
protocol_path = _persist_protocol(run_dir, result)
|
||||||
@@ -330,6 +436,7 @@ def run_mvp_meeting(
|
|||||||
except Exception as exc:
|
except Exception as exc:
|
||||||
metadata_stage = {
|
metadata_stage = {
|
||||||
"preparing": "setup",
|
"preparing": "setup",
|
||||||
|
"audio_preparation": "audio_preparation",
|
||||||
"transcription": "whisper",
|
"transcription": "whisper",
|
||||||
"diarization": "diarization",
|
"diarization": "diarization",
|
||||||
"protocol_generation": "protocol",
|
"protocol_generation": "protocol",
|
||||||
|
|||||||
@@ -3,17 +3,34 @@
|
|||||||
from __future__ import annotations
|
from __future__ import annotations
|
||||||
|
|
||||||
|
|
||||||
DIRECT_PROTOCOL_INSTRUCTION = """Erstelle aus dem vollständigen Transkript und dem Meeting-Kontext ein vollständiges, strukturiertes und professionelles internes Besprechungsprotokoll in deutscher Sprache.
|
DIRECT_PROTOCOL_INSTRUCTION = """Erstelle aus dem vollständigen Transkript und dem Meeting-Kontext ein vollständiges, strukturiertes und professionelles internes Besprechungsprotokoll.
|
||||||
|
|
||||||
Das Protokoll muss themenorientiert sein, nicht chronologisch und nicht nach technischen Kategorien gegliedert. Beginne mit # Meeting Protocol. Verwende für jedes kohärente Thema eine Überschrift ## <Thema> und darunter eine strukturierte Synthese der Diskussion. Bewahre relevante Diskussionsverläufe, unterschiedliche Positionen, offene Punkte und Entscheidungsgrundlagen. Dokumentiere die wesentlichen Inhalte nachvollziehbar und fasse Themenblöcke so zusammen, dass auch Personen, die nicht am Meeting teilgenommen haben, den Kontext und die Entwicklung der Diskussion verstehen können. Nenne Entscheidungen oder abgestimmte Positionen nur, wenn sie tatsächlich belegt sind. Führe Maßnahmen nur auf, wenn eine konkrete zukünftige Handlung gestützt ist; nenne verantwortliche Personen und Fristen ausschließlich bei expliziter Zuweisung, Annahme oder Bestätigung im Transkript. Vorschläge, Einwände, Möglichkeiten und vorläufige Ideen sind keine Entscheidungen oder Verpflichtungen. Bewahre relevante Einschränkungen und ungelöste Meinungsverschiedenheiten. Nenne offene Punkte nur, wenn sie wirklich offen bleiben. Nicht jedes Thema benötigt Entscheidungen, Maßnahmen oder offene Punkte.
|
Das Protokoll muss themenorientiert sein, nicht chronologisch und nicht nach technischen Kategorien gegliedert. Beginne mit # Meeting Protocol. Verwende für jedes kohärente Thema eine Überschrift ## <Thema> und darunter eine strukturierte Synthese der Diskussion. Bewahre relevante Diskussionsverläufe, unterschiedliche Positionen, offene Punkte und Entscheidungsgrundlagen. Dokumentiere die wesentlichen Inhalte nachvollziehbar und fasse Themenblöcke so zusammen, dass auch Personen, die nicht am Meeting teilgenommen haben, den Kontext und die Entwicklung der Diskussion verstehen können. Nenne Entscheidungen oder abgestimmte Positionen nur, wenn sie tatsächlich belegt sind. Führe Maßnahmen nur auf, wenn eine konkrete zukünftige Handlung gestützt ist; nenne verantwortliche Personen und Fristen ausschließlich bei expliziter Zuweisung, Annahme oder Bestätigung im Transkript. Vorschläge, Einwände, Möglichkeiten und vorläufige Ideen sind keine Entscheidungen oder Verpflichtungen. Bewahre relevante Einschränkungen und ungelöste Meinungsverschiedenheiten. Nenne offene Punkte nur, wenn sie wirklich offen bleiben. Nicht jedes Thema benötigt Entscheidungen, Maßnahmen oder offene Punkte.
|
||||||
|
|
||||||
Erzeuge keine reine Wiedergabe des Transkripts und verlängere das Protokoll nicht unnötig durch Wiederholungen. Synthetisiere zusammengehörige Aussagen, entferne Füllwörter und Gesprächsrauschen und erfinde keine Fakten, Entscheidungen, Zustimmungen, Verantwortlichen oder Fristen. Gib kein JSON, keine internen Labels und keine Analyse oder Denkprotokolle aus. Das Ergebnis soll als Markdown-Protokoll nach geringfügiger menschlicher Redaktion intern versendbar sein. Eine kompakte themenübergreifende Maßnahmenliste am Ende ist optional, wenn sie nützlich und vollständig belegt ist."""
|
Erzeuge keine reine Wiedergabe des Transkripts und verlängere das Protokoll nicht unnötig durch Wiederholungen. Synthetisiere zusammengehörige Aussagen, entferne Füllwörter und Gesprächsrauschen und erfinde keine Fakten, Entscheidungen, Zustimmungen, Verantwortlichen oder Fristen. Gib kein JSON, keine internen Labels und keine Analyse oder Denkprotokolle aus. Das Ergebnis soll als Markdown-Protokoll nach geringfügiger menschlicher Redaktion intern versendbar sein. Eine kompakte themenübergreifende Maßnahmenliste am Ende ist optional, wenn sie nützlich und vollständig belegt ist."""
|
||||||
|
|
||||||
|
COMPACT_DIARIZED_PROTOCOL_INSTRUCTION = """Erstelle aus dem vollständigen Transkript und Meeting-Kontext ein vollständiges, professionelles internes Besprechungsprotokoll. Das Transkript ist in aufeinanderfolgende anonyme Sprecherblöcke gegliedert.
|
||||||
|
|
||||||
def build_direct_protocol_prompt(transcript: str, meeting_context: str | None = None) -> str:
|
Beginne mit # Meeting Protocol. Gliedere themenorientiert mit ## <Thema> und synthetisiere je Thema den relevanten Diskussionsverlauf, Kontext, unterschiedliche Positionen, Entscheidungsgrundlagen, Einschränkungen und ungelöste Meinungsverschiedenheiten so, dass Dritte ihn nachvollziehen können. Nenne Entscheidungen nur bei Beleg. Nenne Maßnahmen, Verantwortliche und Fristen nur bei expliziter Zuweisung, Annahme oder Bestätigung; Vorschläge sind keine Verpflichtungen.
|
||||||
|
|
||||||
|
Entferne nur Wiederholungen, Füllwörter und Gesprächsrauschen. Erfinde keine Fakten oder Identitäten. Gib kein JSON, keine Sprecherlabels und kein Denkprotokoll aus. Eine belegte themenübergreifende Maßnahmenliste am Ende ist optional."""
|
||||||
|
|
||||||
|
MAPPED_SPEAKER_ATTRIBUTION_INSTRUCTION = """Nutze die autoritativen SPEAKER_XX-zu-Teilnehmer-Zuordnungen im Meeting-Kontext, um ausdrücklich belegte Aussagen, Positionen, Entscheidungen, Zuweisungen und angenommene persönliche Verpflichtungen namentlich zuzuordnen. Eine ausdrückliche Ich-Zusage eines zugeordneten Sprechers belegt persönliche Verantwortung. Unterscheide stets den Sprecher einer Aussage von darin nur erwähnten Personen. Leite für nicht zugeordnete Sprecher keine Identität ab und erfinde keine persönliche Verantwortung. Gib die technischen SPEAKER_XX-Bezeichnungen nicht im nutzerseitigen Protokoll aus."""
|
||||||
|
|
||||||
|
|
||||||
|
def build_direct_protocol_prompt(
|
||||||
|
transcript: str,
|
||||||
|
meeting_context: str | None = None,
|
||||||
|
*,
|
||||||
|
instruction: str = DIRECT_PROTOCOL_INSTRUCTION,
|
||||||
|
meeting_language: str = "de",
|
||||||
|
) -> str:
|
||||||
context = meeting_context.strip() if meeting_context else "Kein Meeting-Kontext bereitgestellt."
|
context = meeting_context.strip() if meeting_context else "Kein Meeting-Kontext bereitgestellt."
|
||||||
|
language = {"de": "German", "en": "English"}.get(meeting_language, meeting_language)
|
||||||
return (
|
return (
|
||||||
f"{DIRECT_PROTOCOL_INSTRUCTION}\n\n"
|
f"{instruction}\n\n"
|
||||||
|
f"Write the meeting protocol in {language}. "
|
||||||
|
"Preserve speaker and person names exactly as supplied.\n\n"
|
||||||
f"MEETING-KONTEXT:\n{context}\n\n"
|
f"MEETING-KONTEXT:\n{context}\n\n"
|
||||||
f"VOLLSTAENDIGES TRANSKRIPT:\n{transcript.strip()}\n"
|
f"VOLLSTAENDIGES TRANSKRIPT:\n{transcript.strip()}\n"
|
||||||
)
|
)
|
||||||
|
|||||||
@@ -18,13 +18,24 @@ from src.meeting_lab.models.meeting_context import (
|
|||||||
load_meeting_context,
|
load_meeting_context,
|
||||||
render_meeting_context_for_prompt,
|
render_meeting_context_for_prompt,
|
||||||
)
|
)
|
||||||
from src.meeting_lab.protocol.direct_protocol_prompt import build_direct_protocol_prompt
|
from src.meeting_lab.protocol.direct_protocol_prompt import (
|
||||||
|
COMPACT_DIARIZED_PROTOCOL_INSTRUCTION,
|
||||||
|
MAPPED_SPEAKER_ATTRIBUTION_INSTRUCTION,
|
||||||
|
build_direct_protocol_prompt,
|
||||||
|
)
|
||||||
|
from src.meeting_lab.protocol.transcript_input import (
|
||||||
|
TranscriptInputError,
|
||||||
|
compact_diarized_transcript,
|
||||||
|
plain_segment_transcript,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
DEFAULT_MODEL = "qwen3.6:35B-A3B"
|
DEFAULT_MODEL = "qwen3.6:35B-A3B"
|
||||||
DEFAULT_NUM_CTX = 32768
|
DEFAULT_NUM_CTX = 32768
|
||||||
DEFAULT_NUM_PREDICT = 8192
|
DEFAULT_NUM_PREDICT = 8192
|
||||||
DEFAULT_TIMEOUT = 1800
|
DEFAULT_TIMEOUT = 1800
|
||||||
|
DEFAULT_SAFE_INPUT_TOKEN_BUDGET = 29_000
|
||||||
|
ESTIMATED_UTF8_BYTES_PER_TOKEN = 4.4
|
||||||
|
|
||||||
|
|
||||||
class DirectProtocolError(ValueError):
|
class DirectProtocolError(ValueError):
|
||||||
@@ -38,9 +49,29 @@ class DirectProtocolResult:
|
|||||||
model_metadata: dict[str, Any]
|
model_metadata: dict[str, Any]
|
||||||
runtime_metadata: dict[str, Any]
|
runtime_metadata: dict[str, Any]
|
||||||
raw_response: dict[str, Any]
|
raw_response: dict[str, Any]
|
||||||
|
transcript_input: str | None = None
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass(frozen=True)
|
||||||
|
class SelectedTranscriptInput:
|
||||||
|
text: str
|
||||||
|
prompt: str
|
||||||
|
representation: str
|
||||||
|
estimated_input_tokens: int
|
||||||
|
safe_input_token_budget: int
|
||||||
|
fallback_used: bool
|
||||||
|
diarization_enabled: bool
|
||||||
|
|
||||||
|
|
||||||
def load_compact_transcript(path: Path) -> str:
|
def load_compact_transcript(path: Path) -> str:
|
||||||
|
data = _load_transcript_document(path)
|
||||||
|
text = data.get("text")
|
||||||
|
if not isinstance(text, str) or not text.strip():
|
||||||
|
raise DirectProtocolError("Transcript top-level 'text' must be a non-empty string.")
|
||||||
|
return text
|
||||||
|
|
||||||
|
|
||||||
|
def _load_transcript_document(path: Path) -> dict[str, Any]:
|
||||||
if not path.is_file():
|
if not path.is_file():
|
||||||
raise DirectProtocolError(f"Transcript file does not exist: {path}")
|
raise DirectProtocolError(f"Transcript file does not exist: {path}")
|
||||||
try:
|
try:
|
||||||
@@ -51,10 +82,83 @@ def load_compact_transcript(path: Path) -> str:
|
|||||||
raise DirectProtocolError("Transcript JSON must contain a top-level object.")
|
raise DirectProtocolError("Transcript JSON must contain a top-level object.")
|
||||||
if "text" not in data:
|
if "text" not in data:
|
||||||
raise DirectProtocolError("Transcript JSON must contain top-level 'text'.")
|
raise DirectProtocolError("Transcript JSON must contain top-level 'text'.")
|
||||||
text = data["text"]
|
return data
|
||||||
if not isinstance(text, str) or not text.strip():
|
|
||||||
raise DirectProtocolError("Transcript top-level 'text' must be a non-empty string.")
|
|
||||||
return text
|
def estimate_input_tokens(prompt: str) -> int:
|
||||||
|
"""Estimate tokens without adding a model-specific tokenizer dependency."""
|
||||||
|
byte_count = len(prompt.encode("utf-8"))
|
||||||
|
return max(1, int(byte_count / ESTIMATED_UTF8_BYTES_PER_TOKEN + 0.999999))
|
||||||
|
|
||||||
|
|
||||||
|
def select_transcript_input(
|
||||||
|
transcript: dict[str, Any],
|
||||||
|
rendered_context: str | None,
|
||||||
|
*,
|
||||||
|
safe_input_token_budget: int = DEFAULT_SAFE_INPUT_TOKEN_BUDGET,
|
||||||
|
meeting_language: str = "de",
|
||||||
|
) -> SelectedTranscriptInput:
|
||||||
|
"""Select complete prompt input without allowing silent tail truncation."""
|
||||||
|
if safe_input_token_budget <= 0:
|
||||||
|
raise DirectProtocolError("Safe protocol input token budget must be positive.")
|
||||||
|
diarization_enabled = transcript.get("speaker_labels_anonymous") is True
|
||||||
|
if diarization_enabled:
|
||||||
|
try:
|
||||||
|
compact = compact_diarized_transcript(transcript.get("segments"))
|
||||||
|
plain_text = plain_segment_transcript(transcript.get("segments"))
|
||||||
|
except TranscriptInputError as exc:
|
||||||
|
raise DirectProtocolError(str(exc)) from exc
|
||||||
|
instruction = COMPACT_DIARIZED_PROTOCOL_INSTRUCTION
|
||||||
|
if (
|
||||||
|
rendered_context
|
||||||
|
and "Confirmed diarization speaker mappings" in rendered_context
|
||||||
|
):
|
||||||
|
instruction = f"{instruction}\n\n{MAPPED_SPEAKER_ATTRIBUTION_INSTRUCTION}"
|
||||||
|
compact_prompt = build_direct_protocol_prompt(
|
||||||
|
compact.text,
|
||||||
|
rendered_context,
|
||||||
|
instruction=instruction,
|
||||||
|
meeting_language=meeting_language,
|
||||||
|
)
|
||||||
|
compact_estimate = estimate_input_tokens(compact_prompt)
|
||||||
|
if compact_estimate <= safe_input_token_budget:
|
||||||
|
return SelectedTranscriptInput(
|
||||||
|
text=compact.text,
|
||||||
|
prompt=compact_prompt,
|
||||||
|
representation="diarized_compact",
|
||||||
|
estimated_input_tokens=compact_estimate,
|
||||||
|
safe_input_token_budget=safe_input_token_budget,
|
||||||
|
fallback_used=False,
|
||||||
|
diarization_enabled=True,
|
||||||
|
)
|
||||||
|
representation = "plain_transcript_fallback"
|
||||||
|
fallback_used = True
|
||||||
|
else:
|
||||||
|
plain_text = transcript.get("text")
|
||||||
|
if not isinstance(plain_text, str) or not plain_text.strip():
|
||||||
|
raise DirectProtocolError("Transcript top-level 'text' must be a non-empty string.")
|
||||||
|
representation = "plain_transcript"
|
||||||
|
fallback_used = False
|
||||||
|
|
||||||
|
plain_prompt = build_direct_protocol_prompt(
|
||||||
|
plain_text, rendered_context, meeting_language=meeting_language
|
||||||
|
)
|
||||||
|
plain_estimate = estimate_input_tokens(plain_prompt)
|
||||||
|
if plain_estimate > safe_input_token_budget:
|
||||||
|
raise DirectProtocolError(
|
||||||
|
"Protocol prompt/input is too large for the configured safe input budget "
|
||||||
|
f"({plain_estimate} estimated tokens > {safe_input_token_budget}). "
|
||||||
|
"No LLM request was made; silent truncation is not allowed."
|
||||||
|
)
|
||||||
|
return SelectedTranscriptInput(
|
||||||
|
text=plain_text,
|
||||||
|
prompt=plain_prompt,
|
||||||
|
representation=representation,
|
||||||
|
estimated_input_tokens=plain_estimate,
|
||||||
|
safe_input_token_budget=safe_input_token_budget,
|
||||||
|
fallback_used=fallback_used,
|
||||||
|
diarization_enabled=diarization_enabled,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
def generate_direct_protocol(
|
def generate_direct_protocol(
|
||||||
@@ -66,27 +170,38 @@ def generate_direct_protocol(
|
|||||||
timeout: int = DEFAULT_TIMEOUT,
|
timeout: int = DEFAULT_TIMEOUT,
|
||||||
num_ctx: int = DEFAULT_NUM_CTX,
|
num_ctx: int = DEFAULT_NUM_CTX,
|
||||||
num_predict: int = DEFAULT_NUM_PREDICT,
|
num_predict: int = DEFAULT_NUM_PREDICT,
|
||||||
|
safe_input_token_budget: int = DEFAULT_SAFE_INPUT_TOKEN_BUDGET,
|
||||||
model_check: Callable[[str, str, int], dict[str, Any]] = require_model,
|
model_check: Callable[[str, str, int], dict[str, Any]] = require_model,
|
||||||
generation_call: Callable[..., OllamaGeneration] = generate_once,
|
generation_call: Callable[..., OllamaGeneration] = generate_once,
|
||||||
) -> DirectProtocolResult:
|
) -> DirectProtocolResult:
|
||||||
transcript = load_compact_transcript(transcript_path)
|
transcript = _load_transcript_document(transcript_path)
|
||||||
context: MeetingContext | None = (
|
context: MeetingContext | None = (
|
||||||
load_meeting_context(context_path) if context_path is not None else None
|
load_meeting_context(context_path) if context_path is not None else None
|
||||||
)
|
)
|
||||||
rendered_context = render_meeting_context_for_prompt(context) if context else None
|
rendered_context = render_meeting_context_for_prompt(context) if context else None
|
||||||
prompt = build_direct_protocol_prompt(transcript, rendered_context)
|
# Older runs without a meeting language retain the historical German output.
|
||||||
|
meeting_language = (
|
||||||
|
str(context.data["meeting"].get("language") or "de") if context else "de"
|
||||||
|
)
|
||||||
|
selected = select_transcript_input(
|
||||||
|
transcript,
|
||||||
|
rendered_context,
|
||||||
|
safe_input_token_budget=safe_input_token_budget,
|
||||||
|
meeting_language=meeting_language,
|
||||||
|
)
|
||||||
|
|
||||||
model_metadata = model_check(endpoint, model, 10)
|
model_metadata = model_check(endpoint, model, 10)
|
||||||
generation = generation_call(
|
generation = generation_call(
|
||||||
endpoint,
|
endpoint,
|
||||||
model,
|
model,
|
||||||
prompt,
|
selected.prompt,
|
||||||
timeout=timeout,
|
timeout=timeout,
|
||||||
num_ctx=num_ctx,
|
num_ctx=num_ctx,
|
||||||
num_predict=num_predict,
|
num_predict=num_predict,
|
||||||
)
|
)
|
||||||
data = generation.raw_response
|
data = generation.raw_response
|
||||||
runtime_metadata = {
|
runtime_metadata = {
|
||||||
|
"output_language": meeting_language,
|
||||||
"model": model,
|
"model": model,
|
||||||
"prompt_token_count": data.get("prompt_eval_count"),
|
"prompt_token_count": data.get("prompt_eval_count"),
|
||||||
"output_token_count": data.get("eval_count"),
|
"output_token_count": data.get("eval_count"),
|
||||||
@@ -101,11 +216,31 @@ def generate_direct_protocol(
|
|||||||
"think": False,
|
"think": False,
|
||||||
"num_ctx": num_ctx,
|
"num_ctx": num_ctx,
|
||||||
"num_predict": num_predict,
|
"num_predict": num_predict,
|
||||||
|
"selected_transcript_representation": selected.representation,
|
||||||
|
"estimated_input_tokens": selected.estimated_input_tokens,
|
||||||
|
"safe_input_token_budget": selected.safe_input_token_budget,
|
||||||
|
"input_token_estimation_method": "utf8_bytes_divided_by_4.4",
|
||||||
|
"fallback_used": selected.fallback_used,
|
||||||
|
"diarization_enabled": selected.diarization_enabled,
|
||||||
|
"speaker_attribution_available": (
|
||||||
|
True
|
||||||
|
if selected.representation == "diarized_compact"
|
||||||
|
else False
|
||||||
|
if selected.representation == "plain_transcript_fallback"
|
||||||
|
else None
|
||||||
|
),
|
||||||
|
"speaker_attribution_loss_reason": (
|
||||||
|
"plain_transcript_fallback"
|
||||||
|
if selected.representation == "plain_transcript_fallback"
|
||||||
|
else None
|
||||||
|
),
|
||||||
|
"speaker_mapping_count": len(context.speaker_mappings) if context else 0,
|
||||||
}
|
}
|
||||||
return DirectProtocolResult(
|
return DirectProtocolResult(
|
||||||
protocol_text=generation.text,
|
protocol_text=generation.text,
|
||||||
exact_prompt=prompt,
|
exact_prompt=selected.prompt,
|
||||||
model_metadata=model_metadata,
|
model_metadata=model_metadata,
|
||||||
runtime_metadata=runtime_metadata,
|
runtime_metadata=runtime_metadata,
|
||||||
raw_response=data,
|
raw_response=data,
|
||||||
|
transcript_input=selected.text,
|
||||||
)
|
)
|
||||||
|
|||||||
@@ -0,0 +1,97 @@
|
|||||||
|
"""Deterministic transcript representations for one-call protocol prompts."""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
from dataclasses import dataclass
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
|
||||||
|
class TranscriptInputError(ValueError):
|
||||||
|
"""Raised when a transcript cannot be represented without content loss."""
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass(frozen=True)
|
||||||
|
class SpeakerBlock:
|
||||||
|
"""One contiguous run of transcript segments assigned to one speaker."""
|
||||||
|
|
||||||
|
speaker_id: str
|
||||||
|
segment_texts: tuple[str, ...]
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass(frozen=True)
|
||||||
|
class CompactDiarizedTranscript:
|
||||||
|
"""Compact prompt text plus structural evidence of segment preservation."""
|
||||||
|
|
||||||
|
text: str
|
||||||
|
blocks: tuple[SpeakerBlock, ...]
|
||||||
|
source_segment_count: int
|
||||||
|
|
||||||
|
@property
|
||||||
|
def represented_segment_count(self) -> int:
|
||||||
|
return sum(len(block.segment_texts) for block in self.blocks)
|
||||||
|
|
||||||
|
@property
|
||||||
|
def segment_texts(self) -> tuple[str, ...]:
|
||||||
|
return tuple(text for block in self.blocks for text in block.segment_texts)
|
||||||
|
|
||||||
|
|
||||||
|
def normalize_segment_text(value: Any, index: int) -> str:
|
||||||
|
"""Normalize formatting whitespace while retaining all semantic text."""
|
||||||
|
if not isinstance(value, str):
|
||||||
|
raise TranscriptInputError(f"Transcript segment {index} text must be a string.")
|
||||||
|
return " ".join(value.split())
|
||||||
|
|
||||||
|
|
||||||
|
def compact_diarized_transcript(segments: Any) -> CompactDiarizedTranscript:
|
||||||
|
"""Group only adjacent same-speaker segments and omit repeated timestamps."""
|
||||||
|
if not isinstance(segments, list) or not segments:
|
||||||
|
raise TranscriptInputError(
|
||||||
|
"Diarized transcript must contain a non-empty 'segments' list."
|
||||||
|
)
|
||||||
|
|
||||||
|
mutable_blocks: list[tuple[str, list[str]]] = []
|
||||||
|
source_texts: list[str] = []
|
||||||
|
for index, segment in enumerate(segments):
|
||||||
|
if not isinstance(segment, dict):
|
||||||
|
raise TranscriptInputError(f"Transcript segment {index} must be an object.")
|
||||||
|
speaker = segment.get("speaker_id") or "SPEAKER_UNASSIGNED"
|
||||||
|
if not isinstance(speaker, str) or not speaker.startswith("SPEAKER_"):
|
||||||
|
raise TranscriptInputError(
|
||||||
|
f"Transcript segment {index} must use an anonymous SPEAKER_ label."
|
||||||
|
)
|
||||||
|
text = normalize_segment_text(segment.get("text"), index)
|
||||||
|
source_texts.append(text)
|
||||||
|
if mutable_blocks and mutable_blocks[-1][0] == speaker:
|
||||||
|
mutable_blocks[-1][1].append(text)
|
||||||
|
else:
|
||||||
|
mutable_blocks.append((speaker, [text]))
|
||||||
|
|
||||||
|
blocks = tuple(
|
||||||
|
SpeakerBlock(speaker_id=speaker, segment_texts=tuple(texts))
|
||||||
|
for speaker, texts in mutable_blocks
|
||||||
|
)
|
||||||
|
rendered = "\n".join(
|
||||||
|
f"{block.speaker_id}: {' '.join(block.segment_texts)}" for block in blocks
|
||||||
|
)
|
||||||
|
result = CompactDiarizedTranscript(
|
||||||
|
text=rendered + "\n",
|
||||||
|
blocks=blocks,
|
||||||
|
source_segment_count=len(segments),
|
||||||
|
)
|
||||||
|
if result.represented_segment_count != len(segments):
|
||||||
|
raise TranscriptInputError("Compact diarized transcript lost source segments.")
|
||||||
|
if result.segment_texts != tuple(source_texts):
|
||||||
|
raise TranscriptInputError("Compact diarized transcript changed segment order or text.")
|
||||||
|
return result
|
||||||
|
|
||||||
|
|
||||||
|
def plain_segment_transcript(segments: Any) -> str:
|
||||||
|
"""Reconstruct plain transcript text from every segment in source order."""
|
||||||
|
if not isinstance(segments, list) or not segments:
|
||||||
|
raise TranscriptInputError("Transcript must contain a non-empty 'segments' list.")
|
||||||
|
texts = []
|
||||||
|
for index, segment in enumerate(segments):
|
||||||
|
if not isinstance(segment, dict):
|
||||||
|
raise TranscriptInputError(f"Transcript segment {index} must be an object.")
|
||||||
|
texts.append(normalize_segment_text(segment.get("text"), index))
|
||||||
|
return " ".join(texts)
|
||||||
@@ -0,0 +1,230 @@
|
|||||||
|
import subprocess
|
||||||
|
import tempfile
|
||||||
|
import unittest
|
||||||
|
import wave
|
||||||
|
from pathlib import Path
|
||||||
|
from unittest.mock import patch
|
||||||
|
|
||||||
|
from src.meeting_lab.audio.preparation import (
|
||||||
|
DEFAULT_NORMALIZATION_FILTER,
|
||||||
|
DEFAULT_NORMALIZATION_METHOD,
|
||||||
|
AudioPreparationError,
|
||||||
|
prepare_audio,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _write_wav(
|
||||||
|
path: Path, *, channels: int = 1, sample_rate: int = 16_000, sample_width: int = 2
|
||||||
|
) -> None:
|
||||||
|
with wave.open(str(path), "wb") as recording:
|
||||||
|
recording.setnchannels(channels)
|
||||||
|
recording.setsampwidth(sample_width)
|
||||||
|
recording.setframerate(sample_rate)
|
||||||
|
recording.writeframes(b"\x00" * channels * sample_width * 32)
|
||||||
|
|
||||||
|
|
||||||
|
def _successful_runner(commands: list[list[str]]):
|
||||||
|
def run(command, **kwargs):
|
||||||
|
commands.append(list(command))
|
||||||
|
_write_wav(Path(command[-1]))
|
||||||
|
return subprocess.CompletedProcess(command, 0, "", "")
|
||||||
|
|
||||||
|
return run
|
||||||
|
|
||||||
|
|
||||||
|
class AudioPreparationTests(unittest.TestCase):
|
||||||
|
def test_supported_inputs_are_prepared_with_normalization_on_and_off(self) -> None:
|
||||||
|
for suffix in (".wav", ".flac", ".m4a"):
|
||||||
|
for normalization_enabled in (True, False):
|
||||||
|
with (
|
||||||
|
self.subTest(
|
||||||
|
suffix=suffix, normalization_enabled=normalization_enabled
|
||||||
|
),
|
||||||
|
tempfile.TemporaryDirectory() as directory,
|
||||||
|
):
|
||||||
|
root = Path(directory)
|
||||||
|
source = root / f"meeting{suffix}"
|
||||||
|
if suffix == ".wav":
|
||||||
|
_write_wav(source)
|
||||||
|
else:
|
||||||
|
source.write_bytes(b"original encoded audio")
|
||||||
|
original = source.read_bytes()
|
||||||
|
destination = root / "run" / "audio" / "prepared.wav"
|
||||||
|
commands: list[list[str]] = []
|
||||||
|
|
||||||
|
with patch(
|
||||||
|
"src.meeting_lab.audio.preparation.shutil.which",
|
||||||
|
return_value="/usr/bin/ffmpeg",
|
||||||
|
):
|
||||||
|
result = prepare_audio(
|
||||||
|
source,
|
||||||
|
destination,
|
||||||
|
normalization_enabled=normalization_enabled,
|
||||||
|
runner=_successful_runner(commands),
|
||||||
|
)
|
||||||
|
|
||||||
|
self.assertEqual(source.read_bytes(), original)
|
||||||
|
self.assertEqual(result.prepared_path, destination)
|
||||||
|
with wave.open(str(destination), "rb") as recording:
|
||||||
|
self.assertEqual(recording.getnchannels(), 1)
|
||||||
|
self.assertEqual(recording.getframerate(), 16_000)
|
||||||
|
self.assertEqual(recording.getsampwidth(), 2)
|
||||||
|
self.assertEqual(recording.getcomptype(), "NONE")
|
||||||
|
self.assertEqual(commands[0][commands[0].index("-ac") + 1], "1")
|
||||||
|
self.assertEqual(commands[0][commands[0].index("-ar") + 1], "16000")
|
||||||
|
self.assertEqual(
|
||||||
|
commands[0][commands[0].index("-c:a") + 1], "pcm_s16le"
|
||||||
|
)
|
||||||
|
self.assertEqual("-af" in commands[0], normalization_enabled)
|
||||||
|
if normalization_enabled:
|
||||||
|
self.assertEqual(
|
||||||
|
commands[0][commands[0].index("-af") + 1],
|
||||||
|
DEFAULT_NORMALIZATION_FILTER,
|
||||||
|
)
|
||||||
|
self.assertEqual(
|
||||||
|
result.normalization_enabled, normalization_enabled
|
||||||
|
)
|
||||||
|
|
||||||
|
def test_normalization_defaults_to_on_and_explicit_on_matches(self) -> None:
|
||||||
|
with tempfile.TemporaryDirectory() as directory:
|
||||||
|
root = Path(directory)
|
||||||
|
source = root / "meeting.wav"
|
||||||
|
_write_wav(source)
|
||||||
|
commands: list[list[str]] = []
|
||||||
|
with patch(
|
||||||
|
"src.meeting_lab.audio.preparation.shutil.which",
|
||||||
|
return_value="/usr/bin/ffmpeg",
|
||||||
|
):
|
||||||
|
default = prepare_audio(
|
||||||
|
source, root / "default.wav", runner=_successful_runner(commands)
|
||||||
|
)
|
||||||
|
explicit = prepare_audio(
|
||||||
|
source,
|
||||||
|
root / "explicit.wav",
|
||||||
|
normalization_enabled=True,
|
||||||
|
runner=_successful_runner(commands),
|
||||||
|
)
|
||||||
|
|
||||||
|
self.assertTrue(default.normalization_enabled)
|
||||||
|
self.assertTrue(explicit.normalization_enabled)
|
||||||
|
self.assertEqual(
|
||||||
|
commands[0][commands[0].index("-af") + 1],
|
||||||
|
commands[1][commands[1].index("-af") + 1],
|
||||||
|
)
|
||||||
|
|
||||||
|
def test_noncanonical_wav_is_normalized(self) -> None:
|
||||||
|
with tempfile.TemporaryDirectory() as directory:
|
||||||
|
root = Path(directory)
|
||||||
|
source = root / "stereo-48k.wav"
|
||||||
|
_write_wav(source, channels=2, sample_rate=48_000)
|
||||||
|
destination = root / "prepared.wav"
|
||||||
|
|
||||||
|
with patch(
|
||||||
|
"src.meeting_lab.audio.preparation.shutil.which",
|
||||||
|
return_value="/usr/bin/ffmpeg",
|
||||||
|
):
|
||||||
|
prepare_audio(source, destination, runner=_successful_runner([]))
|
||||||
|
|
||||||
|
with wave.open(str(destination), "rb") as recording:
|
||||||
|
self.assertEqual(
|
||||||
|
(recording.getnchannels(), recording.getframerate()), (1, 16_000)
|
||||||
|
)
|
||||||
|
|
||||||
|
def test_ffmpeg_missing_has_actionable_error(self) -> None:
|
||||||
|
with tempfile.TemporaryDirectory() as directory:
|
||||||
|
root = Path(directory)
|
||||||
|
source = root / "meeting.flac"
|
||||||
|
source.write_bytes(b"audio")
|
||||||
|
|
||||||
|
with (
|
||||||
|
patch(
|
||||||
|
"src.meeting_lab.audio.preparation.shutil.which", return_value=None
|
||||||
|
),
|
||||||
|
self.assertRaisesRegex(AudioPreparationError, "not found on PATH"),
|
||||||
|
):
|
||||||
|
prepare_audio(source, root / "prepared.wav")
|
||||||
|
|
||||||
|
def test_ffmpeg_failure_includes_diagnostic_and_preserves_source(self) -> None:
|
||||||
|
with tempfile.TemporaryDirectory() as directory:
|
||||||
|
root = Path(directory)
|
||||||
|
source = root / "meeting.m4a"
|
||||||
|
source.write_bytes(b"original")
|
||||||
|
|
||||||
|
def fail(command, **kwargs):
|
||||||
|
return subprocess.CompletedProcess(command, 1, "", "decoder exploded")
|
||||||
|
|
||||||
|
with (
|
||||||
|
patch(
|
||||||
|
"src.meeting_lab.audio.preparation.shutil.which",
|
||||||
|
return_value="/usr/bin/ffmpeg",
|
||||||
|
),
|
||||||
|
self.assertRaisesRegex(AudioPreparationError, "decoder exploded"),
|
||||||
|
):
|
||||||
|
prepare_audio(source, root / "prepared.wav", runner=fail)
|
||||||
|
|
||||||
|
self.assertEqual(source.read_bytes(), b"original")
|
||||||
|
self.assertFalse((root / "prepared.wav").exists())
|
||||||
|
|
||||||
|
def test_prepared_audio_metadata_is_traceable(self) -> None:
|
||||||
|
with tempfile.TemporaryDirectory() as directory:
|
||||||
|
root = Path(directory)
|
||||||
|
source = root / "unknown_meeting.flac"
|
||||||
|
source.write_bytes(b"source")
|
||||||
|
destination = root / "audio" / "prepared.wav"
|
||||||
|
with patch(
|
||||||
|
"src.meeting_lab.audio.preparation.shutil.which",
|
||||||
|
return_value="/usr/bin/ffmpeg",
|
||||||
|
):
|
||||||
|
result = prepare_audio(
|
||||||
|
source, destination, runner=_successful_runner([])
|
||||||
|
)
|
||||||
|
|
||||||
|
metadata = result.metadata()
|
||||||
|
self.assertEqual(metadata["original_source_name"], "unknown_meeting.flac")
|
||||||
|
self.assertEqual(metadata["original_format"], "flac")
|
||||||
|
self.assertEqual(
|
||||||
|
metadata["prepared_audio_path"], str(destination.resolve())
|
||||||
|
)
|
||||||
|
self.assertEqual(metadata["preparation_method"], "ffmpeg")
|
||||||
|
self.assertTrue(metadata["normalization_enabled"])
|
||||||
|
self.assertEqual(
|
||||||
|
metadata["normalization_method"], DEFAULT_NORMALIZATION_METHOD
|
||||||
|
)
|
||||||
|
self.assertEqual(
|
||||||
|
metadata["normalization_filter"], DEFAULT_NORMALIZATION_FILTER
|
||||||
|
)
|
||||||
|
self.assertEqual(
|
||||||
|
metadata["canonical_output"],
|
||||||
|
{
|
||||||
|
"container": "wav",
|
||||||
|
"codec": "pcm_s16le",
|
||||||
|
"channels": 1,
|
||||||
|
"sample_rate_hz": 16_000,
|
||||||
|
"bits_per_sample": 16,
|
||||||
|
},
|
||||||
|
)
|
||||||
|
|
||||||
|
def test_disabled_normalization_metadata_has_no_method_or_filter(self) -> None:
|
||||||
|
with tempfile.TemporaryDirectory() as directory:
|
||||||
|
root = Path(directory)
|
||||||
|
source = root / "meeting.m4a"
|
||||||
|
source.write_bytes(b"source")
|
||||||
|
with patch(
|
||||||
|
"src.meeting_lab.audio.preparation.shutil.which",
|
||||||
|
return_value="/usr/bin/ffmpeg",
|
||||||
|
):
|
||||||
|
result = prepare_audio(
|
||||||
|
source,
|
||||||
|
root / "prepared.wav",
|
||||||
|
normalization_enabled=False,
|
||||||
|
runner=_successful_runner([]),
|
||||||
|
)
|
||||||
|
|
||||||
|
metadata = result.metadata()
|
||||||
|
self.assertFalse(metadata["normalization_enabled"])
|
||||||
|
self.assertIsNone(metadata["normalization_method"])
|
||||||
|
self.assertIsNone(metadata["normalization_filter"])
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
unittest.main()
|
||||||
@@ -116,6 +116,7 @@ class GeneratorTests(unittest.TestCase):
|
|||||||
self.assertEqual(check.call_count, 1)
|
self.assertEqual(check.call_count, 1)
|
||||||
self.assertEqual(call.call_count, 1)
|
self.assertEqual(call.call_count, 1)
|
||||||
self.assertEqual(call.call_args.args[1], "qwen3.6:35B-A3B")
|
self.assertEqual(call.call_args.args[1], "qwen3.6:35B-A3B")
|
||||||
|
self.assertEqual(call.call_args.kwargs["num_ctx"], 32768)
|
||||||
self.assertEqual(result.runtime_metadata["request_count"], 1)
|
self.assertEqual(result.runtime_metadata["request_count"], 1)
|
||||||
self.assertEqual(result.runtime_metadata["prompt_token_count"], 123)
|
self.assertEqual(result.runtime_metadata["prompt_token_count"], 123)
|
||||||
self.assertFalse(result.runtime_metadata["think"])
|
self.assertFalse(result.runtime_metadata["think"])
|
||||||
@@ -181,7 +182,7 @@ class OllamaTests(unittest.TestCase):
|
|||||||
with patch.object(ollama.requests, "post", return_value=response) as post:
|
with patch.object(ollama.requests, "post", return_value=response) as post:
|
||||||
result = ollama.generate_once(
|
result = ollama.generate_once(
|
||||||
"http://localhost:11434",
|
"http://localhost:11434",
|
||||||
"chosen:model",
|
"qwen3.8:27b",
|
||||||
"prompt",
|
"prompt",
|
||||||
timeout=30,
|
timeout=30,
|
||||||
num_ctx=32768,
|
num_ctx=32768,
|
||||||
@@ -190,8 +191,9 @@ class OllamaTests(unittest.TestCase):
|
|||||||
|
|
||||||
self.assertEqual(post.call_count, 1)
|
self.assertEqual(post.call_count, 1)
|
||||||
payload = post.call_args.kwargs["json"]
|
payload = post.call_args.kwargs["json"]
|
||||||
self.assertEqual(payload["model"], "chosen:model")
|
self.assertEqual(payload["model"], "qwen3.8:27b")
|
||||||
self.assertEqual(payload["options"]["temperature"], 0.0)
|
self.assertEqual(payload["options"]["temperature"], 0.0)
|
||||||
|
self.assertEqual(payload["options"]["num_ctx"], 32768)
|
||||||
self.assertFalse(payload["think"])
|
self.assertFalse(payload["think"])
|
||||||
self.assertFalse(payload["stream"])
|
self.assertFalse(payload["stream"])
|
||||||
self.assertEqual(result.raw_response, raw)
|
self.assertEqual(result.raw_response, raw)
|
||||||
@@ -232,6 +234,7 @@ class DirectProtocolCliTests(unittest.TestCase):
|
|||||||
return_value=type("Result", (), {
|
return_value=type("Result", (), {
|
||||||
"protocol_text": protocol_text,
|
"protocol_text": protocol_text,
|
||||||
"exact_prompt": "exact prompt\n",
|
"exact_prompt": "exact prompt\n",
|
||||||
|
"transcript_input": "selected transcript\n",
|
||||||
"raw_response": {"response": protocol_text},
|
"raw_response": {"response": protocol_text},
|
||||||
"runtime_metadata": {"request_count": 1},
|
"runtime_metadata": {"request_count": 1},
|
||||||
})(),
|
})(),
|
||||||
@@ -245,6 +248,10 @@ class DirectProtocolCliTests(unittest.TestCase):
|
|||||||
(run_dir / "protocol/exact_prompt.txt").read_text(encoding="utf-8"),
|
(run_dir / "protocol/exact_prompt.txt").read_text(encoding="utf-8"),
|
||||||
"exact prompt\n",
|
"exact prompt\n",
|
||||||
)
|
)
|
||||||
|
self.assertEqual(
|
||||||
|
(run_dir / "protocol/transcript_input.txt").read_text(encoding="utf-8"),
|
||||||
|
"selected transcript\n",
|
||||||
|
)
|
||||||
self.assertEqual(
|
self.assertEqual(
|
||||||
json.loads((run_dir / "protocol/raw_response.json").read_text())["response"],
|
json.loads((run_dir / "protocol/raw_response.json").read_text())["response"],
|
||||||
protocol_text,
|
protocol_text,
|
||||||
|
|||||||
@@ -72,6 +72,38 @@ class MeetingContextTests(unittest.TestCase):
|
|||||||
with self.assertRaisesRegex(MeetingContextValidationError, "invalid value"):
|
with self.assertRaisesRegex(MeetingContextValidationError, "invalid value"):
|
||||||
validate_meeting_context(data)
|
validate_meeting_context(data)
|
||||||
|
|
||||||
|
def test_missing_participant_attendance_defaults_to_present(self) -> None:
|
||||||
|
data = copy.deepcopy(self.context.data)
|
||||||
|
del data["participants"][0]["attendance_status"]
|
||||||
|
|
||||||
|
context = create_meeting_context(data)
|
||||||
|
|
||||||
|
self.assertEqual(
|
||||||
|
context.data["participants"][0]["attendance_status"], "present"
|
||||||
|
)
|
||||||
|
|
||||||
|
def test_explicit_mentioned_only_is_preserved(self) -> None:
|
||||||
|
data = copy.deepcopy(self.context.data)
|
||||||
|
data["mentioned_people"][0]["attendance_status"] = "mentioned_only"
|
||||||
|
|
||||||
|
context = create_meeting_context(data)
|
||||||
|
|
||||||
|
self.assertEqual(
|
||||||
|
context.data["mentioned_people"][0]["attendance_status"],
|
||||||
|
"mentioned_only",
|
||||||
|
)
|
||||||
|
self.assertIn(
|
||||||
|
"Mentioned but absent people:", render_meeting_context_for_prompt(context)
|
||||||
|
)
|
||||||
|
|
||||||
|
def test_mentioned_only_person_cannot_be_a_diarized_speaker(self) -> None:
|
||||||
|
data = copy.deepcopy(self.context.data)
|
||||||
|
mentioned_id = data["mentioned_people"][0]["person_id"]
|
||||||
|
data["speaker_mappings"] = {"SPEAKER_00": mentioned_id}
|
||||||
|
|
||||||
|
with self.assertRaisesRegex(MeetingContextValidationError, "unknown participant"):
|
||||||
|
validate_meeting_context(data)
|
||||||
|
|
||||||
def test_prompt_representation_is_deterministic(self) -> None:
|
def test_prompt_representation_is_deterministic(self) -> None:
|
||||||
first = render_meeting_context_for_prompt(self.context)
|
first = render_meeting_context_for_prompt(self.context)
|
||||||
second = render_meeting_context_for_prompt(self.context)
|
second = render_meeting_context_for_prompt(self.context)
|
||||||
|
|||||||
@@ -0,0 +1,164 @@
|
|||||||
|
"""Meeting-language regression tests without media or model execution."""
|
||||||
|
|
||||||
|
import json
|
||||||
|
import tempfile
|
||||||
|
import unittest
|
||||||
|
from functools import partial
|
||||||
|
from pathlib import Path
|
||||||
|
from unittest.mock import Mock, patch
|
||||||
|
|
||||||
|
from src.meeting_lab.llm.ollama import OllamaGeneration
|
||||||
|
from src.meeting_lab.models.meeting_context import load_meeting_context
|
||||||
|
from src.meeting_lab.orchestration.mvp import regenerate_mvp_protocol
|
||||||
|
from src.meeting_lab.protocol.direct_protocol_prompt import build_direct_protocol_prompt
|
||||||
|
from src.meeting_lab.protocol.generate_direct_protocol import (
|
||||||
|
estimate_input_tokens,
|
||||||
|
generate_direct_protocol,
|
||||||
|
select_transcript_input,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
class MeetingLanguageTests(unittest.TestCase):
|
||||||
|
def test_diarized_plain_fallback_retains_english_instruction(self) -> None:
|
||||||
|
segments = [
|
||||||
|
{
|
||||||
|
"start": index,
|
||||||
|
"end": index + 1,
|
||||||
|
"text": "Word.",
|
||||||
|
"speaker_id": f"SPEAKER_{index % 2:02d}",
|
||||||
|
}
|
||||||
|
for index in range(200)
|
||||||
|
]
|
||||||
|
plain = " ".join(segment["text"] for segment in segments)
|
||||||
|
budget = estimate_input_tokens(
|
||||||
|
build_direct_protocol_prompt(plain, meeting_language="en")
|
||||||
|
)
|
||||||
|
selected = select_transcript_input(
|
||||||
|
{"text": plain, "segments": segments, "speaker_labels_anonymous": True},
|
||||||
|
None,
|
||||||
|
meeting_language="en",
|
||||||
|
safe_input_token_budget=budget,
|
||||||
|
)
|
||||||
|
self.assertEqual(selected.representation, "plain_transcript_fallback")
|
||||||
|
self.assertIn("Write the meeting protocol in English.", selected.prompt)
|
||||||
|
self.assertEqual(selected.text, plain)
|
||||||
|
|
||||||
|
def test_generation_and_regeneration_preserve_language_and_inputs(self) -> None:
|
||||||
|
for language, expected in (
|
||||||
|
("de", "German"),
|
||||||
|
("en", "English"),
|
||||||
|
(None, "German"),
|
||||||
|
):
|
||||||
|
for diarized in (False, True):
|
||||||
|
with (
|
||||||
|
self.subTest(language=language, diarized=diarized),
|
||||||
|
tempfile.TemporaryDirectory() as directory,
|
||||||
|
):
|
||||||
|
run = Path(directory)
|
||||||
|
context_path = run / "context" / "meeting_context.yaml"
|
||||||
|
context_path.parent.mkdir()
|
||||||
|
context_text = (
|
||||||
|
'schema_version: "1"\nmeeting:\n'
|
||||||
|
" meeting_id: test\n title: Mixed terminology\n"
|
||||||
|
+ (f" language: {language}\n" if language else "")
|
||||||
|
+ ' notes: "Freigabe für Product X"\n'
|
||||||
|
"participants:\n - participant_id: person\n"
|
||||||
|
' display_name: "Jörg Müller"\n'
|
||||||
|
"speaker_mappings:\n SPEAKER_00: person\n"
|
||||||
|
)
|
||||||
|
context_path.write_text(context_text, encoding="utf-8")
|
||||||
|
transcript = run / ("diarization" if diarized else "transcript")
|
||||||
|
transcript.mkdir()
|
||||||
|
transcript /= (
|
||||||
|
"transcript_diarized.json" if diarized else "transcript.json"
|
||||||
|
)
|
||||||
|
transcript.write_text(
|
||||||
|
json.dumps(
|
||||||
|
{
|
||||||
|
"text": "I will check Product X.",
|
||||||
|
"speaker_labels_anonymous": diarized,
|
||||||
|
"segments": [
|
||||||
|
{
|
||||||
|
"id": 0,
|
||||||
|
"start": 0,
|
||||||
|
"end": 1,
|
||||||
|
"speaker_id": "SPEAKER_00",
|
||||||
|
"text": "I will check Product X.",
|
||||||
|
}
|
||||||
|
],
|
||||||
|
}
|
||||||
|
),
|
||||||
|
encoding="utf-8",
|
||||||
|
)
|
||||||
|
original = transcript.read_bytes()
|
||||||
|
call = Mock(
|
||||||
|
return_value=OllamaGeneration(
|
||||||
|
text="# Meeting Protocol",
|
||||||
|
raw_response={},
|
||||||
|
client_wall_time_seconds=0.1,
|
||||||
|
)
|
||||||
|
)
|
||||||
|
generate = partial(
|
||||||
|
generate_direct_protocol,
|
||||||
|
model_check=Mock(return_value={}),
|
||||||
|
generation_call=call,
|
||||||
|
)
|
||||||
|
result = generate(transcript, context_path)
|
||||||
|
self.assertIn(
|
||||||
|
f"Write the meeting protocol in {expected}.",
|
||||||
|
result.exact_prompt,
|
||||||
|
)
|
||||||
|
self.assertNotIn("in deutscher Sprache", result.exact_prompt)
|
||||||
|
self.assertNotIn("auf Deutsch", result.exact_prompt)
|
||||||
|
self.assertIn("Jörg Müller", result.exact_prompt)
|
||||||
|
self.assertIn("Mixed terminology", result.exact_prompt)
|
||||||
|
self.assertIn("I will check Product X.", result.transcript_input)
|
||||||
|
self.assertEqual(
|
||||||
|
result.runtime_metadata["output_language"], language or "de"
|
||||||
|
)
|
||||||
|
self.assertEqual(
|
||||||
|
context_path.read_text(encoding="utf-8"), context_text
|
||||||
|
)
|
||||||
|
context = load_meeting_context(context_path)
|
||||||
|
with (
|
||||||
|
patch(
|
||||||
|
"src.meeting_lab.orchestration.mvp.generate_direct_protocol",
|
||||||
|
side_effect=generate,
|
||||||
|
),
|
||||||
|
patch(
|
||||||
|
"src.meeting_lab.orchestration.mvp.transcribe_audio",
|
||||||
|
side_effect=AssertionError("Retranscription is forbidden"),
|
||||||
|
),
|
||||||
|
):
|
||||||
|
regenerate_mvp_protocol(run, meeting_context=context)
|
||||||
|
metadata = json.loads(
|
||||||
|
(run / "protocol" / "runtime_metadata.json").read_text()
|
||||||
|
)
|
||||||
|
self.assertEqual(metadata["output_language"], language or "de")
|
||||||
|
self.assertIn(
|
||||||
|
f"Write the meeting protocol in {expected}.",
|
||||||
|
call.call_args.args[2],
|
||||||
|
)
|
||||||
|
self.assertEqual(transcript.read_bytes(), original)
|
||||||
|
self.assertEqual(
|
||||||
|
load_meeting_context(context_path).data, context.data
|
||||||
|
)
|
||||||
|
self.assertEqual(context.speaker_mappings, {"SPEAKER_00": "person"})
|
||||||
|
|
||||||
|
def test_no_context_defaults_to_german(self) -> None:
|
||||||
|
with tempfile.TemporaryDirectory() as directory:
|
||||||
|
transcript = Path(directory) / "transcript.json"
|
||||||
|
transcript.write_text('{"text": "English source text."}')
|
||||||
|
result = generate_direct_protocol(
|
||||||
|
transcript,
|
||||||
|
model_check=Mock(return_value={}),
|
||||||
|
generation_call=Mock(
|
||||||
|
return_value=OllamaGeneration(
|
||||||
|
text="Protocol",
|
||||||
|
raw_response={},
|
||||||
|
client_wall_time_seconds=0.1,
|
||||||
|
)
|
||||||
|
),
|
||||||
|
)
|
||||||
|
self.assertIn("Write the meeting protocol in German.", result.exact_prompt)
|
||||||
|
self.assertEqual(result.runtime_metadata["output_language"], "de")
|
||||||
+137
-2
@@ -1,11 +1,14 @@
|
|||||||
import json
|
import json
|
||||||
|
import shutil
|
||||||
import subprocess
|
import subprocess
|
||||||
import tempfile
|
import tempfile
|
||||||
import unittest
|
import unittest
|
||||||
|
from dataclasses import replace
|
||||||
from pathlib import Path
|
from pathlib import Path
|
||||||
from unittest.mock import patch
|
from unittest.mock import patch
|
||||||
|
|
||||||
from scripts import run_mvp_meeting as cli
|
from scripts import run_mvp_meeting as cli
|
||||||
|
from src.meeting_lab.audio import PreparedAudio
|
||||||
from src.meeting_lab.models.meeting_context import load_meeting_context
|
from src.meeting_lab.models.meeting_context import load_meeting_context
|
||||||
from src.meeting_lab.orchestration import mvp as mvp_api
|
from src.meeting_lab.orchestration import mvp as mvp_api
|
||||||
from src.meeting_lab.orchestration.mvp import MvpMeetingConfig, MvpRunResult
|
from src.meeting_lab.orchestration.mvp import MvpMeetingConfig, MvpRunResult
|
||||||
@@ -71,6 +74,12 @@ def fake_transcribe(audio, model, output, language, **kwargs):
|
|||||||
return TranscriptionResult(output, raw, transcript, text, metadata, 0.1)
|
return TranscriptionResult(output, raw, transcript, text, metadata, 0.1)
|
||||||
|
|
||||||
|
|
||||||
|
def fake_prepare(source, destination, **kwargs):
|
||||||
|
destination.parent.mkdir(parents=True, exist_ok=True)
|
||||||
|
shutil.copyfile(source, destination)
|
||||||
|
return PreparedAudio(source, source.suffix.removeprefix("."), destination, "ffmpeg", "ffmpeg")
|
||||||
|
|
||||||
|
|
||||||
def fake_protocol(transcript, context, **kwargs):
|
def fake_protocol(transcript, context, **kwargs):
|
||||||
rendered_context = load_meeting_context(context)
|
rendered_context = load_meeting_context(context)
|
||||||
assert rendered_context.meeting_id == "programmatic-test"
|
assert rendered_context.meeting_id == "programmatic-test"
|
||||||
@@ -103,9 +112,10 @@ class MvpApiTests(unittest.TestCase):
|
|||||||
events = []
|
events = []
|
||||||
with (
|
with (
|
||||||
patch.object(mvp_api, "transcribe_audio", side_effect=fake_transcribe),
|
patch.object(mvp_api, "transcribe_audio", side_effect=fake_transcribe),
|
||||||
|
patch.object(mvp_api, "prepare_audio", side_effect=fake_prepare),
|
||||||
patch.object(
|
patch.object(
|
||||||
mvp_api, "generate_direct_protocol", side_effect=fake_protocol
|
mvp_api, "generate_direct_protocol", side_effect=fake_protocol
|
||||||
),
|
) as protocol_generator,
|
||||||
patch.object(subprocess, "run") as subprocess_run,
|
patch.object(subprocess, "run") as subprocess_run,
|
||||||
):
|
):
|
||||||
result = mvp_api.run_mvp_meeting(
|
result = mvp_api.run_mvp_meeting(
|
||||||
@@ -132,6 +142,65 @@ class MvpApiTests(unittest.TestCase):
|
|||||||
],
|
],
|
||||||
)
|
)
|
||||||
self.assertTrue(all(event.progress is None for event in events))
|
self.assertTrue(all(event.progress is None for event in events))
|
||||||
|
self.assertEqual(protocol_generator.call_args.kwargs["num_ctx"], 32_768)
|
||||||
|
self.assertEqual(
|
||||||
|
protocol_generator.call_args.kwargs["safe_input_token_budget"],
|
||||||
|
29_000,
|
||||||
|
)
|
||||||
|
|
||||||
|
def test_protocol_only_regeneration_reuses_diarized_artifacts(self):
|
||||||
|
with tempfile.TemporaryDirectory() as directory:
|
||||||
|
root = Path(directory)
|
||||||
|
run_dir = root / "existing-run"
|
||||||
|
diarization_dir = run_dir / "diarization"
|
||||||
|
diarization_dir.mkdir(parents=True)
|
||||||
|
source = diarization_dir / "transcript_diarized.json"
|
||||||
|
source.write_text(
|
||||||
|
json.dumps(
|
||||||
|
{
|
||||||
|
"text": "SPEAKER_00: Existing statement.\n",
|
||||||
|
"segments": [
|
||||||
|
{
|
||||||
|
"start": 0.0,
|
||||||
|
"end": 1.0,
|
||||||
|
"speaker_id": "SPEAKER_00",
|
||||||
|
"text": "Existing statement.",
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"speaker_labels_anonymous": True,
|
||||||
|
}
|
||||||
|
),
|
||||||
|
encoding="utf-8",
|
||||||
|
)
|
||||||
|
source_before = source.read_bytes()
|
||||||
|
mapped_context = context_data()
|
||||||
|
mapped_context["speaker_mappings"] = {"SPEAKER_00": "person-1"}
|
||||||
|
with (
|
||||||
|
patch.object(mvp_api, "prepare_audio") as preparation,
|
||||||
|
patch.object(mvp_api, "transcribe_audio") as transcription,
|
||||||
|
patch.object(mvp_api, "diarize_audio") as diarization,
|
||||||
|
patch.object(
|
||||||
|
mvp_api, "generate_direct_protocol", side_effect=fake_protocol
|
||||||
|
) as protocol,
|
||||||
|
):
|
||||||
|
result = mvp_api.regenerate_mvp_protocol(
|
||||||
|
run_dir,
|
||||||
|
meeting_context=mapped_context,
|
||||||
|
model="qwen3.8:27b",
|
||||||
|
protocol_num_ctx=32_768,
|
||||||
|
protocol_safe_input_token_budget=29_000,
|
||||||
|
)
|
||||||
|
|
||||||
|
self.assertEqual(result.exit_code, 0)
|
||||||
|
self.assertEqual(result.protocol_path, run_dir / "protocol.md")
|
||||||
|
preparation.assert_not_called()
|
||||||
|
transcription.assert_not_called()
|
||||||
|
diarization.assert_not_called()
|
||||||
|
self.assertEqual(protocol.call_args.args[0], source)
|
||||||
|
self.assertEqual(protocol.call_args.kwargs["num_ctx"], 32_768)
|
||||||
|
self.assertEqual(source.read_bytes(), source_before)
|
||||||
|
persisted = load_meeting_context(run_dir / "context/meeting_context.yaml")
|
||||||
|
self.assertEqual(persisted.speaker_mappings, {"SPEAKER_00": "person-1"})
|
||||||
|
|
||||||
def test_failure_emits_terminal_failure_event(self):
|
def test_failure_emits_terminal_failure_event(self):
|
||||||
with tempfile.TemporaryDirectory() as directory:
|
with tempfile.TemporaryDirectory() as directory:
|
||||||
@@ -142,7 +211,7 @@ class MvpApiTests(unittest.TestCase):
|
|||||||
mvp_api,
|
mvp_api,
|
||||||
"transcribe_audio",
|
"transcribe_audio",
|
||||||
side_effect=TranscriptionError("stopped"),
|
side_effect=TranscriptionError("stopped"),
|
||||||
):
|
), patch.object(mvp_api, "prepare_audio", side_effect=fake_prepare):
|
||||||
result = mvp_api.run_mvp_meeting(config, progress_sink=events.append)
|
result = mvp_api.run_mvp_meeting(config, progress_sink=events.append)
|
||||||
|
|
||||||
self.assertEqual(result.exit_code, 2)
|
self.assertEqual(result.exit_code, 2)
|
||||||
@@ -168,8 +237,74 @@ class MvpApiTests(unittest.TestCase):
|
|||||||
self.assertEqual(delegated.diarization, "off")
|
self.assertEqual(delegated.diarization, "off")
|
||||||
self.assertEqual(delegated.language, "de")
|
self.assertEqual(delegated.language, "de")
|
||||||
self.assertEqual(delegated.whisper_executable, "whisper-cli")
|
self.assertEqual(delegated.whisper_executable, "whisper-cli")
|
||||||
|
self.assertEqual(delegated.ffmpeg_executable, "ffmpeg")
|
||||||
|
self.assertTrue(delegated.audio_normalization)
|
||||||
|
self.assertEqual(delegated.protocol_num_ctx, 32_768)
|
||||||
|
self.assertEqual(delegated.protocol_safe_input_token_budget, 29_000)
|
||||||
self.assertEqual(api.call_args.kwargs["meeting_context"], context_data())
|
self.assertEqual(api.call_args.kwargs["meeting_context"], context_data())
|
||||||
|
|
||||||
|
def test_cli_explicit_audio_normalization_values_are_propagated(self):
|
||||||
|
with tempfile.TemporaryDirectory() as directory:
|
||||||
|
root = Path(directory)
|
||||||
|
config = self.config(root)
|
||||||
|
enabled = cli.config_from_args(
|
||||||
|
cli.parse_args(
|
||||||
|
[
|
||||||
|
str(config.audio_file),
|
||||||
|
"--whisper-model",
|
||||||
|
str(config.whisper_model),
|
||||||
|
"--audio-normalization",
|
||||||
|
]
|
||||||
|
)
|
||||||
|
)
|
||||||
|
disabled = cli.config_from_args(
|
||||||
|
cli.parse_args(
|
||||||
|
[
|
||||||
|
str(config.audio_file),
|
||||||
|
"--whisper-model",
|
||||||
|
str(config.whisper_model),
|
||||||
|
"--no-audio-normalization",
|
||||||
|
]
|
||||||
|
)
|
||||||
|
)
|
||||||
|
|
||||||
|
self.assertTrue(enabled.audio_normalization)
|
||||||
|
self.assertFalse(disabled.audio_normalization)
|
||||||
|
|
||||||
|
def test_transcription_receives_prepared_wav_for_encoded_inputs(self):
|
||||||
|
for suffix in (".flac", ".m4a"):
|
||||||
|
with self.subTest(suffix=suffix), tempfile.TemporaryDirectory() as directory:
|
||||||
|
root = Path(directory)
|
||||||
|
config = self.config(root)
|
||||||
|
encoded = config.audio_file.with_suffix(suffix)
|
||||||
|
config.audio_file.rename(encoded)
|
||||||
|
config = replace(config, audio_file=encoded)
|
||||||
|
received = []
|
||||||
|
|
||||||
|
def capture_transcribe(audio, *args, received_paths=received, **kwargs):
|
||||||
|
received_paths.append(audio)
|
||||||
|
return fake_transcribe(audio, *args, **kwargs)
|
||||||
|
|
||||||
|
with (
|
||||||
|
patch.object(mvp_api, "prepare_audio", side_effect=fake_prepare),
|
||||||
|
patch.object(mvp_api, "transcribe_audio", side_effect=capture_transcribe),
|
||||||
|
patch.object(mvp_api, "generate_direct_protocol", side_effect=fake_protocol),
|
||||||
|
):
|
||||||
|
result = mvp_api.run_mvp_meeting(
|
||||||
|
config, meeting_context=context_data()
|
||||||
|
)
|
||||||
|
|
||||||
|
self.assertEqual(result.exit_code, 0)
|
||||||
|
self.assertEqual(received, [result.run_dir / "audio" / "prepared.wav"])
|
||||||
|
manifest = json.loads(
|
||||||
|
(result.run_dir / "audio" / "input_manifest.json").read_text()
|
||||||
|
)
|
||||||
|
self.assertEqual(manifest["format"], suffix.removeprefix("."))
|
||||||
|
self.assertEqual(
|
||||||
|
manifest["prepared_audio"]["prepared_audio_path"],
|
||||||
|
str((result.run_dir / "audio" / "prepared.wav").resolve()),
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
if __name__ == "__main__":
|
if __name__ == "__main__":
|
||||||
unittest.main()
|
unittest.main()
|
||||||
|
|||||||
@@ -1,4 +1,5 @@
|
|||||||
import json
|
import json
|
||||||
|
import shutil
|
||||||
import tempfile
|
import tempfile
|
||||||
import unittest
|
import unittest
|
||||||
from datetime import datetime
|
from datetime import datetime
|
||||||
@@ -6,6 +7,11 @@ from pathlib import Path
|
|||||||
from unittest.mock import Mock, patch
|
from unittest.mock import Mock, patch
|
||||||
|
|
||||||
from scripts import run_mvp_meeting
|
from scripts import run_mvp_meeting
|
||||||
|
from src.meeting_lab.audio import PreparedAudio
|
||||||
|
from src.meeting_lab.audio.preparation import (
|
||||||
|
DEFAULT_NORMALIZATION_FILTER,
|
||||||
|
DEFAULT_NORMALIZATION_METHOD,
|
||||||
|
)
|
||||||
from src.meeting_lab.orchestration import mvp as mvp_api
|
from src.meeting_lab.orchestration import mvp as mvp_api
|
||||||
from src.meeting_lab.protocol.generate_direct_protocol import DirectProtocolResult
|
from src.meeting_lab.protocol.generate_direct_protocol import DirectProtocolResult
|
||||||
from src.meeting_lab.diarization.backend import DiarizationResult
|
from src.meeting_lab.diarization.backend import DiarizationResult
|
||||||
@@ -33,6 +39,7 @@ def protocol_result(model: str = "chosen:model") -> DirectProtocolResult:
|
|||||||
model_metadata={"model": model},
|
model_metadata={"model": model},
|
||||||
runtime_metadata={"model": model, "request_count": 1, "client_wall_time_seconds": 0.5},
|
runtime_metadata={"model": model, "request_count": 1, "client_wall_time_seconds": 0.5},
|
||||||
raw_response={"response": text, "done": True},
|
raw_response={"response": text, "done": True},
|
||||||
|
transcript_input="selected transcript\n",
|
||||||
)
|
)
|
||||||
|
|
||||||
|
|
||||||
@@ -106,7 +113,28 @@ def fake_diarize(audio_path, output_dir, device_mode, **kwargs):
|
|||||||
)
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def fake_prepare(source, destination, **kwargs):
|
||||||
|
destination.parent.mkdir(parents=True, exist_ok=True)
|
||||||
|
shutil.copyfile(source, destination)
|
||||||
|
normalization_enabled = kwargs.get("normalization_enabled", True)
|
||||||
|
return PreparedAudio(
|
||||||
|
source,
|
||||||
|
source.suffix.removeprefix("."),
|
||||||
|
destination,
|
||||||
|
"ffmpeg",
|
||||||
|
kwargs.get("ffmpeg_executable", "ffmpeg"),
|
||||||
|
normalization_enabled,
|
||||||
|
DEFAULT_NORMALIZATION_METHOD if normalization_enabled else None,
|
||||||
|
DEFAULT_NORMALIZATION_FILTER if normalization_enabled else None,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
class MvpOrchestratorTests(unittest.TestCase):
|
class MvpOrchestratorTests(unittest.TestCase):
|
||||||
|
def setUp(self) -> None:
|
||||||
|
patcher = patch.object(mvp_api, "prepare_audio", side_effect=fake_prepare)
|
||||||
|
self.prepare_audio = patcher.start()
|
||||||
|
self.addCleanup(patcher.stop)
|
||||||
|
|
||||||
def create_inputs(self, root: Path) -> tuple[Path, Path, Path]:
|
def create_inputs(self, root: Path) -> tuple[Path, Path, Path]:
|
||||||
audio = root / "team meeting.wav"
|
audio = root / "team meeting.wav"
|
||||||
whisper_model = root / "ggml-model.bin"
|
whisper_model = root / "ggml-model.bin"
|
||||||
@@ -151,12 +179,15 @@ class MvpOrchestratorTests(unittest.TestCase):
|
|||||||
expected = {
|
expected = {
|
||||||
"run_metadata.json",
|
"run_metadata.json",
|
||||||
"audio/input_manifest.json",
|
"audio/input_manifest.json",
|
||||||
|
"audio/prepared.wav",
|
||||||
|
"audio/preparation_metadata.json",
|
||||||
"transcript/whisper_raw.json",
|
"transcript/whisper_raw.json",
|
||||||
"transcript/transcript.json",
|
"transcript/transcript.json",
|
||||||
"transcript/transcript.txt",
|
"transcript/transcript.txt",
|
||||||
"transcript/runtime_metadata.json",
|
"transcript/runtime_metadata.json",
|
||||||
"context/meeting_context.yaml",
|
"context/meeting_context.yaml",
|
||||||
"protocol/exact_prompt.txt",
|
"protocol/exact_prompt.txt",
|
||||||
|
"protocol/transcript_input.txt",
|
||||||
"protocol/raw_response.json",
|
"protocol/raw_response.json",
|
||||||
"protocol/runtime_metadata.json",
|
"protocol/runtime_metadata.json",
|
||||||
"protocol.md",
|
"protocol.md",
|
||||||
@@ -167,14 +198,60 @@ class MvpOrchestratorTests(unittest.TestCase):
|
|||||||
self.assertEqual(metadata["model"], "chosen:model")
|
self.assertEqual(metadata["model"], "chosen:model")
|
||||||
self.assertIsNone(metadata["failure"])
|
self.assertIsNone(metadata["failure"])
|
||||||
self.assertEqual(whisper.call_count, 1)
|
self.assertEqual(whisper.call_count, 1)
|
||||||
|
self.assertEqual(
|
||||||
|
whisper.call_args.args[0], run_dir / "audio" / "prepared.wav"
|
||||||
|
)
|
||||||
self.assertEqual(protocol.call_count, 1)
|
self.assertEqual(protocol.call_count, 1)
|
||||||
|
self.assertTrue(
|
||||||
|
self.prepare_audio.call_args.kwargs["normalization_enabled"]
|
||||||
|
)
|
||||||
|
preparation = json.loads(
|
||||||
|
(run_dir / "audio" / "preparation_metadata.json").read_text()
|
||||||
|
)
|
||||||
|
self.assertTrue(preparation["normalization_enabled"])
|
||||||
|
|
||||||
|
def test_cli_can_disable_audio_normalization_without_bypassing_preparation(self) -> None:
|
||||||
|
with tempfile.TemporaryDirectory() as directory:
|
||||||
|
root = Path(directory)
|
||||||
|
args = self.args(root, ["--no-audio-normalization"])
|
||||||
|
with (
|
||||||
|
patch.object(mvp_api, "transcribe_audio", side_effect=fake_transcribe) as whisper,
|
||||||
|
patch.object(
|
||||||
|
mvp_api,
|
||||||
|
"generate_direct_protocol",
|
||||||
|
return_value=protocol_result(),
|
||||||
|
),
|
||||||
|
):
|
||||||
|
code, run_dir, _ = run_mvp_meeting.run(args)
|
||||||
|
|
||||||
|
self.assertEqual(code, 0)
|
||||||
|
self.prepare_audio.assert_called_once()
|
||||||
|
self.assertFalse(
|
||||||
|
self.prepare_audio.call_args.kwargs["normalization_enabled"]
|
||||||
|
)
|
||||||
|
self.assertEqual(
|
||||||
|
whisper.call_args.args[0], run_dir / "audio" / "prepared.wav"
|
||||||
|
)
|
||||||
|
preparation = json.loads(
|
||||||
|
(run_dir / "audio" / "preparation_metadata.json").read_text()
|
||||||
|
)
|
||||||
|
self.assertFalse(preparation["normalization_enabled"])
|
||||||
|
self.assertIsNone(preparation["normalization_method"])
|
||||||
|
self.assertIsNone(preparation["normalization_filter"])
|
||||||
|
|
||||||
def test_context_model_endpoint_and_whisper_options_are_forwarded(self) -> None:
|
def test_context_model_endpoint_and_whisper_options_are_forwarded(self) -> None:
|
||||||
with tempfile.TemporaryDirectory() as directory:
|
with tempfile.TemporaryDirectory() as directory:
|
||||||
root = Path(directory)
|
root = Path(directory)
|
||||||
args = self.args(
|
args = self.args(
|
||||||
root,
|
root,
|
||||||
["--whisper-executable", "/tools/whisper-cli", "--threads", "4"],
|
[
|
||||||
|
"--whisper-executable",
|
||||||
|
"/tools/whisper-cli",
|
||||||
|
"--ffmpeg-executable",
|
||||||
|
"/tools/ffmpeg",
|
||||||
|
"--threads",
|
||||||
|
"4",
|
||||||
|
],
|
||||||
)
|
)
|
||||||
with (
|
with (
|
||||||
patch.object(mvp_api, "transcribe_audio", side_effect=fake_transcribe) as whisper,
|
patch.object(mvp_api, "transcribe_audio", side_effect=fake_transcribe) as whisper,
|
||||||
@@ -190,6 +267,10 @@ class MvpOrchestratorTests(unittest.TestCase):
|
|||||||
self.assertEqual(whisper.call_args.args[3], "de")
|
self.assertEqual(whisper.call_args.args[3], "de")
|
||||||
self.assertEqual(whisper.call_args.kwargs["executable"], "/tools/whisper-cli")
|
self.assertEqual(whisper.call_args.kwargs["executable"], "/tools/whisper-cli")
|
||||||
self.assertEqual(whisper.call_args.kwargs["threads"], "4")
|
self.assertEqual(whisper.call_args.kwargs["threads"], "4")
|
||||||
|
self.assertEqual(
|
||||||
|
self.prepare_audio.call_args.kwargs["ffmpeg_executable"],
|
||||||
|
"/tools/ffmpeg",
|
||||||
|
)
|
||||||
self.assertEqual(protocol.call_args.args[1], run_dir / "context/meeting_context.yaml")
|
self.assertEqual(protocol.call_args.args[1], run_dir / "context/meeting_context.yaml")
|
||||||
self.assertEqual(protocol.call_args.kwargs["model"], "chosen:model")
|
self.assertEqual(protocol.call_args.kwargs["model"], "chosen:model")
|
||||||
self.assertEqual(
|
self.assertEqual(
|
||||||
@@ -245,6 +326,9 @@ class MvpOrchestratorTests(unittest.TestCase):
|
|||||||
|
|
||||||
self.assertEqual(code, 0)
|
self.assertEqual(code, 0)
|
||||||
self.assertEqual(diarization.call_args.args[2], "gpu")
|
self.assertEqual(diarization.call_args.args[2], "gpu")
|
||||||
|
self.assertEqual(
|
||||||
|
diarization.call_args.args[0], run_dir / "audio" / "prepared.wav"
|
||||||
|
)
|
||||||
self.assertEqual(diarization.call_args.kwargs["runtime"], "container")
|
self.assertEqual(diarization.call_args.kwargs["runtime"], "container")
|
||||||
self.assertEqual(
|
self.assertEqual(
|
||||||
diarization.call_args.kwargs["container_args"], ("--device=/dev/kfd",)
|
diarization.call_args.kwargs["container_args"], ("--device=/dev/kfd",)
|
||||||
|
|||||||
@@ -0,0 +1,315 @@
|
|||||||
|
import json
|
||||||
|
import tempfile
|
||||||
|
import unittest
|
||||||
|
from pathlib import Path
|
||||||
|
from unittest.mock import Mock, patch
|
||||||
|
|
||||||
|
from src.meeting_lab.diarization.alignment import diarized_transcript_text
|
||||||
|
from src.meeting_lab.llm import ollama
|
||||||
|
from src.meeting_lab.llm.ollama import OllamaGeneration
|
||||||
|
from src.meeting_lab.protocol.direct_protocol_prompt import (
|
||||||
|
COMPACT_DIARIZED_PROTOCOL_INSTRUCTION,
|
||||||
|
build_direct_protocol_prompt,
|
||||||
|
)
|
||||||
|
from src.meeting_lab.protocol.generate_direct_protocol import (
|
||||||
|
DirectProtocolError,
|
||||||
|
estimate_input_tokens,
|
||||||
|
generate_direct_protocol,
|
||||||
|
)
|
||||||
|
from src.meeting_lab.protocol.transcript_input import compact_diarized_transcript
|
||||||
|
|
||||||
|
|
||||||
|
def segments() -> list[dict[str, object]]:
|
||||||
|
return [
|
||||||
|
{"id": 0, "start": 0.0, "end": 1.0, "text": "First.", "speaker_id": "SPEAKER_01"},
|
||||||
|
{"id": 1, "start": 1.0, "end": 2.0, "text": "Second.", "speaker_id": "SPEAKER_01"},
|
||||||
|
{"id": 2, "start": 2.0, "end": 3.0, "text": "Third.", "speaker_id": "SPEAKER_04"},
|
||||||
|
{"id": 3, "start": 3.0, "end": 4.0, "text": "Unassigned.", "speaker_id": None},
|
||||||
|
{"id": 4, "start": 4.0, "end": 5.0, "text": "Last.", "speaker_id": "SPEAKER_01"},
|
||||||
|
]
|
||||||
|
|
||||||
|
|
||||||
|
def diarized_document(repetitions: int = 1) -> dict[str, object]:
|
||||||
|
source = segments() * repetitions
|
||||||
|
return {
|
||||||
|
"text": diarized_transcript_text(source),
|
||||||
|
"segments": source,
|
||||||
|
"speaker_labels_anonymous": True,
|
||||||
|
"alignment_source": "exclusive_diarization",
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def completed_generation() -> OllamaGeneration:
|
||||||
|
text = "# Meeting Protocol\n\nComplete."
|
||||||
|
return OllamaGeneration(
|
||||||
|
raw_response={
|
||||||
|
"response": text,
|
||||||
|
"done": True,
|
||||||
|
"done_reason": "stop",
|
||||||
|
"prompt_eval_count": 100,
|
||||||
|
"eval_count": 10,
|
||||||
|
},
|
||||||
|
text=text,
|
||||||
|
client_wall_time_seconds=0.1,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
class CompactDiarizedTranscriptTests(unittest.TestCase):
|
||||||
|
def test_adjacent_segments_group_and_transitions_remain_separate(self) -> None:
|
||||||
|
compact = compact_diarized_transcript(segments())
|
||||||
|
|
||||||
|
self.assertEqual(
|
||||||
|
[block.speaker_id for block in compact.blocks],
|
||||||
|
["SPEAKER_01", "SPEAKER_04", "SPEAKER_UNASSIGNED", "SPEAKER_01"],
|
||||||
|
)
|
||||||
|
self.assertEqual(compact.blocks[0].segment_texts, ("First.", "Second."))
|
||||||
|
self.assertEqual(compact.blocks[-1].segment_texts, ("Last.",))
|
||||||
|
self.assertEqual(compact.text.count("SPEAKER_01:"), 2)
|
||||||
|
|
||||||
|
def test_every_segment_text_and_order_are_preserved(self) -> None:
|
||||||
|
source = segments()
|
||||||
|
compact = compact_diarized_transcript(source)
|
||||||
|
|
||||||
|
self.assertEqual(compact.source_segment_count, len(source))
|
||||||
|
self.assertEqual(compact.represented_segment_count, len(source))
|
||||||
|
self.assertEqual(
|
||||||
|
compact.segment_texts,
|
||||||
|
tuple(str(segment["text"]) for segment in source),
|
||||||
|
)
|
||||||
|
self.assertEqual(compact.segment_texts[0], "First.")
|
||||||
|
self.assertEqual(compact.segment_texts[-1], "Last.")
|
||||||
|
self.assertIn("SPEAKER_UNASSIGNED: Unassigned.", compact.text)
|
||||||
|
|
||||||
|
def test_compact_form_is_materially_smaller_than_per_segment_format(self) -> None:
|
||||||
|
source = [
|
||||||
|
{
|
||||||
|
"start": index,
|
||||||
|
"end": index + 1,
|
||||||
|
"text": "Repeated transcript content.",
|
||||||
|
"speaker_id": "SPEAKER_01",
|
||||||
|
}
|
||||||
|
for index in range(100)
|
||||||
|
]
|
||||||
|
|
||||||
|
compact = compact_diarized_transcript(source).text
|
||||||
|
verbose = diarized_transcript_text(source)
|
||||||
|
|
||||||
|
self.assertLess(len(compact), len(verbose) * 0.6)
|
||||||
|
|
||||||
|
|
||||||
|
class ProtocolInputBudgetTests(unittest.TestCase):
|
||||||
|
def _write(self, root: Path, document: dict[str, object]) -> Path:
|
||||||
|
path = root / "transcript.json"
|
||||||
|
path.write_text(json.dumps(document), encoding="utf-8")
|
||||||
|
return path
|
||||||
|
|
||||||
|
def test_token_estimate_uses_utf8_bytes_for_non_ascii_safety(self) -> None:
|
||||||
|
self.assertEqual(estimate_input_tokens("ä" * 44), 20)
|
||||||
|
|
||||||
|
def test_compact_diarized_representation_selected_within_budget(self) -> None:
|
||||||
|
with tempfile.TemporaryDirectory() as directory:
|
||||||
|
root = Path(directory)
|
||||||
|
document = diarized_document()
|
||||||
|
transcript = self._write(root, document)
|
||||||
|
compact = compact_diarized_transcript(document["segments"]).text
|
||||||
|
budget = estimate_input_tokens(
|
||||||
|
build_direct_protocol_prompt(
|
||||||
|
compact,
|
||||||
|
instruction=COMPACT_DIARIZED_PROTOCOL_INSTRUCTION,
|
||||||
|
)
|
||||||
|
)
|
||||||
|
call = Mock(return_value=completed_generation())
|
||||||
|
|
||||||
|
result = generate_direct_protocol(
|
||||||
|
transcript,
|
||||||
|
safe_input_token_budget=budget,
|
||||||
|
model_check=Mock(return_value={}),
|
||||||
|
generation_call=call,
|
||||||
|
)
|
||||||
|
|
||||||
|
self.assertEqual(
|
||||||
|
result.runtime_metadata["selected_transcript_representation"],
|
||||||
|
"diarized_compact",
|
||||||
|
)
|
||||||
|
self.assertFalse(result.runtime_metadata["fallback_used"])
|
||||||
|
self.assertTrue(result.runtime_metadata["diarization_enabled"])
|
||||||
|
self.assertEqual(result.runtime_metadata["safe_input_token_budget"], budget)
|
||||||
|
self.assertEqual(result.transcript_input, compact)
|
||||||
|
self.assertEqual(call.call_count, 1)
|
||||||
|
|
||||||
|
def test_mapped_speakers_and_statements_reach_final_ollama_payload(self) -> None:
|
||||||
|
diarized = {
|
||||||
|
"text": "",
|
||||||
|
"segments": [
|
||||||
|
{
|
||||||
|
"start": 0.0,
|
||||||
|
"end": 1.0,
|
||||||
|
"speaker_id": "SPEAKER_00",
|
||||||
|
"text": "We will run the trial on Wednesday.",
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"start": 1.0,
|
||||||
|
"end": 2.0,
|
||||||
|
"speaker_id": "SPEAKER_01",
|
||||||
|
"text": "I will prepare the raw materials before then.",
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"start": 2.0,
|
||||||
|
"end": 3.0,
|
||||||
|
"speaker_id": "SPEAKER_00",
|
||||||
|
"text": "Good. Anna owns the material preparation.",
|
||||||
|
},
|
||||||
|
],
|
||||||
|
"speaker_labels_anonymous": True,
|
||||||
|
"alignment_source": "exclusive_diarization",
|
||||||
|
}
|
||||||
|
context = {
|
||||||
|
"schema_version": "1",
|
||||||
|
"meeting": {
|
||||||
|
"meeting_id": "speaker-test",
|
||||||
|
"title": "Speaker test",
|
||||||
|
"language": "en",
|
||||||
|
},
|
||||||
|
"participants": [
|
||||||
|
{"participant_id": "martin", "display_name": "Martin"},
|
||||||
|
{"participant_id": "anna", "display_name": "Anna"},
|
||||||
|
],
|
||||||
|
"speaker_mappings": {"SPEAKER_00": "martin", "SPEAKER_01": "anna"},
|
||||||
|
"mentioned_people": [],
|
||||||
|
"organization": {"departments": []},
|
||||||
|
"known_entities": {},
|
||||||
|
}
|
||||||
|
response = Mock()
|
||||||
|
response.raise_for_status.return_value = None
|
||||||
|
response.json.return_value = {"response": "# Meeting Protocol\n", "done": True}
|
||||||
|
|
||||||
|
with tempfile.TemporaryDirectory() as directory:
|
||||||
|
root = Path(directory)
|
||||||
|
transcript = self._write(root, diarized)
|
||||||
|
context_path = root / "context.yaml"
|
||||||
|
context_path.write_text(json.dumps(context), encoding="utf-8")
|
||||||
|
with patch.object(ollama.requests, "post", return_value=response) as post:
|
||||||
|
result = generate_direct_protocol(
|
||||||
|
transcript,
|
||||||
|
context_path,
|
||||||
|
model="qwen3.8:27b",
|
||||||
|
model_check=Mock(return_value={}),
|
||||||
|
)
|
||||||
|
|
||||||
|
prompt = post.call_args.kwargs["json"]["prompt"]
|
||||||
|
self.assertEqual(result.exact_prompt, prompt)
|
||||||
|
self.assertIn("- SPEAKER_00: Martin (participant_id: martin)", prompt)
|
||||||
|
self.assertIn("- SPEAKER_01: Anna (participant_id: anna)", prompt)
|
||||||
|
self.assertIn("SPEAKER_00: We will run the trial on Wednesday.", prompt)
|
||||||
|
self.assertIn(
|
||||||
|
"SPEAKER_01: I will prepare the raw materials before then.", prompt
|
||||||
|
)
|
||||||
|
self.assertIn("SPEAKER_00: Good. Anna owns the material preparation.", prompt)
|
||||||
|
self.assertNotIn("Martin: We will run the trial on Wednesday.", prompt)
|
||||||
|
self.assertNotIn("Anna: I will prepare the raw materials before then.", prompt)
|
||||||
|
self.assertIn("autoritativen SPEAKER_XX-zu-Teilnehmer-Zuordnungen", prompt)
|
||||||
|
self.assertIn("Ich-Zusage", prompt)
|
||||||
|
self.assertIn("nur erwähnten Personen", prompt)
|
||||||
|
self.assertIn("keine persönliche Verantwortung", prompt)
|
||||||
|
|
||||||
|
def test_plain_fallback_selected_when_diarized_compact_exceeds_budget(self) -> None:
|
||||||
|
with tempfile.TemporaryDirectory() as directory:
|
||||||
|
root = Path(directory)
|
||||||
|
alternating_segments = [
|
||||||
|
{
|
||||||
|
"id": index,
|
||||||
|
"start": float(index),
|
||||||
|
"end": float(index + 1),
|
||||||
|
"text": "Word.",
|
||||||
|
"speaker_id": f"SPEAKER_{index % 2:02d}",
|
||||||
|
}
|
||||||
|
for index in range(200)
|
||||||
|
]
|
||||||
|
document = {
|
||||||
|
"text": diarized_transcript_text(alternating_segments),
|
||||||
|
"segments": alternating_segments,
|
||||||
|
"speaker_labels_anonymous": True,
|
||||||
|
"alignment_source": "exclusive_diarization",
|
||||||
|
}
|
||||||
|
transcript = self._write(root, document)
|
||||||
|
compact = compact_diarized_transcript(document["segments"]).text
|
||||||
|
plain = " ".join(str(segment["text"]) for segment in document["segments"])
|
||||||
|
compact_estimate = estimate_input_tokens(
|
||||||
|
build_direct_protocol_prompt(
|
||||||
|
compact,
|
||||||
|
instruction=COMPACT_DIARIZED_PROTOCOL_INSTRUCTION,
|
||||||
|
)
|
||||||
|
)
|
||||||
|
plain_estimate = estimate_input_tokens(build_direct_protocol_prompt(plain))
|
||||||
|
self.assertLess(plain_estimate, compact_estimate)
|
||||||
|
|
||||||
|
result = generate_direct_protocol(
|
||||||
|
transcript,
|
||||||
|
safe_input_token_budget=plain_estimate,
|
||||||
|
model_check=Mock(return_value={}),
|
||||||
|
generation_call=Mock(return_value=completed_generation()),
|
||||||
|
)
|
||||||
|
|
||||||
|
self.assertEqual(
|
||||||
|
result.runtime_metadata["selected_transcript_representation"],
|
||||||
|
"plain_transcript_fallback",
|
||||||
|
)
|
||||||
|
self.assertTrue(result.runtime_metadata["fallback_used"])
|
||||||
|
self.assertEqual(result.transcript_input, plain)
|
||||||
|
self.assertNotIn("SPEAKER_00", result.exact_prompt)
|
||||||
|
self.assertNotIn("SPEAKER_01", result.exact_prompt)
|
||||||
|
self.assertFalse(result.runtime_metadata["speaker_attribution_available"])
|
||||||
|
self.assertEqual(
|
||||||
|
result.runtime_metadata["speaker_attribution_loss_reason"],
|
||||||
|
"plain_transcript_fallback",
|
||||||
|
)
|
||||||
|
|
||||||
|
def test_oversized_plain_transcript_fails_before_any_network_call(self) -> None:
|
||||||
|
with tempfile.TemporaryDirectory() as directory:
|
||||||
|
root = Path(directory)
|
||||||
|
transcript = self._write(
|
||||||
|
root,
|
||||||
|
{"text": "large input " * 1000, "segments": []},
|
||||||
|
)
|
||||||
|
model_check = Mock()
|
||||||
|
generation_call = Mock()
|
||||||
|
|
||||||
|
with self.assertRaisesRegex(
|
||||||
|
DirectProtocolError,
|
||||||
|
"No LLM request was made; silent truncation is not allowed",
|
||||||
|
):
|
||||||
|
generate_direct_protocol(
|
||||||
|
transcript,
|
||||||
|
safe_input_token_budget=1,
|
||||||
|
model_check=model_check,
|
||||||
|
generation_call=generation_call,
|
||||||
|
)
|
||||||
|
|
||||||
|
model_check.assert_not_called()
|
||||||
|
generation_call.assert_not_called()
|
||||||
|
|
||||||
|
def test_existing_plain_path_and_metadata_remain_direct(self) -> None:
|
||||||
|
with tempfile.TemporaryDirectory() as directory:
|
||||||
|
root = Path(directory)
|
||||||
|
transcript = self._write(
|
||||||
|
root,
|
||||||
|
{"text": "Plain original transcript.", "segments": []},
|
||||||
|
)
|
||||||
|
result = generate_direct_protocol(
|
||||||
|
transcript,
|
||||||
|
model_check=Mock(return_value={}),
|
||||||
|
generation_call=Mock(return_value=completed_generation()),
|
||||||
|
)
|
||||||
|
|
||||||
|
self.assertEqual(result.transcript_input, "Plain original transcript.")
|
||||||
|
self.assertEqual(
|
||||||
|
result.runtime_metadata["selected_transcript_representation"],
|
||||||
|
"plain_transcript",
|
||||||
|
)
|
||||||
|
self.assertFalse(result.runtime_metadata["fallback_used"])
|
||||||
|
self.assertFalse(result.runtime_metadata["diarization_enabled"])
|
||||||
|
self.assertIsNone(result.runtime_metadata["speaker_attribution_available"])
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
unittest.main()
|
||||||
Reference in New Issue
Block a user