Add GTM Hub real-world regression case
This commit is contained in:
@@ -0,0 +1,69 @@
|
||||
# GTM-Hub meeting — 2026-09-07
|
||||
|
||||
Durable real-world regression case for a German GTM-Hub meeting. This is a
|
||||
private real-meeting sample; do not publish or redistribute it outside the
|
||||
intended development context.
|
||||
|
||||
## Case metadata
|
||||
|
||||
- **Case ID:** `gtm_hub_2026-09-07`
|
||||
- **Meeting:** GTM-Hub Meeting, 2026-09-07
|
||||
- **Duration:** about 78 min 36 s (4,715.52 s from existing diarization metadata)
|
||||
- **Context:** a biweekly cross-functional board meeting about coordinating
|
||||
projects for new markets, products, and applications; this discussion focused
|
||||
initially on administrative questions and process definition.
|
||||
- **Participants:** five confirmed participants: Malte Schnau, Henning
|
||||
Ehrenberg, Björn-Erik Falkenau, Martin Tazl, and Jovana Husemann.
|
||||
|
||||
## Contents and provenance
|
||||
|
||||
`transcript/transcript.txt` is the exact compact diarized transcript
|
||||
representation supplied to both comparable Meeting Assistant protocol calls;
|
||||
it is the evidentiary source for evaluation. It is byte-identical to both
|
||||
source runs' `protocol/transcript_input.txt` files (SHA-256
|
||||
`42ac52223cac36eeeb7c01e6d6673b2fb50973c4365b5ddaafdd12927b2f12eb`).
|
||||
The source plain Whisper transcripts were also identical between the two runs,
|
||||
but are not copied because they were not the exact representation supplied to
|
||||
protocol generation.
|
||||
|
||||
| Case file | Meeting Assistant source | Generation | Speaker status |
|
||||
| --- | --- | --- | --- |
|
||||
| `references/meeting_assistant_anonymous_speakers.md` | `data/meetings/gtm-hub-meeting-2026-09-07/runs/b895cda2_2026-09-07_11-18-54_GTM-HUB_20260909_025239/protocol.md` | 2026-09-09 01:00:46 UTC; `qwen3.8:27b` | Diarized with anonymous `SPEAKER_XX` labels; zero confirmed speaker-to-person mappings. |
|
||||
| `references/meeting_assistant_speaker.md` | `data/meetings/gtm-hub-meeting-2026-09-07/runs/284a47be_2026-09-07_11-18-54_GTM-HUB_20260909_031106/protocol.md` | 2026-09-09 01:30:44 UTC; `qwen3.8:27b` | Diarized with the same anonymous labels plus five confirmed speaker-to-person mappings in the prompt/context. |
|
||||
| `references/human_protocol_malte.txt` | copied human colleague reference supplied for this case | not generated | author identified as Malte by filename; preserve unchanged. |
|
||||
|
||||
Both runs have meeting ID `gtm-hub-meeting-2026-09-07`, title/date
|
||||
`GTM-Hub Meeting 2026-09-07`, the same source recording name and size
|
||||
(154,620,897-byte FLAC), the same 4,715.52-second prepared-audio duration, the
|
||||
same Whisper model (`ggml-large-v3-turbo.bin`), and the same protocol model,
|
||||
temperature (0), context window (32,768), and compact diarized transcript
|
||||
selection. Their Meeting Context V1 files identify the same meeting and four
|
||||
shared confirmed participants; the speaker-aware context additionally includes
|
||||
Jovana Husemann and the five confirmed mappings. Exact run IDs, hashes, flags,
|
||||
and paths are in `manifest.json`.
|
||||
|
||||
The source runs were separate full-pipeline executions: their audio,
|
||||
transcription, diarization, and exact protocol-input artifacts are
|
||||
byte-identical, but the mapped-speaker run is not operational evidence that it
|
||||
reused the anonymous-speaker run's Whisper or diarization artifacts. This case
|
||||
therefore compares equivalent diarized input conditions with different speaker
|
||||
identity mappings only.
|
||||
|
||||
## Evaluation role
|
||||
|
||||
The transcript is the evidentiary source. The human colleague protocol is
|
||||
**not** a gold protocol: it is a human reference showing what an experienced
|
||||
participant considered worth preserving in a concise working protocol. The two
|
||||
Meeting Assistant files are comparison variants. This case must not force a
|
||||
system to imitate the human protocol verbatim.
|
||||
|
||||
This is a valuable regression case because a long, five-speaker discussion
|
||||
contains competing mental models, repeated disagreement and clarification, and
|
||||
only partial convergence. It distinguishes strategic collection/evaluation and
|
||||
coordination from QMS/process documentation and operational project work. It
|
||||
also discusses responsibility repeatedly without always creating a concrete
|
||||
personal commitment.
|
||||
|
||||
See `evaluation.md` for the qualitative comparison and explicit regression
|
||||
expectations. No audio, raw model response, container log, or other unrelated
|
||||
run artifacts are intentionally included.
|
||||
Reference in New Issue
Block a user