Add GTM Hub real-world regression case

This commit is contained in:
2026-09-12 11:23:18 +02:00
parent 6fc07690d9
commit 21082e66b3
8 changed files with 565 additions and 0 deletions
+29
View File
@@ -2259,3 +2259,32 @@ the basis of this prototype. Further work should first analyze whether the
failure comes from the schema/prompt representation, the model's sparse-output
reliability, or the boundary between subject grouping and semantic synthesis.
It should not proceed through repeated prompt tuning against these nine cases.
## EXP-0027 — GTM-Hub real-world protocol regression case
Status: Accepted as a qualitative regression case
Date: 2026-09-09
The private `samples/real_live/gtm_hub_2026-09-07/` case preserves an existing
78-minute German GTM-Hub discussion, its exact protocol input transcript, a
human colleague's concise reference, and two existing Meeting Assistant
protocols (with anonymous speaker labels and with five confirmed mappings).
It is an evaluation dataset, not a prompt/model experiment; no inference was
run to create it.
The evaluation establishes product and architecture learning without changing
the current architecture:
- the detailed/working protocol remains the primary product artifact; a short
distribution protocol is a later separate renderer, not a replacement;
- diarization has particular value for long, multi-speaker,
disagreement-heavy meetings because it improves discourse attribution;
- diarization does not solve semantic commitment extraction: speaker
attribution and normative interpretation are separate problems;
- responsibility and action extraction require stricter evidence than ordinary
discussion summarization, especially when a named speaker is involved.
Evidence and explicit qualitative regression expectations are recorded in
`samples/real_live/gtm_hub_2026-09-07/evaluation.md`. The case does not define
a synthetic gold protocol or a benchmark metric.