Add GTM Hub real-world regression case
This commit is contained in:
@@ -0,0 +1,113 @@
|
||||
# Qualitative evaluation — GTM-Hub 2026-09-07
|
||||
|
||||
## Method and status
|
||||
|
||||
This is a qualitative regression case, not a measured accuracy benchmark.
|
||||
Assess future outputs against `transcript/transcript.txt`, the evidentiary
|
||||
source. The human reference is a concise colleague-authored view, not gold
|
||||
labels. The two Meeting Assistant protocols are comparison variants; their
|
||||
wording is not itself evidence.
|
||||
|
||||
The transcript shows a long attempt to align several interpretations of the
|
||||
GTM-Hub: a cross-functional place to collect, evaluate, coordinate, and then
|
||||
assign ideas; separate QMS/process-description work; and later operational
|
||||
project execution. It records repeated requests for clarification and limited
|
||||
common understanding, not immediate detailed closure.
|
||||
|
||||
## Comparison
|
||||
|
||||
| Dimension | Human reference | Assistant with anonymous speaker labels | Assistant with confirmed speaker mappings |
|
||||
| --- | --- | --- | --- |
|
||||
| Core meeting intent | Captures the organizational outcome and the three aims: responsibility clarity, information flow, described project flow. | Strong GTM-Hub coverage as strategic collection, evaluation, coordination. | Equally strong, with competing framings retained. |
|
||||
| Topic coverage | Selective and highly compressed. | Strong: purpose, process/QMS, criteria, roles. | Strong; additionally separates threads and proponents. |
|
||||
| Discussion reconstruction | Loses much of why discussion evolved. | Reconstructs broad process logic but smooths argumentative turns. | Best reconstruction of structure and repeated clarification. |
|
||||
| Differing viewpoints | Notes different perspectives, not their content or ownership. | Weak at retaining who held which position. | Best preservation: differing positions remain tied to confirmed speakers. |
|
||||
| Speaker/person attribution | Initials only; not a complete attribution record. | No confirmed person mapping in input context. | Valuable where mappings are confirmed; names make unsupported interpretation more consequential. |
|
||||
| Position / proposal / agreement / consensus / decision | Compresses them into a basic shared understanding. | Sometimes promotes discussion to “es wurde vereinbart” or “die Teilnehmer waren sich einig.” | Preserves positions better, but still overstates some proposals/open details as agreement, decision, or obligation. |
|
||||
| Action-item precision | One concise genuine next step: comment on/give feedback on the BD presentation. | Creates broadly attributed implied actions despite absent explicit individual assignment. | Turns suggestions into named commitments, including delegation by Martin Tazl. |
|
||||
| Responsibility attribution | Focuses on later project-responsibility clarity rather than naming it prematurely. | Describes future assignment reasonably, but implies responsibility in action wording. | Better attribution does not establish responsibility; the named false action is high risk. |
|
||||
| Open-question preservation | Retains only the core unresolved process issue. | Lists open points, some synthesized rather than clearly preserved as open. | Keeps several open themes but turns others into settled next steps. |
|
||||
| Useful detail | Good concise working/distribution reference; insufficient alone for later argument reconstruction. | Detailed enough for overall logic, though attribution and restraint are weak. | Most useful detailed reference for complex multi-speaker discussion, provided commitments are evidence-checked. |
|
||||
| Risk of overinterpretation | Low through editorial compression, at cost of omitted context. | Moderate: consensus and implied collective actions strengthened. | Highest impact when a semantic overinterpretation is attached to a named person. |
|
||||
| Product suitability | Strong concise working/distribution protocol; not detailed evidentiary reconstruction. | Better detailed reference than short distribution protocol, but needs restraint. | Best detailed reference of the comparison outputs; a separate later renderer should create the concise distribution protocol. |
|
||||
|
||||
### Overall conclusion
|
||||
|
||||
The human reference is very strongly editorially compressed. It captures the
|
||||
main organizational outcome well and works as a concise working/distribution
|
||||
protocol, but loses much argumentative context and can be insufficient weeks
|
||||
later when reconstructing why the discussion evolved.
|
||||
|
||||
The anonymous-speaker Assistant output has strong topic coverage and
|
||||
reconstructs overall process logic well. It tends to smooth disagreement into
|
||||
stronger consensus than the transcript supports, is weaker at retaining who
|
||||
held which position, and can overstate phrases such as “es wurde vereinbart”
|
||||
or “die Teilnehmer waren sich einig.”
|
||||
|
||||
The mapped-speaker Assistant output best preserves differing viewpoints and
|
||||
discussion structure. It is substantially better for reconstructing who argued
|
||||
what in this complex multi-speaker discussion; speaker information increases
|
||||
the value of a detailed protocol. It does not solve semantic interpretation of
|
||||
commitments/responsibility. A false action is more dangerous when confidently
|
||||
attached to a named person.
|
||||
|
||||
## Explicit regression expectations
|
||||
|
||||
These are qualitative expectations, not a synthetic gold protocol.
|
||||
|
||||
### R1 — Core purpose
|
||||
|
||||
The detailed protocol should preserve that the GTM-Hub is a cross-functional
|
||||
collection, evaluation, and coordination point for ideas, markets, products,
|
||||
or initiatives.
|
||||
|
||||
### R2 — Different perspectives
|
||||
|
||||
Preserve that participants approached the subject from different
|
||||
perspectives/mental models; do not claim immediate detailed agreement.
|
||||
|
||||
### R3 — Strategic vs. operational distinction
|
||||
|
||||
Preserve the distinction among strategic evaluation/coordination, formal
|
||||
process description/QMS, and operational project execution.
|
||||
|
||||
### R4 — Responsibility timing
|
||||
|
||||
Do not imply concrete project responsibility before group evaluation and
|
||||
assignment. The transcript explicitly says responsibility does not yet exist at
|
||||
that stage and is assigned later.
|
||||
|
||||
### R5 — Consensus strength
|
||||
|
||||
It is valid to report a basic/common understanding around collection,
|
||||
discussion, and assignment of ideas. Do not turn unresolved process details
|
||||
into final consensus.
|
||||
|
||||
### R6 — Speaker-aware value
|
||||
|
||||
For diarized processing, differing positions should remain attributable to
|
||||
separate speakers/persons where mappings are confirmed.
|
||||
|
||||
### R7 — No responsibility invention
|
||||
|
||||
Speaker-aware processing must not turn an organizational suggestion into a
|
||||
personal commitment. In particular, the statement that formal completeness
|
||||
checking could be done by an administrative office/secretariat must **not**
|
||||
become “Martin Tazl will delegate the formal review to an administrative
|
||||
office” without explicit commitment evidence.
|
||||
|
||||
### R8 — Proposal vs. action
|
||||
|
||||
A suggestion to structure discussion into two parts must not automatically
|
||||
become a personal action item for the person proposing it.
|
||||
|
||||
### R9 — Human-reference role
|
||||
|
||||
Detailed-protocol mode must not be forced to become as short as the human
|
||||
reference.
|
||||
|
||||
### R10 — Detailed-protocol product direction
|
||||
|
||||
The main protocol should retain enough detail for a participant to reconstruct
|
||||
the substance several weeks later. A future short/distribution protocol may
|
||||
deliberately compress much more strongly.
|
||||
Reference in New Issue
Block a user