Files
meeting-lab/samples/real_live/gtm_hub_2026-09-07/evaluation.md
T

114 lines
7.1 KiB
Markdown

# Qualitative evaluation — GTM-Hub 2026-09-07
## Method and status
This is a qualitative regression case, not a measured accuracy benchmark.
Assess future outputs against `transcript/transcript.txt`, the evidentiary
source. The human reference is a concise colleague-authored view, not gold
labels. The two Meeting Assistant protocols are comparison variants; their
wording is not itself evidence.
The transcript shows a long attempt to align several interpretations of the
GTM-Hub: a cross-functional place to collect, evaluate, coordinate, and then
assign ideas; separate QMS/process-description work; and later operational
project execution. It records repeated requests for clarification and limited
common understanding, not immediate detailed closure.
## Comparison
| Dimension | Human reference | Assistant with anonymous speaker labels | Assistant with confirmed speaker mappings |
| --- | --- | --- | --- |
| Core meeting intent | Captures the organizational outcome and the three aims: responsibility clarity, information flow, described project flow. | Strong GTM-Hub coverage as strategic collection, evaluation, coordination. | Equally strong, with competing framings retained. |
| Topic coverage | Selective and highly compressed. | Strong: purpose, process/QMS, criteria, roles. | Strong; additionally separates threads and proponents. |
| Discussion reconstruction | Loses much of why discussion evolved. | Reconstructs broad process logic but smooths argumentative turns. | Best reconstruction of structure and repeated clarification. |
| Differing viewpoints | Notes different perspectives, not their content or ownership. | Weak at retaining who held which position. | Best preservation: differing positions remain tied to confirmed speakers. |
| Speaker/person attribution | Initials only; not a complete attribution record. | No confirmed person mapping in input context. | Valuable where mappings are confirmed; names make unsupported interpretation more consequential. |
| Position / proposal / agreement / consensus / decision | Compresses them into a basic shared understanding. | Sometimes promotes discussion to “es wurde vereinbart” or “die Teilnehmer waren sich einig.” | Preserves positions better, but still overstates some proposals/open details as agreement, decision, or obligation. |
| Action-item precision | One concise genuine next step: comment on/give feedback on the BD presentation. | Creates broadly attributed implied actions despite absent explicit individual assignment. | Turns suggestions into named commitments, including delegation by Martin Tazl. |
| Responsibility attribution | Focuses on later project-responsibility clarity rather than naming it prematurely. | Describes future assignment reasonably, but implies responsibility in action wording. | Better attribution does not establish responsibility; the named false action is high risk. |
| Open-question preservation | Retains only the core unresolved process issue. | Lists open points, some synthesized rather than clearly preserved as open. | Keeps several open themes but turns others into settled next steps. |
| Useful detail | Good concise working/distribution reference; insufficient alone for later argument reconstruction. | Detailed enough for overall logic, though attribution and restraint are weak. | Most useful detailed reference for complex multi-speaker discussion, provided commitments are evidence-checked. |
| Risk of overinterpretation | Low through editorial compression, at cost of omitted context. | Moderate: consensus and implied collective actions strengthened. | Highest impact when a semantic overinterpretation is attached to a named person. |
| Product suitability | Strong concise working/distribution protocol; not detailed evidentiary reconstruction. | Better detailed reference than short distribution protocol, but needs restraint. | Best detailed reference of the comparison outputs; a separate later renderer should create the concise distribution protocol. |
### Overall conclusion
The human reference is very strongly editorially compressed. It captures the
main organizational outcome well and works as a concise working/distribution
protocol, but loses much argumentative context and can be insufficient weeks
later when reconstructing why the discussion evolved.
The anonymous-speaker Assistant output has strong topic coverage and
reconstructs overall process logic well. It tends to smooth disagreement into
stronger consensus than the transcript supports, is weaker at retaining who
held which position, and can overstate phrases such as “es wurde vereinbart”
or “die Teilnehmer waren sich einig.”
The mapped-speaker Assistant output best preserves differing viewpoints and
discussion structure. It is substantially better for reconstructing who argued
what in this complex multi-speaker discussion; speaker information increases
the value of a detailed protocol. It does not solve semantic interpretation of
commitments/responsibility. A false action is more dangerous when confidently
attached to a named person.
## Explicit regression expectations
These are qualitative expectations, not a synthetic gold protocol.
### R1 — Core purpose
The detailed protocol should preserve that the GTM-Hub is a cross-functional
collection, evaluation, and coordination point for ideas, markets, products,
or initiatives.
### R2 — Different perspectives
Preserve that participants approached the subject from different
perspectives/mental models; do not claim immediate detailed agreement.
### R3 — Strategic vs. operational distinction
Preserve the distinction among strategic evaluation/coordination, formal
process description/QMS, and operational project execution.
### R4 — Responsibility timing
Do not imply concrete project responsibility before group evaluation and
assignment. The transcript explicitly says responsibility does not yet exist at
that stage and is assigned later.
### R5 — Consensus strength
It is valid to report a basic/common understanding around collection,
discussion, and assignment of ideas. Do not turn unresolved process details
into final consensus.
### R6 — Speaker-aware value
For diarized processing, differing positions should remain attributable to
separate speakers/persons where mappings are confirmed.
### R7 — No responsibility invention
Speaker-aware processing must not turn an organizational suggestion into a
personal commitment. In particular, the statement that formal completeness
checking could be done by an administrative office/secretariat must **not**
become “Martin Tazl will delegate the formal review to an administrative
office” without explicit commitment evidence.
### R8 — Proposal vs. action
A suggestion to structure discussion into two parts must not automatically
become a personal action item for the person proposing it.
### R9 — Human-reference role
Detailed-protocol mode must not be forced to become as short as the human
reference.
### R10 — Detailed-protocol product direction
The main protocol should retain enough detail for a participant to reconstruct
the substance several weeks later. A future short/distribution protocol may
deliberately compress much more strongly.