114 lines
7.1 KiB
Markdown
114 lines
7.1 KiB
Markdown
# Qualitative evaluation — GTM-Hub 2026-09-07
|
|
|
|
## Method and status
|
|
|
|
This is a qualitative regression case, not a measured accuracy benchmark.
|
|
Assess future outputs against `transcript/transcript.txt`, the evidentiary
|
|
source. The human reference is a concise colleague-authored view, not gold
|
|
labels. The two Meeting Assistant protocols are comparison variants; their
|
|
wording is not itself evidence.
|
|
|
|
The transcript shows a long attempt to align several interpretations of the
|
|
GTM-Hub: a cross-functional place to collect, evaluate, coordinate, and then
|
|
assign ideas; separate QMS/process-description work; and later operational
|
|
project execution. It records repeated requests for clarification and limited
|
|
common understanding, not immediate detailed closure.
|
|
|
|
## Comparison
|
|
|
|
| Dimension | Human reference | Assistant with anonymous speaker labels | Assistant with confirmed speaker mappings |
|
|
| --- | --- | --- | --- |
|
|
| Core meeting intent | Captures the organizational outcome and the three aims: responsibility clarity, information flow, described project flow. | Strong GTM-Hub coverage as strategic collection, evaluation, coordination. | Equally strong, with competing framings retained. |
|
|
| Topic coverage | Selective and highly compressed. | Strong: purpose, process/QMS, criteria, roles. | Strong; additionally separates threads and proponents. |
|
|
| Discussion reconstruction | Loses much of why discussion evolved. | Reconstructs broad process logic but smooths argumentative turns. | Best reconstruction of structure and repeated clarification. |
|
|
| Differing viewpoints | Notes different perspectives, not their content or ownership. | Weak at retaining who held which position. | Best preservation: differing positions remain tied to confirmed speakers. |
|
|
| Speaker/person attribution | Initials only; not a complete attribution record. | No confirmed person mapping in input context. | Valuable where mappings are confirmed; names make unsupported interpretation more consequential. |
|
|
| Position / proposal / agreement / consensus / decision | Compresses them into a basic shared understanding. | Sometimes promotes discussion to “es wurde vereinbart” or “die Teilnehmer waren sich einig.” | Preserves positions better, but still overstates some proposals/open details as agreement, decision, or obligation. |
|
|
| Action-item precision | One concise genuine next step: comment on/give feedback on the BD presentation. | Creates broadly attributed implied actions despite absent explicit individual assignment. | Turns suggestions into named commitments, including delegation by Martin Tazl. |
|
|
| Responsibility attribution | Focuses on later project-responsibility clarity rather than naming it prematurely. | Describes future assignment reasonably, but implies responsibility in action wording. | Better attribution does not establish responsibility; the named false action is high risk. |
|
|
| Open-question preservation | Retains only the core unresolved process issue. | Lists open points, some synthesized rather than clearly preserved as open. | Keeps several open themes but turns others into settled next steps. |
|
|
| Useful detail | Good concise working/distribution reference; insufficient alone for later argument reconstruction. | Detailed enough for overall logic, though attribution and restraint are weak. | Most useful detailed reference for complex multi-speaker discussion, provided commitments are evidence-checked. |
|
|
| Risk of overinterpretation | Low through editorial compression, at cost of omitted context. | Moderate: consensus and implied collective actions strengthened. | Highest impact when a semantic overinterpretation is attached to a named person. |
|
|
| Product suitability | Strong concise working/distribution protocol; not detailed evidentiary reconstruction. | Better detailed reference than short distribution protocol, but needs restraint. | Best detailed reference of the comparison outputs; a separate later renderer should create the concise distribution protocol. |
|
|
|
|
### Overall conclusion
|
|
|
|
The human reference is very strongly editorially compressed. It captures the
|
|
main organizational outcome well and works as a concise working/distribution
|
|
protocol, but loses much argumentative context and can be insufficient weeks
|
|
later when reconstructing why the discussion evolved.
|
|
|
|
The anonymous-speaker Assistant output has strong topic coverage and
|
|
reconstructs overall process logic well. It tends to smooth disagreement into
|
|
stronger consensus than the transcript supports, is weaker at retaining who
|
|
held which position, and can overstate phrases such as “es wurde vereinbart”
|
|
or “die Teilnehmer waren sich einig.”
|
|
|
|
The mapped-speaker Assistant output best preserves differing viewpoints and
|
|
discussion structure. It is substantially better for reconstructing who argued
|
|
what in this complex multi-speaker discussion; speaker information increases
|
|
the value of a detailed protocol. It does not solve semantic interpretation of
|
|
commitments/responsibility. A false action is more dangerous when confidently
|
|
attached to a named person.
|
|
|
|
## Explicit regression expectations
|
|
|
|
These are qualitative expectations, not a synthetic gold protocol.
|
|
|
|
### R1 — Core purpose
|
|
|
|
The detailed protocol should preserve that the GTM-Hub is a cross-functional
|
|
collection, evaluation, and coordination point for ideas, markets, products,
|
|
or initiatives.
|
|
|
|
### R2 — Different perspectives
|
|
|
|
Preserve that participants approached the subject from different
|
|
perspectives/mental models; do not claim immediate detailed agreement.
|
|
|
|
### R3 — Strategic vs. operational distinction
|
|
|
|
Preserve the distinction among strategic evaluation/coordination, formal
|
|
process description/QMS, and operational project execution.
|
|
|
|
### R4 — Responsibility timing
|
|
|
|
Do not imply concrete project responsibility before group evaluation and
|
|
assignment. The transcript explicitly says responsibility does not yet exist at
|
|
that stage and is assigned later.
|
|
|
|
### R5 — Consensus strength
|
|
|
|
It is valid to report a basic/common understanding around collection,
|
|
discussion, and assignment of ideas. Do not turn unresolved process details
|
|
into final consensus.
|
|
|
|
### R6 — Speaker-aware value
|
|
|
|
For diarized processing, differing positions should remain attributable to
|
|
separate speakers/persons where mappings are confirmed.
|
|
|
|
### R7 — No responsibility invention
|
|
|
|
Speaker-aware processing must not turn an organizational suggestion into a
|
|
personal commitment. In particular, the statement that formal completeness
|
|
checking could be done by an administrative office/secretariat must **not**
|
|
become “Martin Tazl will delegate the formal review to an administrative
|
|
office” without explicit commitment evidence.
|
|
|
|
### R8 — Proposal vs. action
|
|
|
|
A suggestion to structure discussion into two parts must not automatically
|
|
become a personal action item for the person proposing it.
|
|
|
|
### R9 — Human-reference role
|
|
|
|
Detailed-protocol mode must not be forced to become as short as the human
|
|
reference.
|
|
|
|
### R10 — Detailed-protocol product direction
|
|
|
|
The main protocol should retain enough detail for a participant to reconstruct
|
|
the substance several weeks later. A future short/distribution protocol may
|
|
deliberately compress much more strongly.
|