Files
meeting-lab/docs/protocol_generation_decision.md
T

230 lines
14 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Protocol Generation Decision
## Executive Summary
Meeting Lab tested ten protocol-generation and runtime variants against the
same 93.5-minute reference meeting, `project_process_meeting`. The current
evidence does not support a fully automatic protocol. The best practical local
baseline remains one direct call to `qwen3.6:35B-A3B`, followed by informed
human review. It is fast and produces readable, broadly useful Markdown, but it
still overstates consensus, compresses unresolved process boundaries, misses
some challenge and follow-up paths, and can infer unsafe ownership. Its current
quality verdict is **C — promising but insufficient**.
Additional prompting, review, Meeting Map, hierarchical, diarized, dense-model
and 70B-scale variants did not produce a reliable step change. Some improved
individual dimensions, but none reached a stable B result or removed the need
for substantive review. This is a current evidence-based product choice, not a
permanent architecture decision.
## Experiments Compared
All protocol rows used the complete cleaned transcript and meeting context
unless stated otherwise. A dash means that the artifact did not record the
number; it is not an estimate. Verdicts marked “assessment” are comparative
assessments of saved outputs because those older artifact directories contain
no formal `quality_review.json`.
| Experiment | Model | Architecture | LLM calls | Prompt tokens | Runtime | Human editing | Verdict | Main strength | Main failure | MVP | Research |
| --- | --- | --- | ---: | ---: | ---: | --- | --- | --- | --- | --- | --- |
| Direct one-shot baseline | Qwen3.6 35B-A3B | Direct full-context protocol | 1 | 18,385 | 72.5 s cold; about 38 s inference | Not recorded | C | Fast, readable, broad topic outline | Consensus and ownership overpromotion; missing boundaries and follow-up | **Yes, with review** | Baseline |
| Conservative one-shot | Qwen3.6 35B-A3B | Direct with stronger safety instructions | 1 | 19,120 | 39.5 s warm | Not recorded | C (assessment) | Better uncertainty and pending-feedback language | Still invents or upgrades named follow-up actions | No | Limited |
| Draft → review | Qwen3.6 35B-A3B | Conservative draft plus review call | 2 | 39,071 total | About 110.4 s summed | Not recorded | C (assessment) | Removes some unsafe named attribution | Does not reliably restore omitted content; empty/weak action sections remain | No | Limited |
| Meeting Map → protocol | Qwen3.6 35B-A3B | Semantic map followed by rendering | 2 | 39,129 total | 134.8 s | Not recorded | C (assessment) | Explicit intermediate structure | Map errors propagate: false consensus and named ownership remain | No | Yes |
| Hierarchical notes → protocol | Qwen3.6 35B-A3B | Five chunk-note calls plus synthesis | 6 | 32,145 total | 251.1 s | Not recorded | C (assessment) | Highest recall in several detailed/open topics | Amplifies unsupported speaker/name interpretations and confirmed actions | No | Yes |
| Segment-level anonymous diarization | Qwen3.6 35B-A3B | One-shot over 1,154 labeled segments | 1 | 39,308 | 125.4 s | 25–35 min | C | Preserves some filtered-idea challenge structure | Token count more than doubled; actions and deadlines became less safe | No | No further protocol tests |
| Turn-merged anonymous diarization | Qwen3.6 35B-A3B | One-shot over 435 merged turns | 1 | 22,490 | 101.2 s | 25–35 min | C | Corrected token inflation; recovered some topic and feedback detail | Still did not beat raw input; unsafe actions/deadlines persisted | No | UI/search only |
| Dense one-shot | Qwen3.5 27B | Direct full-context protocol | 1 | 18,385 | 181.2 s | Not recorded | C (assessment) | Somewhat better recall of process details | Much slower; no material overall quality gain | No | No |
| 70B scale one-shot | Llama 3.3 70B Q3_K_S | Dense, 64% CPU / 36% GPU | 1 | 21,156 | 470.1 s warm | 45–60 min | D | Technically proved a 70B hybrid load can run | Severe coverage loss, invented governance, internal contradiction | No | Negative scale result |
| Ollama vs native llama.cpp | Qwen3.6 35B-A3B | Same Q4_K_M GGUF; ROCm/Vulkan servers | 1 per backend | 18,385 | ROCm 38.4 s; Vulkan 42.3 s | N/A | Runtime only | Native ROCm reached 51.38 generated tok/s | No meaningful end-to-end advantage; more operational complexity | Ollama | Runtime reference |
The draft-review total combines the saved conservative draft call and the
saved review call. Its review metadata itself reports only the one new review
call (19,951 prompt tokens and 70.8 seconds). The direct baseline's 72.5-second
wall time includes a 34.5-second cold load; its measured prompt evaluation plus
generation was 37.8 seconds. These distinctions explain apparent runtime
differences between otherwise similar Qwen3.6 calls.
### Recurring quality patterns
- **Topic coverage and factual accuracy:** Direct Qwen3.6 captures the main
process but misses the second review after enrichment, project reporting and
parts of the filtered-idea challenge path. Hierarchical processing recalls
more detail but introduces too many unsupported interpretations. Llama 3.3
loses most of the meeting and invents a governance role for the
Geschäftsführung.
- **Consensus and unresolved boundaries:** Every broad one-shot family remains
vulnerable to turning discussion or a working direction into agreement. The
unresolved boundary between central coordination and autonomous department
work, and the uncertainty around universal filter criteria, are especially
fragile.
- **Visibility, veto and reconsideration:** No approach consistently preserves
initial filtering, later cross-functional challenge, reconsideration after
enrichment and the return through the project cycle together.
- **Stakeholder feedback:** Pending Jovana and Björn feedback is an important
quality probe. Some variants preserve both; segment-level diarization drops
Björn, while Llama 3.3 drops both.
- **Actions and attribution:** Added structure does not guarantee safety.
Conservative, reviewed, Meeting Map, hierarchical and diarized outputs still
promote proposals or expected work into confirmed actions, infer owners from
roles or conversational context, or invent deadlines. Human review remains
mandatory.
## Model Findings
### Qwen3.6:35B-A3B
Qwen3.6 is the best overall local practical baseline. Its Q4_K_M model is
operationally fast on the RX 9070/CPU hybrid setup, follows the requested
Markdown form and usually provides a useful first draft. It remains verdict C:
larger context and fluent synthesis do not reliably protect evidence strength,
responsibility attribution or unresolved process boundaries.
### Qwen3.5:27b dense
The dense 27B run recalled some process details better than the MoE baseline,
but took 181.2 seconds and generated at 8.45 tokens/s. The gains did not amount
to a material overall quality improvement. This result does not prove that
dense models are generally inferior; it shows that this dense model is not a
better product choice on this hardware and meeting.
### Llama 3.3 70B Q3_K_S
Llama 3.3 70B was technically runnable at 32k context with a 42 GB loaded
footprint and a 64% CPU / 36% GPU split. Its 7m50s warm meeting run produced a
very short, materially worse protocol: one critical invented governance claim,
five new major errors and an estimated 45–60 minutes of editing. Raw parameter
count alone is therefore insufficient. The older model generation and
aggressive Q3 quantization are plausible contributors, but this experiment
does not isolate or prove either cause.
## Runtime Findings
The native comparison reused the exact 23,938,321,664-byte Qwen3.6 Q4_K_M GGUF
that Ollama uses. The tested `llama-server` binary was the llama.cpp runtime
shipped with the installed Ollama distribution, not an independent source
build.
Native ROCm processed the reference request in 38.4 seconds and generated at
51.38 tokens/s. Vulkan took 42.3 seconds and generated at 46.73 tokens/s. The
comparable Ollama baseline generated at 44.68 tokens/s, with about 37.8 seconds
of prompt evaluation plus generation when load time is excluded. Output token
counts differed, so generation throughput alone is not an end-to-end quality or
latency comparison.
Native ROCm gained some generation throughput, but did not provide a meaningful
end-to-end advantage for this workload. Vulkan required more host spill and was
not preferable. Ollama already provides the relevant llama.cpp runtime
components, model lifecycle and API integration; it remains the preferred
routine Meeting Lab runtime.
## Diarization Findings
### Technical feasibility
Pyannote `speaker-diarization-community-1` successfully processed the
93.5-minute meeting on CPU in 1,647 seconds (about 27m27s), an RTF of 0.293. It
detected four anonymous clusters and assigned 1,145 of 1,154 Whisper segments
(99.2%). Peak RSS was about 3.3 GiB. This establishes technical feasibility; it
does not establish speaker identity or diarization accuracy against labeled
ground truth.
### Protocol-quality impact
Annotating every Whisper segment increased the Qwen prompt from 18,385 to
39,308 tokens, confounding speaker structure with fragmentation and token
inflation. Deterministic turn merging reduced 1,154 segments to 435 turns and
the complete prompt to 22,490 tokens. That controlled the main representation
confound, but the resulting protocol still did not materially outperform the
raw transcript and remained verdict C.
Anonymous diarization is therefore not justified as a mandatory MVP
protocol-quality feature. This does **not** mean diarization is generally
useless. It may remain valuable for speaker-aware UI, navigation and search,
participation statistics, traceability, or later carefully validated real-name
mapping.
## Semantic Research Findings
The semantic experiments provide architectural evidence, but should not
dominate the product decision:
- **Evidence Observation V3** is a strong evidence-near candidate stage. With
Qwen3.5 9B it achieved 8 PASS, 1 PARTIAL and 0 FAIL while preserving hedges,
alternatives, requests, commitments and boundaries in natural language.
- **Request/Acceptance** and **Collective Commitment** show that narrow semantic
recognition followed by deterministic provenance, ordering, addressee,
negation and deadline gates can safely derive limited consequences. Model
recognition errors were contained without inventing individual ownership.
- **Explicit Rejection** failed when reduced to a coarse binary recognition
problem: semantically valid false positives passed structural gates.
- **Negative Act Form** worked better by distinguishing non-pursuit, personal
preference, recommendation and temporary non-action before any normative
derivation. All eight form classifications matched Gold, although normalized
action text was imperfect in three cases.
- **Target Resolution V0** failed because prompt examples and a weak JSON
boundary encouraged the string `"null"` instead of typed linkage.
**Target Resolution V1** fixed all linkage/ID failures with deterministic
self-linkage, closed ID lists and true JSON Schema, but normalization remained
incomplete.
- **Target Normalization V0** improved polarity and scope preservation to 3/4
PASS, but still lost continuation meaning in the collaboration case.
These findings support Meeting Lab as a research and validation track. They do
not yet justify placing a multi-stage semantic pipeline on the MVP critical
path.
## Current MVP Decision
The current product path is:
```text
Audio
-> transcription
-> direct qwen3.6:35B-A3B protocol generation through Ollama
-> informed human review
-> final protocol
```
The first MVP should treat the generated protocol as an editable draft, not an
authoritative semantic record. Human review must specifically check consensus,
unresolved boundaries, competing positions, action status, owners, deadlines
and pending stakeholder feedback.
Diarization is optional and deferred. The semantic research pipeline remains
in Meeting Lab, outside the MVP critical path. Ollama remains the default local
runtime.
## Rejected / Deferred Directions
- Do not continue Qwen3.6 prompt variants as the main quality strategy.
- Do not add draft-review, Meeting Map or hierarchical generation to the MVP;
their added calls and complexity did not deliver reliable quality gains.
- Do not continue anonymous-diarization protocol experiments. Revisit
diarization for UI, search, statistics or traceability instead.
- Do not use Qwen3.5 27B or Llama 3.3 70B Q3_K_S as the routine protocol model.
- Do not replace Ollama with a manually managed native llama.cpp service for
this workload.
- Retain semantic experiments, but defer production integration and broad
semantic consolidation.
## Open Questions
1. How does a genuinely newer, materially stronger model perform when a useful
quantization fits the available RAM/VRAM without severe swap?
2. If project policy permits, what quality ceiling does the unchanged reference
prompt achieve with a commercial frontier model?
3. What is the measured reviewer time and correction distribution once the
direct Qwen3.6 draft path is exercised in an end-to-end MVP workflow?
4. Which non-protocol product benefits justify revisiting diarization later?
No further Qwen3.6 prompt variants, anonymous-diarization protocol runs or old
70B Q3 scale tests are recommended.
## Recommended Next Product Step
Build the practical end-to-end MVP around direct Qwen3.6 generation and an
explicit human review handoff. Measure reviewer time and correction categories
in real use. Keep the experiment artifacts and semantic Gold work as validation
evidence, but do not block the first product loop on broader research stages.