Files
meeting-lab/docs/protocol_generation_decision.md

14 KiB
Raw Permalink Blame History

Protocol Generation Decision

Executive Summary

Meeting Lab tested ten protocol-generation and runtime variants against the same 93.5-minute reference meeting, project_process_meeting. The current evidence does not support a fully automatic protocol. The best practical local baseline remains one direct call to qwen3.6:35B-A3B, followed by informed human review. It is fast and produces readable, broadly useful Markdown, but it still overstates consensus, compresses unresolved process boundaries, misses some challenge and follow-up paths, and can infer unsafe ownership. Its current quality verdict is C — promising but insufficient.

Additional prompting, review, Meeting Map, hierarchical, diarized, dense-model and 70B-scale variants did not produce a reliable step change. Some improved individual dimensions, but none reached a stable B result or removed the need for substantive review. This is a current evidence-based product choice, not a permanent architecture decision.

Experiments Compared

All protocol rows used the complete cleaned transcript and meeting context unless stated otherwise. A dash means that the artifact did not record the number; it is not an estimate. Verdicts marked “assessment” are comparative assessments of saved outputs because those older artifact directories contain no formal quality_review.json.

Experiment Model Architecture LLM calls Prompt tokens Runtime Human editing Verdict Main strength Main failure MVP Research
Direct one-shot baseline Qwen3.6 35B-A3B Direct full-context protocol 1 18,385 72.5 s cold; about 38 s inference Not recorded C Fast, readable, broad topic outline Consensus and ownership overpromotion; missing boundaries and follow-up Yes, with review Baseline
Conservative one-shot Qwen3.6 35B-A3B Direct with stronger safety instructions 1 19,120 39.5 s warm Not recorded C (assessment) Better uncertainty and pending-feedback language Still invents or upgrades named follow-up actions No Limited
Draft → review Qwen3.6 35B-A3B Conservative draft plus review call 2 39,071 total About 110.4 s summed Not recorded C (assessment) Removes some unsafe named attribution Does not reliably restore omitted content; empty/weak action sections remain No Limited
Meeting Map → protocol Qwen3.6 35B-A3B Semantic map followed by rendering 2 39,129 total 134.8 s Not recorded C (assessment) Explicit intermediate structure Map errors propagate: false consensus and named ownership remain No Yes
Hierarchical notes → protocol Qwen3.6 35B-A3B Five chunk-note calls plus synthesis 6 32,145 total 251.1 s Not recorded C (assessment) Highest recall in several detailed/open topics Amplifies unsupported speaker/name interpretations and confirmed actions No Yes
Segment-level anonymous diarization Qwen3.6 35B-A3B One-shot over 1,154 labeled segments 1 39,308 125.4 s 25–35 min C Preserves some filtered-idea challenge structure Token count more than doubled; actions and deadlines became less safe No No further protocol tests
Turn-merged anonymous diarization Qwen3.6 35B-A3B One-shot over 435 merged turns 1 22,490 101.2 s 25–35 min C Corrected token inflation; recovered some topic and feedback detail Still did not beat raw input; unsafe actions/deadlines persisted No UI/search only
Dense one-shot Qwen3.5 27B Direct full-context protocol 1 18,385 181.2 s Not recorded C (assessment) Somewhat better recall of process details Much slower; no material overall quality gain No No
70B scale one-shot Llama 3.3 70B Q3_K_S Dense, 64% CPU / 36% GPU 1 21,156 470.1 s warm 45–60 min D Technically proved a 70B hybrid load can run Severe coverage loss, invented governance, internal contradiction No Negative scale result
Ollama vs native llama.cpp Qwen3.6 35B-A3B Same Q4_K_M GGUF; ROCm/Vulkan servers 1 per backend 18,385 ROCm 38.4 s; Vulkan 42.3 s N/A Runtime only Native ROCm reached 51.38 generated tok/s No meaningful end-to-end advantage; more operational complexity Ollama Runtime reference

The draft-review total combines the saved conservative draft call and the saved review call. Its review metadata itself reports only the one new review call (19,951 prompt tokens and 70.8 seconds). The direct baseline's 72.5-second wall time includes a 34.5-second cold load; its measured prompt evaluation plus generation was 37.8 seconds. These distinctions explain apparent runtime differences between otherwise similar Qwen3.6 calls.

Recurring quality patterns

  • Topic coverage and factual accuracy: Direct Qwen3.6 captures the main process but misses the second review after enrichment, project reporting and parts of the filtered-idea challenge path. Hierarchical processing recalls more detail but introduces too many unsupported interpretations. Llama 3.3 loses most of the meeting and invents a governance role for the Geschäftsführung.
  • Consensus and unresolved boundaries: Every broad one-shot family remains vulnerable to turning discussion or a working direction into agreement. The unresolved boundary between central coordination and autonomous department work, and the uncertainty around universal filter criteria, are especially fragile.
  • Visibility, veto and reconsideration: No approach consistently preserves initial filtering, later cross-functional challenge, reconsideration after enrichment and the return through the project cycle together.
  • Stakeholder feedback: Pending Jovana and Björn feedback is an important quality probe. Some variants preserve both; segment-level diarization drops Björn, while Llama 3.3 drops both.
  • Actions and attribution: Added structure does not guarantee safety. Conservative, reviewed, Meeting Map, hierarchical and diarized outputs still promote proposals or expected work into confirmed actions, infer owners from roles or conversational context, or invent deadlines. Human review remains mandatory.

Model Findings

Qwen3.6:35B-A3B

Qwen3.6 is the best overall local practical baseline. Its Q4_K_M model is operationally fast on the RX 9070/CPU hybrid setup, follows the requested Markdown form and usually provides a useful first draft. It remains verdict C: larger context and fluent synthesis do not reliably protect evidence strength, responsibility attribution or unresolved process boundaries.

Qwen3.5:27b dense

The dense 27B run recalled some process details better than the MoE baseline, but took 181.2 seconds and generated at 8.45 tokens/s. The gains did not amount to a material overall quality improvement. This result does not prove that dense models are generally inferior; it shows that this dense model is not a better product choice on this hardware and meeting.

Llama 3.3 70B Q3_K_S

Llama 3.3 70B was technically runnable at 32k context with a 42 GB loaded footprint and a 64% CPU / 36% GPU split. Its 7m50s warm meeting run produced a very short, materially worse protocol: one critical invented governance claim, five new major errors and an estimated 45–60 minutes of editing. Raw parameter count alone is therefore insufficient. The older model generation and aggressive Q3 quantization are plausible contributors, but this experiment does not isolate or prove either cause.

Runtime Findings

The native comparison reused the exact 23,938,321,664-byte Qwen3.6 Q4_K_M GGUF that Ollama uses. The tested llama-server binary was the llama.cpp runtime shipped with the installed Ollama distribution, not an independent source build.

Native ROCm processed the reference request in 38.4 seconds and generated at 51.38 tokens/s. Vulkan took 42.3 seconds and generated at 46.73 tokens/s. The comparable Ollama baseline generated at 44.68 tokens/s, with about 37.8 seconds of prompt evaluation plus generation when load time is excluded. Output token counts differed, so generation throughput alone is not an end-to-end quality or latency comparison.

Native ROCm gained some generation throughput, but did not provide a meaningful end-to-end advantage for this workload. Vulkan required more host spill and was not preferable. Ollama already provides the relevant llama.cpp runtime components, model lifecycle and API integration; it remains the preferred routine Meeting Lab runtime.

Diarization Findings

Technical feasibility

Pyannote speaker-diarization-community-1 successfully processed the 93.5-minute meeting on CPU in 1,647 seconds (about 27m27s), an RTF of 0.293. It detected four anonymous clusters and assigned 1,145 of 1,154 Whisper segments (99.2%). Peak RSS was about 3.3 GiB. This establishes technical feasibility; it does not establish speaker identity or diarization accuracy against labeled ground truth.

Protocol-quality impact

Annotating every Whisper segment increased the Qwen prompt from 18,385 to 39,308 tokens, confounding speaker structure with fragmentation and token inflation. Deterministic turn merging reduced 1,154 segments to 435 turns and the complete prompt to 22,490 tokens. That controlled the main representation confound, but the resulting protocol still did not materially outperform the raw transcript and remained verdict C.

Anonymous diarization is therefore not justified as a mandatory MVP protocol-quality feature. This does not mean diarization is generally useless. It may remain valuable for speaker-aware UI, navigation and search, participation statistics, traceability, or later carefully validated real-name mapping.

Semantic Research Findings

The semantic experiments provide architectural evidence, but should not dominate the product decision:

  • Evidence Observation V3 is a strong evidence-near candidate stage. With Qwen3.5 9B it achieved 8 PASS, 1 PARTIAL and 0 FAIL while preserving hedges, alternatives, requests, commitments and boundaries in natural language.
  • Request/Acceptance and Collective Commitment show that narrow semantic recognition followed by deterministic provenance, ordering, addressee, negation and deadline gates can safely derive limited consequences. Model recognition errors were contained without inventing individual ownership.
  • Explicit Rejection failed when reduced to a coarse binary recognition problem: semantically valid false positives passed structural gates.
  • Negative Act Form worked better by distinguishing non-pursuit, personal preference, recommendation and temporary non-action before any normative derivation. All eight form classifications matched Gold, although normalized action text was imperfect in three cases.
  • Target Resolution V0 failed because prompt examples and a weak JSON boundary encouraged the string "null" instead of typed linkage. Target Resolution V1 fixed all linkage/ID failures with deterministic self-linkage, closed ID lists and true JSON Schema, but normalization remained incomplete.
  • Target Normalization V0 improved polarity and scope preservation to 3/4 PASS, but still lost continuation meaning in the collaboration case.

These findings support Meeting Lab as a research and validation track. They do not yet justify placing a multi-stage semantic pipeline on the MVP critical path.

Current MVP Decision

The current product path is:

Audio
-> transcription
-> direct qwen3.6:35B-A3B protocol generation through Ollama
-> informed human review
-> final protocol

The first MVP should treat the generated protocol as an editable draft, not an authoritative semantic record. Human review must specifically check consensus, unresolved boundaries, competing positions, action status, owners, deadlines and pending stakeholder feedback.

Diarization is optional and deferred. The semantic research pipeline remains in Meeting Lab, outside the MVP critical path. Ollama remains the default local runtime.

Rejected / Deferred Directions

  • Do not continue Qwen3.6 prompt variants as the main quality strategy.
  • Do not add draft-review, Meeting Map or hierarchical generation to the MVP; their added calls and complexity did not deliver reliable quality gains.
  • Do not continue anonymous-diarization protocol experiments. Revisit diarization for UI, search, statistics or traceability instead.
  • Do not use Qwen3.5 27B or Llama 3.3 70B Q3_K_S as the routine protocol model.
  • Do not replace Ollama with a manually managed native llama.cpp service for this workload.
  • Retain semantic experiments, but defer production integration and broad semantic consolidation.

Open Questions

  1. How does a genuinely newer, materially stronger model perform when a useful quantization fits the available RAM/VRAM without severe swap?
  2. If project policy permits, what quality ceiling does the unchanged reference prompt achieve with a commercial frontier model?
  3. What is the measured reviewer time and correction distribution once the direct Qwen3.6 draft path is exercised in an end-to-end MVP workflow?
  4. Which non-protocol product benefits justify revisiting diarization later?

No further Qwen3.6 prompt variants, anonymous-diarization protocol runs or old 70B Q3 scale tests are recommended.

Build the practical end-to-end MVP around direct Qwen3.6 generation and an explicit human review handoff. Measure reviewer time and correction categories in real use. Keep the experiment artifacts and semantic Gold work as validation evidence, but do not block the first product loop on broader research stages.