# Protocol Generation Decision ## Executive Summary Meeting Lab tested ten protocol-generation and runtime variants against the same 93.5-minute reference meeting, `project_process_meeting`. The current evidence does not support a fully automatic protocol. The best practical local baseline remains one direct call to `qwen3.6:35B-A3B`, followed by informed human review. It is fast and produces readable, broadly useful Markdown, but it still overstates consensus, compresses unresolved process boundaries, misses some challenge and follow-up paths, and can infer unsafe ownership. Its current quality verdict is **C — promising but insufficient**. Additional prompting, review, Meeting Map, hierarchical, diarized, dense-model and 70B-scale variants did not produce a reliable step change. Some improved individual dimensions, but none reached a stable B result or removed the need for substantive review. This is a current evidence-based product choice, not a permanent architecture decision. ## Experiments Compared All protocol rows used the complete cleaned transcript and meeting context unless stated otherwise. A dash means that the artifact did not record the number; it is not an estimate. Verdicts marked “assessment” are comparative assessments of saved outputs because those older artifact directories contain no formal `quality_review.json`. | Experiment | Model | Architecture | LLM calls | Prompt tokens | Runtime | Human editing | Verdict | Main strength | Main failure | MVP | Research | | --- | --- | --- | ---: | ---: | ---: | --- | --- | --- | --- | --- | --- | | Direct one-shot baseline | Qwen3.6 35B-A3B | Direct full-context protocol | 1 | 18,385 | 72.5 s cold; about 38 s inference | Not recorded | C | Fast, readable, broad topic outline | Consensus and ownership overpromotion; missing boundaries and follow-up | **Yes, with review** | Baseline | | Conservative one-shot | Qwen3.6 35B-A3B | Direct with stronger safety instructions | 1 | 19,120 | 39.5 s warm | Not recorded | C (assessment) | Better uncertainty and pending-feedback language | Still invents or upgrades named follow-up actions | No | Limited | | Draft → review | Qwen3.6 35B-A3B | Conservative draft plus review call | 2 | 39,071 total | About 110.4 s summed | Not recorded | C (assessment) | Removes some unsafe named attribution | Does not reliably restore omitted content; empty/weak action sections remain | No | Limited | | Meeting Map → protocol | Qwen3.6 35B-A3B | Semantic map followed by rendering | 2 | 39,129 total | 134.8 s | Not recorded | C (assessment) | Explicit intermediate structure | Map errors propagate: false consensus and named ownership remain | No | Yes | | Hierarchical notes → protocol | Qwen3.6 35B-A3B | Five chunk-note calls plus synthesis | 6 | 32,145 total | 251.1 s | Not recorded | C (assessment) | Highest recall in several detailed/open topics | Amplifies unsupported speaker/name interpretations and confirmed actions | No | Yes | | Segment-level anonymous diarization | Qwen3.6 35B-A3B | One-shot over 1,154 labeled segments | 1 | 39,308 | 125.4 s | 25–35 min | C | Preserves some filtered-idea challenge structure | Token count more than doubled; actions and deadlines became less safe | No | No further protocol tests | | Turn-merged anonymous diarization | Qwen3.6 35B-A3B | One-shot over 435 merged turns | 1 | 22,490 | 101.2 s | 25–35 min | C | Corrected token inflation; recovered some topic and feedback detail | Still did not beat raw input; unsafe actions/deadlines persisted | No | UI/search only | | Dense one-shot | Qwen3.5 27B | Direct full-context protocol | 1 | 18,385 | 181.2 s | Not recorded | C (assessment) | Somewhat better recall of process details | Much slower; no material overall quality gain | No | No | | 70B scale one-shot | Llama 3.3 70B Q3_K_S | Dense, 64% CPU / 36% GPU | 1 | 21,156 | 470.1 s warm | 45–60 min | D | Technically proved a 70B hybrid load can run | Severe coverage loss, invented governance, internal contradiction | No | Negative scale result | | Ollama vs native llama.cpp | Qwen3.6 35B-A3B | Same Q4_K_M GGUF; ROCm/Vulkan servers | 1 per backend | 18,385 | ROCm 38.4 s; Vulkan 42.3 s | N/A | Runtime only | Native ROCm reached 51.38 generated tok/s | No meaningful end-to-end advantage; more operational complexity | Ollama | Runtime reference | The draft-review total combines the saved conservative draft call and the saved review call. Its review metadata itself reports only the one new review call (19,951 prompt tokens and 70.8 seconds). The direct baseline's 72.5-second wall time includes a 34.5-second cold load; its measured prompt evaluation plus generation was 37.8 seconds. These distinctions explain apparent runtime differences between otherwise similar Qwen3.6 calls. ### Recurring quality patterns - **Topic coverage and factual accuracy:** Direct Qwen3.6 captures the main process but misses the second review after enrichment, project reporting and parts of the filtered-idea challenge path. Hierarchical processing recalls more detail but introduces too many unsupported interpretations. Llama 3.3 loses most of the meeting and invents a governance role for the Geschäftsführung. - **Consensus and unresolved boundaries:** Every broad one-shot family remains vulnerable to turning discussion or a working direction into agreement. The unresolved boundary between central coordination and autonomous department work, and the uncertainty around universal filter criteria, are especially fragile. - **Visibility, veto and reconsideration:** No approach consistently preserves initial filtering, later cross-functional challenge, reconsideration after enrichment and the return through the project cycle together. - **Stakeholder feedback:** Pending Jovana and Björn feedback is an important quality probe. Some variants preserve both; segment-level diarization drops Björn, while Llama 3.3 drops both. - **Actions and attribution:** Added structure does not guarantee safety. Conservative, reviewed, Meeting Map, hierarchical and diarized outputs still promote proposals or expected work into confirmed actions, infer owners from roles or conversational context, or invent deadlines. Human review remains mandatory. ## Model Findings ### Qwen3.6:35B-A3B Qwen3.6 is the best overall local practical baseline. Its Q4_K_M model is operationally fast on the RX 9070/CPU hybrid setup, follows the requested Markdown form and usually provides a useful first draft. It remains verdict C: larger context and fluent synthesis do not reliably protect evidence strength, responsibility attribution or unresolved process boundaries. ### Qwen3.5:27b dense The dense 27B run recalled some process details better than the MoE baseline, but took 181.2 seconds and generated at 8.45 tokens/s. The gains did not amount to a material overall quality improvement. This result does not prove that dense models are generally inferior; it shows that this dense model is not a better product choice on this hardware and meeting. ### Llama 3.3 70B Q3_K_S Llama 3.3 70B was technically runnable at 32k context with a 42 GB loaded footprint and a 64% CPU / 36% GPU split. Its 7m50s warm meeting run produced a very short, materially worse protocol: one critical invented governance claim, five new major errors and an estimated 45–60 minutes of editing. Raw parameter count alone is therefore insufficient. The older model generation and aggressive Q3 quantization are plausible contributors, but this experiment does not isolate or prove either cause. ## Runtime Findings The native comparison reused the exact 23,938,321,664-byte Qwen3.6 Q4_K_M GGUF that Ollama uses. The tested `llama-server` binary was the llama.cpp runtime shipped with the installed Ollama distribution, not an independent source build. Native ROCm processed the reference request in 38.4 seconds and generated at 51.38 tokens/s. Vulkan took 42.3 seconds and generated at 46.73 tokens/s. The comparable Ollama baseline generated at 44.68 tokens/s, with about 37.8 seconds of prompt evaluation plus generation when load time is excluded. Output token counts differed, so generation throughput alone is not an end-to-end quality or latency comparison. Native ROCm gained some generation throughput, but did not provide a meaningful end-to-end advantage for this workload. Vulkan required more host spill and was not preferable. Ollama already provides the relevant llama.cpp runtime components, model lifecycle and API integration; it remains the preferred routine Meeting Lab runtime. ## Diarization Findings ### Technical feasibility Pyannote `speaker-diarization-community-1` successfully processed the 93.5-minute meeting on CPU in 1,647 seconds (about 27m27s), an RTF of 0.293. It detected four anonymous clusters and assigned 1,145 of 1,154 Whisper segments (99.2%). Peak RSS was about 3.3 GiB. This establishes technical feasibility; it does not establish speaker identity or diarization accuracy against labeled ground truth. ### Protocol-quality impact Annotating every Whisper segment increased the Qwen prompt from 18,385 to 39,308 tokens, confounding speaker structure with fragmentation and token inflation. Deterministic turn merging reduced 1,154 segments to 435 turns and the complete prompt to 22,490 tokens. That controlled the main representation confound, but the resulting protocol still did not materially outperform the raw transcript and remained verdict C. Anonymous diarization is therefore not justified as a mandatory MVP protocol-quality feature. This does **not** mean diarization is generally useless. It may remain valuable for speaker-aware UI, navigation and search, participation statistics, traceability, or later carefully validated real-name mapping. ## Semantic Research Findings The semantic experiments provide architectural evidence, but should not dominate the product decision: - **Evidence Observation V3** is a strong evidence-near candidate stage. With Qwen3.5 9B it achieved 8 PASS, 1 PARTIAL and 0 FAIL while preserving hedges, alternatives, requests, commitments and boundaries in natural language. - **Request/Acceptance** and **Collective Commitment** show that narrow semantic recognition followed by deterministic provenance, ordering, addressee, negation and deadline gates can safely derive limited consequences. Model recognition errors were contained without inventing individual ownership. - **Explicit Rejection** failed when reduced to a coarse binary recognition problem: semantically valid false positives passed structural gates. - **Negative Act Form** worked better by distinguishing non-pursuit, personal preference, recommendation and temporary non-action before any normative derivation. All eight form classifications matched Gold, although normalized action text was imperfect in three cases. - **Target Resolution V0** failed because prompt examples and a weak JSON boundary encouraged the string `"null"` instead of typed linkage. **Target Resolution V1** fixed all linkage/ID failures with deterministic self-linkage, closed ID lists and true JSON Schema, but normalization remained incomplete. - **Target Normalization V0** improved polarity and scope preservation to 3/4 PASS, but still lost continuation meaning in the collaboration case. These findings support Meeting Lab as a research and validation track. They do not yet justify placing a multi-stage semantic pipeline on the MVP critical path. ## Current MVP Decision The current product path is: ```text Audio -> transcription -> direct qwen3.6:35B-A3B protocol generation through Ollama -> informed human review -> final protocol ``` The first MVP should treat the generated protocol as an editable draft, not an authoritative semantic record. Human review must specifically check consensus, unresolved boundaries, competing positions, action status, owners, deadlines and pending stakeholder feedback. Diarization is optional and deferred. The semantic research pipeline remains in Meeting Lab, outside the MVP critical path. Ollama remains the default local runtime. ## Rejected / Deferred Directions - Do not continue Qwen3.6 prompt variants as the main quality strategy. - Do not add draft-review, Meeting Map or hierarchical generation to the MVP; their added calls and complexity did not deliver reliable quality gains. - Do not continue anonymous-diarization protocol experiments. Revisit diarization for UI, search, statistics or traceability instead. - Do not use Qwen3.5 27B or Llama 3.3 70B Q3_K_S as the routine protocol model. - Do not replace Ollama with a manually managed native llama.cpp service for this workload. - Retain semantic experiments, but defer production integration and broad semantic consolidation. ## Open Questions 1. How does a genuinely newer, materially stronger model perform when a useful quantization fits the available RAM/VRAM without severe swap? 2. If project policy permits, what quality ceiling does the unchanged reference prompt achieve with a commercial frontier model? 3. What is the measured reviewer time and correction distribution once the direct Qwen3.6 draft path is exercised in an end-to-end MVP workflow? 4. Which non-protocol product benefits justify revisiting diarization later? No further Qwen3.6 prompt variants, anonymous-diarization protocol runs or old 70B Q3 scale tests are recommended. ## Recommended Next Product Step Build the practical end-to-end MVP around direct Qwen3.6 generation and an explicit human review handoff. Measure reviewer time and correction categories in real use. Keep the experiment artifacts and semantic Gold work as validation evidence, but do not block the first product loop on broader research stages.