diff --git a/docs/protocol_generation_decision.md b/docs/protocol_generation_decision.md new file mode 100644 index 0000000..d755b2d --- /dev/null +++ b/docs/protocol_generation_decision.md @@ -0,0 +1,229 @@ +# Protocol Generation Decision + +## Executive Summary + +Meeting Lab tested ten protocol-generation and runtime variants against the +same 93.5-minute reference meeting, `project_process_meeting`. The current +evidence does not support a fully automatic protocol. The best practical local +baseline remains one direct call to `qwen3.6:35B-A3B`, followed by informed +human review. It is fast and produces readable, broadly useful Markdown, but it +still overstates consensus, compresses unresolved process boundaries, misses +some challenge and follow-up paths, and can infer unsafe ownership. Its current +quality verdict is **C — promising but insufficient**. + +Additional prompting, review, Meeting Map, hierarchical, diarized, dense-model +and 70B-scale variants did not produce a reliable step change. Some improved +individual dimensions, but none reached a stable B result or removed the need +for substantive review. This is a current evidence-based product choice, not a +permanent architecture decision. + +## Experiments Compared + +All protocol rows used the complete cleaned transcript and meeting context +unless stated otherwise. A dash means that the artifact did not record the +number; it is not an estimate. Verdicts marked “assessment” are comparative +assessments of saved outputs because those older artifact directories contain +no formal `quality_review.json`. + +| Experiment | Model | Architecture | LLM calls | Prompt tokens | Runtime | Human editing | Verdict | Main strength | Main failure | MVP | Research | +| --- | --- | --- | ---: | ---: | ---: | --- | --- | --- | --- | --- | --- | +| Direct one-shot baseline | Qwen3.6 35B-A3B | Direct full-context protocol | 1 | 18,385 | 72.5 s cold; about 38 s inference | Not recorded | C | Fast, readable, broad topic outline | Consensus and ownership overpromotion; missing boundaries and follow-up | **Yes, with review** | Baseline | +| Conservative one-shot | Qwen3.6 35B-A3B | Direct with stronger safety instructions | 1 | 19,120 | 39.5 s warm | Not recorded | C (assessment) | Better uncertainty and pending-feedback language | Still invents or upgrades named follow-up actions | No | Limited | +| Draft → review | Qwen3.6 35B-A3B | Conservative draft plus review call | 2 | 39,071 total | About 110.4 s summed | Not recorded | C (assessment) | Removes some unsafe named attribution | Does not reliably restore omitted content; empty/weak action sections remain | No | Limited | +| Meeting Map → protocol | Qwen3.6 35B-A3B | Semantic map followed by rendering | 2 | 39,129 total | 134.8 s | Not recorded | C (assessment) | Explicit intermediate structure | Map errors propagate: false consensus and named ownership remain | No | Yes | +| Hierarchical notes → protocol | Qwen3.6 35B-A3B | Five chunk-note calls plus synthesis | 6 | 32,145 total | 251.1 s | Not recorded | C (assessment) | Highest recall in several detailed/open topics | Amplifies unsupported speaker/name interpretations and confirmed actions | No | Yes | +| Segment-level anonymous diarization | Qwen3.6 35B-A3B | One-shot over 1,154 labeled segments | 1 | 39,308 | 125.4 s | 25–35 min | C | Preserves some filtered-idea challenge structure | Token count more than doubled; actions and deadlines became less safe | No | No further protocol tests | +| Turn-merged anonymous diarization | Qwen3.6 35B-A3B | One-shot over 435 merged turns | 1 | 22,490 | 101.2 s | 25–35 min | C | Corrected token inflation; recovered some topic and feedback detail | Still did not beat raw input; unsafe actions/deadlines persisted | No | UI/search only | +| Dense one-shot | Qwen3.5 27B | Direct full-context protocol | 1 | 18,385 | 181.2 s | Not recorded | C (assessment) | Somewhat better recall of process details | Much slower; no material overall quality gain | No | No | +| 70B scale one-shot | Llama 3.3 70B Q3_K_S | Dense, 64% CPU / 36% GPU | 1 | 21,156 | 470.1 s warm | 45–60 min | D | Technically proved a 70B hybrid load can run | Severe coverage loss, invented governance, internal contradiction | No | Negative scale result | +| Ollama vs native llama.cpp | Qwen3.6 35B-A3B | Same Q4_K_M GGUF; ROCm/Vulkan servers | 1 per backend | 18,385 | ROCm 38.4 s; Vulkan 42.3 s | N/A | Runtime only | Native ROCm reached 51.38 generated tok/s | No meaningful end-to-end advantage; more operational complexity | Ollama | Runtime reference | + +The draft-review total combines the saved conservative draft call and the +saved review call. Its review metadata itself reports only the one new review +call (19,951 prompt tokens and 70.8 seconds). The direct baseline's 72.5-second +wall time includes a 34.5-second cold load; its measured prompt evaluation plus +generation was 37.8 seconds. These distinctions explain apparent runtime +differences between otherwise similar Qwen3.6 calls. + +### Recurring quality patterns + +- **Topic coverage and factual accuracy:** Direct Qwen3.6 captures the main + process but misses the second review after enrichment, project reporting and + parts of the filtered-idea challenge path. Hierarchical processing recalls + more detail but introduces too many unsupported interpretations. Llama 3.3 + loses most of the meeting and invents a governance role for the + Geschäftsführung. +- **Consensus and unresolved boundaries:** Every broad one-shot family remains + vulnerable to turning discussion or a working direction into agreement. The + unresolved boundary between central coordination and autonomous department + work, and the uncertainty around universal filter criteria, are especially + fragile. +- **Visibility, veto and reconsideration:** No approach consistently preserves + initial filtering, later cross-functional challenge, reconsideration after + enrichment and the return through the project cycle together. +- **Stakeholder feedback:** Pending Jovana and Björn feedback is an important + quality probe. Some variants preserve both; segment-level diarization drops + Björn, while Llama 3.3 drops both. +- **Actions and attribution:** Added structure does not guarantee safety. + Conservative, reviewed, Meeting Map, hierarchical and diarized outputs still + promote proposals or expected work into confirmed actions, infer owners from + roles or conversational context, or invent deadlines. Human review remains + mandatory. + +## Model Findings + +### Qwen3.6:35B-A3B + +Qwen3.6 is the best overall local practical baseline. Its Q4_K_M model is +operationally fast on the RX 9070/CPU hybrid setup, follows the requested +Markdown form and usually provides a useful first draft. It remains verdict C: +larger context and fluent synthesis do not reliably protect evidence strength, +responsibility attribution or unresolved process boundaries. + +### Qwen3.5:27b dense + +The dense 27B run recalled some process details better than the MoE baseline, +but took 181.2 seconds and generated at 8.45 tokens/s. The gains did not amount +to a material overall quality improvement. This result does not prove that +dense models are generally inferior; it shows that this dense model is not a +better product choice on this hardware and meeting. + +### Llama 3.3 70B Q3_K_S + +Llama 3.3 70B was technically runnable at 32k context with a 42 GB loaded +footprint and a 64% CPU / 36% GPU split. Its 7m50s warm meeting run produced a +very short, materially worse protocol: one critical invented governance claim, +five new major errors and an estimated 45–60 minutes of editing. Raw parameter +count alone is therefore insufficient. The older model generation and +aggressive Q3 quantization are plausible contributors, but this experiment +does not isolate or prove either cause. + +## Runtime Findings + +The native comparison reused the exact 23,938,321,664-byte Qwen3.6 Q4_K_M GGUF +that Ollama uses. The tested `llama-server` binary was the llama.cpp runtime +shipped with the installed Ollama distribution, not an independent source +build. + +Native ROCm processed the reference request in 38.4 seconds and generated at +51.38 tokens/s. Vulkan took 42.3 seconds and generated at 46.73 tokens/s. The +comparable Ollama baseline generated at 44.68 tokens/s, with about 37.8 seconds +of prompt evaluation plus generation when load time is excluded. Output token +counts differed, so generation throughput alone is not an end-to-end quality or +latency comparison. + +Native ROCm gained some generation throughput, but did not provide a meaningful +end-to-end advantage for this workload. Vulkan required more host spill and was +not preferable. Ollama already provides the relevant llama.cpp runtime +components, model lifecycle and API integration; it remains the preferred +routine Meeting Lab runtime. + +## Diarization Findings + +### Technical feasibility + +Pyannote `speaker-diarization-community-1` successfully processed the +93.5-minute meeting on CPU in 1,647 seconds (about 27m27s), an RTF of 0.293. It +detected four anonymous clusters and assigned 1,145 of 1,154 Whisper segments +(99.2%). Peak RSS was about 3.3 GiB. This establishes technical feasibility; it +does not establish speaker identity or diarization accuracy against labeled +ground truth. + +### Protocol-quality impact + +Annotating every Whisper segment increased the Qwen prompt from 18,385 to +39,308 tokens, confounding speaker structure with fragmentation and token +inflation. Deterministic turn merging reduced 1,154 segments to 435 turns and +the complete prompt to 22,490 tokens. That controlled the main representation +confound, but the resulting protocol still did not materially outperform the +raw transcript and remained verdict C. + +Anonymous diarization is therefore not justified as a mandatory MVP +protocol-quality feature. This does **not** mean diarization is generally +useless. It may remain valuable for speaker-aware UI, navigation and search, +participation statistics, traceability, or later carefully validated real-name +mapping. + +## Semantic Research Findings + +The semantic experiments provide architectural evidence, but should not +dominate the product decision: + +- **Evidence Observation V3** is a strong evidence-near candidate stage. With + Qwen3.5 9B it achieved 8 PASS, 1 PARTIAL and 0 FAIL while preserving hedges, + alternatives, requests, commitments and boundaries in natural language. +- **Request/Acceptance** and **Collective Commitment** show that narrow semantic + recognition followed by deterministic provenance, ordering, addressee, + negation and deadline gates can safely derive limited consequences. Model + recognition errors were contained without inventing individual ownership. +- **Explicit Rejection** failed when reduced to a coarse binary recognition + problem: semantically valid false positives passed structural gates. +- **Negative Act Form** worked better by distinguishing non-pursuit, personal + preference, recommendation and temporary non-action before any normative + derivation. All eight form classifications matched Gold, although normalized + action text was imperfect in three cases. +- **Target Resolution V0** failed because prompt examples and a weak JSON + boundary encouraged the string `"null"` instead of typed linkage. + **Target Resolution V1** fixed all linkage/ID failures with deterministic + self-linkage, closed ID lists and true JSON Schema, but normalization remained + incomplete. +- **Target Normalization V0** improved polarity and scope preservation to 3/4 + PASS, but still lost continuation meaning in the collaboration case. + +These findings support Meeting Lab as a research and validation track. They do +not yet justify placing a multi-stage semantic pipeline on the MVP critical +path. + +## Current MVP Decision + +The current product path is: + +```text +Audio +-> transcription +-> direct qwen3.6:35B-A3B protocol generation through Ollama +-> informed human review +-> final protocol +``` + +The first MVP should treat the generated protocol as an editable draft, not an +authoritative semantic record. Human review must specifically check consensus, +unresolved boundaries, competing positions, action status, owners, deadlines +and pending stakeholder feedback. + +Diarization is optional and deferred. The semantic research pipeline remains +in Meeting Lab, outside the MVP critical path. Ollama remains the default local +runtime. + +## Rejected / Deferred Directions + +- Do not continue Qwen3.6 prompt variants as the main quality strategy. +- Do not add draft-review, Meeting Map or hierarchical generation to the MVP; + their added calls and complexity did not deliver reliable quality gains. +- Do not continue anonymous-diarization protocol experiments. Revisit + diarization for UI, search, statistics or traceability instead. +- Do not use Qwen3.5 27B or Llama 3.3 70B Q3_K_S as the routine protocol model. +- Do not replace Ollama with a manually managed native llama.cpp service for + this workload. +- Retain semantic experiments, but defer production integration and broad + semantic consolidation. + +## Open Questions + +1. How does a genuinely newer, materially stronger model perform when a useful + quantization fits the available RAM/VRAM without severe swap? +2. If project policy permits, what quality ceiling does the unchanged reference + prompt achieve with a commercial frontier model? +3. What is the measured reviewer time and correction distribution once the + direct Qwen3.6 draft path is exercised in an end-to-end MVP workflow? +4. Which non-protocol product benefits justify revisiting diarization later? + +No further Qwen3.6 prompt variants, anonymous-diarization protocol runs or old +70B Q3 scale tests are recommended. + +## Recommended Next Product Step + +Build the practical end-to-end MVP around direct Qwen3.6 generation and an +explicit human review handoff. Measure reviewer time and correction categories +in real use. Keep the experiment artifacts and semantic Gold work as validation +evidence, but do not block the first product loop on broader research stages.