Document protocol generation decision
This commit is contained in:
@@ -0,0 +1,229 @@
|
||||
# Protocol Generation Decision
|
||||
|
||||
## Executive Summary
|
||||
|
||||
Meeting Lab tested ten protocol-generation and runtime variants against the
|
||||
same 93.5-minute reference meeting, `project_process_meeting`. The current
|
||||
evidence does not support a fully automatic protocol. The best practical local
|
||||
baseline remains one direct call to `qwen3.6:35B-A3B`, followed by informed
|
||||
human review. It is fast and produces readable, broadly useful Markdown, but it
|
||||
still overstates consensus, compresses unresolved process boundaries, misses
|
||||
some challenge and follow-up paths, and can infer unsafe ownership. Its current
|
||||
quality verdict is **C — promising but insufficient**.
|
||||
|
||||
Additional prompting, review, Meeting Map, hierarchical, diarized, dense-model
|
||||
and 70B-scale variants did not produce a reliable step change. Some improved
|
||||
individual dimensions, but none reached a stable B result or removed the need
|
||||
for substantive review. This is a current evidence-based product choice, not a
|
||||
permanent architecture decision.
|
||||
|
||||
## Experiments Compared
|
||||
|
||||
All protocol rows used the complete cleaned transcript and meeting context
|
||||
unless stated otherwise. A dash means that the artifact did not record the
|
||||
number; it is not an estimate. Verdicts marked “assessment” are comparative
|
||||
assessments of saved outputs because those older artifact directories contain
|
||||
no formal `quality_review.json`.
|
||||
|
||||
| Experiment | Model | Architecture | LLM calls | Prompt tokens | Runtime | Human editing | Verdict | Main strength | Main failure | MVP | Research |
|
||||
| --- | --- | --- | ---: | ---: | ---: | --- | --- | --- | --- | --- | --- |
|
||||
| Direct one-shot baseline | Qwen3.6 35B-A3B | Direct full-context protocol | 1 | 18,385 | 72.5 s cold; about 38 s inference | Not recorded | C | Fast, readable, broad topic outline | Consensus and ownership overpromotion; missing boundaries and follow-up | **Yes, with review** | Baseline |
|
||||
| Conservative one-shot | Qwen3.6 35B-A3B | Direct with stronger safety instructions | 1 | 19,120 | 39.5 s warm | Not recorded | C (assessment) | Better uncertainty and pending-feedback language | Still invents or upgrades named follow-up actions | No | Limited |
|
||||
| Draft → review | Qwen3.6 35B-A3B | Conservative draft plus review call | 2 | 39,071 total | About 110.4 s summed | Not recorded | C (assessment) | Removes some unsafe named attribution | Does not reliably restore omitted content; empty/weak action sections remain | No | Limited |
|
||||
| Meeting Map → protocol | Qwen3.6 35B-A3B | Semantic map followed by rendering | 2 | 39,129 total | 134.8 s | Not recorded | C (assessment) | Explicit intermediate structure | Map errors propagate: false consensus and named ownership remain | No | Yes |
|
||||
| Hierarchical notes → protocol | Qwen3.6 35B-A3B | Five chunk-note calls plus synthesis | 6 | 32,145 total | 251.1 s | Not recorded | C (assessment) | Highest recall in several detailed/open topics | Amplifies unsupported speaker/name interpretations and confirmed actions | No | Yes |
|
||||
| Segment-level anonymous diarization | Qwen3.6 35B-A3B | One-shot over 1,154 labeled segments | 1 | 39,308 | 125.4 s | 25–35 min | C | Preserves some filtered-idea challenge structure | Token count more than doubled; actions and deadlines became less safe | No | No further protocol tests |
|
||||
| Turn-merged anonymous diarization | Qwen3.6 35B-A3B | One-shot over 435 merged turns | 1 | 22,490 | 101.2 s | 25–35 min | C | Corrected token inflation; recovered some topic and feedback detail | Still did not beat raw input; unsafe actions/deadlines persisted | No | UI/search only |
|
||||
| Dense one-shot | Qwen3.5 27B | Direct full-context protocol | 1 | 18,385 | 181.2 s | Not recorded | C (assessment) | Somewhat better recall of process details | Much slower; no material overall quality gain | No | No |
|
||||
| 70B scale one-shot | Llama 3.3 70B Q3_K_S | Dense, 64% CPU / 36% GPU | 1 | 21,156 | 470.1 s warm | 45–60 min | D | Technically proved a 70B hybrid load can run | Severe coverage loss, invented governance, internal contradiction | No | Negative scale result |
|
||||
| Ollama vs native llama.cpp | Qwen3.6 35B-A3B | Same Q4_K_M GGUF; ROCm/Vulkan servers | 1 per backend | 18,385 | ROCm 38.4 s; Vulkan 42.3 s | N/A | Runtime only | Native ROCm reached 51.38 generated tok/s | No meaningful end-to-end advantage; more operational complexity | Ollama | Runtime reference |
|
||||
|
||||
The draft-review total combines the saved conservative draft call and the
|
||||
saved review call. Its review metadata itself reports only the one new review
|
||||
call (19,951 prompt tokens and 70.8 seconds). The direct baseline's 72.5-second
|
||||
wall time includes a 34.5-second cold load; its measured prompt evaluation plus
|
||||
generation was 37.8 seconds. These distinctions explain apparent runtime
|
||||
differences between otherwise similar Qwen3.6 calls.
|
||||
|
||||
### Recurring quality patterns
|
||||
|
||||
- **Topic coverage and factual accuracy:** Direct Qwen3.6 captures the main
|
||||
process but misses the second review after enrichment, project reporting and
|
||||
parts of the filtered-idea challenge path. Hierarchical processing recalls
|
||||
more detail but introduces too many unsupported interpretations. Llama 3.3
|
||||
loses most of the meeting and invents a governance role for the
|
||||
Geschäftsführung.
|
||||
- **Consensus and unresolved boundaries:** Every broad one-shot family remains
|
||||
vulnerable to turning discussion or a working direction into agreement. The
|
||||
unresolved boundary between central coordination and autonomous department
|
||||
work, and the uncertainty around universal filter criteria, are especially
|
||||
fragile.
|
||||
- **Visibility, veto and reconsideration:** No approach consistently preserves
|
||||
initial filtering, later cross-functional challenge, reconsideration after
|
||||
enrichment and the return through the project cycle together.
|
||||
- **Stakeholder feedback:** Pending Jovana and Björn feedback is an important
|
||||
quality probe. Some variants preserve both; segment-level diarization drops
|
||||
Björn, while Llama 3.3 drops both.
|
||||
- **Actions and attribution:** Added structure does not guarantee safety.
|
||||
Conservative, reviewed, Meeting Map, hierarchical and diarized outputs still
|
||||
promote proposals or expected work into confirmed actions, infer owners from
|
||||
roles or conversational context, or invent deadlines. Human review remains
|
||||
mandatory.
|
||||
|
||||
## Model Findings
|
||||
|
||||
### Qwen3.6:35B-A3B
|
||||
|
||||
Qwen3.6 is the best overall local practical baseline. Its Q4_K_M model is
|
||||
operationally fast on the RX 9070/CPU hybrid setup, follows the requested
|
||||
Markdown form and usually provides a useful first draft. It remains verdict C:
|
||||
larger context and fluent synthesis do not reliably protect evidence strength,
|
||||
responsibility attribution or unresolved process boundaries.
|
||||
|
||||
### Qwen3.5:27b dense
|
||||
|
||||
The dense 27B run recalled some process details better than the MoE baseline,
|
||||
but took 181.2 seconds and generated at 8.45 tokens/s. The gains did not amount
|
||||
to a material overall quality improvement. This result does not prove that
|
||||
dense models are generally inferior; it shows that this dense model is not a
|
||||
better product choice on this hardware and meeting.
|
||||
|
||||
### Llama 3.3 70B Q3_K_S
|
||||
|
||||
Llama 3.3 70B was technically runnable at 32k context with a 42 GB loaded
|
||||
footprint and a 64% CPU / 36% GPU split. Its 7m50s warm meeting run produced a
|
||||
very short, materially worse protocol: one critical invented governance claim,
|
||||
five new major errors and an estimated 45–60 minutes of editing. Raw parameter
|
||||
count alone is therefore insufficient. The older model generation and
|
||||
aggressive Q3 quantization are plausible contributors, but this experiment
|
||||
does not isolate or prove either cause.
|
||||
|
||||
## Runtime Findings
|
||||
|
||||
The native comparison reused the exact 23,938,321,664-byte Qwen3.6 Q4_K_M GGUF
|
||||
that Ollama uses. The tested `llama-server` binary was the llama.cpp runtime
|
||||
shipped with the installed Ollama distribution, not an independent source
|
||||
build.
|
||||
|
||||
Native ROCm processed the reference request in 38.4 seconds and generated at
|
||||
51.38 tokens/s. Vulkan took 42.3 seconds and generated at 46.73 tokens/s. The
|
||||
comparable Ollama baseline generated at 44.68 tokens/s, with about 37.8 seconds
|
||||
of prompt evaluation plus generation when load time is excluded. Output token
|
||||
counts differed, so generation throughput alone is not an end-to-end quality or
|
||||
latency comparison.
|
||||
|
||||
Native ROCm gained some generation throughput, but did not provide a meaningful
|
||||
end-to-end advantage for this workload. Vulkan required more host spill and was
|
||||
not preferable. Ollama already provides the relevant llama.cpp runtime
|
||||
components, model lifecycle and API integration; it remains the preferred
|
||||
routine Meeting Lab runtime.
|
||||
|
||||
## Diarization Findings
|
||||
|
||||
### Technical feasibility
|
||||
|
||||
Pyannote `speaker-diarization-community-1` successfully processed the
|
||||
93.5-minute meeting on CPU in 1,647 seconds (about 27m27s), an RTF of 0.293. It
|
||||
detected four anonymous clusters and assigned 1,145 of 1,154 Whisper segments
|
||||
(99.2%). Peak RSS was about 3.3 GiB. This establishes technical feasibility; it
|
||||
does not establish speaker identity or diarization accuracy against labeled
|
||||
ground truth.
|
||||
|
||||
### Protocol-quality impact
|
||||
|
||||
Annotating every Whisper segment increased the Qwen prompt from 18,385 to
|
||||
39,308 tokens, confounding speaker structure with fragmentation and token
|
||||
inflation. Deterministic turn merging reduced 1,154 segments to 435 turns and
|
||||
the complete prompt to 22,490 tokens. That controlled the main representation
|
||||
confound, but the resulting protocol still did not materially outperform the
|
||||
raw transcript and remained verdict C.
|
||||
|
||||
Anonymous diarization is therefore not justified as a mandatory MVP
|
||||
protocol-quality feature. This does **not** mean diarization is generally
|
||||
useless. It may remain valuable for speaker-aware UI, navigation and search,
|
||||
participation statistics, traceability, or later carefully validated real-name
|
||||
mapping.
|
||||
|
||||
## Semantic Research Findings
|
||||
|
||||
The semantic experiments provide architectural evidence, but should not
|
||||
dominate the product decision:
|
||||
|
||||
- **Evidence Observation V3** is a strong evidence-near candidate stage. With
|
||||
Qwen3.5 9B it achieved 8 PASS, 1 PARTIAL and 0 FAIL while preserving hedges,
|
||||
alternatives, requests, commitments and boundaries in natural language.
|
||||
- **Request/Acceptance** and **Collective Commitment** show that narrow semantic
|
||||
recognition followed by deterministic provenance, ordering, addressee,
|
||||
negation and deadline gates can safely derive limited consequences. Model
|
||||
recognition errors were contained without inventing individual ownership.
|
||||
- **Explicit Rejection** failed when reduced to a coarse binary recognition
|
||||
problem: semantically valid false positives passed structural gates.
|
||||
- **Negative Act Form** worked better by distinguishing non-pursuit, personal
|
||||
preference, recommendation and temporary non-action before any normative
|
||||
derivation. All eight form classifications matched Gold, although normalized
|
||||
action text was imperfect in three cases.
|
||||
- **Target Resolution V0** failed because prompt examples and a weak JSON
|
||||
boundary encouraged the string `"null"` instead of typed linkage.
|
||||
**Target Resolution V1** fixed all linkage/ID failures with deterministic
|
||||
self-linkage, closed ID lists and true JSON Schema, but normalization remained
|
||||
incomplete.
|
||||
- **Target Normalization V0** improved polarity and scope preservation to 3/4
|
||||
PASS, but still lost continuation meaning in the collaboration case.
|
||||
|
||||
These findings support Meeting Lab as a research and validation track. They do
|
||||
not yet justify placing a multi-stage semantic pipeline on the MVP critical
|
||||
path.
|
||||
|
||||
## Current MVP Decision
|
||||
|
||||
The current product path is:
|
||||
|
||||
```text
|
||||
Audio
|
||||
-> transcription
|
||||
-> direct qwen3.6:35B-A3B protocol generation through Ollama
|
||||
-> informed human review
|
||||
-> final protocol
|
||||
```
|
||||
|
||||
The first MVP should treat the generated protocol as an editable draft, not an
|
||||
authoritative semantic record. Human review must specifically check consensus,
|
||||
unresolved boundaries, competing positions, action status, owners, deadlines
|
||||
and pending stakeholder feedback.
|
||||
|
||||
Diarization is optional and deferred. The semantic research pipeline remains
|
||||
in Meeting Lab, outside the MVP critical path. Ollama remains the default local
|
||||
runtime.
|
||||
|
||||
## Rejected / Deferred Directions
|
||||
|
||||
- Do not continue Qwen3.6 prompt variants as the main quality strategy.
|
||||
- Do not add draft-review, Meeting Map or hierarchical generation to the MVP;
|
||||
their added calls and complexity did not deliver reliable quality gains.
|
||||
- Do not continue anonymous-diarization protocol experiments. Revisit
|
||||
diarization for UI, search, statistics or traceability instead.
|
||||
- Do not use Qwen3.5 27B or Llama 3.3 70B Q3_K_S as the routine protocol model.
|
||||
- Do not replace Ollama with a manually managed native llama.cpp service for
|
||||
this workload.
|
||||
- Retain semantic experiments, but defer production integration and broad
|
||||
semantic consolidation.
|
||||
|
||||
## Open Questions
|
||||
|
||||
1. How does a genuinely newer, materially stronger model perform when a useful
|
||||
quantization fits the available RAM/VRAM without severe swap?
|
||||
2. If project policy permits, what quality ceiling does the unchanged reference
|
||||
prompt achieve with a commercial frontier model?
|
||||
3. What is the measured reviewer time and correction distribution once the
|
||||
direct Qwen3.6 draft path is exercised in an end-to-end MVP workflow?
|
||||
4. Which non-protocol product benefits justify revisiting diarization later?
|
||||
|
||||
No further Qwen3.6 prompt variants, anonymous-diarization protocol runs or old
|
||||
70B Q3 scale tests are recommended.
|
||||
|
||||
## Recommended Next Product Step
|
||||
|
||||
Build the practical end-to-end MVP around direct Qwen3.6 generation and an
|
||||
explicit human review handoff. Measure reviewer time and correction categories
|
||||
in real use. Keep the experiment artifacts and semantic Gold work as validation
|
||||
evidence, but do not block the first product loop on broader research stages.
|
||||
Reference in New Issue
Block a user