Document and preserve the controlled Progeo benchmark series. Whisper comparison: - Compare whisper.cpp large-v3 and large-v3-turbo - Identify and verify BUG-012: extraction stability depends on chunk size - Repeat both transcription variants with identical reduced chunk budgets - Select large-v3-turbo as the current production transcription model LLM comparison: - Compare qwen3.5:9b with qwen3.5:35B-A3B - Preserve identical transcript, Meeting Context, prompts and chunking - Retain qwen3.5:9b as the production recommendation - Record runtime, repair burden and semantic-quality findings Current production benchmark configuration: - whisper.cpp large-v3-turbo - target_chars=4500 - max_chars=5500 - min_chars=2500 - overlap_blocks=0 - qwen3.5:9b - num_ctx=32768 - think=false - temperature=0 BUG-012 remains verified but not yet fixed in the production chunker.
2.3 KiB
2.3 KiB
Experiment 2 LLM Comparison Report
Benchmark Directories
- qwen3.5:9b:
samples\benchmarks\progeo_qwen35_9b_20260805_133942 - qwen3.5:35B-A3B:
samples\benchmarks\progeo_qwen35_35b_a3b_20260805_133942
Configuration
- Input:
samples/real_live/progeo_meeting/progeo_whispercpp_vulkan_turbo_trimmed_converted.json - Meeting Context:
samples/real_live/progeo_meeting/meeting_context.yaml - Chunking:
target_chars=4500,max_chars=5500,min_chars=2500,overlap_blocks=0 - Model parameters:
think=false,temperature=0,num_ctx=32768; committed generation limits and adaptive consolidator sizing; no retries; no manual intervention.
Runtime
| Stage | qwen3.5:9b | qwen3.5:35B-A3B |
|---|---|---|
| chunking | 0.271 | 0.276 |
| normalization | 1.366 | 1.344 |
| extraction | 389.75 | 953.042 |
| canonicalizer | 0.139 | 0.316 |
| semantic_consolidator | 103.788 | 119.612 |
| renderer | 64.85 | 87.72 |
- Total qwen3.5:9b: 560.164 seconds
- Total qwen3.5:35B-A3B: 1162.31 seconds
Structural Counts
- qwen3.5:9b extraction: {'facts': 106, 'decisions': 10, 'todos': 40, 'questions': 26, 'positions': 0, 'technical': 63}
- qwen3.5:35B-A3B extraction: {'facts': 125, 'decisions': 9, 'todos': 36, 'questions': 37, 'positions': 0, 'technical': 79}
- qwen3.5:9b canonical: {'fact': 106, 'decision': 10, 'action_item': 35, 'open_question': 26, 'position': 0, 'technical_detail': 63}
- qwen3.5:35B-A3B canonical: {'fact': 125, 'decision': 9, 'action_item': 36, 'open_question': 37, 'position': 0, 'technical_detail': 79}
Validation And Repair
- qwen3.5:9b: validator before valid=False, violations=11; repair_count=14; validator after valid=True.
- qwen3.5:35B-A3B: validator before valid=False, violations=78; repair_count=78; validator after valid=True.
Renderer
- qwen3.5:9b: valid=False; done_reason=length; eval_count=4096; response_text_length=16429.
- qwen3.5:35B-A3B: valid=False; done_reason=stop; eval_count=1861; response_text_length=7594.
Result
- Selected winner:
qwen3.5:9b WINS - Confidence: medium
- Engineering recommendation: do not switch production to
qwen3.5:35B-A3Bon this evidence. It was slower, required much more deterministic repair, still failed the renderer contract, and did not reduce core extraction overclassification enough to justify the runtime cost.