# Experiment 2 LLM Comparison Report ## Benchmark Directories - qwen3.5:9b: `samples\benchmarks\progeo_qwen35_9b_20260805_133942` - qwen3.5:35B-A3B: `samples\benchmarks\progeo_qwen35_35b_a3b_20260805_133942` ## Configuration - Input: `samples/real_live/progeo_meeting/progeo_whispercpp_vulkan_turbo_trimmed_converted.json` - Meeting Context: `samples/real_live/progeo_meeting/meeting_context.yaml` - Chunking: `target_chars=4500`, `max_chars=5500`, `min_chars=2500`, `overlap_blocks=0` - Model parameters: `think=false`, `temperature=0`, `num_ctx=32768`; committed generation limits and adaptive consolidator sizing; no retries; no manual intervention. ## Runtime | Stage | qwen3.5:9b | qwen3.5:35B-A3B | |---|---:|---:| | chunking | 0.271 | 0.276 | | normalization | 1.366 | 1.344 | | extraction | 389.75 | 953.042 | | canonicalizer | 0.139 | 0.316 | | semantic_consolidator | 103.788 | 119.612 | | renderer | 64.85 | 87.72 | - Total qwen3.5:9b: 560.164 seconds - Total qwen3.5:35B-A3B: 1162.31 seconds ## Structural Counts - qwen3.5:9b extraction: {'facts': 106, 'decisions': 10, 'todos': 40, 'questions': 26, 'positions': 0, 'technical': 63} - qwen3.5:35B-A3B extraction: {'facts': 125, 'decisions': 9, 'todos': 36, 'questions': 37, 'positions': 0, 'technical': 79} - qwen3.5:9b canonical: {'fact': 106, 'decision': 10, 'action_item': 35, 'open_question': 26, 'position': 0, 'technical_detail': 63} - qwen3.5:35B-A3B canonical: {'fact': 125, 'decision': 9, 'action_item': 36, 'open_question': 37, 'position': 0, 'technical_detail': 79} ## Validation And Repair - qwen3.5:9b: validator before valid=False, violations=11; repair_count=14; validator after valid=True. - qwen3.5:35B-A3B: validator before valid=False, violations=78; repair_count=78; validator after valid=True. ## Renderer - qwen3.5:9b: valid=False; done_reason=length; eval_count=4096; response_text_length=16429. - qwen3.5:35B-A3B: valid=False; done_reason=stop; eval_count=1861; response_text_length=7594. ## Result - Selected winner: `qwen3.5:9b WINS` - Confidence: medium - Engineering recommendation: do not switch production to `qwen3.5:35B-A3B` on this evidence. It was slower, required much more deterministic repair, still failed the renderer contract, and did not reduce core extraction overclassification enough to justify the runtime cost.