Files
meeting-lab/samples/benchmarks/progeo_llm_comparison_20260805_133942_report.md
admin 03bc6b1d90 Add controlled Whisper and LLM benchmark results
Document and preserve the controlled Progeo benchmark series.

Whisper comparison:
- Compare whisper.cpp large-v3 and large-v3-turbo
- Identify and verify BUG-012: extraction stability depends on chunk size
- Repeat both transcription variants with identical reduced chunk budgets
- Select large-v3-turbo as the current production transcription model

LLM comparison:
- Compare qwen3.5:9b with qwen3.5:35B-A3B
- Preserve identical transcript, Meeting Context, prompts and chunking
- Retain qwen3.5:9b as the production recommendation
- Record runtime, repair burden and semantic-quality findings

Current production benchmark configuration:
- whisper.cpp large-v3-turbo
- target_chars=4500
- max_chars=5500
- min_chars=2500
- overlap_blocks=0
- qwen3.5:9b
- num_ctx=32768
- think=false
- temperature=0

BUG-012 remains verified but not yet fixed in the production chunker.
2026-08-05 14:19:40 +02:00

2.3 KiB

Experiment 2 LLM Comparison Report

Benchmark Directories

  • qwen3.5:9b: samples\benchmarks\progeo_qwen35_9b_20260805_133942
  • qwen3.5:35B-A3B: samples\benchmarks\progeo_qwen35_35b_a3b_20260805_133942

Configuration

  • Input: samples/real_live/progeo_meeting/progeo_whispercpp_vulkan_turbo_trimmed_converted.json
  • Meeting Context: samples/real_live/progeo_meeting/meeting_context.yaml
  • Chunking: target_chars=4500, max_chars=5500, min_chars=2500, overlap_blocks=0
  • Model parameters: think=false, temperature=0, num_ctx=32768; committed generation limits and adaptive consolidator sizing; no retries; no manual intervention.

Runtime

Stage qwen3.5:9b qwen3.5:35B-A3B
chunking 0.271 0.276
normalization 1.366 1.344
extraction 389.75 953.042
canonicalizer 0.139 0.316
semantic_consolidator 103.788 119.612
renderer 64.85 87.72
  • Total qwen3.5:9b: 560.164 seconds
  • Total qwen3.5:35B-A3B: 1162.31 seconds

Structural Counts

  • qwen3.5:9b extraction: {'facts': 106, 'decisions': 10, 'todos': 40, 'questions': 26, 'positions': 0, 'technical': 63}
  • qwen3.5:35B-A3B extraction: {'facts': 125, 'decisions': 9, 'todos': 36, 'questions': 37, 'positions': 0, 'technical': 79}
  • qwen3.5:9b canonical: {'fact': 106, 'decision': 10, 'action_item': 35, 'open_question': 26, 'position': 0, 'technical_detail': 63}
  • qwen3.5:35B-A3B canonical: {'fact': 125, 'decision': 9, 'action_item': 36, 'open_question': 37, 'position': 0, 'technical_detail': 79}

Validation And Repair

  • qwen3.5:9b: validator before valid=False, violations=11; repair_count=14; validator after valid=True.
  • qwen3.5:35B-A3B: validator before valid=False, violations=78; repair_count=78; validator after valid=True.

Renderer

  • qwen3.5:9b: valid=False; done_reason=length; eval_count=4096; response_text_length=16429.
  • qwen3.5:35B-A3B: valid=False; done_reason=stop; eval_count=1861; response_text_length=7594.

Result

  • Selected winner: qwen3.5:9b WINS
  • Confidence: medium
  • Engineering recommendation: do not switch production to qwen3.5:35B-A3B on this evidence. It was slower, required much more deterministic repair, still failed the renderer contract, and did not reduce core extraction overclassification enough to justify the runtime cost.