Add controlled Whisper and LLM benchmark results

Document and preserve the controlled Progeo benchmark series.

Whisper comparison:
- Compare whisper.cpp large-v3 and large-v3-turbo
- Identify and verify BUG-012: extraction stability depends on chunk size
- Repeat both transcription variants with identical reduced chunk budgets
- Select large-v3-turbo as the current production transcription model

LLM comparison:
- Compare qwen3.5:9b with qwen3.5:35B-A3B
- Preserve identical transcript, Meeting Context, prompts and chunking
- Retain qwen3.5:9b as the production recommendation
- Record runtime, repair burden and semantic-quality findings

Current production benchmark configuration:
- whisper.cpp large-v3-turbo
- target_chars=4500
- max_chars=5500
- min_chars=2500
- overlap_blocks=0
- qwen3.5:9b
- num_ctx=32768
- think=false
- temperature=0

BUG-012 remains verified but not yet fixed in the production chunker.
This commit is contained in:
2026-08-05 14:19:40 +02:00
parent dcc3a5c734
commit 03bc6b1d90
544 changed files with 464774 additions and 0 deletions
@@ -0,0 +1,50 @@
# Experiment 2 LLM Comparison Report
## Benchmark Directories
- qwen3.5:9b: `samples\benchmarks\progeo_qwen35_9b_20260805_133942`
- qwen3.5:35B-A3B: `samples\benchmarks\progeo_qwen35_35b_a3b_20260805_133942`
## Configuration
- Input: `samples/real_live/progeo_meeting/progeo_whispercpp_vulkan_turbo_trimmed_converted.json`
- Meeting Context: `samples/real_live/progeo_meeting/meeting_context.yaml`
- Chunking: `target_chars=4500`, `max_chars=5500`, `min_chars=2500`, `overlap_blocks=0`
- Model parameters: `think=false`, `temperature=0`, `num_ctx=32768`; committed generation limits and adaptive consolidator sizing; no retries; no manual intervention.
## Runtime
| Stage | qwen3.5:9b | qwen3.5:35B-A3B |
|---|---:|---:|
| chunking | 0.271 | 0.276 |
| normalization | 1.366 | 1.344 |
| extraction | 389.75 | 953.042 |
| canonicalizer | 0.139 | 0.316 |
| semantic_consolidator | 103.788 | 119.612 |
| renderer | 64.85 | 87.72 |
- Total qwen3.5:9b: 560.164 seconds
- Total qwen3.5:35B-A3B: 1162.31 seconds
## Structural Counts
- qwen3.5:9b extraction: {'facts': 106, 'decisions': 10, 'todos': 40, 'questions': 26, 'positions': 0, 'technical': 63}
- qwen3.5:35B-A3B extraction: {'facts': 125, 'decisions': 9, 'todos': 36, 'questions': 37, 'positions': 0, 'technical': 79}
- qwen3.5:9b canonical: {'fact': 106, 'decision': 10, 'action_item': 35, 'open_question': 26, 'position': 0, 'technical_detail': 63}
- qwen3.5:35B-A3B canonical: {'fact': 125, 'decision': 9, 'action_item': 36, 'open_question': 37, 'position': 0, 'technical_detail': 79}
## Validation And Repair
- qwen3.5:9b: validator before valid=False, violations=11; repair_count=14; validator after valid=True.
- qwen3.5:35B-A3B: validator before valid=False, violations=78; repair_count=78; validator after valid=True.
## Renderer
- qwen3.5:9b: valid=False; done_reason=length; eval_count=4096; response_text_length=16429.
- qwen3.5:35B-A3B: valid=False; done_reason=stop; eval_count=1861; response_text_length=7594.
## Result
- Selected winner: `qwen3.5:9b WINS`
- Confidence: medium
- Engineering recommendation: do not switch production to `qwen3.5:35B-A3B` on this evidence. It was slower, required much more deterministic repair, still failed the renderer contract, and did not reduce core extraction overclassification enough to justify the runtime cost.