Add controlled Whisper and LLM benchmark results
Document and preserve the controlled Progeo benchmark series. Whisper comparison: - Compare whisper.cpp large-v3 and large-v3-turbo - Identify and verify BUG-012: extraction stability depends on chunk size - Repeat both transcription variants with identical reduced chunk budgets - Select large-v3-turbo as the current production transcription model LLM comparison: - Compare qwen3.5:9b with qwen3.5:35B-A3B - Preserve identical transcript, Meeting Context, prompts and chunking - Retain qwen3.5:9b as the production recommendation - Record runtime, repair burden and semantic-quality findings Current production benchmark configuration: - whisper.cpp large-v3-turbo - target_chars=4500 - max_chars=5500 - min_chars=2500 - overlap_blocks=0 - qwen3.5:9b - num_ctx=32768 - think=false - temperature=0 BUG-012 remains verified but not yet fixed in the production chunker.
This commit is contained in:
@@ -0,0 +1,50 @@
|
||||
# Experiment 2 LLM Comparison Report
|
||||
|
||||
## Benchmark Directories
|
||||
|
||||
- qwen3.5:9b: `samples\benchmarks\progeo_qwen35_9b_20260805_133942`
|
||||
- qwen3.5:35B-A3B: `samples\benchmarks\progeo_qwen35_35b_a3b_20260805_133942`
|
||||
|
||||
## Configuration
|
||||
|
||||
- Input: `samples/real_live/progeo_meeting/progeo_whispercpp_vulkan_turbo_trimmed_converted.json`
|
||||
- Meeting Context: `samples/real_live/progeo_meeting/meeting_context.yaml`
|
||||
- Chunking: `target_chars=4500`, `max_chars=5500`, `min_chars=2500`, `overlap_blocks=0`
|
||||
- Model parameters: `think=false`, `temperature=0`, `num_ctx=32768`; committed generation limits and adaptive consolidator sizing; no retries; no manual intervention.
|
||||
|
||||
## Runtime
|
||||
|
||||
| Stage | qwen3.5:9b | qwen3.5:35B-A3B |
|
||||
|---|---:|---:|
|
||||
| chunking | 0.271 | 0.276 |
|
||||
| normalization | 1.366 | 1.344 |
|
||||
| extraction | 389.75 | 953.042 |
|
||||
| canonicalizer | 0.139 | 0.316 |
|
||||
| semantic_consolidator | 103.788 | 119.612 |
|
||||
| renderer | 64.85 | 87.72 |
|
||||
|
||||
- Total qwen3.5:9b: 560.164 seconds
|
||||
- Total qwen3.5:35B-A3B: 1162.31 seconds
|
||||
|
||||
## Structural Counts
|
||||
|
||||
- qwen3.5:9b extraction: {'facts': 106, 'decisions': 10, 'todos': 40, 'questions': 26, 'positions': 0, 'technical': 63}
|
||||
- qwen3.5:35B-A3B extraction: {'facts': 125, 'decisions': 9, 'todos': 36, 'questions': 37, 'positions': 0, 'technical': 79}
|
||||
- qwen3.5:9b canonical: {'fact': 106, 'decision': 10, 'action_item': 35, 'open_question': 26, 'position': 0, 'technical_detail': 63}
|
||||
- qwen3.5:35B-A3B canonical: {'fact': 125, 'decision': 9, 'action_item': 36, 'open_question': 37, 'position': 0, 'technical_detail': 79}
|
||||
|
||||
## Validation And Repair
|
||||
|
||||
- qwen3.5:9b: validator before valid=False, violations=11; repair_count=14; validator after valid=True.
|
||||
- qwen3.5:35B-A3B: validator before valid=False, violations=78; repair_count=78; validator after valid=True.
|
||||
|
||||
## Renderer
|
||||
|
||||
- qwen3.5:9b: valid=False; done_reason=length; eval_count=4096; response_text_length=16429.
|
||||
- qwen3.5:35B-A3B: valid=False; done_reason=stop; eval_count=1861; response_text_length=7594.
|
||||
|
||||
## Result
|
||||
|
||||
- Selected winner: `qwen3.5:9b WINS`
|
||||
- Confidence: medium
|
||||
- Engineering recommendation: do not switch production to `qwen3.5:35B-A3B` on this evidence. It was slower, required much more deterministic repair, still failed the renderer contract, and did not reduce core extraction overclassification enough to justify the runtime cost.
|
||||
Reference in New Issue
Block a user