Add controlled Whisper and LLM benchmark results

Document and preserve the controlled Progeo benchmark series.

Whisper comparison:
- Compare whisper.cpp large-v3 and large-v3-turbo
- Identify and verify BUG-012: extraction stability depends on chunk size
- Repeat both transcription variants with identical reduced chunk budgets
- Select large-v3-turbo as the current production transcription model

LLM comparison:
- Compare qwen3.5:9b with qwen3.5:35B-A3B
- Preserve identical transcript, Meeting Context, prompts and chunking
- Retain qwen3.5:9b as the production recommendation
- Record runtime, repair burden and semantic-quality findings

Current production benchmark configuration:
- whisper.cpp large-v3-turbo
- target_chars=4500
- max_chars=5500
- min_chars=2500
- overlap_blocks=0
- qwen3.5:9b
- num_ctx=32768
- think=false
- temperature=0

BUG-012 remains verified but not yet fixed in the production chunker.
This commit is contained in:
2026-08-05 14:19:40 +02:00
parent dcc3a5c734
commit 03bc6b1d90
544 changed files with 464774 additions and 0 deletions
+92
View File
@@ -839,3 +839,95 @@ Current status:
Verified. The renderer contract is now enforced deterministically. This does
not improve semantic quality of the generated prose; it prevents invalid
renderer output from being accepted as a final Working Protocol.
## BUG-012
ID: BUG-012
Title: Chunking sensitivity to Whisper segmentation density
Pipeline stage: Chunking / Extraction
Severity: High
Status: Verified hypothesis
Date discovered: 2026-08-05
Version first observed: `progeo_whisper_turbo_20260805_120903`
Description:
Experiment 1 compared two valid trimmed whisper.cpp transcriptions while
keeping Meeting Context, source code, prompts, LLM, model parameters, pipeline
stages, validation/repair behavior and renderer configuration identical.
The large-v3 transcription produced 3,822 input segments, 216,500 segment-text
characters and 25 chunks. It completed extraction, canonicalization, Semantic
Consolidator V0 and renderer invocation under the identical downstream
settings.
The large-v3-turbo transcription produced 1,211 input segments, 83,473
segment-text characters and 10 chunks. Extraction failed at `chunk_04`. The
preserved raw response repeated the same fact many times and truncated
mid-string, producing invalid JSON. Because the extraction output for
`chunk_04` was invalid, canonicalization, semantic consolidation and rendering
could not proceed.
Expected behaviour:
Chunk construction should provide stable extraction inputs across reasonable
transcription segment-density differences. A valid, shorter Turbo transcript
should not fail extraction solely because its segment boundaries create
different chunk shapes.
Actual behaviour:
The pipeline was materially more compatible with the large-v3 transcript than
with the Turbo transcript. The Turbo run failed before downstream stages even
though the downstream settings were unchanged.
Important interpretation:
This does not prove that whisper.cpp large-v3 is intrinsically better than
large-v3-turbo. It proves only that large-v3 was more compatible with the
current chunking/extraction path in this controlled run.
Suspected root cause:
The current chunker treats Whisper JSON segments as atomic blocks and builds
chunks by character budget around those source segments. Different Whisper
segment density therefore changes the effective extraction input shape. The
suspected failure mode is chunk/token budgeting and extraction stability, not
merely transcription readability or Whisper quality.
Related files:
- `samples/real_live/progeo_meeting/progeo_whispercpp_vulkan_large_v3_trimmed_converted.json`
- `samples/real_live/progeo_meeting/progeo_whispercpp_vulkan_turbo_trimmed_converted.json`
- `samples/benchmarks/progeo_whisper_large_v3_20260805_120903/`
- `samples/benchmarks/progeo_whisper_turbo_20260805_120903/`
- `samples/benchmarks/progeo_whisper_turbo_small_chunks_20260805_124324/`
- `samples/benchmarks/progeo_whisper_turbo_20260805_120903/extractions/chunk_04_extraction.raw.txt`
- `samples/benchmarks/progeo_whisper_comparison_20260805_120903_metrics.json`
- `src/meeting_lab/chunking/chunk_transcript.py`
Regression test available (yes/no): no
Current status:
Verified hypothesis. Experiment 1B reran the same Turbo transcription through
the same downstream settings while changing only the chunk budget from the
default `target_chars=9000`, `max_chars=11000`, `min_chars=5000`,
`overlap_blocks=0` to `target_chars=4500`, `max_chars=5500`,
`min_chars=2500`, `overlap_blocks=0`.
The reduced-size Turbo run produced 19 chunks instead of 10. All extraction
chunks produced valid JSON, canonicalization completed, Semantic Consolidator
V0 completed with existing deterministic source-coverage repair, and the
renderer was invoked. Renderer contract validation still failed because the
candidate did not start with `# Working Protocol`; that remains covered by
BUG-011 and does not invalidate this chunking hypothesis.
This verifies the hypothesis for this transcript/configuration only. It does
not claim the general chunker is fixed.