Add controlled Whisper and LLM benchmark results
Document and preserve the controlled Progeo benchmark series. Whisper comparison: - Compare whisper.cpp large-v3 and large-v3-turbo - Identify and verify BUG-012: extraction stability depends on chunk size - Repeat both transcription variants with identical reduced chunk budgets - Select large-v3-turbo as the current production transcription model LLM comparison: - Compare qwen3.5:9b with qwen3.5:35B-A3B - Preserve identical transcript, Meeting Context, prompts and chunking - Retain qwen3.5:9b as the production recommendation - Record runtime, repair burden and semantic-quality findings Current production benchmark configuration: - whisper.cpp large-v3-turbo - target_chars=4500 - max_chars=5500 - min_chars=2500 - overlap_blocks=0 - qwen3.5:9b - num_ctx=32768 - think=false - temperature=0 BUG-012 remains verified but not yet fixed in the production chunker.
This commit is contained in:
@@ -839,3 +839,95 @@ Current status:
|
||||
Verified. The renderer contract is now enforced deterministically. This does
|
||||
not improve semantic quality of the generated prose; it prevents invalid
|
||||
renderer output from being accepted as a final Working Protocol.
|
||||
|
||||
## BUG-012
|
||||
|
||||
ID: BUG-012
|
||||
|
||||
Title: Chunking sensitivity to Whisper segmentation density
|
||||
|
||||
Pipeline stage: Chunking / Extraction
|
||||
|
||||
Severity: High
|
||||
|
||||
Status: Verified hypothesis
|
||||
|
||||
Date discovered: 2026-08-05
|
||||
|
||||
Version first observed: `progeo_whisper_turbo_20260805_120903`
|
||||
|
||||
Description:
|
||||
|
||||
Experiment 1 compared two valid trimmed whisper.cpp transcriptions while
|
||||
keeping Meeting Context, source code, prompts, LLM, model parameters, pipeline
|
||||
stages, validation/repair behavior and renderer configuration identical.
|
||||
|
||||
The large-v3 transcription produced 3,822 input segments, 216,500 segment-text
|
||||
characters and 25 chunks. It completed extraction, canonicalization, Semantic
|
||||
Consolidator V0 and renderer invocation under the identical downstream
|
||||
settings.
|
||||
|
||||
The large-v3-turbo transcription produced 1,211 input segments, 83,473
|
||||
segment-text characters and 10 chunks. Extraction failed at `chunk_04`. The
|
||||
preserved raw response repeated the same fact many times and truncated
|
||||
mid-string, producing invalid JSON. Because the extraction output for
|
||||
`chunk_04` was invalid, canonicalization, semantic consolidation and rendering
|
||||
could not proceed.
|
||||
|
||||
Expected behaviour:
|
||||
|
||||
Chunk construction should provide stable extraction inputs across reasonable
|
||||
transcription segment-density differences. A valid, shorter Turbo transcript
|
||||
should not fail extraction solely because its segment boundaries create
|
||||
different chunk shapes.
|
||||
|
||||
Actual behaviour:
|
||||
|
||||
The pipeline was materially more compatible with the large-v3 transcript than
|
||||
with the Turbo transcript. The Turbo run failed before downstream stages even
|
||||
though the downstream settings were unchanged.
|
||||
|
||||
Important interpretation:
|
||||
|
||||
This does not prove that whisper.cpp large-v3 is intrinsically better than
|
||||
large-v3-turbo. It proves only that large-v3 was more compatible with the
|
||||
current chunking/extraction path in this controlled run.
|
||||
|
||||
Suspected root cause:
|
||||
|
||||
The current chunker treats Whisper JSON segments as atomic blocks and builds
|
||||
chunks by character budget around those source segments. Different Whisper
|
||||
segment density therefore changes the effective extraction input shape. The
|
||||
suspected failure mode is chunk/token budgeting and extraction stability, not
|
||||
merely transcription readability or Whisper quality.
|
||||
|
||||
Related files:
|
||||
|
||||
- `samples/real_live/progeo_meeting/progeo_whispercpp_vulkan_large_v3_trimmed_converted.json`
|
||||
- `samples/real_live/progeo_meeting/progeo_whispercpp_vulkan_turbo_trimmed_converted.json`
|
||||
- `samples/benchmarks/progeo_whisper_large_v3_20260805_120903/`
|
||||
- `samples/benchmarks/progeo_whisper_turbo_20260805_120903/`
|
||||
- `samples/benchmarks/progeo_whisper_turbo_small_chunks_20260805_124324/`
|
||||
- `samples/benchmarks/progeo_whisper_turbo_20260805_120903/extractions/chunk_04_extraction.raw.txt`
|
||||
- `samples/benchmarks/progeo_whisper_comparison_20260805_120903_metrics.json`
|
||||
- `src/meeting_lab/chunking/chunk_transcript.py`
|
||||
|
||||
Regression test available (yes/no): no
|
||||
|
||||
Current status:
|
||||
|
||||
Verified hypothesis. Experiment 1B reran the same Turbo transcription through
|
||||
the same downstream settings while changing only the chunk budget from the
|
||||
default `target_chars=9000`, `max_chars=11000`, `min_chars=5000`,
|
||||
`overlap_blocks=0` to `target_chars=4500`, `max_chars=5500`,
|
||||
`min_chars=2500`, `overlap_blocks=0`.
|
||||
|
||||
The reduced-size Turbo run produced 19 chunks instead of 10. All extraction
|
||||
chunks produced valid JSON, canonicalization completed, Semantic Consolidator
|
||||
V0 completed with existing deterministic source-coverage repair, and the
|
||||
renderer was invoked. Renderer contract validation still failed because the
|
||||
candidate did not start with `# Working Protocol`; that remains covered by
|
||||
BUG-011 and does not invalidate this chunking hypothesis.
|
||||
|
||||
This verifies the hypothesis for this transcript/configuration only. It does
|
||||
not claim the general chunker is fixed.
|
||||
|
||||
Reference in New Issue
Block a user