Document and preserve the controlled Progeo benchmark series. Whisper comparison: - Compare whisper.cpp large-v3 and large-v3-turbo - Identify and verify BUG-012: extraction stability depends on chunk size - Repeat both transcription variants with identical reduced chunk budgets - Select large-v3-turbo as the current production transcription model LLM comparison: - Compare qwen3.5:9b with qwen3.5:35B-A3B - Preserve identical transcript, Meeting Context, prompts and chunking - Retain qwen3.5:9b as the production recommendation - Record runtime, repair burden and semantic-quality findings Current production benchmark configuration: - whisper.cpp large-v3-turbo - target_chars=4500 - max_chars=5500 - min_chars=2500 - overlap_blocks=0 - qwen3.5:9b - num_ctx=32768 - think=false - temperature=0 BUG-012 remains verified but not yet fixed in the production chunker.
18 lines
529 B
JSON
18 lines
529 B
JSON
{
|
|
"normalization": {
|
|
"exit_code": 0,
|
|
"log": "samples\\benchmarks\\progeo_whisper_turbo_20260805_120903\\normalization_stdout_stderr.txt",
|
|
"runtime_seconds": 0.788
|
|
},
|
|
"chunking": {
|
|
"exit_code": 0,
|
|
"log": "samples\\benchmarks\\progeo_whisper_turbo_20260805_120903\\chunking_stdout_stderr.txt",
|
|
"runtime_seconds": 0.269
|
|
},
|
|
"extraction": {
|
|
"runtime_seconds": 198.349,
|
|
"exit_code": 1,
|
|
"log": "samples\\benchmarks\\progeo_whisper_turbo_20260805_120903\\extraction_stdout_stderr.txt"
|
|
}
|
|
}
|