Add controlled Whisper and LLM benchmark results
Document and preserve the controlled Progeo benchmark series. Whisper comparison: - Compare whisper.cpp large-v3 and large-v3-turbo - Identify and verify BUG-012: extraction stability depends on chunk size - Repeat both transcription variants with identical reduced chunk budgets - Select large-v3-turbo as the current production transcription model LLM comparison: - Compare qwen3.5:9b with qwen3.5:35B-A3B - Preserve identical transcript, Meeting Context, prompts and chunking - Retain qwen3.5:9b as the production recommendation - Record runtime, repair burden and semantic-quality findings Current production benchmark configuration: - whisper.cpp large-v3-turbo - target_chars=4500 - max_chars=5500 - min_chars=2500 - overlap_blocks=0 - qwen3.5:9b - num_ctx=32768 - think=false - temperature=0 BUG-012 remains verified but not yet fixed in the production chunker.
This commit is contained in:
@@ -0,0 +1,17 @@
|
||||
{
|
||||
"normalization": {
|
||||
"exit_code": 0,
|
||||
"log": "samples\\benchmarks\\progeo_whisper_turbo_20260805_120903\\normalization_stdout_stderr.txt",
|
||||
"runtime_seconds": 0.788
|
||||
},
|
||||
"chunking": {
|
||||
"exit_code": 0,
|
||||
"log": "samples\\benchmarks\\progeo_whisper_turbo_20260805_120903\\chunking_stdout_stderr.txt",
|
||||
"runtime_seconds": 0.269
|
||||
},
|
||||
"extraction": {
|
||||
"runtime_seconds": 198.349,
|
||||
"exit_code": 1,
|
||||
"log": "samples\\benchmarks\\progeo_whisper_turbo_20260805_120903\\extraction_stdout_stderr.txt"
|
||||
}
|
||||
}
|
||||
Reference in New Issue
Block a user