Document and preserve the controlled Progeo benchmark series.
Whisper comparison:
- Compare whisper.cpp large-v3 and large-v3-turbo
- Identify and verify BUG-012: extraction stability depends on chunk size
- Repeat both transcription variants with identical reduced chunk budgets
- Select large-v3-turbo as the current production transcription model
LLM comparison:
- Compare qwen3.5:9b with qwen3.5:35B-A3B
- Preserve identical transcript, Meeting Context, prompts and chunking
- Retain qwen3.5:9b as the production recommendation
- Record runtime, repair burden and semantic-quality findings
Current production benchmark configuration:
- whisper.cpp large-v3-turbo
- target_chars=4500
- max_chars=5500
- min_chars=2500
- overlap_blocks=0
- qwen3.5:9b
- num_ctx=32768
- think=false
- temperature=0
BUG-012 remains verified but not yet fixed in the production chunker.
Add benchmark artifacts created during the first RC1 robustness evaluation of the Meeting Lab pipeline.
Included benchmark sets:
- hardware_experimental_qwen35_9b
- meeting_context_v1
- progeo_meeting_20260804_083849
- progeo_meeting_context_v1_20260804_110913
- progeo_meeting_rc1_20260804_121502
These artifacts document the evolution of the pipeline during the implementation
and verification of BUG-009, BUG-010 and BUG-011.
The benchmark data provide reproducible real-life regression cases for future
development and allow quality comparisons across pipeline revisions.
Current benchmark policy:
During the active development phase, representative benchmark artifacts are
intentionally versioned to preserve reproducibility and simplify regression
analysis.
Benchmark artifacts are considered part of the engineering evidence rather than
temporary build output. Benchmark retention strategy will be revisited once the
Meeting Assistant reaches production maturity.
Introduce Meeting Context V1 with YAML schema, validation and template.
Support optional --meeting-context during chunk extraction.
Inject authoritative Meeting Context into extraction prompts.
Record Meeting Context provenance in extraction output.
Activate todos.md in shared prompt assembly.
Strengthen responsibility attribution and decision/todo boundaries.
Add focused Gold scenarios and validation tests.
Update architecture and pipeline documentation.
- add deterministic canonicalization support for extraction items
- add facts-only semantic consolidation using local Ollama
- preserve source evidence and validate complete fact coverage
- add conservative merge rules and non-LLM tests
- record the first validated real-life consolidation benchmark
- document current scope, limitations and next evaluation step