Add a repository-native runner for reproducible Meeting Lab benchmark runs on machines without Codex. The runner: - validates normalized Whisper input and Meeting Context - checks Ollama availability and the requested model - rejects preloaded Ollama models by default for clean benchmarks - supports an explicit --allow-loaded-models override - uses the selected production configuration: - qwen3.5:9b - target_chars=4500 - max_chars=5500 - min_chars=2500 - overlap_blocks=0 - think=false - temperature=0 - num_ctx=32768 - executes the complete current pipeline - creates unique benchmark output directories - preserves artifacts up to failure - records runtime, environment and validation metadata - writes working_protocol.md only when the renderer contract passes Add focused mocked tests and Linux-first setup documentation for the AI-PC. The runner does not include Whisper execution.
4.2 KiB
Running A Meeting Benchmark
This page documents the repository-native runner for the current Meeting Lab production benchmark on the Linux AI-PC. It does not run Whisper. It starts from a converted, normalized Whisper JSON file.
Prerequisites
- Git
- Python 3.11 or newer
- Ollama running locally
- Required Ollama model installed
- Meeting Lab dependencies installed in a virtual environment
Setup
From the repository root on Linux:
python3 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -e .
Install the reference model in Ollama if it is not already present:
ollama pull qwen3.5:9b
Confirm Ollama sees the model:
ollama list
ollama ps
ollama ps must show no loaded model before a reference benchmark run. By
default, the runner fails pre-flight if Ollama reports a loaded model. Use
--allow-loaded-models only for a deliberately non-clean run; the default
reference command does not use that override.
Progeo Reference Command
Run from the repository root:
python scripts/run_meeting.py \
--input samples/real_live/progeo_meeting/progeo_whispercpp_vulkan_turbo_trimmed_converted.json \
--context samples/real_live/progeo_meeting/meeting_context.yaml \
--model qwen3.5:9b \
--benchmark-label progeo_ai_pc
Windows / PowerShell Example
For local Windows development, activate the virtual environment through PowerShell and run the same benchmark:
python -m venv .venv
.\.venv\Scripts\Activate.ps1
python -m pip install --upgrade pip
python -m pip install -e .
ollama pull qwen3.5:9b
python scripts\run_meeting.py `
--input samples\real_live\progeo_meeting\progeo_whispercpp_vulkan_turbo_trimmed_converted.json `
--context samples\real_live\progeo_meeting\meeting_context.yaml `
--model qwen3.5:9b `
--benchmark-label progeo_ai_pc
Production Defaults
The Progeo reference benchmark defaults match the latest controlled experiment selection:
- Chunking:
target_chars=4500,max_chars=5500,min_chars=2500,overlap_blocks=0 - LLM:
qwen3.5:9b,think=false,temperature=0,num_ctx=32768 - Extraction: context-aware chunk extraction with current committed prompts
- Consolidation: Canonicalizer V1, Semantic Consolidator V0 adaptive sizing, strict source-coverage validation and deterministic source-coverage repair
- Rendering: Working Protocol Renderer V2
- No semantic retries and no manual intervention
Use --help to see explicit overrides.
Output
The runner creates a unique directory:
samples/benchmarks/<benchmark-label>_<timestamp>/
It never overwrites an existing run. Artifacts include:
chunks/normalized_chunks/extractions/, including parsed extraction JSON and raw successful model responsescanonicalizer/canonicalized_extractions.jsonsemantic_consolidator/, including raw response, repaired groups, validator reports, repair metadata and consolidated outputworking_protocol/, including raw renderer response and validation reportworking_protocol.mdonly when the renderer contract is validrun_metadata.jsonbenchmark_report.md
If a stage fails, artifacts produced up to the failure are preserved and the report records the failure.
Comparing Benchmarks
For AI-PC comparison, compare the new progeo_ai_pc_<timestamp> directory
against:
samples/benchmarks/progeo_qwen35_9b_20260805_133942/
Key fields are in run_metadata.json and summarized in benchmark_report.md:
- branch and commit
- operating system, CPU, RAM, Python version and Ollama version
- model tag, observed Ollama active models and whether
--allow-loaded-modelswas used - chunk count and chunk-size distribution
- runtime per stage and total wall-clock runtime
- average extraction time per chunk
- validator violations and deterministic repair count
- renderer contract result
Known Limitation
The current Working Protocol Renderer V2 often produces contract-invalid output
that does not start with # Working Protocol. The runner preserves the raw
renderer response and validation report, and writes working_protocol.md only
when the contract validator passes. This is the known BUG-011 renderer
limitation, not a runner failure by itself.