Files
admin e04e2533fc Add one-command Meeting Lab benchmark runner
Add a repository-native runner for reproducible Meeting Lab benchmark runs on machines without Codex.

The runner:

- validates normalized Whisper input and Meeting Context
- checks Ollama availability and the requested model
- rejects preloaded Ollama models by default for clean benchmarks
- supports an explicit --allow-loaded-models override
- uses the selected production configuration:
  - qwen3.5:9b
  - target_chars=4500
  - max_chars=5500
  - min_chars=2500
  - overlap_blocks=0
  - think=false
  - temperature=0
  - num_ctx=32768
- executes the complete current pipeline
- creates unique benchmark output directories
- preserves artifacts up to failure
- records runtime, environment and validation metadata
- writes working_protocol.md only when the renderer contract passes

Add focused mocked tests and Linux-first setup documentation for the AI-PC.
The runner does not include Whisper execution.
2026-08-05 14:48:42 +02:00

143 lines
4.2 KiB
Markdown

# Running A Meeting Benchmark
This page documents the repository-native runner for the current Meeting Lab
production benchmark on the Linux AI-PC. It does not run Whisper. It starts
from a converted, normalized Whisper JSON file.
## Prerequisites
- Git
- Python 3.11 or newer
- Ollama running locally
- Required Ollama model installed
- Meeting Lab dependencies installed in a virtual environment
## Setup
From the repository root on Linux:
```bash
python3 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -e .
```
Install the reference model in Ollama if it is not already present:
```bash
ollama pull qwen3.5:9b
```
Confirm Ollama sees the model:
```bash
ollama list
ollama ps
```
`ollama ps` must show no loaded model before a reference benchmark run. By
default, the runner fails pre-flight if Ollama reports a loaded model. Use
`--allow-loaded-models` only for a deliberately non-clean run; the default
reference command does not use that override.
## Progeo Reference Command
Run from the repository root:
```bash
python scripts/run_meeting.py \
--input samples/real_live/progeo_meeting/progeo_whispercpp_vulkan_turbo_trimmed_converted.json \
--context samples/real_live/progeo_meeting/meeting_context.yaml \
--model qwen3.5:9b \
--benchmark-label progeo_ai_pc
```
## Windows / PowerShell Example
For local Windows development, activate the virtual environment through
PowerShell and run the same benchmark:
```powershell
python -m venv .venv
.\.venv\Scripts\Activate.ps1
python -m pip install --upgrade pip
python -m pip install -e .
ollama pull qwen3.5:9b
python scripts\run_meeting.py `
--input samples\real_live\progeo_meeting\progeo_whispercpp_vulkan_turbo_trimmed_converted.json `
--context samples\real_live\progeo_meeting\meeting_context.yaml `
--model qwen3.5:9b `
--benchmark-label progeo_ai_pc
```
## Production Defaults
The Progeo reference benchmark defaults match the latest controlled experiment
selection:
- Chunking: `target_chars=4500`, `max_chars=5500`, `min_chars=2500`,
`overlap_blocks=0`
- LLM: `qwen3.5:9b`, `think=false`, `temperature=0`, `num_ctx=32768`
- Extraction: context-aware chunk extraction with current committed prompts
- Consolidation: Canonicalizer V1, Semantic Consolidator V0 adaptive sizing,
strict source-coverage validation and deterministic source-coverage repair
- Rendering: Working Protocol Renderer V2
- No semantic retries and no manual intervention
Use `--help` to see explicit overrides.
## Output
The runner creates a unique directory:
```text
samples/benchmarks/<benchmark-label>_<timestamp>/
```
It never overwrites an existing run. Artifacts include:
- `chunks/`
- `normalized_chunks/`
- `extractions/`, including parsed extraction JSON and raw successful model
responses
- `canonicalizer/canonicalized_extractions.json`
- `semantic_consolidator/`, including raw response, repaired groups,
validator reports, repair metadata and consolidated output
- `working_protocol/`, including raw renderer response and validation report
- `working_protocol.md` only when the renderer contract is valid
- `run_metadata.json`
- `benchmark_report.md`
If a stage fails, artifacts produced up to the failure are preserved and the
report records the failure.
## Comparing Benchmarks
For AI-PC comparison, compare the new `progeo_ai_pc_<timestamp>` directory
against:
```text
samples/benchmarks/progeo_qwen35_9b_20260805_133942/
```
Key fields are in `run_metadata.json` and summarized in `benchmark_report.md`:
- branch and commit
- operating system, CPU, RAM, Python version and Ollama version
- model tag, observed Ollama active models and whether
`--allow-loaded-models` was used
- chunk count and chunk-size distribution
- runtime per stage and total wall-clock runtime
- average extraction time per chunk
- validator violations and deterministic repair count
- renderer contract result
## Known Limitation
The current Working Protocol Renderer V2 often produces contract-invalid output
that does not start with `# Working Protocol`. The runner preserves the raw
renderer response and validation report, and writes `working_protocol.md` only
when the contract validator passes. This is the known BUG-011 renderer
limitation, not a runner failure by itself.