Add a repository-native runner for reproducible Meeting Lab benchmark runs on machines without Codex. The runner: - validates normalized Whisper input and Meeting Context - checks Ollama availability and the requested model - rejects preloaded Ollama models by default for clean benchmarks - supports an explicit --allow-loaded-models override - uses the selected production configuration: - qwen3.5:9b - target_chars=4500 - max_chars=5500 - min_chars=2500 - overlap_blocks=0 - think=false - temperature=0 - num_ctx=32768 - executes the complete current pipeline - creates unique benchmark output directories - preserves artifacts up to failure - records runtime, environment and validation metadata - writes working_protocol.md only when the renderer contract passes Add focused mocked tests and Linux-first setup documentation for the AI-PC. The runner does not include Whisper execution.
143 lines
4.2 KiB
Markdown
143 lines
4.2 KiB
Markdown
# Running A Meeting Benchmark
|
|
|
|
This page documents the repository-native runner for the current Meeting Lab
|
|
production benchmark on the Linux AI-PC. It does not run Whisper. It starts
|
|
from a converted, normalized Whisper JSON file.
|
|
|
|
## Prerequisites
|
|
|
|
- Git
|
|
- Python 3.11 or newer
|
|
- Ollama running locally
|
|
- Required Ollama model installed
|
|
- Meeting Lab dependencies installed in a virtual environment
|
|
|
|
## Setup
|
|
|
|
From the repository root on Linux:
|
|
|
|
```bash
|
|
python3 -m venv .venv
|
|
source .venv/bin/activate
|
|
python -m pip install --upgrade pip
|
|
python -m pip install -e .
|
|
```
|
|
|
|
Install the reference model in Ollama if it is not already present:
|
|
|
|
```bash
|
|
ollama pull qwen3.5:9b
|
|
```
|
|
|
|
Confirm Ollama sees the model:
|
|
|
|
```bash
|
|
ollama list
|
|
ollama ps
|
|
```
|
|
|
|
`ollama ps` must show no loaded model before a reference benchmark run. By
|
|
default, the runner fails pre-flight if Ollama reports a loaded model. Use
|
|
`--allow-loaded-models` only for a deliberately non-clean run; the default
|
|
reference command does not use that override.
|
|
|
|
## Progeo Reference Command
|
|
|
|
Run from the repository root:
|
|
|
|
```bash
|
|
python scripts/run_meeting.py \
|
|
--input samples/real_live/progeo_meeting/progeo_whispercpp_vulkan_turbo_trimmed_converted.json \
|
|
--context samples/real_live/progeo_meeting/meeting_context.yaml \
|
|
--model qwen3.5:9b \
|
|
--benchmark-label progeo_ai_pc
|
|
```
|
|
|
|
## Windows / PowerShell Example
|
|
|
|
For local Windows development, activate the virtual environment through
|
|
PowerShell and run the same benchmark:
|
|
|
|
```powershell
|
|
python -m venv .venv
|
|
.\.venv\Scripts\Activate.ps1
|
|
python -m pip install --upgrade pip
|
|
python -m pip install -e .
|
|
ollama pull qwen3.5:9b
|
|
python scripts\run_meeting.py `
|
|
--input samples\real_live\progeo_meeting\progeo_whispercpp_vulkan_turbo_trimmed_converted.json `
|
|
--context samples\real_live\progeo_meeting\meeting_context.yaml `
|
|
--model qwen3.5:9b `
|
|
--benchmark-label progeo_ai_pc
|
|
```
|
|
|
|
## Production Defaults
|
|
|
|
The Progeo reference benchmark defaults match the latest controlled experiment
|
|
selection:
|
|
|
|
- Chunking: `target_chars=4500`, `max_chars=5500`, `min_chars=2500`,
|
|
`overlap_blocks=0`
|
|
- LLM: `qwen3.5:9b`, `think=false`, `temperature=0`, `num_ctx=32768`
|
|
- Extraction: context-aware chunk extraction with current committed prompts
|
|
- Consolidation: Canonicalizer V1, Semantic Consolidator V0 adaptive sizing,
|
|
strict source-coverage validation and deterministic source-coverage repair
|
|
- Rendering: Working Protocol Renderer V2
|
|
- No semantic retries and no manual intervention
|
|
|
|
Use `--help` to see explicit overrides.
|
|
|
|
## Output
|
|
|
|
The runner creates a unique directory:
|
|
|
|
```text
|
|
samples/benchmarks/<benchmark-label>_<timestamp>/
|
|
```
|
|
|
|
It never overwrites an existing run. Artifacts include:
|
|
|
|
- `chunks/`
|
|
- `normalized_chunks/`
|
|
- `extractions/`, including parsed extraction JSON and raw successful model
|
|
responses
|
|
- `canonicalizer/canonicalized_extractions.json`
|
|
- `semantic_consolidator/`, including raw response, repaired groups,
|
|
validator reports, repair metadata and consolidated output
|
|
- `working_protocol/`, including raw renderer response and validation report
|
|
- `working_protocol.md` only when the renderer contract is valid
|
|
- `run_metadata.json`
|
|
- `benchmark_report.md`
|
|
|
|
If a stage fails, artifacts produced up to the failure are preserved and the
|
|
report records the failure.
|
|
|
|
## Comparing Benchmarks
|
|
|
|
For AI-PC comparison, compare the new `progeo_ai_pc_<timestamp>` directory
|
|
against:
|
|
|
|
```text
|
|
samples/benchmarks/progeo_qwen35_9b_20260805_133942/
|
|
```
|
|
|
|
Key fields are in `run_metadata.json` and summarized in `benchmark_report.md`:
|
|
|
|
- branch and commit
|
|
- operating system, CPU, RAM, Python version and Ollama version
|
|
- model tag, observed Ollama active models and whether
|
|
`--allow-loaded-models` was used
|
|
- chunk count and chunk-size distribution
|
|
- runtime per stage and total wall-clock runtime
|
|
- average extraction time per chunk
|
|
- validator violations and deterministic repair count
|
|
- renderer contract result
|
|
|
|
## Known Limitation
|
|
|
|
The current Working Protocol Renderer V2 often produces contract-invalid output
|
|
that does not start with `# Working Protocol`. The runner preserves the raw
|
|
renderer response and validation report, and writes `working_protocol.md` only
|
|
when the contract validator passes. This is the known BUG-011 renderer
|
|
limitation, not a runner failure by itself.
|