Files
meeting-lab/docs/run-meeting.md
T
admin e04e2533fc Add one-command Meeting Lab benchmark runner
Add a repository-native runner for reproducible Meeting Lab benchmark runs on machines without Codex.

The runner:

- validates normalized Whisper input and Meeting Context
- checks Ollama availability and the requested model
- rejects preloaded Ollama models by default for clean benchmarks
- supports an explicit --allow-loaded-models override
- uses the selected production configuration:
  - qwen3.5:9b
  - target_chars=4500
  - max_chars=5500
  - min_chars=2500
  - overlap_blocks=0
  - think=false
  - temperature=0
  - num_ctx=32768
- executes the complete current pipeline
- creates unique benchmark output directories
- preserves artifacts up to failure
- records runtime, environment and validation metadata
- writes working_protocol.md only when the renderer contract passes

Add focused mocked tests and Linux-first setup documentation for the AI-PC.
The runner does not include Whisper execution.
2026-08-05 14:48:42 +02:00

4.2 KiB

Running A Meeting Benchmark

This page documents the repository-native runner for the current Meeting Lab production benchmark on the Linux AI-PC. It does not run Whisper. It starts from a converted, normalized Whisper JSON file.

Prerequisites

  • Git
  • Python 3.11 or newer
  • Ollama running locally
  • Required Ollama model installed
  • Meeting Lab dependencies installed in a virtual environment

Setup

From the repository root on Linux:

python3 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -e .

Install the reference model in Ollama if it is not already present:

ollama pull qwen3.5:9b

Confirm Ollama sees the model:

ollama list
ollama ps

ollama ps must show no loaded model before a reference benchmark run. By default, the runner fails pre-flight if Ollama reports a loaded model. Use --allow-loaded-models only for a deliberately non-clean run; the default reference command does not use that override.

Progeo Reference Command

Run from the repository root:

python scripts/run_meeting.py \
  --input samples/real_live/progeo_meeting/progeo_whispercpp_vulkan_turbo_trimmed_converted.json \
  --context samples/real_live/progeo_meeting/meeting_context.yaml \
  --model qwen3.5:9b \
  --benchmark-label progeo_ai_pc

Windows / PowerShell Example

For local Windows development, activate the virtual environment through PowerShell and run the same benchmark:

python -m venv .venv
.\.venv\Scripts\Activate.ps1
python -m pip install --upgrade pip
python -m pip install -e .
ollama pull qwen3.5:9b
python scripts\run_meeting.py `
  --input samples\real_live\progeo_meeting\progeo_whispercpp_vulkan_turbo_trimmed_converted.json `
  --context samples\real_live\progeo_meeting\meeting_context.yaml `
  --model qwen3.5:9b `
  --benchmark-label progeo_ai_pc

Production Defaults

The Progeo reference benchmark defaults match the latest controlled experiment selection:

  • Chunking: target_chars=4500, max_chars=5500, min_chars=2500, overlap_blocks=0
  • LLM: qwen3.5:9b, think=false, temperature=0, num_ctx=32768
  • Extraction: context-aware chunk extraction with current committed prompts
  • Consolidation: Canonicalizer V1, Semantic Consolidator V0 adaptive sizing, strict source-coverage validation and deterministic source-coverage repair
  • Rendering: Working Protocol Renderer V2
  • No semantic retries and no manual intervention

Use --help to see explicit overrides.

Output

The runner creates a unique directory:

samples/benchmarks/<benchmark-label>_<timestamp>/

It never overwrites an existing run. Artifacts include:

  • chunks/
  • normalized_chunks/
  • extractions/, including parsed extraction JSON and raw successful model responses
  • canonicalizer/canonicalized_extractions.json
  • semantic_consolidator/, including raw response, repaired groups, validator reports, repair metadata and consolidated output
  • working_protocol/, including raw renderer response and validation report
  • working_protocol.md only when the renderer contract is valid
  • run_metadata.json
  • benchmark_report.md

If a stage fails, artifacts produced up to the failure are preserved and the report records the failure.

Comparing Benchmarks

For AI-PC comparison, compare the new progeo_ai_pc_<timestamp> directory against:

samples/benchmarks/progeo_qwen35_9b_20260805_133942/

Key fields are in run_metadata.json and summarized in benchmark_report.md:

  • branch and commit
  • operating system, CPU, RAM, Python version and Ollama version
  • model tag, observed Ollama active models and whether --allow-loaded-models was used
  • chunk count and chunk-size distribution
  • runtime per stage and total wall-clock runtime
  • average extraction time per chunk
  • validator violations and deterministic repair count
  • renderer contract result

Known Limitation

The current Working Protocol Renderer V2 often produces contract-invalid output that does not start with # Working Protocol. The runner preserves the raw renderer response and validation report, and writes working_protocol.md only when the contract validator passes. This is the known BUG-011 renderer limitation, not a runner failure by itself.