Add a repository-native runner for reproducible Meeting Lab benchmark runs on machines without Codex.
The runner:
- validates normalized Whisper input and Meeting Context
- checks Ollama availability and the requested model
- rejects preloaded Ollama models by default for clean benchmarks
- supports an explicit --allow-loaded-models override
- uses the selected production configuration:
- qwen3.5:9b
- target_chars=4500
- max_chars=5500
- min_chars=2500
- overlap_blocks=0
- think=false
- temperature=0
- num_ctx=32768
- executes the complete current pipeline
- creates unique benchmark output directories
- preserves artifacts up to failure
- records runtime, environment and validation metadata
- writes working_protocol.md only when the renderer contract passes
Add focused mocked tests and Linux-first setup documentation for the AI-PC.
The runner does not include Whisper execution.
Introduce Meeting Context V1 with YAML schema, validation and template.
Support optional --meeting-context during chunk extraction.
Inject authoritative Meeting Context into extraction prompts.
Record Meeting Context provenance in extraction output.
Activate todos.md in shared prompt assembly.
Strengthen responsibility attribution and decision/todo boundaries.
Add focused Gold scenarios and validation tests.
Update architecture and pipeline documentation.
- document responsibility attribution as a project-wide invariant
- prevent inferred ownership in protocol rendering
- add negative gold regression for false responsibility assignment
- document future responsibility evidence and attribution model
- record the real-life benchmark finding
- add deterministic canonicalization support for extraction items
- add facts-only semantic consolidation using local Ollama
- preserve source evidence and validate complete fact coverage
- add conservative merge rules and non-LLM tests
- record the first validated real-life consolidation benchmark
- document current scope, limitations and next evaluation step