# AGENTS.md Practical instructions for coding agents working in Meeting Lab. ## Project Purpose Meeting Lab extracts and structures organizational knowledge from meeting recordings. It is an experimental local discussion analyzer, not merely a one-step meeting-protocol generator. Successful approaches may later move into the Meeting Assistant project. ## Current Pipeline Current and intended flow: ```text Audio -> Whisper -> cleanup -> normalization -> chunking -> local chunk extraction -> consolidation -> Canonical Meeting Knowledge -> Output Views ``` Status: - Implemented: Whisper JSON cleanup script, normalization, technical chunking, local chunk extraction, interim Markdown protocol builder. - Experimental/prototype: topic segmentation and review tooling. - Planned: consolidation, Canonical Meeting Knowledge implementation, final Output Views. ## Architectural Principles - Canonical Meeting Knowledge is the intended semantic source of truth. - Working Protocol / Arbeitsprotokoll, Distribution Protocol / Verteilerprotokoll and Knowledge Objects / Wissensdatenbankeintrag are parallel output views. - Output views must not silently change meaning. They may select, condense or render information for an audience, but not invent new semantics. - Extraction, consolidation, synthesis and rendering are separate concerns. - Prefer small, testable processing stages over one monolithic LLM prompt. - Current extraction strategy is one normalized chunk per LLM call. - Do not expand context windows or redesign the extraction strategy without an explicit experiment. - Deterministic stages should remain deterministic where possible. ## Prompt Engineering Rules The Gold Standard corpus is the reference specification. Follow Rules 1-11 from `tests/gold/PROMPT_ENGINEERING_METHODOLOGY.md`: 1. Make only one prompt change per iteration. 2. Optimize only one target gold test case at a time. 3. Validate every prompt modification immediately. 4. Accept a prompt change only if it improves the target and causes no regressions in previously passing gold tests. 5. Never modify `expected.json` merely to make a prompt pass. 6. Prompt engineering edits prompt files only; Python code changes require a separate explicit task. 7. Maintain a prompt evolution log for every iteration. 8. Stop arbitrary iterations if small changes do not improve the test; analyze the root cause. 9. Avoid gold-test overfitting. Prompt changes must generalize and must not special-case one transcript. 10. Stop after two consecutive non-improving prompt iterations and classify the root cause. 11. Verify whether the target gold test has objectively unique ground truth before changing a prompt for unexpected behavior. Current documented Prompt Version 2 decision baseline: - `decision_simple`: passing - `decision_deferred`: passing - `decision_none`: passing ## LLM Execution Safety - Never start a full multi-chunk LLM run unless explicitly requested. - Before any LLM run, state the model, inputs, expected LLM-call count and output location. - Do not retry LLM calls automatically unless explicitly allowed. - Do not download models automatically. - Prefer small-scope validation runs. - Never use generated output as committed source data. - Preserve raw model responses when diagnosing parser or truncation failures. - Do not run Ollama from unit tests. ## Development Rules - Make small, focused changes. - Preserve the existing architecture unless a redesign is explicitly requested. - Add regression tests for bugs. - Run non-LLM tests before committing when code changes are made. - Do not commit generated transcripts, audio, extraction JSON, protocol output or temporary files. - Report files changed, tests run and assumptions. - Do not commit or push unless explicitly requested. ## Repository Conventions - `README.md`: project overview and current high-level status. - `docs/`: architecture, pipeline, data model and output-view documentation. - `prompts/`: extraction and segmentation prompts. Treat prompt edits as controlled experiments. - `tests/gold/`: Gold Standard corpus and semantic specification for extraction behavior. - `scripts/`: command-line support scripts such as Whisper cleanup and gold test execution. - `src/meeting_lab/normalization/`: deterministic transcript cleanup. - `src/meeting_lab/chunking/`: technical chunk creation; chunks are not topics. - `src/meeting_lab/segmentation/`: experimental topic segmentation tooling. - `src/meeting_lab/extraction/`: local LLM extraction flow and category extractor modules. - `src/meeting_lab/consolidation/`: planned consolidation area. - `src/meeting_lab/protocol/`: interim protocol rendering. - `src/meeting_lab/models/`: current lightweight data models. - `samples/`: sample inputs and generated/experimental artifacts; do not treat sample output as canonical source data.