# AGENTS.md Practical instructions for coding agents working in Meeting Lab. ## Project Purpose Meeting Lab extracts and structures organizational knowledge from meeting recordings. It is an experimental local discussion analyzer, not merely a one-step meeting-protocol generator. Successful approaches may later move into the Meeting Assistant project. ## Current Pipeline Current and intended flow: ```text Audio -> Whisper -> cleanup -> normalization -> chunking -> local chunk extraction -> deterministic canonicalization -> semantic consolidation -> Canonical Meeting Knowledge -> Output Views ``` Status: - Implemented: Whisper JSON cleanup script, normalization, technical chunking, local chunk extraction, Canonicalizer V1, Semantic Consolidator V0 facts-only duplicate detection, interim Markdown protocol builder. - Experimental/prototype: topic segmentation and review tooling. - Planned: broader semantic consolidation, Canonical Meeting Knowledge implementation, final Output Views. ## Architectural Principles - Canonical Meeting Knowledge is the intended semantic source of truth. - Working Protocol / Arbeitsprotokoll, Distribution Protocol / Verteilerprotokoll and Knowledge Objects / Wissensdatenbankeintrag are parallel output views. - Output views must not silently change meaning. They may select, condense or render information for an audience, but not invent new semantics. - Extraction, consolidation, synthesis and rendering are separate concerns. - Deterministic canonicalization and semantic consolidation are separate concerns. - Canonicalizer V1 is implemented Python code. It validates and normalizes extraction objects, assigns stable source references and IDs, normalizes category names and basic field structure, performs only safe deterministic cleanup, may group exact duplicates, and must preserve all source evidence. It must not perform uncertain semantic merging. - Semantic Consolidator V0 is implemented as local-LLM facts-only duplicate detection. It merges semantically equivalent fact items conservatively, preserves source references and evidence, and validates complete source fact coverage. It is not a summarizer, topic grouper, protocol renderer or complete Canonical Meeting Knowledge stage. - Broader semantic consolidation remains planned. It should group content by topic, mark contradictions and uncertainty, separate durable information from transient discussion, and prepare Canonical Meeting Knowledge. It does not directly write a protocol. - Prefer small, testable processing stages over one monolithic LLM prompt. - Current extraction strategy is one normalized chunk per LLM call. - Do not expand context windows or redesign the extraction strategy without an explicit experiment. - Deterministic stages should remain deterministic where possible. - Rendered protocol output language should normally match the dominant language of the source transcript or consolidated meeting knowledge unless an explicit output language is requested. ## Prompt Engineering Rules The Gold Standard corpus is the reference specification. Follow Rules 1-11 from `tests/gold/PROMPT_ENGINEERING_METHODOLOGY.md`: 1. Make only one prompt change per iteration. 2. Optimize only one target gold test case at a time. 3. Validate every prompt modification immediately. 4. Accept a prompt change only if it improves the target and causes no regressions in previously passing gold tests. 5. Never modify `expected.json` merely to make a prompt pass. 6. Prompt engineering edits prompt files only; Python code changes require a separate explicit task. 7. Maintain a prompt evolution log for every iteration. 8. Stop arbitrary iterations if small changes do not improve the test; analyze the root cause. 9. Avoid gold-test overfitting. Prompt changes must generalize and must not special-case one transcript. 10. Stop after two consecutive non-improving prompt iterations and classify the root cause. 11. Verify whether the target gold test has objectively unique ground truth before changing a prompt for unexpected behavior. Current documented Prompt Version 2 decision baseline: - `decision_simple`: passing - `decision_deferred`: passing - `decision_none`: passing ## LLM Execution Safety - Never start a full multi-chunk LLM run unless explicitly requested. - Before any LLM run, state the model, inputs, expected LLM-call count and output location. - Do not retry LLM calls automatically unless explicitly allowed. - Do not download models automatically. - Prefer small-scope validation runs. - Never use generated output as committed source data. - Preserve raw model responses when diagnosing parser or truncation failures. - Do not run Ollama from unit tests. ## Development Rules - Make small, focused changes. - Preserve the existing architecture unless a redesign is explicitly requested. - Add regression tests for bugs. - Run non-LLM tests before committing when code changes are made. - Do not commit generated transcripts, audio, extraction JSON, protocol output or temporary files. - Report files changed, tests run and assumptions. - Do not commit or push unless explicitly requested. ## Repository Conventions - `README.md`: project overview and current high-level status. - `docs/`: architecture, pipeline, data model and output-view documentation. - `prompts/`: extraction and segmentation prompts. Treat prompt edits as controlled experiments. - `tests/gold/`: Gold Standard corpus and semantic specification for extraction behavior. - `scripts/`: command-line support scripts such as Whisper cleanup and gold test execution. - `src/meeting_lab/normalization/`: deterministic transcript cleanup. - `src/meeting_lab/chunking/`: technical chunk creation; chunks are not topics. - `src/meeting_lab/segmentation/`: experimental topic segmentation tooling. - `src/meeting_lab/extraction/`: local LLM extraction flow and category extractor modules. - `src/meeting_lab/consolidation/`: Canonicalizer V1 and Semantic Consolidator V0. - `src/meeting_lab/protocol/`: interim protocol rendering. - `src/meeting_lab/models/`: current lightweight data models. - `samples/`: sample inputs and generated/experimental artifacts; do not treat sample output as canonical source data.