Files
meeting-lab/AGENTS.md
T
admin 63075eaca9 Enforce explicit responsibility attribution
- document responsibility attribution as a project-wide invariant
- prevent inferred ownership in protocol rendering
- add negative gold regression for false responsibility assignment
- document future responsibility evidence and attribution model
- record the real-life benchmark finding
2026-07-31 13:01:11 +02:00

164 lines
7.0 KiB
Markdown

# AGENTS.md
Practical instructions for coding agents working in Meeting Lab.
## Project Purpose
Meeting Lab extracts and structures organizational knowledge from meeting
recordings. It is an experimental local discussion analyzer, not merely a
one-step meeting-protocol generator.
Successful approaches may later move into the Meeting Assistant project.
## Current Pipeline
Current and intended flow:
```text
Audio
-> Whisper
-> cleanup
-> normalization
-> chunking
-> local chunk extraction
-> deterministic canonicalization
-> semantic consolidation
-> Canonical Meeting Knowledge
-> Output Views
```
Status:
- Implemented: Whisper JSON cleanup script, normalization, technical chunking,
local chunk extraction, Canonicalizer V1, Semantic Consolidator V0
facts-only duplicate detection, interim Markdown protocol builder.
- Experimental/prototype: topic segmentation and review tooling.
- Planned: broader semantic consolidation, Canonical Meeting Knowledge
implementation, final Output Views.
## Architectural Principles
- Canonical Meeting Knowledge is the intended semantic source of truth.
- Working Protocol / Arbeitsprotokoll, Distribution Protocol /
Verteilerprotokoll and Knowledge Objects / Wissensdatenbankeintrag are
parallel output views.
- Output views must not silently change meaning. They may select, condense or
render information for an audience, but not invent new semantics.
- Extraction, consolidation, synthesis and rendering are separate concerns.
- Deterministic canonicalization and semantic consolidation are separate
concerns.
- Canonicalizer V1 is implemented Python code. It validates and normalizes
extraction objects, assigns stable source references and IDs, normalizes
category names and basic field structure, performs only safe deterministic
cleanup, may group exact duplicates, and must preserve all source evidence.
It must not perform uncertain semantic merging.
- Semantic Consolidator V0 is implemented as local-LLM facts-only duplicate
detection. It merges semantically equivalent fact items conservatively,
preserves source references and evidence, and validates complete source fact
coverage. It is not a summarizer, topic grouper, protocol renderer or
complete Canonical Meeting Knowledge stage.
- Broader semantic consolidation remains planned. It should group content by
topic, mark contradictions and uncertainty, separate durable information from
transient discussion, and prepare Canonical Meeting Knowledge. It does not
directly write a protocol.
- Prefer small, testable processing stages over one monolithic LLM prompt.
- Current extraction strategy is one normalized chunk per LLM call.
- Do not expand context windows or redesign the extraction strategy without an
explicit experiment.
- Deterministic stages should remain deterministic where possible.
- Rendered protocol output language should normally match the dominant language
of the source transcript or consolidated meeting knowledge unless an explicit
output language is requested.
## Responsibility Attribution Invariant
A person, team or department may be recorded as responsible only when the
meeting evidence explicitly assigns, accepts or confirms that responsibility.
Discussion, expertise, objection, suggestion, thematic proximity, speaker
adjacency, organizational assumptions, likely job roles or mere mention do not
establish ownership.
When support is incomplete or ambiguous, leave the responsible person unset,
mark the item as unclear where the current schema supports it, and preserve the
supporting evidence. Never guess.
This invariant applies to extraction, canonicalization, semantic consolidation,
Canonical Meeting Knowledge, Working Protocol / Arbeitsprotokoll, Distribution
Protocol / Verteilerprotokoll and Knowledge Objects /
Wissensdatenbankeintrag.
## Prompt Engineering Rules
The Gold Standard corpus is the reference specification. Follow Rules 1-11 from
`tests/gold/PROMPT_ENGINEERING_METHODOLOGY.md`:
1. Make only one prompt change per iteration.
2. Optimize only one target gold test case at a time.
3. Validate every prompt modification immediately.
4. Accept a prompt change only if it improves the target and causes no
regressions in previously passing gold tests.
5. Never modify `expected.json` merely to make a prompt pass.
6. Prompt engineering edits prompt files only; Python code changes require a
separate explicit task.
7. Maintain a prompt evolution log for every iteration.
8. Stop arbitrary iterations if small changes do not improve the test; analyze
the root cause.
9. Avoid gold-test overfitting. Prompt changes must generalize and must not
special-case one transcript.
10. Stop after two consecutive non-improving prompt iterations and classify the
root cause.
11. Verify whether the target gold test has objectively unique ground truth
before changing a prompt for unexpected behavior.
Current documented Prompt Version 2 decision baseline:
- `decision_simple`: passing
- `decision_deferred`: passing
- `decision_none`: passing
## LLM Execution Safety
- Never start a full multi-chunk LLM run unless explicitly requested.
- Before any LLM run, state the model, inputs, expected LLM-call count and
output location.
- Do not retry LLM calls automatically unless explicitly allowed.
- Do not download models automatically.
- Prefer small-scope validation runs.
- Never use generated output as committed source data.
- Preserve raw model responses when diagnosing parser or truncation failures.
- Do not run Ollama from unit tests.
## Development Rules
- Make small, focused changes.
- Preserve the existing architecture unless a redesign is explicitly requested.
- Add regression tests for bugs.
- Run non-LLM tests before committing when code changes are made.
- Do not commit generated transcripts, audio, extraction JSON, protocol output
or temporary files.
- Report files changed, tests run and assumptions.
- Do not commit or push unless explicitly requested.
## Repository Conventions
- `README.md`: project overview and current high-level status.
- `docs/`: architecture, pipeline, data model and output-view documentation.
- `prompts/`: extraction and segmentation prompts. Treat prompt edits as
controlled experiments.
- `tests/gold/`: Gold Standard corpus and semantic specification for extraction
behavior.
- `scripts/`: command-line support scripts such as Whisper cleanup and gold
test execution.
- `src/meeting_lab/normalization/`: deterministic transcript cleanup.
- `src/meeting_lab/chunking/`: technical chunk creation; chunks are not topics.
- `src/meeting_lab/segmentation/`: experimental topic segmentation tooling.
- `src/meeting_lab/extraction/`: local LLM extraction flow and category
extractor modules.
- `src/meeting_lab/consolidation/`: Canonicalizer V1 and Semantic
Consolidator V0.
- `src/meeting_lab/protocol/`: interim protocol rendering.
- `src/meeting_lab/models/`: current lightweight data models.
- `samples/`: sample inputs and generated/experimental artifacts; do not treat
sample output as canonical source data.