- document responsibility attribution as a project-wide invariant - prevent inferred ownership in protocol rendering - add negative gold regression for false responsibility assignment - document future responsibility evidence and attribution model - record the real-life benchmark finding
164 lines
7.0 KiB
Markdown
164 lines
7.0 KiB
Markdown
# AGENTS.md
|
|
|
|
Practical instructions for coding agents working in Meeting Lab.
|
|
|
|
## Project Purpose
|
|
|
|
Meeting Lab extracts and structures organizational knowledge from meeting
|
|
recordings. It is an experimental local discussion analyzer, not merely a
|
|
one-step meeting-protocol generator.
|
|
|
|
Successful approaches may later move into the Meeting Assistant project.
|
|
|
|
## Current Pipeline
|
|
|
|
Current and intended flow:
|
|
|
|
```text
|
|
Audio
|
|
-> Whisper
|
|
-> cleanup
|
|
-> normalization
|
|
-> chunking
|
|
-> local chunk extraction
|
|
-> deterministic canonicalization
|
|
-> semantic consolidation
|
|
-> Canonical Meeting Knowledge
|
|
-> Output Views
|
|
```
|
|
|
|
Status:
|
|
|
|
- Implemented: Whisper JSON cleanup script, normalization, technical chunking,
|
|
local chunk extraction, Canonicalizer V1, Semantic Consolidator V0
|
|
facts-only duplicate detection, interim Markdown protocol builder.
|
|
- Experimental/prototype: topic segmentation and review tooling.
|
|
- Planned: broader semantic consolidation, Canonical Meeting Knowledge
|
|
implementation, final Output Views.
|
|
|
|
## Architectural Principles
|
|
|
|
- Canonical Meeting Knowledge is the intended semantic source of truth.
|
|
- Working Protocol / Arbeitsprotokoll, Distribution Protocol /
|
|
Verteilerprotokoll and Knowledge Objects / Wissensdatenbankeintrag are
|
|
parallel output views.
|
|
- Output views must not silently change meaning. They may select, condense or
|
|
render information for an audience, but not invent new semantics.
|
|
- Extraction, consolidation, synthesis and rendering are separate concerns.
|
|
- Deterministic canonicalization and semantic consolidation are separate
|
|
concerns.
|
|
- Canonicalizer V1 is implemented Python code. It validates and normalizes
|
|
extraction objects, assigns stable source references and IDs, normalizes
|
|
category names and basic field structure, performs only safe deterministic
|
|
cleanup, may group exact duplicates, and must preserve all source evidence.
|
|
It must not perform uncertain semantic merging.
|
|
- Semantic Consolidator V0 is implemented as local-LLM facts-only duplicate
|
|
detection. It merges semantically equivalent fact items conservatively,
|
|
preserves source references and evidence, and validates complete source fact
|
|
coverage. It is not a summarizer, topic grouper, protocol renderer or
|
|
complete Canonical Meeting Knowledge stage.
|
|
- Broader semantic consolidation remains planned. It should group content by
|
|
topic, mark contradictions and uncertainty, separate durable information from
|
|
transient discussion, and prepare Canonical Meeting Knowledge. It does not
|
|
directly write a protocol.
|
|
- Prefer small, testable processing stages over one monolithic LLM prompt.
|
|
- Current extraction strategy is one normalized chunk per LLM call.
|
|
- Do not expand context windows or redesign the extraction strategy without an
|
|
explicit experiment.
|
|
- Deterministic stages should remain deterministic where possible.
|
|
- Rendered protocol output language should normally match the dominant language
|
|
of the source transcript or consolidated meeting knowledge unless an explicit
|
|
output language is requested.
|
|
|
|
## Responsibility Attribution Invariant
|
|
|
|
A person, team or department may be recorded as responsible only when the
|
|
meeting evidence explicitly assigns, accepts or confirms that responsibility.
|
|
|
|
Discussion, expertise, objection, suggestion, thematic proximity, speaker
|
|
adjacency, organizational assumptions, likely job roles or mere mention do not
|
|
establish ownership.
|
|
|
|
When support is incomplete or ambiguous, leave the responsible person unset,
|
|
mark the item as unclear where the current schema supports it, and preserve the
|
|
supporting evidence. Never guess.
|
|
|
|
This invariant applies to extraction, canonicalization, semantic consolidation,
|
|
Canonical Meeting Knowledge, Working Protocol / Arbeitsprotokoll, Distribution
|
|
Protocol / Verteilerprotokoll and Knowledge Objects /
|
|
Wissensdatenbankeintrag.
|
|
|
|
## Prompt Engineering Rules
|
|
|
|
The Gold Standard corpus is the reference specification. Follow Rules 1-11 from
|
|
`tests/gold/PROMPT_ENGINEERING_METHODOLOGY.md`:
|
|
|
|
1. Make only one prompt change per iteration.
|
|
2. Optimize only one target gold test case at a time.
|
|
3. Validate every prompt modification immediately.
|
|
4. Accept a prompt change only if it improves the target and causes no
|
|
regressions in previously passing gold tests.
|
|
5. Never modify `expected.json` merely to make a prompt pass.
|
|
6. Prompt engineering edits prompt files only; Python code changes require a
|
|
separate explicit task.
|
|
7. Maintain a prompt evolution log for every iteration.
|
|
8. Stop arbitrary iterations if small changes do not improve the test; analyze
|
|
the root cause.
|
|
9. Avoid gold-test overfitting. Prompt changes must generalize and must not
|
|
special-case one transcript.
|
|
10. Stop after two consecutive non-improving prompt iterations and classify the
|
|
root cause.
|
|
11. Verify whether the target gold test has objectively unique ground truth
|
|
before changing a prompt for unexpected behavior.
|
|
|
|
Current documented Prompt Version 2 decision baseline:
|
|
|
|
- `decision_simple`: passing
|
|
- `decision_deferred`: passing
|
|
- `decision_none`: passing
|
|
|
|
## LLM Execution Safety
|
|
|
|
- Never start a full multi-chunk LLM run unless explicitly requested.
|
|
- Before any LLM run, state the model, inputs, expected LLM-call count and
|
|
output location.
|
|
- Do not retry LLM calls automatically unless explicitly allowed.
|
|
- Do not download models automatically.
|
|
- Prefer small-scope validation runs.
|
|
- Never use generated output as committed source data.
|
|
- Preserve raw model responses when diagnosing parser or truncation failures.
|
|
- Do not run Ollama from unit tests.
|
|
|
|
## Development Rules
|
|
|
|
- Make small, focused changes.
|
|
- Preserve the existing architecture unless a redesign is explicitly requested.
|
|
- Add regression tests for bugs.
|
|
- Run non-LLM tests before committing when code changes are made.
|
|
- Do not commit generated transcripts, audio, extraction JSON, protocol output
|
|
or temporary files.
|
|
- Report files changed, tests run and assumptions.
|
|
- Do not commit or push unless explicitly requested.
|
|
|
|
## Repository Conventions
|
|
|
|
- `README.md`: project overview and current high-level status.
|
|
- `docs/`: architecture, pipeline, data model and output-view documentation.
|
|
- `prompts/`: extraction and segmentation prompts. Treat prompt edits as
|
|
controlled experiments.
|
|
- `tests/gold/`: Gold Standard corpus and semantic specification for extraction
|
|
behavior.
|
|
- `scripts/`: command-line support scripts such as Whisper cleanup and gold
|
|
test execution.
|
|
- `src/meeting_lab/normalization/`: deterministic transcript cleanup.
|
|
- `src/meeting_lab/chunking/`: technical chunk creation; chunks are not topics.
|
|
- `src/meeting_lab/segmentation/`: experimental topic segmentation tooling.
|
|
- `src/meeting_lab/extraction/`: local LLM extraction flow and category
|
|
extractor modules.
|
|
- `src/meeting_lab/consolidation/`: Canonicalizer V1 and Semantic
|
|
Consolidator V0.
|
|
- `src/meeting_lab/protocol/`: interim protocol rendering.
|
|
- `src/meeting_lab/models/`: current lightweight data models.
|
|
- `samples/`: sample inputs and generated/experimental artifacts; do not treat
|
|
sample output as canonical source data.
|