- add AGENTS.md with development and prompt-engineering rules - add PROJECT_KNOWLEDGE.md summarizing current architecture and findings - add CHANGELOG.md - add ROADMAP.md - establish experiments.md as the project's experiment log - document Canonical Meeting Knowledge architecture - document Output Views and Knowledge Objects - capture accepted experimental results and engineering methodology
125 lines
4.8 KiB
Markdown
125 lines
4.8 KiB
Markdown
# AGENTS.md
|
|
|
|
Practical instructions for coding agents working in Meeting Lab.
|
|
|
|
## Project Purpose
|
|
|
|
Meeting Lab extracts and structures organizational knowledge from meeting
|
|
recordings. It is an experimental local discussion analyzer, not merely a
|
|
one-step meeting-protocol generator.
|
|
|
|
Successful approaches may later move into the Meeting Assistant project.
|
|
|
|
## Current Pipeline
|
|
|
|
Current and intended flow:
|
|
|
|
```text
|
|
Audio
|
|
-> Whisper
|
|
-> cleanup
|
|
-> normalization
|
|
-> chunking
|
|
-> local chunk extraction
|
|
-> consolidation
|
|
-> Canonical Meeting Knowledge
|
|
-> Output Views
|
|
```
|
|
|
|
Status:
|
|
|
|
- Implemented: Whisper JSON cleanup script, normalization, technical chunking,
|
|
local chunk extraction, interim Markdown protocol builder.
|
|
- Experimental/prototype: topic segmentation and review tooling.
|
|
- Planned: consolidation, Canonical Meeting Knowledge implementation, final
|
|
Output Views.
|
|
|
|
## Architectural Principles
|
|
|
|
- Canonical Meeting Knowledge is the intended semantic source of truth.
|
|
- Working Protocol / Arbeitsprotokoll, Distribution Protocol /
|
|
Verteilerprotokoll and Knowledge Objects / Wissensdatenbankeintrag are
|
|
parallel output views.
|
|
- Output views must not silently change meaning. They may select, condense or
|
|
render information for an audience, but not invent new semantics.
|
|
- Extraction, consolidation, synthesis and rendering are separate concerns.
|
|
- Prefer small, testable processing stages over one monolithic LLM prompt.
|
|
- Current extraction strategy is one normalized chunk per LLM call.
|
|
- Do not expand context windows or redesign the extraction strategy without an
|
|
explicit experiment.
|
|
- Deterministic stages should remain deterministic where possible.
|
|
|
|
## Prompt Engineering Rules
|
|
|
|
The Gold Standard corpus is the reference specification. Follow Rules 1-11 from
|
|
`tests/gold/PROMPT_ENGINEERING_METHODOLOGY.md`:
|
|
|
|
1. Make only one prompt change per iteration.
|
|
2. Optimize only one target gold test case at a time.
|
|
3. Validate every prompt modification immediately.
|
|
4. Accept a prompt change only if it improves the target and causes no
|
|
regressions in previously passing gold tests.
|
|
5. Never modify `expected.json` merely to make a prompt pass.
|
|
6. Prompt engineering edits prompt files only; Python code changes require a
|
|
separate explicit task.
|
|
7. Maintain a prompt evolution log for every iteration.
|
|
8. Stop arbitrary iterations if small changes do not improve the test; analyze
|
|
the root cause.
|
|
9. Avoid gold-test overfitting. Prompt changes must generalize and must not
|
|
special-case one transcript.
|
|
10. Stop after two consecutive non-improving prompt iterations and classify the
|
|
root cause.
|
|
11. Verify whether the target gold test has objectively unique ground truth
|
|
before changing a prompt for unexpected behavior.
|
|
|
|
Current documented Prompt Version 2 decision baseline:
|
|
|
|
- `decision_simple`: passing
|
|
- `decision_deferred`: passing
|
|
- `decision_none`: passing
|
|
|
|
## LLM Execution Safety
|
|
|
|
- Never start a full multi-chunk LLM run unless explicitly requested.
|
|
- Before any LLM run, state the model, inputs, expected LLM-call count and
|
|
output location.
|
|
- Do not retry LLM calls automatically unless explicitly allowed.
|
|
- Do not download models automatically.
|
|
- Prefer small-scope validation runs.
|
|
- Never use generated output as committed source data.
|
|
- Preserve raw model responses when diagnosing parser or truncation failures.
|
|
- Do not run Ollama from unit tests.
|
|
|
|
## Development Rules
|
|
|
|
- Make small, focused changes.
|
|
- Preserve the existing architecture unless a redesign is explicitly requested.
|
|
- Add regression tests for bugs.
|
|
- Run non-LLM tests before committing when code changes are made.
|
|
- Do not commit generated transcripts, audio, extraction JSON, protocol output
|
|
or temporary files.
|
|
- Report files changed, tests run and assumptions.
|
|
- Do not commit or push unless explicitly requested.
|
|
|
|
## Repository Conventions
|
|
|
|
- `README.md`: project overview and current high-level status.
|
|
- `docs/`: architecture, pipeline, data model and output-view documentation.
|
|
- `prompts/`: extraction and segmentation prompts. Treat prompt edits as
|
|
controlled experiments.
|
|
- `tests/gold/`: Gold Standard corpus and semantic specification for extraction
|
|
behavior.
|
|
- `scripts/`: command-line support scripts such as Whisper cleanup and gold
|
|
test execution.
|
|
- `src/meeting_lab/normalization/`: deterministic transcript cleanup.
|
|
- `src/meeting_lab/chunking/`: technical chunk creation; chunks are not topics.
|
|
- `src/meeting_lab/segmentation/`: experimental topic segmentation tooling.
|
|
- `src/meeting_lab/extraction/`: local LLM extraction flow and category
|
|
extractor modules.
|
|
- `src/meeting_lab/consolidation/`: planned consolidation area.
|
|
- `src/meeting_lab/protocol/`: interim protocol rendering.
|
|
- `src/meeting_lab/models/`: current lightweight data models.
|
|
- `samples/`: sample inputs and generated/experimental artifacts; do not treat
|
|
sample output as canonical source data.
|
|
|