- introduce Gold Standard evaluation corpus - document decision taxonomy - define prompt-engineering methodology - add regression workflow - establish Prompt Version 2 baseline - validate decision_simple, decision_deferred and decision_none
35 lines
1.4 KiB
Markdown
35 lines
1.4 KiB
Markdown
# Prompt Engineering Methodology
|
|
|
|
Prompt engineering uses the Gold Standard corpus as the reference. The corpus is
|
|
not adjusted to make a prompt pass.
|
|
|
|
Rules:
|
|
|
|
1. Only one prompt change is allowed per iteration.
|
|
2. Only one gold test case may be optimized at a time.
|
|
3. Every prompt modification must be validated immediately.
|
|
4. A prompt modification is acceptable only if it improves the current target
|
|
and does not degrade any previously passing gold test.
|
|
5. Never modify `expected.json` to make a prompt pass.
|
|
6. Prompt engineering edits prompt files only. Python code changes require a
|
|
separate explicit task.
|
|
7. Maintain a prompt evolution log for every iteration.
|
|
8. If a prompt cannot improve a test after several small iterations, stop and
|
|
analyze the root cause.
|
|
9. Prompt changes must be generally applicable and must not special-case one
|
|
transcript.
|
|
10. If two consecutive prompt iterations fail to improve the current target,
|
|
stop further prompt modifications and classify the root cause.
|
|
11. If a prompt produces unexpected behavior, first verify whether the targeted
|
|
gold test has an objectively unique ground truth.
|
|
|
|
Prompt Version 2 baseline:
|
|
|
|
- `decision_simple`: passing
|
|
- `decision_deferred`: passing
|
|
- `decision_none`: passing
|
|
|
|
Prompt Version 2 adds explicit support for process decisions where the group
|
|
agrees to defer a substantive decision until additional information is
|
|
available.
|