Establish prompt engineering baseline with Gold Standard tests
- introduce Gold Standard evaluation corpus - document decision taxonomy - define prompt-engineering methodology - add regression workflow - establish Prompt Version 2 baseline - validate decision_simple, decision_deferred and decision_none
This commit is contained in:
@@ -0,0 +1,34 @@
|
||||
# Prompt Engineering Methodology
|
||||
|
||||
Prompt engineering uses the Gold Standard corpus as the reference. The corpus is
|
||||
not adjusted to make a prompt pass.
|
||||
|
||||
Rules:
|
||||
|
||||
1. Only one prompt change is allowed per iteration.
|
||||
2. Only one gold test case may be optimized at a time.
|
||||
3. Every prompt modification must be validated immediately.
|
||||
4. A prompt modification is acceptable only if it improves the current target
|
||||
and does not degrade any previously passing gold test.
|
||||
5. Never modify `expected.json` to make a prompt pass.
|
||||
6. Prompt engineering edits prompt files only. Python code changes require a
|
||||
separate explicit task.
|
||||
7. Maintain a prompt evolution log for every iteration.
|
||||
8. If a prompt cannot improve a test after several small iterations, stop and
|
||||
analyze the root cause.
|
||||
9. Prompt changes must be generally applicable and must not special-case one
|
||||
transcript.
|
||||
10. If two consecutive prompt iterations fail to improve the current target,
|
||||
stop further prompt modifications and classify the root cause.
|
||||
11. If a prompt produces unexpected behavior, first verify whether the targeted
|
||||
gold test has an objectively unique ground truth.
|
||||
|
||||
Prompt Version 2 baseline:
|
||||
|
||||
- `decision_simple`: passing
|
||||
- `decision_deferred`: passing
|
||||
- `decision_none`: passing
|
||||
|
||||
Prompt Version 2 adds explicit support for process decisions where the group
|
||||
agrees to defer a substantive decision until additional information is
|
||||
available.
|
||||
Reference in New Issue
Block a user