Establish prompt engineering baseline with Gold Standard tests
- introduce Gold Standard evaluation corpus - document decision taxonomy - define prompt-engineering methodology - add regression workflow - establish Prompt Version 2 baseline - validate decision_simple, decision_deferred and decision_none
This commit is contained in:
@@ -0,0 +1,11 @@
|
||||
# decision_simple
|
||||
|
||||
Tests one explicit decision with clear agreement language.
|
||||
|
||||
The difficult part is separating the decision from nearby rationale about user confusion and from the non-decision statement that the copy can stay unchanged for now.
|
||||
|
||||
Typical LLM mistakes:
|
||||
|
||||
- Extracting the rationale as a separate decision.
|
||||
- Treating "copy can stay as it is" as a formal decision.
|
||||
- Losing the evidence that shows explicit agreement.
|
||||
@@ -0,0 +1,13 @@
|
||||
{
|
||||
"facts": [],
|
||||
"decisions": [
|
||||
{
|
||||
"decision": "The welcome email will be sent after account activation.",
|
||||
"evidence": "So are we agreed that the welcome email moves to after activation? Ben: Agreed. Cara: Yes, let's do that."
|
||||
}
|
||||
],
|
||||
"todos": [],
|
||||
"questions": [],
|
||||
"positions": [],
|
||||
"technical": []
|
||||
}
|
||||
@@ -0,0 +1,19 @@
|
||||
Anna: Before we leave the onboarding flow, can we settle the email step?
|
||||
|
||||
Ben: I still think the welcome email should go out after account activation, not before.
|
||||
|
||||
Cara: Yes, before activation it keeps confusing people.
|
||||
|
||||
Anna: So are we agreed that the welcome email moves to after activation?
|
||||
|
||||
Ben: Agreed.
|
||||
|
||||
Cara: Yes, let's do that.
|
||||
|
||||
Anna: Good. Then that is the decision for this release.
|
||||
|
||||
Ben: Separate note on the copy: I am not proposing any wording decision today.
|
||||
|
||||
Cara: Same here, no wording proposal from me.
|
||||
|
||||
Anna: Okay, then the only decision is the timing after activation.
|
||||
Reference in New Issue
Block a user