Establish prompt engineering baseline with Gold Standard tests

- introduce Gold Standard evaluation corpus
- document decision taxonomy
- define prompt-engineering methodology
- add regression workflow
- establish Prompt Version 2 baseline
- validate decision_simple, decision_deferred and decision_none
This commit is contained in:
2026-07-30 12:13:10 +02:00
parent 07b0d80113
commit f7ad9ba51f
43 changed files with 1288 additions and 45 deletions
+11
View File
@@ -0,0 +1,11 @@
# todo_simple
Tests one clear action item with responsible person and deadline.
The difficult part is not adding extra scope: Nina explicitly says she will not touch the layout.
Typical LLM mistakes:
- Adding a layout update as a task.
- Dropping the deadline.
- Turning Omar's request into the task evidence instead of Nina's commitment.
+20
View File
@@ -0,0 +1,20 @@
{
"facts": [
{
"fact": "The beta signup page points to the old privacy note.",
"evidence": "Omar: The beta signup page still points to the old privacy note."
}
],
"decisions": [],
"todos": [
{
"task": "Update the privacy link on the beta signup page.",
"responsible": "Nina",
"deadline": "Thursday noon",
"evidence": "Nina: Yes, I will update the privacy link by Thursday noon."
}
],
"questions": [],
"positions": [],
"technical": []
}
+19
View File
@@ -0,0 +1,19 @@
Omar: The beta signup page still points to the old privacy note.
Nina: Yes, I saw that yesterday.
Omar: Can you update the link before the partner demo?
Nina: Yes, I will update the privacy link by Thursday noon.
Omar: Great. Nothing else on that page from my side.
Nina: I will only touch the link, not the layout.
Omar: Fine.
Nina: Then I have what I need.
Omar: We can move on.
Nina: Yes.