Establish prompt engineering baseline with Gold Standard tests
- introduce Gold Standard evaluation corpus - document decision taxonomy - define prompt-engineering methodology - add regression workflow - establish Prompt Version 2 baseline - validate decision_simple, decision_deferred and decision_none
This commit is contained in:
@@ -0,0 +1,21 @@
|
||||
# decision_deferred
|
||||
|
||||
Tests that a process decision to defer a substantive decision is still extracted
|
||||
as a decision.
|
||||
|
||||
The transcript contains several candidate options for the weekly dashboard, but
|
||||
the group does not choose any of them. Instead, Mira explicitly says not to
|
||||
decide today and Jonas agrees that more input is needed.
|
||||
|
||||
Ground truth:
|
||||
|
||||
- No substantive dashboard schedule decision is reached.
|
||||
- One process decision is reached: the substantive decision is deferred until
|
||||
more information is available.
|
||||
|
||||
Typical LLM mistakes:
|
||||
|
||||
- Extracting Monday, Tuesday, or Friday as the chosen dashboard day.
|
||||
- Treating "I like shorter" as an approval.
|
||||
- Missing the deferral because the substantive decision is unresolved.
|
||||
- Treating "no decision today" as equivalent to no decision at all.
|
||||
Reference in New Issue
Block a user