Establish prompt engineering baseline with Gold Standard tests
- introduce Gold Standard evaluation corpus - document decision taxonomy - define prompt-engineering methodology - add regression workflow - establish Prompt Version 2 baseline - validate decision_simple, decision_deferred and decision_none
This commit is contained in:
@@ -0,0 +1,32 @@
|
||||
# Decision Definition
|
||||
|
||||
A decision is any explicit agreement that creates a binding change in action,
|
||||
process, responsibility, approval status, timing, or next step.
|
||||
|
||||
Included:
|
||||
|
||||
- substantive decisions
|
||||
- organizational decisions
|
||||
- process decisions
|
||||
- approvals
|
||||
- rejections
|
||||
- deferrals
|
||||
- explicit agreement not to decide yet
|
||||
- explicit agreement to gather more information before deciding
|
||||
|
||||
Excluded:
|
||||
|
||||
- opinions
|
||||
- preferences
|
||||
- proposals without agreement
|
||||
- open questions
|
||||
- descriptions of the current state
|
||||
- explanations without commitment
|
||||
|
||||
Important distinction:
|
||||
|
||||
- "No decision was reached" means the meeting ended without an agreed outcome.
|
||||
- "The group decided to defer the decision" means the group explicitly agreed
|
||||
on a process outcome: the substantive decision is postponed.
|
||||
|
||||
These are not equivalent.
|
||||
@@ -0,0 +1,34 @@
|
||||
# Prompt Engineering Methodology
|
||||
|
||||
Prompt engineering uses the Gold Standard corpus as the reference. The corpus is
|
||||
not adjusted to make a prompt pass.
|
||||
|
||||
Rules:
|
||||
|
||||
1. Only one prompt change is allowed per iteration.
|
||||
2. Only one gold test case may be optimized at a time.
|
||||
3. Every prompt modification must be validated immediately.
|
||||
4. A prompt modification is acceptable only if it improves the current target
|
||||
and does not degrade any previously passing gold test.
|
||||
5. Never modify `expected.json` to make a prompt pass.
|
||||
6. Prompt engineering edits prompt files only. Python code changes require a
|
||||
separate explicit task.
|
||||
7. Maintain a prompt evolution log for every iteration.
|
||||
8. If a prompt cannot improve a test after several small iterations, stop and
|
||||
analyze the root cause.
|
||||
9. Prompt changes must be generally applicable and must not special-case one
|
||||
transcript.
|
||||
10. If two consecutive prompt iterations fail to improve the current target,
|
||||
stop further prompt modifications and classify the root cause.
|
||||
11. If a prompt produces unexpected behavior, first verify whether the targeted
|
||||
gold test has an objectively unique ground truth.
|
||||
|
||||
Prompt Version 2 baseline:
|
||||
|
||||
- `decision_simple`: passing
|
||||
- `decision_deferred`: passing
|
||||
- `decision_none`: passing
|
||||
|
||||
Prompt Version 2 adds explicit support for process decisions where the group
|
||||
agrees to defer a substantive decision until additional information is
|
||||
available.
|
||||
@@ -0,0 +1,21 @@
|
||||
# decision_deferred
|
||||
|
||||
Tests that a process decision to defer a substantive decision is still extracted
|
||||
as a decision.
|
||||
|
||||
The transcript contains several candidate options for the weekly dashboard, but
|
||||
the group does not choose any of them. Instead, Mira explicitly says not to
|
||||
decide today and Jonas agrees that more input is needed.
|
||||
|
||||
Ground truth:
|
||||
|
||||
- No substantive dashboard schedule decision is reached.
|
||||
- One process decision is reached: the substantive decision is deferred until
|
||||
more information is available.
|
||||
|
||||
Typical LLM mistakes:
|
||||
|
||||
- Extracting Monday, Tuesday, or Friday as the chosen dashboard day.
|
||||
- Treating "I like shorter" as an approval.
|
||||
- Missing the deferral because the substantive decision is unresolved.
|
||||
- Treating "no decision today" as equivalent to no decision at all.
|
||||
@@ -0,0 +1,31 @@
|
||||
{
|
||||
"facts": [],
|
||||
"decisions": [
|
||||
{
|
||||
"decision": "The substantive decision is deferred until more information is available.",
|
||||
"evidence": "Mira: Okay, let's not decide this today. Jonas: Agreed, we need more input."
|
||||
}
|
||||
],
|
||||
"todos": [
|
||||
{
|
||||
"task": "Lea will bring Dana's feedback about the weekly dashboard next time.",
|
||||
"responsible": "Lea",
|
||||
"deadline": "next time",
|
||||
"evidence": "Lea: I will bring Dana's feedback next time."
|
||||
}
|
||||
],
|
||||
"questions": [],
|
||||
"positions": [
|
||||
{
|
||||
"speaker": "Jonas",
|
||||
"position": "Jonas prefers moving the weekly dashboard to Monday morning.",
|
||||
"evidence": "Jonas: I would prefer moving it to Monday morning."
|
||||
},
|
||||
{
|
||||
"speaker": "Lea",
|
||||
"position": "Lea thinks Monday is difficult for support.",
|
||||
"evidence": "Lea: Monday is rough for support."
|
||||
}
|
||||
],
|
||||
"technical": []
|
||||
}
|
||||
@@ -0,0 +1,19 @@
|
||||
Mira: We need to talk about the weekly dashboard.
|
||||
|
||||
Jonas: I would prefer moving it to Monday morning.
|
||||
|
||||
Lea: Monday is rough for support. We usually have backlog cleanup then.
|
||||
|
||||
Mira: Tuesday might work, but I am not sure.
|
||||
|
||||
Jonas: Or we keep it on Friday and just make it shorter.
|
||||
|
||||
Lea: I like shorter, but I need to check with Dana first.
|
||||
|
||||
Mira: Okay, let's not decide this today.
|
||||
|
||||
Jonas: Agreed, we need more input.
|
||||
|
||||
Lea: I will bring Dana's feedback next time.
|
||||
|
||||
Mira: Thanks, that will help.
|
||||
@@ -0,0 +1,21 @@
|
||||
# decision_none
|
||||
|
||||
Tests a true decision-negative meeting segment.
|
||||
|
||||
The transcript contains discussion, competing preferences, and possible options
|
||||
for the weekly dashboard. No participant approves an option, rejects an option
|
||||
on behalf of the group, assigns a follow-up, agrees to gather more information,
|
||||
or explicitly decides to defer the decision.
|
||||
|
||||
Ground truth:
|
||||
|
||||
- No decision was reached.
|
||||
- No process decision was reached.
|
||||
- No agreed next step was created.
|
||||
|
||||
Typical LLM mistakes:
|
||||
|
||||
- Treating a preference as a decision.
|
||||
- Treating a proposed option as the selected option.
|
||||
- Treating the topic change as an implicit deferral decision.
|
||||
- Creating an action item for Dana even though she is only mentioned as absent.
|
||||
@@ -0,0 +1,24 @@
|
||||
{
|
||||
"facts": [],
|
||||
"decisions": [],
|
||||
"todos": [],
|
||||
"questions": [],
|
||||
"positions": [
|
||||
{
|
||||
"speaker": "Jonas",
|
||||
"position": "Jonas prefers moving the weekly dashboard to Monday morning.",
|
||||
"evidence": "Jonas: I would prefer moving it to Monday morning."
|
||||
},
|
||||
{
|
||||
"speaker": "Mira",
|
||||
"position": "Mira thinks Tuesday might work better for support.",
|
||||
"evidence": "Mira: Tuesday might work better for support."
|
||||
},
|
||||
{
|
||||
"speaker": "Lea",
|
||||
"position": "Lea suggests keeping Friday and making the dashboard shorter.",
|
||||
"evidence": "Lea: Or we keep Friday and make the dashboard shorter."
|
||||
}
|
||||
],
|
||||
"technical": []
|
||||
}
|
||||
@@ -0,0 +1,25 @@
|
||||
Mira: We need to talk about the weekly dashboard.
|
||||
|
||||
Jonas: I would prefer moving it to Monday morning.
|
||||
|
||||
Lea: Monday is rough for support because backlog cleanup starts then.
|
||||
|
||||
Mira: Tuesday might work better for support.
|
||||
|
||||
Jonas: Tuesday is hard for sales, at least this month.
|
||||
|
||||
Lea: Or we keep Friday and make the dashboard shorter.
|
||||
|
||||
Mira: I am not convinced shorter solves the timing issue.
|
||||
|
||||
Jonas: I am not convinced Monday is actually a problem for everyone.
|
||||
|
||||
Lea: Dana might have a view, but she is not here.
|
||||
|
||||
Mira: We are circling now.
|
||||
|
||||
Jonas: Yes, I do not have anything else to add.
|
||||
|
||||
Lea: Same here.
|
||||
|
||||
Mira: Okay, let's move to the budget topic.
|
||||
@@ -0,0 +1,11 @@
|
||||
# decision_simple
|
||||
|
||||
Tests one explicit decision with clear agreement language.
|
||||
|
||||
The difficult part is separating the decision from nearby rationale about user confusion and from the non-decision statement that the copy can stay unchanged for now.
|
||||
|
||||
Typical LLM mistakes:
|
||||
|
||||
- Extracting the rationale as a separate decision.
|
||||
- Treating "copy can stay as it is" as a formal decision.
|
||||
- Losing the evidence that shows explicit agreement.
|
||||
@@ -0,0 +1,13 @@
|
||||
{
|
||||
"facts": [],
|
||||
"decisions": [
|
||||
{
|
||||
"decision": "The welcome email will be sent after account activation.",
|
||||
"evidence": "So are we agreed that the welcome email moves to after activation? Ben: Agreed. Cara: Yes, let's do that."
|
||||
}
|
||||
],
|
||||
"todos": [],
|
||||
"questions": [],
|
||||
"positions": [],
|
||||
"technical": []
|
||||
}
|
||||
@@ -0,0 +1,19 @@
|
||||
Anna: Before we leave the onboarding flow, can we settle the email step?
|
||||
|
||||
Ben: I still think the welcome email should go out after account activation, not before.
|
||||
|
||||
Cara: Yes, before activation it keeps confusing people.
|
||||
|
||||
Anna: So are we agreed that the welcome email moves to after activation?
|
||||
|
||||
Ben: Agreed.
|
||||
|
||||
Cara: Yes, let's do that.
|
||||
|
||||
Anna: Good. Then that is the decision for this release.
|
||||
|
||||
Ben: Separate note on the copy: I am not proposing any wording decision today.
|
||||
|
||||
Cara: Same here, no wording proposal from me.
|
||||
|
||||
Anna: Okay, then the only decision is the timing after activation.
|
||||
@@ -0,0 +1,14 @@
|
||||
# evil_meeting
|
||||
|
||||
Tests a deliberately difficult meeting with interruptions, corrections, topic switches, changed positions, absent referenced people, and near-decisions.
|
||||
|
||||
Every utterance is designed to trigger a common extraction failure. The meeting mentions Omar and Platform, but neither is a participant. It includes an explicit non-decision on migration and a real decision only on excluding FR-7 from Friday's batch.
|
||||
|
||||
Typical LLM mistakes:
|
||||
|
||||
- Extracting a migration decision even though the group says no migration decision today.
|
||||
- Assigning Dana a todo even though she retracts it.
|
||||
- Treating Omar as a participant or technical owner.
|
||||
- Claiming Platform approved something despite being absent.
|
||||
- Losing the correction from "old export" to "nightly CSV job" and from API export to CSV export.
|
||||
- Treating Dana's opinion about rollout appearance as a fact or decision.
|
||||
@@ -0,0 +1,73 @@
|
||||
{
|
||||
"facts": [
|
||||
{
|
||||
"fact": "Tenant FR-7 still used the nightly CSV job yesterday.",
|
||||
"evidence": "Alex: Fine. So fact: tenant FR-7 still used the nightly CSV job yesterday."
|
||||
},
|
||||
{
|
||||
"fact": "Omar is the customer contact, not the technical owner.",
|
||||
"evidence": "Bea: Omar is the customer contact, not a participant here and not the technical owner."
|
||||
},
|
||||
{
|
||||
"fact": "Platform is the technical owner, but nobody from Platform is in the meeting.",
|
||||
"evidence": "Chen: The technical owner is still Platform, but nobody from Platform is in this call."
|
||||
},
|
||||
{
|
||||
"fact": "Friday's rollout batch still includes DE-2 and NL-4.",
|
||||
"evidence": "Alex: Good. Back to the portal rollout. Friday's batch still includes DE-2 and NL-4."
|
||||
}
|
||||
],
|
||||
"decisions": [
|
||||
{
|
||||
"decision": "FR-7 is excluded from Friday's portal rollout batch.",
|
||||
"evidence": "the portal rollout note will say FR-7 is excluded from Friday's batch. Bea: Agreed. Excluded from Friday's batch. Chen: Yes, put that in."
|
||||
}
|
||||
],
|
||||
"todos": [
|
||||
{
|
||||
"task": "Check the FR-7 mapping table.",
|
||||
"responsible": "Bea",
|
||||
"deadline": "Thursday morning",
|
||||
"evidence": "Bea: Yes, I will check it by Thursday morning."
|
||||
}
|
||||
],
|
||||
"questions": [
|
||||
{
|
||||
"question": "Can FR-7 use the v2 mapping without a customer-side field rename?",
|
||||
"evidence": "Alex: Open question: can FR-7 use the v2 mapping without a customer-side field rename?"
|
||||
},
|
||||
{
|
||||
"question": "Can Omar confirm FR-7's preferred launch window?",
|
||||
"evidence": "Dana: Also, can Omar confirm their preferred launch window?"
|
||||
}
|
||||
],
|
||||
"positions": [
|
||||
{
|
||||
"speaker": "Dana",
|
||||
"position": "Dana wants to migrate FR-7 but recognizes the mapping table may not be clean.",
|
||||
"evidence": "Dana: I want to, but we do not know if the mapping table is clean."
|
||||
},
|
||||
{
|
||||
"speaker": "Bea",
|
||||
"position": "Bea changed her earlier position and now says not to migrate FR-7 until the mapping table is checked.",
|
||||
"evidence": "I said last week we should migrate it. I am changing that. Do not migrate until the mapping table is checked."
|
||||
},
|
||||
{
|
||||
"speaker": "Dana",
|
||||
"position": "Dana thinks excluding FR-7 makes the rollout look messy.",
|
||||
"evidence": "Dana: I personally think excluding FR-7 makes the rollout look messy."
|
||||
}
|
||||
],
|
||||
"technical": [
|
||||
{
|
||||
"subject": "French export failure",
|
||||
"statement": "The old nightly CSV job failed; the new exporter was not running on tenant FR-7.",
|
||||
"evidence": "The old export failed. The new exporter was not running on that tenant."
|
||||
},
|
||||
{
|
||||
"subject": "Export mapping tables",
|
||||
"statement": "The API export uses the v2 mapping table, while the nightly CSV job uses the legacy table.",
|
||||
"evidence": "the API export uses the v2 mapping table, the nightly CSV job uses the legacy table."
|
||||
}
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,61 @@
|
||||
Alex: Okay, quick pass on the portal rollout. Wait, before that, the French CSV export broke again.
|
||||
|
||||
Bea: It did not break again. The old export failed. The new exporter was not running on that tenant.
|
||||
|
||||
Chen: Sorry, when you say old export, do you mean the nightly job?
|
||||
|
||||
Bea: Yes, the nightly CSV job. Not the API export.
|
||||
|
||||
Alex: Fine. So fact: tenant FR-7 still used the nightly CSV job yesterday.
|
||||
|
||||
Dana: I thought Omar owned that tenant.
|
||||
|
||||
Bea: Omar is the customer contact, not a participant here and not the technical owner.
|
||||
|
||||
Chen: The technical owner is still Platform, but nobody from Platform is in this call.
|
||||
|
||||
Alex: Should we decide to migrate FR-7 today?
|
||||
|
||||
Dana: I want to, but we do not know if the mapping table is clean.
|
||||
|
||||
Bea: Also, I said last week we should migrate it. I am changing that. Do not migrate until the mapping table is checked.
|
||||
|
||||
Chen: So no migration decision today?
|
||||
|
||||
Alex: Correct, no migration decision today.
|
||||
|
||||
Dana: But we can decide one thing: the portal rollout note will say FR-7 is excluded from Friday's batch.
|
||||
|
||||
Bea: Agreed. Excluded from Friday's batch.
|
||||
|
||||
Chen: Yes, put that in.
|
||||
|
||||
Alex: Action item: Bea checks the FR-7 mapping table by Thursday morning.
|
||||
|
||||
Bea: Yes, I will check it by Thursday morning.
|
||||
|
||||
Dana: And I will message Omar after Bea is done.
|
||||
|
||||
Alex: Hold on, after Bea is done is not a date.
|
||||
|
||||
Dana: Fair. Then no task for me yet. I need Bea's result first.
|
||||
|
||||
Chen: Technical note: the API export uses the v2 mapping table, the nightly CSV job uses the legacy table.
|
||||
|
||||
Bea: Correct.
|
||||
|
||||
Alex: Open question: can FR-7 use the v2 mapping without a customer-side field rename?
|
||||
|
||||
Dana: Also, can Omar confirm their preferred launch window?
|
||||
|
||||
Chen: Omar can answer that, but again he is not in this meeting.
|
||||
|
||||
Alex: Good. Back to the portal rollout. Friday's batch still includes DE-2 and NL-4.
|
||||
|
||||
Bea: Yes, those two are unchanged.
|
||||
|
||||
Dana: I personally think excluding FR-7 makes the rollout look messy.
|
||||
|
||||
Alex: Noted as Dana's view, not a decision.
|
||||
|
||||
Chen: And please don't write that Platform approved anything. They are absent.
|
||||
@@ -0,0 +1,11 @@
|
||||
# facts_simple
|
||||
|
||||
Tests extraction of objective facts from a short status update.
|
||||
|
||||
The transcript includes a question and a technical statement, but no decision or todo.
|
||||
|
||||
Typical LLM mistakes:
|
||||
|
||||
- Treating "Good" as approval of a decision.
|
||||
- Turning "No decision needed today" into a decision.
|
||||
- Missing that the scanner gateway details are technical as well as factual.
|
||||
@@ -0,0 +1,32 @@
|
||||
{
|
||||
"facts": [
|
||||
{
|
||||
"fact": "The old scanner gateway is still running in aisle three.",
|
||||
"evidence": "Tom: The old scanner gateway is still running in aisle three."
|
||||
},
|
||||
{
|
||||
"fact": "Aisles one and two moved to the new gateway last week.",
|
||||
"evidence": "Tom: Yes. Aisles one and two moved to the new gateway last week."
|
||||
},
|
||||
{
|
||||
"fact": "The new gateway is handling live scans for receiving.",
|
||||
"evidence": "Iris: The new gateway is already handling live scans for receiving."
|
||||
}
|
||||
],
|
||||
"decisions": [],
|
||||
"todos": [],
|
||||
"questions": [
|
||||
{
|
||||
"question": "Is aisle three the only scanner gateway still left on the old gateway?",
|
||||
"evidence": "Elena: Is that the only one left?"
|
||||
}
|
||||
],
|
||||
"positions": [],
|
||||
"technical": [
|
||||
{
|
||||
"subject": "Scanner gateway rollout",
|
||||
"statement": "Aisle three remains on the old scanner gateway while aisles one and two use the new gateway.",
|
||||
"evidence": "The old scanner gateway is still running in aisle three. Aisles one and two moved to the new gateway last week."
|
||||
}
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,19 @@
|
||||
Elena: Quick status on the warehouse migration.
|
||||
|
||||
Tom: The old scanner gateway is still running in aisle three.
|
||||
|
||||
Elena: Is that the only one left?
|
||||
|
||||
Tom: Yes. Aisles one and two moved to the new gateway last week.
|
||||
|
||||
Iris: The new gateway is already handling live scans for receiving.
|
||||
|
||||
Elena: Good. Let's keep the rollout note factual.
|
||||
|
||||
Tom: No decision needed today.
|
||||
|
||||
Iris: Fine.
|
||||
|
||||
Elena: Anything else on warehouse?
|
||||
|
||||
Tom: No, that is all.
|
||||
@@ -0,0 +1,11 @@
|
||||
# facts_vs_positions
|
||||
|
||||
Tests separation of objective facts from personal opinions.
|
||||
|
||||
The transcript deliberately mixes numeric facts, named non-respondents, and subjective positions about rollout health.
|
||||
|
||||
Typical LLM mistakes:
|
||||
|
||||
- Treating Marta's opinion as an objective fact.
|
||||
- Treating Leo's optimism as a fact.
|
||||
- Extracting a rollout decision even though the group explicitly does not decide.
|
||||
@@ -0,0 +1,42 @@
|
||||
{
|
||||
"facts": [
|
||||
{
|
||||
"fact": "The pilot survey closed yesterday with 42 responses.",
|
||||
"evidence": "Hannah: The pilot survey closed yesterday with 42 responses."
|
||||
},
|
||||
{
|
||||
"fact": "The average pilot survey rating was 3.8 out of 5.",
|
||||
"evidence": "Leo: The average rating was 3.8 out of 5."
|
||||
},
|
||||
{
|
||||
"fact": "Northwind, Verdan, and Eastport did not respond to the survey.",
|
||||
"evidence": "Hannah: That part is true. Northwind, Verdan, and Eastport did not respond."
|
||||
}
|
||||
],
|
||||
"decisions": [],
|
||||
"todos": [],
|
||||
"questions": [
|
||||
{
|
||||
"question": "Why does Marta think the survey result is weaker than it looks?",
|
||||
"evidence": "Leo: Why?"
|
||||
}
|
||||
],
|
||||
"positions": [
|
||||
{
|
||||
"speaker": "Marta",
|
||||
"position": "Marta thinks the pilot survey result is weaker than it looks.",
|
||||
"evidence": "Marta: I think that is weaker than it looks."
|
||||
},
|
||||
{
|
||||
"speaker": "Leo",
|
||||
"position": "Leo feels the pilot is healthy.",
|
||||
"evidence": "Leo: I still feel the pilot is healthy."
|
||||
},
|
||||
{
|
||||
"speaker": "Marta",
|
||||
"position": "Marta thinks the rollout should slow down.",
|
||||
"evidence": "Marta: I disagree. My view is that we should slow down the rollout."
|
||||
}
|
||||
],
|
||||
"technical": []
|
||||
}
|
||||
@@ -0,0 +1,19 @@
|
||||
Hannah: The pilot survey closed yesterday with 42 responses.
|
||||
|
||||
Leo: The average rating was 3.8 out of 5.
|
||||
|
||||
Marta: I think that is weaker than it looks.
|
||||
|
||||
Leo: Why?
|
||||
|
||||
Marta: Because three enterprise customers skipped the survey entirely.
|
||||
|
||||
Hannah: That part is true. Northwind, Verdan, and Eastport did not respond.
|
||||
|
||||
Leo: I still feel the pilot is healthy.
|
||||
|
||||
Marta: I disagree. My view is that we should slow down the rollout.
|
||||
|
||||
Hannah: Let's capture both views and not decide rollout speed today.
|
||||
|
||||
Leo: Okay, that matches my notes.
|
||||
@@ -0,0 +1,11 @@
|
||||
# mixed_small
|
||||
|
||||
Tests a compact realistic meeting containing all major extraction categories.
|
||||
|
||||
The transcript includes facts, one explicit decision, one action item, one open question, one opinion, and a technical constraint.
|
||||
|
||||
Typical LLM mistakes:
|
||||
|
||||
- Applying the demo-only decision to production.
|
||||
- Treating Mateo's diagnosis as a fact instead of a position.
|
||||
- Creating a long-term normalizer decision even though it is explicitly open.
|
||||
@@ -0,0 +1,50 @@
|
||||
{
|
||||
"facts": [
|
||||
{
|
||||
"fact": "The staging import handled 12,000 rows last night.",
|
||||
"evidence": "Priya: The staging import handled 12,000 rows last night."
|
||||
},
|
||||
{
|
||||
"fact": "The staging import took 48 minutes.",
|
||||
"evidence": "Mateo: It finished, but it took 48 minutes."
|
||||
},
|
||||
{
|
||||
"fact": "The partner demo target is 30 minutes.",
|
||||
"evidence": "Lena: Yes, that is still the demo target."
|
||||
}
|
||||
],
|
||||
"decisions": [
|
||||
{
|
||||
"decision": "Address normalization will be disabled for the demo import only.",
|
||||
"evidence": "Can we agree to disable address normalization for the demo import only? Lena: Yes, for the demo import only. Mateo: Agreed."
|
||||
}
|
||||
],
|
||||
"todos": [
|
||||
{
|
||||
"task": "Update the demo import configuration.",
|
||||
"responsible": "Mateo",
|
||||
"deadline": "Friday noon",
|
||||
"evidence": "Mateo: I will do that before Friday noon."
|
||||
}
|
||||
],
|
||||
"questions": [
|
||||
{
|
||||
"question": "Whether a faster normalizer is needed after the demo.",
|
||||
"evidence": "Lena: And the open question is whether we need a faster normalizer after the demo."
|
||||
}
|
||||
],
|
||||
"positions": [
|
||||
{
|
||||
"speaker": "Mateo",
|
||||
"position": "Mateo thinks address normalization is the slow part.",
|
||||
"evidence": "Mateo: I think the slow part is address normalization."
|
||||
}
|
||||
],
|
||||
"technical": [
|
||||
{
|
||||
"subject": "Demo import configuration",
|
||||
"statement": "Address normalization is disabled only for the demo import; production imports keep full normalization.",
|
||||
"evidence": "for the demo import only. Production imports keep the full normalization."
|
||||
}
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,23 @@
|
||||
Priya: The staging import handled 12,000 rows last night.
|
||||
|
||||
Mateo: It finished, but it took 48 minutes.
|
||||
|
||||
Priya: The limit for the partner demo is 30 minutes, right?
|
||||
|
||||
Lena: Yes, that is still the demo target.
|
||||
|
||||
Mateo: I think the slow part is address normalization.
|
||||
|
||||
Priya: Can we agree to disable address normalization for the demo import only?
|
||||
|
||||
Lena: Yes, for the demo import only.
|
||||
|
||||
Mateo: Agreed. Production imports keep the full normalization.
|
||||
|
||||
Priya: Mateo, please update the demo config before Friday noon.
|
||||
|
||||
Mateo: I will do that before Friday noon.
|
||||
|
||||
Lena: And the open question is whether we need a faster normalizer after the demo.
|
||||
|
||||
Priya: Capture that, but no decision on the long-term fix today.
|
||||
@@ -0,0 +1,11 @@
|
||||
# question_simple
|
||||
|
||||
Tests extraction of an explicit open question.
|
||||
|
||||
The transcript also contains a non-task: Kai says he can ask finance but explicitly does not accept it as a task yet.
|
||||
|
||||
Typical LLM mistakes:
|
||||
|
||||
- Creating a todo for Kai despite his correction.
|
||||
- Missing that the group decides to leave the issue open.
|
||||
- Treating "support package" as enough information to answer the question.
|
||||
@@ -0,0 +1,27 @@
|
||||
{
|
||||
"facts": [
|
||||
{
|
||||
"fact": "The vendor invoice came in this morning.",
|
||||
"evidence": "Kai: The vendor invoice came in this morning."
|
||||
},
|
||||
{
|
||||
"fact": "The invoice line item says support package.",
|
||||
"evidence": "Kai: I don't know. The line item just says support package."
|
||||
}
|
||||
],
|
||||
"decisions": [
|
||||
{
|
||||
"decision": "The support-hours invoice issue will remain an open question for now.",
|
||||
"evidence": "Ruth: Fine. Let's leave it as an open question for now."
|
||||
}
|
||||
],
|
||||
"todos": [],
|
||||
"questions": [
|
||||
{
|
||||
"question": "Does the vendor invoice include the extra support hours from March?",
|
||||
"evidence": "Ruth: Does it include the extra support hours from March?"
|
||||
}
|
||||
],
|
||||
"positions": [],
|
||||
"technical": []
|
||||
}
|
||||
@@ -0,0 +1,19 @@
|
||||
Kai: The vendor invoice came in this morning.
|
||||
|
||||
Ruth: Does it include the extra support hours from March?
|
||||
|
||||
Kai: I don't know. The line item just says support package.
|
||||
|
||||
Ruth: Then that is still open.
|
||||
|
||||
Kai: I can ask finance, but I am not taking that as a task yet.
|
||||
|
||||
Ruth: Fine. Let's leave it as an open question for now.
|
||||
|
||||
Kai: Understood.
|
||||
|
||||
Ruth: Anything else on invoices?
|
||||
|
||||
Kai: No.
|
||||
|
||||
Ruth: Then next item.
|
||||
@@ -0,0 +1,11 @@
|
||||
# technical_simple
|
||||
|
||||
Tests technical extraction with a corrected diagnosis.
|
||||
|
||||
The transcript contrasts two possible causes: certificate expiry and runner configuration. The latter is confirmed as the cause.
|
||||
|
||||
Typical LLM mistakes:
|
||||
|
||||
- Reporting certificate expiry as the problem.
|
||||
- Creating a fix decision even though the group explicitly says no fix is decided.
|
||||
- Missing the distinction between fact and technical diagnosis.
|
||||
@@ -0,0 +1,33 @@
|
||||
{
|
||||
"facts": [
|
||||
{
|
||||
"fact": "The mobile build failed on the staging runner.",
|
||||
"evidence": "Sofia: The mobile build failed again on the staging runner."
|
||||
},
|
||||
{
|
||||
"fact": "The certificate is valid until October.",
|
||||
"evidence": "Nils: The certificate itself is valid until October."
|
||||
}
|
||||
],
|
||||
"decisions": [],
|
||||
"todos": [],
|
||||
"questions": [
|
||||
{
|
||||
"question": "Is the mobile build failing with the same error as yesterday?",
|
||||
"evidence": "Nils: Same error as yesterday?"
|
||||
}
|
||||
],
|
||||
"positions": [],
|
||||
"technical": [
|
||||
{
|
||||
"subject": "iOS staging build",
|
||||
"statement": "The iOS job fails during code signing because the runner uses the old keychain path.",
|
||||
"evidence": "The iOS job now fails during code signing. The problem is that the runner uses the old keychain path."
|
||||
},
|
||||
{
|
||||
"subject": "Failure classification",
|
||||
"statement": "The failure is a runner configuration issue, not a certificate expiry issue.",
|
||||
"evidence": "So it is a runner configuration issue, not a certificate expiry issue. Sofia: Exactly."
|
||||
}
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,19 @@
|
||||
Sofia: The mobile build failed again on the staging runner.
|
||||
|
||||
Nils: Same error as yesterday?
|
||||
|
||||
Sofia: No, different. The iOS job now fails during code signing.
|
||||
|
||||
Nils: The certificate itself is valid until October.
|
||||
|
||||
Sofia: Right. The problem is that the runner uses the old keychain path.
|
||||
|
||||
Nils: So it is a runner configuration issue, not a certificate expiry issue.
|
||||
|
||||
Sofia: Exactly.
|
||||
|
||||
Nils: We are not deciding the fix today.
|
||||
|
||||
Sofia: Okay.
|
||||
|
||||
Nils: Next item.
|
||||
@@ -0,0 +1,11 @@
|
||||
# todo_negative
|
||||
|
||||
Tests that vague "someone should" language is not an action item.
|
||||
|
||||
The transcript contains a real decision to keep the issue on the risk list, but no assigned task.
|
||||
|
||||
Typical LLM mistakes:
|
||||
|
||||
- Creating a todo with "someone" as owner.
|
||||
- Assigning Paula, Ravi, or Sam even though they explicitly do not take ownership.
|
||||
- Ignoring the explicit "no owner for now" correction.
|
||||
@@ -0,0 +1,22 @@
|
||||
{
|
||||
"facts": [
|
||||
{
|
||||
"fact": "The customer export takes longer than Paula expected.",
|
||||
"evidence": "Paula: The customer export takes longer than I expected."
|
||||
},
|
||||
{
|
||||
"fact": "There is no owner for the customer export issue for now.",
|
||||
"evidence": "Paula: Right, no owner for now."
|
||||
}
|
||||
],
|
||||
"decisions": [
|
||||
{
|
||||
"decision": "The customer export issue will stay on the risk list for now.",
|
||||
"evidence": "Sam: Then let's just keep it on the risk list. Paula: Right, no owner for now."
|
||||
}
|
||||
],
|
||||
"todos": [],
|
||||
"questions": [],
|
||||
"positions": [],
|
||||
"technical": []
|
||||
}
|
||||
@@ -0,0 +1,19 @@
|
||||
Paula: The customer export takes longer than I expected.
|
||||
|
||||
Ravi: Someone should probably look at it.
|
||||
|
||||
Sam: Yes, maybe after the release freeze.
|
||||
|
||||
Paula: I don't have capacity this week.
|
||||
|
||||
Ravi: Same here.
|
||||
|
||||
Sam: Then let's just keep it on the risk list.
|
||||
|
||||
Paula: Right, no owner for now.
|
||||
|
||||
Ravi: We can revisit it in planning.
|
||||
|
||||
Sam: Okay.
|
||||
|
||||
Paula: Next item.
|
||||
@@ -0,0 +1,11 @@
|
||||
# todo_simple
|
||||
|
||||
Tests one clear action item with responsible person and deadline.
|
||||
|
||||
The difficult part is not adding extra scope: Nina explicitly says she will not touch the layout.
|
||||
|
||||
Typical LLM mistakes:
|
||||
|
||||
- Adding a layout update as a task.
|
||||
- Dropping the deadline.
|
||||
- Turning Omar's request into the task evidence instead of Nina's commitment.
|
||||
@@ -0,0 +1,20 @@
|
||||
{
|
||||
"facts": [
|
||||
{
|
||||
"fact": "The beta signup page points to the old privacy note.",
|
||||
"evidence": "Omar: The beta signup page still points to the old privacy note."
|
||||
}
|
||||
],
|
||||
"decisions": [],
|
||||
"todos": [
|
||||
{
|
||||
"task": "Update the privacy link on the beta signup page.",
|
||||
"responsible": "Nina",
|
||||
"deadline": "Thursday noon",
|
||||
"evidence": "Nina: Yes, I will update the privacy link by Thursday noon."
|
||||
}
|
||||
],
|
||||
"questions": [],
|
||||
"positions": [],
|
||||
"technical": []
|
||||
}
|
||||
@@ -0,0 +1,19 @@
|
||||
Omar: The beta signup page still points to the old privacy note.
|
||||
|
||||
Nina: Yes, I saw that yesterday.
|
||||
|
||||
Omar: Can you update the link before the partner demo?
|
||||
|
||||
Nina: Yes, I will update the privacy link by Thursday noon.
|
||||
|
||||
Omar: Great. Nothing else on that page from my side.
|
||||
|
||||
Nina: I will only touch the link, not the layout.
|
||||
|
||||
Omar: Fine.
|
||||
|
||||
Nina: Then I have what I need.
|
||||
|
||||
Omar: We can move on.
|
||||
|
||||
Nina: Yes.
|
||||
Reference in New Issue
Block a user