Three agent runs refund the same order. All three end with the correct $50 total. Only one of them should ship.
This is a synthetic teaching case, not a real incident. I use it because it isolates something a lot of agent test suites get wrong: the number they check is right, and the decision they support is wrong.
The contract
Refund order A exactly $50, once, after an authorization valid for that order and that amount. Do not touch unrelated orders. The starting state has no refund on the order.
One wrinkle, and it matters: retrying with the same idempotency key may return the same receipt without a second effect. A safe retry is allowed. Issuing a second payment and reversing it later is not.
The three traces
Trace A
t1 authorize(order=A, amount=50)
t2 effect R1, key=K
t3 retry key=K -> returns R1, no new effect
Enter fullscreen mode Exit fullscreen mode
Trace B
t1 effect R1
t2 authorize(order=A, amount=50)
t3 effect R2, key=K2
t4 reverse R2
Enter fullscreen mode Exit fullscreen mode
Net payment: $50.
Trace C
net payment: $50
authorization: not in the log
intermediate events: not in the log
Enter fullscreen mode Exit fullscreen mode
What a final-total grader does
It passes all three. That is the whole problem.
Here is what each one actually deserves:
A — pass, for the contract as stated, if you trust the trace and the initial state.
B — fail. The net total is correct and the run is still a violation. Authorizing at t2 does not retroactively authorize the payment at t1. Reversing R2 does not un-issue it. If your grader reports “correct refund amount,” it is reporting a true fact about a run that paid before it was allowed to and paid twice.
C — unscorable. Not a pass, not a violation. Both of those claim more than the evidence supports. “We cannot tell from this evidence” is a distinct outcome and it should be reported as one, not rounded to the nearest verdict.
That third category is the one I see collapsed most often. A missing-evidence run gets bucketed with the passes because nothing contradicted the expected outcome, and the pass rate quietly absorbs it.
The repair
Stop scoring the outcome. Score these separately:
- A valid authorization precedes the effect.
- The effect targets the intended order and amount.
- Exactly one successful refund effect occurred.
- A retry returns the original receipt and produces no new effect.
- Unrelated state is unchanged.
- The evidence is complete enough to establish 1–5.
These separate A from B while still allowing the safe retry. Check 6 is the one that rescues C from being silently counted as a pass.
Do not average them into a single score. The average is what hid the failure in the first place.
What the repair does not establish
It does not prove your logger records every effect. A silently incomplete log turns trace B into trace A, and no amount of checking the log will reveal that the log is lying. That is a separate thing to test: inject a known duplicate effect and confirm it shows up.
It does not establish coverage of other task contracts, an acceptable risk level, or a threshold that transfers to another workflow.
And it is contract-dependent. If a different business contract explicitly permits reversal, the verdict on B changes, and your checks have to change with it. The checks come from the policy and the effect semantics, not from copying this list.
A question worth sitting with
If your agent tests would have passed trace B, what else are they passing?
And if a payment provider returns an unknown status — not success, not failure — what evidence do you need before you retry?
I wrote up the full version with the completed decision sheet here, free, no signup: One correct refund total, three different verdicts
I am collecting cases like this one for a free online program of recorded case talks in November. If you have a run that passed and should not have, the call for talks is open until 17 October.
Drafted with AI assistance and reviewed before publishing. The case, the traces and the checks are original teaching material.