The bug shouldn’t need you twice
One customer buys 100 credits. The notification arrives twice. Their balance becomes 200.
Imagine the payment notification for that purchase is delivered a second time, and the handler grants on every delivery. Someone has to notice, work out the required behaviour and explain the fix. Without a check that captures it, the next builder can need the same explanation again.
The valuable part of this fix is the behaviour you establish. Keep a check the next builder can run, instead of explaining the same failure again.
This is Part 2 of a field note on agent work and human review. The examples are authored programs that run locally; nothing touches Stripe or a real account.
The demo hid the failure
Every implementation passes the first delivery. The failure shows up only when the same purchase is delivered again, which is exactly the case a quick demo never exercises.
One purchase, buy_42, paid once for 100 credits. The same event, evt_01, is delivered again to a faulty handler that grants on every delivery. The balance is checked against the agreed 100 credits.
evt_01 → evt_01
Expected 100 · actual 200 · does not match
An authored program runs locally. This is a duplicate credit grant, not a duplicate charge. Nothing touches Stripe or a real account, and no real purchase is made.
A different message can still be the same purchase
Deduplicating event IDs is a useful first thought. But this business promise is about granting a purchase once, and an event ID is not necessarily the purchase being fulfilled. Three authored implementations show the difference:
- Every delivery treats another notification as another grant.
- Remember the event stops an identical replay, but not a second event for the same purchase.
- Remember the purchase keeps the account and purchase identity together.
Three authored implementations, not three model results. All check a mocked verification flag. Each column runs the same six declared cases in your browser; each cell shows the actual balance against the expected one.
| Required case | Every delivery | Remember the event | Remember the purchase |
|---|---|---|---|
| A first purchase | Pass100 / 100 | Pass100 / 100 | Pass100 / 100 |
| The same event arrives again | Fail200 / 100 | Pass100 / 100 | Pass100 / 100 |
| A new event ID, the same purchase | Fail200 / 100 | Fail200 / 100 | Pass100 / 100 |
| A genuinely different purchase | Pass200 / 200 | Pass200 / 200 | Pass200 / 200 |
| Verification failed | Pass0 / 0 | Pass0 / 0 | Pass0 / 0 |
| Same purchase, a conflicting amount | Fail300 / 100 | Fail300 / 100 | Pass100 / 100 |
The purchase-keyed version passes these six cases, including the conflicting amount.
Run only the first case and all three implementations pass. The other five required cases stay unrun for each version, so all three are incomplete. Unrun means unknown.
Passing these cases cannot prove a payment integration, the authenticity of inputs or the adequacy of the requirement (Stripe, Receive Stripe events in your webhook endpoint).
This is the question worth keeping: which identity names the thing we promised the customer? A fast model can patch the wrong identity very efficiently. The expectation needs to survive the model change.
The atomic operation still matters
The browser example is synchronous in-memory code. In a service, checking whether the purchase exists and then writing a grant must be coordinated. A separate “check then write” can race. The included SQLite example tests a transaction and uniqueness constraint with local connections and process replay (PostgreSQL 18, Transaction isolation).
BEGIN TRANSACTION claim (account_id, purchase_id) uniquely if claimed: record the credit grant if duplicate: return the existing outcome if conflict: stop and inspect COMMIT
No effect happens outside the transaction in this teaching model. Real work also needs authentic signature verification, authoritative purchase information, rollback, durable operations, observability and deployment review. The kit’s verified flag is not a signature verifier (Stripe webhook documentation).
The next builder should inherit the hard part
You already paid attention once to work out the required behaviour. Keep the input, expectation and result together. That is what can make a later repair cheaper.
- Save the case, so the mistake can happen again. Keep the same-purchase input sequence,
evt_01/buy_42thenevt_02/buy_42, and the expected 100-credit result: different events, one entitlement. A screenshot of a green first purchase cannot express this requirement. - Fix the implementation and leave the expectation alone. Let the agent repair the candidate while the acceptance expectation stays independent. The purchase-keyed function returns 100 for the same case, and the expectation has not changed. If the expectation is wrong, that is a separate decision rather than a quiet test edit.
- Check the other direction: a new purchase must still work. Rejecting every notification can stop duplicate grants. It also stops valid customers receiving credits.
buy_43is a new purchase, so the balance must become 200; rejecting everything is not a fix. Keep the positive case and the exception together. - Hand it on, then compare another builder. Another builder inherits the same cases, and another model can attempt the same task with the same check. Charge the fallback, repeated runs and your review time before calling the cheaper model cheaper.
acceptance / purchase-grant.test one purchase + repeat notification → expected balance: 100 another purchase → expected balance: 200 not: every event grants not: block every new purchase
A check is a maintained asset only while its contract remains right. This starter cannot prevent an agent with write access from rewriting its own judge. Keep the result you care about independent of whoever writes the next version.
What should the actual handoff contain?
| Keep | So you can answer |
|---|---|
| Candidate revision and exact input | Which work was checked? |
| Command, exit status and result | Can this run be repeated? |
| Passed, failed, skipped and unrun cases | What does green leave out? |
| Missing decision, authority or evidence | What specifically needs another person? |
| Complete effort and model cost | Was it actually worth delegating? |
The cheaper model might still cost more
The goal is less repeated discovery, not free checking. The test makes outcomes comparable; it does not guarantee savings. Use the same hypothetical batch model as Part 1 and count the work that comes back to you.
A hypothetical batch. Every input is invented. Equal accepted quality is assumed, not demonstrated. This batch is separate from the six-task opening in Part 1.
Time setup repays after 6 changes in this illustration.
CU = invented cost units. Valuing your time does not turn it into collected cash or salary savings.
The arithmetic, and how to make the advantage disappear
Let N be changes, b the old review minutes per change, q the share with useful repeatable checks, a the review still needed on those, e the minutes for an exception, h coordination per change and S setup minutes.
Hold = Nb; Hnew = S + N[qa + (1 − q)e + h]
At the defaults: 20 × 15 = 300 minutes. The alternative is 45 + 20 × (0.75 × 2 + 0.25 × 18 + 1) = 185. No part of this calculation establishes that the checks will achieve those review times.
Setup pays back in time only when b > qa + (1−q)e + h. If it does not, doing more of the same makes the proposal worse.
Call inputs per change: old 1.2 CU; alternative 0.2 + 0.02 for checks + fallback × 1.2 CU. These are not current vendor prices. The old review stays at 15 minutes in this model.
Both curves use the same axes. Solid = old route; dashed = alternative. Lines interpolate the arithmetic; tasks are counted in whole numbers.
| Route | Model / check calls | Ongoing valued time | Setup valued time | Total |
|---|---|---|---|---|
| Old | 24 | 300 | 0 | 324 |
| Alternative | 9.2 | 140 | 45 | 194.2 |
A sensitivity model with invented inputs, not measured savings.
A runnable failure is not a universal verifier
A useful check can be incomplete. Extra instructions can also make an agent do more work without improving the result (Gloaguen et al., Evaluating AGENTS.md). Keep the contract small enough to inspect.
The opportunity is to move repeated explanation and rediscovery into a check, give another builder a concrete result to reach, and spend your attention on a changed requirement, an important trade-off or a failure the system cannot resolve.
The countercase matters as much. If the task rarely repeats, the check is expensive to maintain, or the agent can optimise around a weak test, this can add process without value. Keep no-improvement runs in the evaluation.
Recent agent-engineering work makes reusable feedback worth testing. It does not make regression testing new, or this kit uniquely capable. The practical moment is the next bug you already understand, before you move on and the context disappears.
Keep the check, and keep its scope with it
Three invented handoff messages, “Implemented.”, “Looks good.” and “Works in the demo.”, say nothing about how the handler behaved on a retry. The check now shows a particular failure and a particular repair. Keep its scope with it. The next model gets to run it too.
Part 1, Your agents gave you homework, puts the human review bill back in the picture.
For coding agents: give the next builder the failure
The useful handoff is the input, the expected result and the command that distinguishes a fix from another confident answer. The instruction below is editable and needs no account or sign-up.
Nothing is sent or collected. Edit this before reusing it.
The kit is local teaching code plus instructions. It does not install a host skill, evaluate arbitrary code or protect a check the builder can rewrite.
Sources and further discussion
Each note says what the source shows and what it does not establish here.
- Stripe, Receive Stripe events in your webhook endpoint. Official documentation, checked 20 September 2026. Documents retries, duplicate handling and signature verification for actual endpoints. The local demo has a mocked verified flag, not Stripe integration. Its business purchase key is an explicit teaching requirement, not a universal Stripe rule. Documentation body and duplicate/verification guidance inspected.
- PostgreSQL 18 documentation, Transaction isolation. Version-pinned official technical documentation. Transaction and concurrency concepts relevant to coordinated validation and change. Browser state and the included SQLite example are not a PostgreSQL deployment or a proof of production correctness. Versioned documentation inspected.
- OpenAI, Harness engineering: leveraging Codex in an agent-first world. First-party team account, 11 February 2026. Describes environment, feedback and controls around a team’s agent-written internal product. Not a controlled trial of this article’s handoff or a general productivity multiplier. Page body inspected; no empirical replication.
- Gloaguen et al., Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents? Versioned research preprint, arXiv, 23 June 2026, version 2. Reports no general success improvement and higher inference cost in its tested repository-context settings. This is not a test of executable acceptance checks or the supplied starter. Results concern its tested agents, tasks and files. Version-2 abstract and revision record inspected; not rerun.
- Wikipedia, Signs of AI writing. Editorial advice from WikiProject AI Cleanup; advice page checked 20 September 2026. Promotional generality, repetitive rhetoric and broken sourcing as editing prompts. Not an authorship detector or a ban on every listed stylistic feature. Caveats and language sections reviewed.
- Nielsen Norman Group, The Layer-Cake Pattern of Scanning Content on the Web. UX research and practice account. Descriptive headings can support scanning. Does not establish these titles’ conversion rate or this article’s comprehension time. Article page inspected.
- Nielsen Norman Group, Progressive Disclosure. UX design guidance. Separate immediately needed information from secondary detail. Essential conditions remain local; collapsing prose is not evidence of easier learning. Article page inspected.
- Bret Victor, Explorable Explanations. Original design essay, 2011, postscript 2024. Make the author’s assumptions and model open to reader challenge. A design approach, not measured efficacy for this article. Original essay and postscript reviewed.
- W3C WAI, Animation from Interactions. Official WCAG 2.2 accessibility guidance. Allow nonessential interaction-triggered motion to be disabled. Reduced-motion checks alone do not certify accessibility. Official page inspected.
- W3C WAI, Pause, Stop, Hide. Official WCAG accessibility guidance. Control moving or updating content while other material is present. No universal reader-comfort claim; real users and assistive technology remain untested. Official page inspected.
- Eduardo Calvo / SmoothUI, AI Design Slop: Why AI-Generated UI Looks Generic — and the Fix. Vendor-authored practitioner critique, 24 June 2026. Prompt to challenge generic composition and one-shot review. Commercial opinion; its quoted statistics and causal claims are not adopted as findings here. Full public post reviewed for design critique, not an empirical quality benchmark.
Execution and attribution. AI assisted the prose and code. The figures are drawn from editable HTML, CSS and SVG; no image generation was used. All tasks, model-like answers, timings and costs are authored examples. No live LLM, Jev or payment service is connected.
These source notes are editorial summaries, not screenshots or endorsements. External page capture was blocked; no social post or engagement count was fabricated. Examples and source records are included in the downloadable pack.