← Back to writing

The check you didn’t know to ask for

Your agent can pass every test you wrote. Who notices the test you forgot? A working local example of keeping the check; discovery and human benefit are unmeasured.

Your agent can pass every test you wrote. Who notices the test you forgot? One customer books a workshop. The write succeeds, the reply disappears, and the screen says “try again”. The first booking already exists.

I want an agent that notices the awkward case before I spell it out. If the lesson transfers, the next project could benefit without someone having to make the same mistake first. The experiments below show both why that is worth testing and where the current machinery breaks.

Good answers need better questions. What this article offers is a runnable local example and a proposed learning loop. It does not establish expert-level AI review; discovery and human benefit are unmeasured.

A retry is not a second intention

The important question can be absent. A green checklist of UI, test and happy path says nothing about a concern it never included. Here is the one this example is built around.

  1. First attempt: the booking exists, and the customer doesn’t know. The database has one row. A timeout says nothing about whether that write happened.
  2. Retry: the same request creates a second booking. The naïve implementation repeats the effect. Your customer wanted one workshop, not two.
  3. The fix: recognise the request and return the original booking. Use the same request identity within the same account. Check the row count after reopening storage.
  4. The exception: another account’s booking must still work. A global duplicate key drops legitimate work. The fix has to preserve the account boundary, and a green test alone does not show that.
A lost reply hides a successful write

An animated replay of the supplied Python fixtures, one booking request at a time. Step through the four saved outcomes.

01 / WRITE COMMITTED

Request and retry sequenceThe first request creates a row. Its reply is lost. The client cannot infer failure from the timeout. ClientAPIDatabase team-a / r-1INSERTreply lost×
Observed rows1 booking
01team-a / r-1workshop

Committed, not acknowledged.The successful write survives a new connection.

Authored example; not a live HTTP call. The browser animates saved outcomes from the starter; it does not run SQLite.

Two explanations fit the second row. One intent ran twice, or these are independent requests. To distinguish them, compare account, request identity and stored effect.

The starter executes isolated SQLite operations. It reopens the database, repeats a request, tries a new request, crosses accounts, and rejects a changed payload reusing an existing identity.

Freshly rerun supplied implementations · local fixture only
ImplementationRetryOther accountResult
Naïve insertTwo rows (fails)Separate rows (holds)BLOCK
Global keyOne row (holds)Booking dropped (fails)BLOCK
Account-scoped keyOne row (holds)Separate rows (holds)PASS_SCOPED_CHECKS
Unknown contextNot testedNot testedDEFER

This is not an exactly-once guarantee for remote services, concurrent production workloads, payment effects or crash recovery. The supplied implementation and contract are authored examples, not independent model outputs. A hash identifies bytes; it does not authenticate truth.

The scoped receipt
{
  "candidate": "scoped",
  "checks": [
    {
      "expected": {
        "rows": 1,
        "same_id": true
      },
      "id": "retry_survives_reopen",
      "observed": {
        "rows": 1,
        "same_id": true
      }
    },
    {
      "expected": {
        "different_id": true,
        "rows": 2
      },
      "id": "new_intent_is_not_dropped",
      "observed": {
        "different_id": true,
        "rows": 2
      }
    },
    {
      "expected": {
        "different_id": true,
        "other_tenant_rows": 1,
        "rows": 2
      },
      "id": "tenant_scope",
      "observed": {
        "different_id": true,
        "other_tenant_rows": 1,
        "rows": 2
      }
    },
    {
      "expected": {
        "rejected": true,
        "rows": 1
      },
      "id": "conflicting_reuse_is_rejected",
      "observed": {
        "rejected": true,
        "rows": 1
      }
    }
  ],
  "contract_version": "lesson-1",
  "evidence_origin": "executed-local-sqlite-fixtures",
  "fingerprint": "96aa0eb3f5c3744c420281605ede5267b5595802d1d74920f2cb0b2401e21847",
  "production_release": false,
  "reason": "four local fixture invariants observed; not production approval",
  "scope": "shared-booking/local-sqlite",
  "status": "PASS_SCOPED_CHECKS"
}

Turn the hunch into something we can challenge

A hunch becomes useful when we can say what would change our mind. That means keeping the rival explanation and finding the discriminating observation.

A proposed learning loop keeps the question, its rival and the evidence
  1. Notice“What happens on a retry?” A proposed concern, not yet a defect.
  2. Challenge“Could these be different accounts?” Keep the legitimate exception.
  3. ObserveRepeat it. Reopen storage. Obtain evidence that distinguishes them.
  4. RetainKeep the test. Record the scope. Revise when later evidence disagrees.

A contradiction sends us back to the question rather than straight to another score.

The loop describes a design, not a demonstrated AI capability. The supplied code exercises a known lesson. The open research is whether an agent can propose an important one we never supplied.

The loop moves from meaning to observation to measurement. The meaning is “A retry shouldn’t book twice”, which raises its own questions: whose request, which operation, what legitimate exception? The observation is the same account, the same request, and a row count. Two rows support the concern; a different account changes it. The measurement asks whether the check helped on later cases, counting missed faults, false alarms, effort and regressions.

A person would put the distinction as “Retry this request. Don’t erase somebody else’s booking.” Context gives the rule its meaning. The retained lesson records it:

lessons/retry-one-intent.json

Concern
One intent, two stored bookings.
Exception
A different account is independent.
Observe
Reopen storage. Count rows and IDs.
Invalidate
The operation or scope changes.

The next agent runs python3 review.py --candidate scoped and gets four local checks passing. That is an inspectable result rather than “I think it works”.

That is the mixed qualitative–quantitative model. Context defines the question. Code tests explicit relationships. Statistics needs appropriate outcome data. A simulation can expose a bad assumption before a complete product exists; it cannot tell us how real users or models perform without observing them.

A prompt asks. A protected check can block. The host must require the approved check. The article cannot.

The possible payoff is useful judgement kept beyond this one review: fewer things only you remember to check, an agent that brings a question, a rival explanation and an observable test, and a repository of lessons that can change when the context does. The practical step is a retained check. Discovering new lessons is the research.

Another opinion is not another observation

Sixteen different states can display the same success message. Take four checks: durable save, one effect, correct account and usable recovery. Each either holds or fails, and all sixteen combinations can show a screen that says “Saved”.

The message alone has ruled out none of these possibilities. Each known-good observation halves what remains within the model: inspecting stored state leaves eight, counting stored effects leaves four, inspecting account scope leaves two, and inspecting recovery leaves one. Repeating the same Saved message leaves sixteen. Ask for evidence that separates the remaining explanations.

This is a noiseless finite teaching model. A remaining state is not proof that unmodelled faults, dishonest observations or human misunderstandings do not exist.

Model, assumptions and the full 16-state table
H(E) = { h in the declared model : h matches evidence E }
A claim is identified within this model only if
it has the same truth value in every h remaining in H(E).

Known-good observations in the worked trace: 16 → 8 → 4 → 2 → 1 compatible states. Repeating the same Saved message leaves 16.

No probability distribution is needed for this counting argument. The inputs are explicitly modelled and trusted. Empty compatibility should trigger model conflict, not universal correctness. This is a new exact teaching model, not a live model result.

Declared finite worlds; not sampled production data.
StateSaveOne effectAccountRecovery
1FailsFailsFailsFails
2HoldsFailsFailsFails
3FailsHoldsFailsFails
4HoldsHoldsFailsFails
5FailsFailsHoldsFails
6HoldsFailsHoldsFails
7FailsHoldsHoldsFails
8HoldsHoldsHoldsFails
9FailsFailsFailsHolds
10HoldsFailsFailsHolds
11FailsHoldsFailsHolds
12HoldsHoldsFailsHolds
13FailsFailsHoldsHolds
14HoldsFailsHoldsHolds
15FailsHoldsHoldsHolds
16HoldsHoldsHoldsHolds

Our learner knew the rule, not its missing context

Learning a rule from a supplied representation is not discovering a missing concept. In the retained v6 E601 counterexample, 0 of 64 answers are wrong in the original context. Flip the omitted context bit and every answer changes: 64 of 64 are wrong. The original-context control reuses training configurations. This is an adversarial finite example, not typical model performance.

A supplied scope guard defers all 64; a supplied correct relation handles them. Neither demonstrates autonomous discovery. Those are authored finite cases, not measured human or LLM error rates. The complete history and shared vocabulary remain in the archive. Auto-Rubric supplies related work on learning explicit criteria; it does not validate this workflow.

A working checker isn’t a discovery engine

EvidenceClaim it supports
Local tests runThese declared checks behave as tested.
Synthetic counterexampleThis failure can occur under these assumptions.
No live model studyDiscovery performance remains unknown.
No human outcome studyLess work for you remains a hypothesis.

Experienced reviewers are a comparator, not infallible ground truth. Compare ordinary strong-model review, the proposed process and human review using independently assessed outcomes.

The research unit is a cue, a contextual concern, a rival explanation, a discriminating observation and an outcome. The qualitative work defines the distinction. Mathematics makes its assumptions explicit. Statistics evaluates appropriate observations. Simulations test consequences before a complete product exists.

Test the full system against a strong ordinary agent and experienced reviewers. First hide the concern list. Then compare judgements on identical evidence. Finally test new cases after feedback. Measure useful discoveries, false alarms, regressions, human effort and completion, not just agreement.

A better judge cannot answer an unasked question

Start with 100 genuine issues in a hypothetical workload. Each probability is conditional on the previous stage, and evidence is obtained 90% of the time.

Expected issues out of 100. Costs and false positives are not represented.
StageStarting assumptions: 35% considered, judge 95%Only improve the judge, to 99%Improve discovery instead, to 80%
Concern considered353580
Evidence obtained31.531.572
Correctly confirmed29.931.268.4

The chain is 100 × P(D | I) × P(O | D,I) × P(C | O,D,I), where I is a genuine issue, D that the concern is considered, O that adequate evidence is obtained and C that it is correctly confirmed. This is the chain rule, not an assumption of independent stages. Changing an input does not demonstrate that the corresponding improvement is achievable or free. Counts are expectations, so they can be fractional.

Spend on the unknown, reuse the known

A cheaper question is useful only if the whole review costs less. At the checked rate, 1,000 packed questions carry an estimated metered input fee of $0.0021. One needless five-minute review has an illustrative time value of $3.33 at US$40 an hour, though that is not cash spent. Token prices are only one term; spend effort on useful evidence.

  • Discover (“What could this flow be missing?”): a strong model with tools.
  • Interpret (“Does this message imply completion?”): a validated semantic judge.
  • Verify (“How many rows survived the retry?”): a query or executable test.

Jev, an SLM, a classifier or a frontier model may fit interpretation. Compare the actual task, not model size alone (R1).

Send the context once instead of 1,000 times. In US dollars at the price checked on 20 September 2026, repeating the context costs $0.42168 and packing it costs $0.00210 for 1,000 questions:

(10,000 + 1,000 × 40) × $0.042 / 1,000,000
= $0.00210

That is 200.8× fewer billed input tokens, not better judgement. The assumed state is 10,000 tokens and the assumed complete question 40 tokens. The published Jev input rate is $0.042 per million tokens, and output is free. Its documented total and state-plus-longest-question limits are 64,000 and 32,000 tokens. Packing repeats state when a new batch is required. These estimates exclude real tokenisation, payload overhead, retries, observation and human work (R1).

No live Jev call occurred. The comparison is two input-packaging strategies at one provider’s rate, not Jev versus a frontier model, and not evidence that 1,000 worthwhile questions exist. A zero count means no semantic calls, not that code can solve any semantic problem.

The upkeep still has to earn its place

Keeping the check costs 55 minutes of setup and upkeep. In the starting case each reuse saves 4.85 minutes net, so time turns positive after 12 uses. Over 30 relevant future reviews, in this illustrative and unmeasured model:

ScenarioHuman minutes over 30 reviewsMetered spend
Starting case90.5 potentially saved$8.01 less
Model included: the original model’s marginal charge is zero90.5 potentially saved$0.09 more
Too many flags: 1.6 false flags per item at 6 minutes each, 4 minutes of review after, 160 minutes of setup388.0 extra required$8.01 less
Faster, less reliable: 0.1 minutes of review after, escaped faults up from 3% to 4%81.5 potentially saved$8.01 less

The time model is Time saved = N[(r₀ − r₁) − f·t − (e₁ − e₀)·k] − F. N is uses, r is human review time, f is false flags per item, t is minutes per flag, e is escaped-fault frequency, k is rework time and F is setup and upkeep. The starting case uses 6 minutes of review before and 0.7 after, 0.15 false flags per item at 3 minutes each, e₀ = e₁ = 3% and k = 90 minutes. Categories must not double-count work.

Metered spend is separate from time value. The example uses $0.30 per original model review, and $0.0031 per replacement pass plus a 10% return to the original model. Those are invented workflow prices, not published frontier pricing. When the original model’s marginal charge is set to zero, the extra service costs more cash. Fixed subscriptions, quotas, hosting and harms beyond rework are not valued here.

A positive time result is not acceptance. In the faster scenario the escaped faults rise, so it is not a reliability improvement. Task mix, consequences and the real comparison need measurement.

A cheaper building block, not a closing window

These are dated sources, not a countdown, and not screenshots, endorsements or evidence that this workflow is faster.

Use your next costly correction as the test. That is the timing argument. I have no evidence this opportunity expires. If useful lessons transfer, a small team could keep more of its hard-earned judgement. A growing pile of noisy checks would do the opposite.

These sources describe available components and practices: tool-using agents, executable feedback and cheap bounded evaluation. They do not establish that our workflow improves productivity, learns human expertise or creates a time-limited commercial advantage. No market deadline is claimed.

Our work is strongest on authored, repeatable distinctions. The research challenge is discovering new ones from independent evidence, preserving their exceptions, and improving later outcomes. Novelty for rubric learning is not claimed. Model subscriptions and local constraints can overturn any small-model cost advantage.

A social post or a news screenshot would establish what was posted, not the claimed effect. The source-capture failure log is included in the package; no fake tweets, engagement counts, news mastheads or publisher screenshots were constructed.

If the lesson transfers, you don’t have to find it first

An agent checking a booking might learn to question premature success elsewhere: an upload, a queued job, a tool saying it finished. That could move some experienced lines of questioning into the workflow.

The condition matters: a familiar pattern is a hypothesis. Each new context needs evidence and its own exceptions. The current lab has not measured autonomous discovery, expert equivalence or that human payoff.

Give the agent a question worth keeping

  1. Find one consequential concern the current tests miss.
  2. Keep a plausible reason the behaviour is acceptable.
  3. Run the observation that distinguishes those possibilities.
  4. Save the scope, exception, check, and observed result.
  5. Report unresolved questions. Don’t rewrite the acceptance contract.

Local demo; Python standard library; no API key.

No merge, deployment, spending or permission changes are authorised. What should come back is a fixture result, not a release certificate:

Changed
Account-scoped retry handling.
Observed
Four local checks; current source fingerprint.
Unresolved
Remote effects, production concurrency, unseen concerns.
Read the full brief, command and enforcement boundary
# Keep the lesson, not just the patch

Use this on the next relevant coding task. Work in a branch or sandbox; do not deploy, merge, spend money, change permissions, or weaken existing checks.

1. Read the task and existing acceptance criteria. State the user outcome in one sentence.
2. Identify one consequential contextual concern that the current checks might miss. Use the actual task, not a generic list. Write a plausible reason the behavior could be acceptable. If there is no useful additional concern, say so.
3. Propose the cheapest observation that distinguishes those explanations. Run exact checks in code. Use a semantic model only where interpretation is needed; keep its version and uncertainty. Missing access means "not observed", not "passed".
4. On a confirmed recurring issue, propose a small lesson file: context, concern, exception, runnable check, expected observation and invalidation condition. Preserve valid alternatives. Never quietly rewrite the acceptance contract to make the implementation pass.
5. Run the agreed check against the current revision. Return the actual command, observed result and scope. A passing statement in your own prose is not a test result. Leave undiscovered mechanisms and user-comprehension claims explicitly untested.
6. End with: changed / observed / unresolved / time and metered usage. Ask for human review of a changed contract or a consequential unresolved decision.

For the supplied local booking example only:
`python3 review.py --candidate scoped --output receipt.json`

For a real repository, propose the smallest adaptation to its existing tests. Do not assume the demo checks verify that repository. Have the host make the approved check required and protect the checker/contract; this instruction alone cannot enforce anything.

python3 review.py --candidate scoped --output receipt.json

Run inside the extracted starter directory. For your own project, propose the smallest adaptation of its real acceptance tests. Protect the checker and contract separately; validate the observation channel; require the current result at the host. Giving an agent text alone cannot enforce this policy.

None of them proved how the booking survived a retry

Imagine three reviews approving the button, the copy and the happy path. Each could be right about the thing it checked. None of them proved how the booking survived a retry.

I want to keep the missing question when I find it, and test whether the agent can learn the next one. Those three reviews are imagined. The research question is real; the full capability remains unproved.

Sources and further discussion
  1. TypeSafe, model reference. R1 · Provider documentation · current price checked 20 Sep 2026. Text-only bounded evaluation; input pricing and request limits inform the packing calculation. Provider capabilities and fees do not establish our task quality, actual token counts or full cost.
  2. OpenAI, Harness engineering. R2 · First-party engineering account · 11 Feb 2026. Describes making repository knowledge, constraints and feedback useful to coding agents. An account of one development setting; not a causal productivity evaluation of this article’s workflow.
  3. Anthropic, Demystifying evals for AI agents. R3 · First-party engineering guidance · Jan 2026. Supports distinguishing outcomes and relevant constraints from an agent’s narrative or one canonical trace. Guidance is not comparative evidence of the supplied starter’s real-world benefit.
  4. Auto-Rubric, arXiv 2510.17314v2. R4 · Primary research preprint · version posted 5 Feb 2026. Related work on learning explicit criteria from contrastive examples; motivates testing discovery and transfer separately. Related rubric-learning results are not validation of this workflow. No local reproduction or live-model comparison was performed.

These informed the original interactive edition’s format:

  1. NN/g, The layer-cake pattern of scanning. D1 · Original practitioner research synthesis · 4 Aug 2019. Headings carry the conclusions; readers can scan for meaning before committing to methods. No claim that this edition has been eye-tracked or proven faster to comprehend.
  2. NN/g, Progressive disclosure. D2 · Practitioner HCI guidance · checked 20 Sep 2026. Keep scope and consequences visible; put derivations and complete evidence behind explicit labels. Hiding essential caveats would undermine the argument. Disclosure is not an excuse to remove them.
  3. Bret Victor, Explorable explanations. D3 · Original design essay · checked 20 Sep 2026. Manipulations expose consequences of assumptions. Reading does not require clicking every control. Design precedent, not measured comprehension improvement on this page.
  4. Wikipedia, Signs of AI writing. D4 · Community editorial guidance · checked 20 Sep 2026. Remove unsupported significance claims, generic restatements and decorative rhetoric. Keep AI assistance disclosed. This is a descriptive project advice page, not an authorship detector or a universal prose rule. It notes limitations for recent models.
  5. Interrogating Design Homogenization in Web Vibe Coding. D5 · Primary research preprint · 13 Mar 2026. Motivates task-specific visual relationships rather than a generic generated dashboard. A sociotechnical analysis, not a score or detector proving an interface is original.
  6. Paul Bakaus, Impeccable by Design. D6 · Practitioner essay; author has commercial interests · checked 20 Sep 2026. Removing fashionable AI design patterns does not create good design; inspect actual hierarchy, states and craft. Opinion and practitioner experience, not a controlled evaluation. The current page is not certified by the author.
  7. W3C, Animation from Interactions. D7 · Official accessibility explanation · checked 20 Sep 2026. Provide reduced motion, a manual motion control, finite playback and static equivalents. These controls alone do not establish WCAG conformance.
  8. W3C, Reflow. D8 · Official accessibility explanation · checked 20 Sep 2026. Test narrow widths, expanded disclosures and increased text size; contain wide tables. Automated browser checks do not replace assistive-technology and human evaluation.
  9. Communicating with Interactive Articles. HCI · Research synthesis / primary article · 11 Sep 2020. Use state changes and controllable models; include manual navigation and a static reading route. Animation can hinder as well as help. Desktop studies do not establish mobile performance here.

Evidence status and inherited research. The v8 publishing archive, including its earlier scientific history, is preserved unchanged. The Python starter and its four local candidate runs were rerun for this article. The calculations here are scenarios, not model benchmarks. No live Jev/SLM/frontier evaluation, human study or production deployment occurred.

The supplied finite demonstrations are not externally validated expertise. Real user comprehension, useful discovery, transfer, reduced supervision and complete-cost improvements remain unmeasured. Publishing checks do not establish those outcomes.

Six terms, only where the distinction matters. The 156-term inherited glossary remains in the research package; these six are enough for the main argument.

Concern
A possible problem, not a finding.
Rival explanation
Why the same behaviour might be fine.
Discriminating observation
Evidence separating those explanations.
Scope
Conditions under which the check applies.
Verification
Conformance to the declared rules.
Validation
Usefulness for the intended real task.
Published 2026-09-19 · Updated 2026-09-27 · Calvin Kennedy · AI-assisted research and writing. Prior scientific results preserved · No live-model or human-performance claim.