← Back to writing

Your agents gave you homework

Six changes, all “done”. Before any can ship, you still have to work out what passed, what failed and what nobody checked.

Take an invented morning: six small SaaS changes from your agents, each marked “done”. A credit-purchase handler, a CSV export, an account check, an empty-state metric, a daily reminder and a cancellation message. Before any of them can ship, someone has to work out what passed, what failed and what nobody checked. That work comes back to you.

The opportunity I see is to make more of the checking happen before the handoff, and keep the decisions that really need you. A finished generation is not a finished task. The useful system does more of the repeatable work before asking for your attention.

This is a field note for people already building with agents. The examples are original teaching code, and no live model is called.

“Done” left six things to investigate

Each change arrived with the same completion message. Those messages are invented for this example, and on their own they don’t establish whether any required behaviour holds. A declared check route makes the next step explicit.

The failing case goes back to the builder, and the missing requirements come to you

An invented morning of six small SaaS changes, each reported as “done”. The four exact checks run in your browser when this page loads. The two missing requirements are authored fixtures.

ChangeDeclared checkResultNext step
Credit-purchase handlerRun the same verified purchase twice.Needs repair200 credits; expected 100.returned 200, expected 100Return the failing case to the builder.
CSV exportEscape a quoted customer name.Check passedComma and quote remain one field.returned "Lee, ""Jo""", expected "Lee, ""Jo"""Review the reported scope before release.
Account accessDeny a different account identifier.Check passedOther account → 403.returned 403, expected 403Review the reported scope before release.
Empty-state metricAn empty denominator stays undefined.Check passedNo observations → no percentage.returned null, expected nullReview the reported scope before release.
Daily reminderNo authoritative timezone provided.Needs a decision“9am” has no agreed timezone.Ask which timezone the requirement uses.
Cancellation messageNo permission to invent a refund promise.Needs a decisionRefund entitlement is unspecified.Ask what the approved policy promises.

3 narrow checks pass, 1 fails and 2 questions remain. No release is approved here.

Four tiny functions run in this browser. Two missing requirements are authored fixtures. No agents, users or performance study are connected.

Don’t buy the same kind of thinking for every job

A full-stack build mixes construction, interpretation, exact facts and authority. Sending every part to a large model, or every unresolved part to you, is a choice rather than a requirement. Don’t pay a model to do declared arithmetic.

One possible allocation: a design example, not a model ranking or a live routing model. A correct JSON shape, a semantic answer, a calibrated probability and permission are separate properties (Anthropic, Building effective agents; OpenAI, Structured Outputs).
QuestionWhere it goesBoundary
Did we grant twice?Exact check: code and authoritative records.Run the same verified purchase twice and compare the balance with the agreed 100 credits. A fluent explanation cannot substitute for that check.Code can repeat the wrong rule perfectly. Someone still has to establish the intended behaviour.
Does this draft promise a refund?Bounded judgement: Jev, a classifier or a constrained LLM.Ask a typed judge such as Jev, another classifier or a constrained LLM to compare the draft with the approved policy. Keep “unclear” as a possible answer when the evidence is missing.This is a proposed allocation, not a Jev result. An answer probability needs task-level calibration; it does not create authority.
Can you write the handler?Construction and explanation: an LLM with tools and feedback.An LLM can construct candidate code and explanations using tools. Give it the failing input, expected result, actual run feedback and permission boundaries, rather than only a request to “make it good”.A passing fixture is narrower than a working product. Keep untested scope and independent acceptance visible.
May this go live?Authority: application policy and the owner.The application enforces the allowed operation. A responsible person resolves missing requirements and reviews the scope when required. A confident model does not acquire that permission.Do not hand every exact check to a person either. Return a specific missing decision, with the evidence already assembled.
Where Jev fits, and what it does not buy you

Jev is TypeSafe’s bounded decision model. Vercel describes a contract with supplied state, declared questions and typed Choice, Score and Boolean answers with probabilities. That makes it a candidate for repeated semantic judgements, not the answer to every engineering task.

A constrained LLM or conventional classifier may already do the job. Compare them on your labels, permitted inputs, abstentions, errors, latency and full fallback cost. No comparative speed or quality claim is made for this kit.

An illustrative question contract, not Jev SDK syntax:

question: Does the draft promise a refund?
input: draft + the approved policy excerpt
answers: supported | contradicted | unclear
next step: review when evidence is missing

A judge can identify a possible contradiction. It cannot create an approved policy or certify its own evidence. Native TypeSafe documentation could not be opened in this research pass.

A failed check is more useful than an unnamed worry

The point is to stop making you rediscover what the agent can already test, and to keep the question precise when the answer really needs you. Here is the credit handler taken through four kinds of work:

  1. Task contract, before the build. Say what finishing means: outcome, data, permissions and missing decisions. For the credit handler, a verified purchase grants once and a new purchase still grants credits, so two notifications must not double the customer’s credits. The agent gets the outcome, not a vague request for “robust code”.
  2. Builder and executable checks, before the handoff. Let failure return to the builder with its input retained. The replay check on a repeated buy_42 returns 200 instead of the expected 100. The agent can inspect that exact case, repair it and rerun within the agreed budget. You do not need to repeat the same investigation first.
  3. Bounded judgement or clarification, when the rule is missing. “Daily at 9am” is missing a timezone, and a test cannot discover the business’s intended answer by itself. Ask whose timezone matters instead of inventing one and calling it verified. “Unclear” is a useful result when evidence is absent.
  4. An inspectable handoff, at the boundary. Give me the revision, the check command, the actual result and what remains open. If another builder takes over, it starts from that record. A proposal is not a deployed operation: approval and deployment are still separate decisions.

This route does not assume that an agent can resolve every failure, or that the author of a check should control its release gate.

Where it holds on your tasks, the next morning changes: fewer investigations from scratch, clearer decisions and checks another builder can inherit. That could create room for another feature, a customer conversation or an evening without repair work. It is an opportunity to measure, not a result this demonstration has established.

The available parts changed, and the acceptance problem didn’t

Better construction tools and specialised decision interfaces make more allocations possible. The useful question is which allocation removes work on your actual task. Three dated sources frame it:

  • 11 February 2026: build the feedback environment. OpenAI’s team account puts repository knowledge, checks and control around agent construction. It is a first-party account and evidence of a working approach in that team, not a productivity promise for ours.
  • 23 June 2026: more instructions can add cost without helping. The revised AGENTS.md study, a version-2 preprint and counterevidence here, found no general task-success improvement in its tested settings, with higher inference cost. Do not turn this article into a huge compulsory context file.
  • 16 September 2026: a bounded judgement need not be a long answer. Vercel introduces Jev’s typed decision interface, which adds a candidate for routing and assessment. It is a provider announcement, not a local benchmark. Whether it beats your existing model after review and fallback is still a test.

The next expensive failure is a chance to stop paying for the same investigation. Capture its input and expected behaviour while you understand them.

Can this give you part of the day back?

Move the assumptions. Count setup, the review still needed and the difficult cases. Lower token spend is not a win if it quietly buys more of your time.

Setup pays back only when each change needs less of your time than the old review

A hypothetical batch. Every input is invented. Equal accepted quality is assumed, not demonstrated. This batch is separate from the six-task opening.

Review every handoff from scratch300 min
Checks + remaining review + setup185 min

Old reviewRemaining workSetup

Your time across the whole batch115 minreleased under these assumptions
Calls + time valued at 60 CU/hour129.8 CU less

Time setup repays after 6 changes in this illustration.

CU = invented cost units. Valuing your time does not turn it into collected cash or salary savings.

The arithmetic, and how to make the advantage disappear

Let N be changes, b the old review minutes per change, q the share with useful repeatable checks, a the review still needed on those, e the minutes for an exception, h coordination per change and S setup minutes.

Hold = Nb; Hnew = S + N[qa + (1 − q)e + h]

At the defaults: 20 × 15 = 300 minutes. The alternative is 45 + 20 × (0.75 × 2 + 0.25 × 18 + 1) = 185. No part of this calculation establishes that the checks will achieve those review times.

Setup pays back in time only when b > qa + (1−q)e + h. If it does not, doing more of the same makes the proposal worse.

Call inputs per change: old 1.2 CU; alternative 0.2 + 0.02 for checks + fallback × 1.2 CU. These are not current vendor prices. The old review stays at 15 minutes in this model.

Cost as the number of changes increasesThe alternative begins with setup cost. It becomes cheaper only if its per-change total is lower. Values are invented.012224336548601530CU · calls + valued timeNumber of changes

Both curves use the same axes. Solid = old route; dashed = alternative. Lines interpolate the arithmetic; tasks are counted in whole numbers.

Current complete-cost breakdown · cost units
RouteModel / check callsOngoing valued timeSetup valued timeTotal
Old243000324
Alternative9.214045194.2

A sensitivity model with invented inputs, not measured savings.

For you, the gain would be less reconstruction: you open the result and the failed case, rather than search for what “done” was supposed to mean. For your agent, it is a useful next move: retry a known failure or ask one grounded question, with a budget and stopping rule to prevent an endless loop. For the business, it means comparing complete routes. Value model calls, setup and your work separately. Released capacity pays off only if it is useful.

A check is exact; a judgement has a distribution

The cost argument depends on which part is reliable. These two optional worked examples keep uncertainty, arithmetic and permission separate.

When should a semantic answer go to review?

These ten scores and outcomes are invented. A threshold selects some answers and sends the rest to review; it does not improve any individual answer. Calibration must be evaluated on relevant held-out records (Guo et al., On Calibration of Modern Neural Networks).

Selected6 / 10
Wrong among selected1 / 6
Sent to review4 / 10
Ten invented answers, their scores and whether each was correct
AnswerScoreOutcomeAt this threshold
A0.99correctAccepted
B0.97correctAccepted
C0.95correctAccepted
D0.92correctAccepted
E0.88wrongAccepted
F0.85correctAccepted
G0.80wrongReview
H0.70correctReview
I0.65wrongReview
J0.55wrongReview

In a separate decision model, wrong output costs 100 CU and review costs 2 CU. Assume the model’s correctness probability is calibrated and the reviewer is correct 98% of the time.

automatic: (1 − p) × 100
review: 2 + (1 − a) × 100
automatic is no worse when p ≥ a − 2 / 100

10.0 CU automatic expected loss versus 4.0 CU with review. Automatic is no worse at p ≥ 0.96 under these assumptions.

Wrongness has one loss here; real outcomes need differentiated costs. No probability authorises an otherwise forbidden action.

One attention calculation, without confusing it with truth

Positive masses [1, m, 1] become weights by dividing by their sum. With m = 2, weights are [0.25, 0.50, 0.25]. They mix the value vectors [1,0], [0,2], [3,0] into [1,1] (Vaswani et al., Attention Is All You Need).

This isolates an attention operation’s arithmetic. It is not a probability that the model’s sentence is true, nor a model of Jev’s architecture.

The original AI-engineering research has not been replaced by the webhook example. This notebook retains the distinctions used to decide which computations and judgements belong in a system.

Weighted sum: [1.000, 1.000]

One weighted contribution from each value
ValueWeightContribution
[1, 0]0.250[0.250, 0.000]
[0, 2]0.500[0.000, 1.000]
[3, 0]0.250[0.750, 0.000]

Measure the work left at the end

I care less about how many agents ran than how much useful work is left at the end. A smaller pile of unresolved decisions would be worth more than a longer list of confident completions.

Start with one task. Measure all the work before and after. Keep the route only if the result and the economics hold up.

Part 2, The bug shouldn’t need you twice, makes one recurring failure reusable.

For coding agents: a handoff you can inspect

You get the work, the evidence and the decision that still needs you. Your agent gets a route to check and repair its own output before asking for help. The instruction below is editable and needs no account or sign-up.

Nothing is sent or collected. Edit this before reusing it.

The kit is local teaching code plus instructions. It does not install a host skill, evaluate arbitrary code or protect a check the builder can rewrite.

Sources and further discussion

Each note says what the source shows and what it does not establish here.

  1. OpenAI, Harness engineering: leveraging Codex in an agent-first world. First-party team account, 11 February 2026. Describes environment, feedback and controls around a team’s agent-written internal product. Not a controlled trial of this article’s handoff or a general productivity multiplier. Page body inspected; no empirical replication.
  2. Vercel, TypeSafe AI’s Jev now available on AI Gateway. Primary hosting announcement, 16 September 2026. Describes state plus declared questions and Choice, Score and Boolean outputs with probabilities. Provider description, not an independent comparison. Direct TypeSafe docs and direct page opening were unavailable. No native fields, version, price or speed ratio asserted here. Primary publisher indexed text retrieved by web search. Direct open failed.
  3. Gloaguen et al., Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents? Versioned research preprint, arXiv, 23 June 2026, version 2. Reports no general success improvement and higher inference cost in its tested repository-context settings. This is not a test of executable acceptance checks or the supplied starter. Results concern its tested agents, tasks and files. Version-2 abstract and revision record inspected; not rerun.
  4. Anthropic, Building effective agents. First-party engineering guidance, 19 December 2024, current page accessed. Discusses workflows, routing, evaluation loops, tool interfaces and bounded autonomy. Practitioner patterns, not a universal optimum or demonstrated results for these examples. Page body inspected.
  5. Guo et al., On Calibration of Modern Neural Networks. Original research, PMLR, 2017. Distinguishes confidence estimates from empirical correctness and studies calibration. Classifier study, not evidence of Jev/LLM calibration on the proposed task. Numerical illustrations here are invented. Abstract and publication record inspected.
  6. OpenAI, Introducing Structured Outputs in the API. Primary product explanation, 6 August 2024. Schema-constrained output and value-level limitations. Historical product announcement; no current model availability or reliability score is repeated. Mechanism and limitations reviewed; no API calls.
  7. Vaswani et al., Attention Is All You Need. Original architecture paper, arXiv, 2017. Scaled dot-product attention and vector mixing. Toy weights in these articles are authored arithmetic, not a trained model or current architecture comparison. Publication record inspected; formula retained from the supplied atlas.
  8. Wikipedia, Signs of AI writing. Editorial advice from WikiProject AI Cleanup; advice page checked 20 September 2026. Promotional generality, repetitive rhetoric and broken sourcing as editing prompts. Not an authorship detector or a ban on every listed stylistic feature. Caveats and language sections reviewed.
  9. Nielsen Norman Group, The Layer-Cake Pattern of Scanning Content on the Web. UX research and practice account. Descriptive headings can support scanning. Does not establish these titles’ conversion rate or this article’s comprehension time. Article page inspected.
  10. Nielsen Norman Group, Progressive Disclosure. UX design guidance. Separate immediately needed information from secondary detail. Essential conditions remain local; collapsing prose is not evidence of easier learning. Article page inspected.
  11. Bret Victor, Explorable Explanations. Original design essay, 2011, postscript 2024. Make the author’s assumptions and model open to reader challenge. A design approach, not measured efficacy for this article. Original essay and postscript reviewed.
  12. W3C WAI, Animation from Interactions. Official WCAG 2.2 accessibility guidance. Allow nonessential interaction-triggered motion to be disabled. Reduced-motion checks alone do not certify accessibility. Official page inspected.
  13. W3C WAI, Pause, Stop, Hide. Official WCAG accessibility guidance. Control moving or updating content while other material is present. No universal reader-comfort claim; real users and assistive technology remain untested. Official page inspected.
  14. Eduardo Calvo / SmoothUI, AI Design Slop: Why AI-Generated UI Looks Generic — and the Fix. Vendor-authored practitioner critique, 24 June 2026. Prompt to challenge generic composition and one-shot review. Commercial opinion; its quoted statistics and causal claims are not adopted as findings here. Full public post reviewed for design critique, not an empirical quality benchmark.

Execution and attribution. AI assisted the prose and code. The figures are drawn from editable HTML, CSS and SVG; no image generation was used. All tasks, model-like answers, timings and costs are authored examples. No live LLM, Jev or payment service is connected.

These source notes are editorial summaries, not screenshots or endorsements. External page capture was blocked; no social post or engagement count was fabricated. Examples and source records are included in the downloadable pack.

Published 2026-09-12 · Updated 2026-09-28 · Source edition v6