It says A$290.
You asked for A$29.
Check the displayed total against an independent requirement.
Design agents / the work after “done”
Give it the checks. Get back the page, the evidence, and the exceptions.
A reusable brief and a runnable example. Not another instruction to “make it better”.
You sell an invoice template.
One purchase. A phone buyer. No subscription.
“Done”
Payment confirmed.
Even when declined.
You still have to find the mistakes.
Checked against the brief
Payment declined.
Try another card.
The agent returns the result and the checks.
01 / The problem
The price is wrong. The pay button runs off the phone. A declined card gets a success message.
I do not want to spend the evening checking whether a generated checkout obeys the brief. I want the agent to bring back the answer to those questions. Start with the page a customer would actually use.
A reusable template for your next invoice.
A$29. One purchase.
Your download follows a confirmed payment.
Compare the amount. No model judgment is needed.
Check the displayed total against an independent requirement.
Open the narrow layout. Check the page and the action.
Exercise the failure state. Keep the next step usable.
Keep the approved visual direction. Attach the current result.
Rule-driven teaching illustration. No card is charged. The separate local fixture in the kit can actually be run.
02 / The steps
Write the requirements once. Make the agent return current evidence, not another confident paragraph.
A$29 once.
Fits the phone.
Decline means retry.
Exact amount.
Narrow viewport.
Failed payment.
Keep the style.
Bound the attempts.
Run checks again.
What changed.
What passed.
What is still unknown.
This is not automatic publishing.
Local review is the end of this loop. Production access and release approval remain separate.
The project began with mathematical design models, then tested browser evidence, qualitative findings, mock judges, failure cases and bounded repair tools. The checkout is one way into that work, not the whole research programme. The broader question is whether checks can catch important failures without erasing legitimate design freedom.
Each result belongs to an artifact, requirement set, observer and context. Missing or unrun checks are not passing checks. A later unsuccessful recheck blocks reuse of an older success until a new valid evaluation is available.
The v7 prototype exposes nine local JSON-lines tools, not an MCP server. Deterministic policies exercised its interface; no real model integration or human study was run. The small checkout kit is a separate fixture, not the entire broker.
Hard requirements define feasibility. Typography, hierarchy, brand suitability and preference still need appropriate models and evidence. A high aesthetic score cannot compensate for a failed required payment-state check.
03 / Use the maths
Calculate what fits. Use the model for the decisions the calculation does not settle.
8 px short. Stack the row.
Compare A$29 with the requirement.
Exercise the actual layout and state.
Propose a change. Recheck its consequences.
Observe understanding, preference and use.
Live numerical model; the diagram is scaled to fit this article. It is not a browser measurement of your site. There is no universal “good design” equation.
W is the viewport width, m the side margin, p the price block’s minimum width, g the gap and b the button’s minimum width. At 327px a rigid row is one pixel too wide. At 328px it fits under these assumptions. Real text, focus treatments and localisation still need rendering.
For equal-width cards, n selects the greatest permitted column count that fits. A zero means infeasible under the chosen assumptions; it is not permission to shrink text until the page passes.
CSS variables or Tailwind theme variables can represent the same semantic tokens. Repeated components then share spacing, colours and control dimensions. A shared mistake can also spread everywhere. Updating a token should therefore trigger checks on its consumers. Tailwind supplies a mechanism, not proof that the choices are right. [T]
The embedded notebook retains interactive line-breaking optimisation, Fitts’ pointing index, correlated-judge errors and review economics. These are different models with different certainty. Exact geometry is not measured usability, and mock judgment is not evidence of human preference.
04 / Keep the freedom
A strict checker can catch faults by rejecting everything unfamiliar. That is not the result I want.
Approved style. Faint action.
The action is visible. The style stays.
Authored styling illustration. In the earlier v7 replay, targeted policies preserved 34/34 permitted focus styles; a reset control preserved 32/34. Those are project cases, not general model performance.
The distance must represent meaningful changes, not reward the smallest edit at any cost. Necessary behaviour changes can justify a larger patch. A narrow colour change that hides a symptom without fixing its cause is not a good repair.
The research found that an invisible parent could make a valid focus treatment invisible too. Repairing the shared cause before replacing the focus style preserved the approved design. The rule was authored; it was not a generally learned causal model.
v6 also found implementation-specific checks that blocked legitimate alternatives. The failure corpus and counterexamples remain in the archive. The objective is to detect consequential failures without erasing legitimate design freedom.
05 / What comes back
Which page? Which tests? What failed? A later edit cannot borrow an earlier pass.
This is evidence for revision 7. Not authorisation to publish.
STALECurrent page: revision 7.
The recorded checks apply to this version.
Ready for local review.
Illustrative state model. Signatures and hashes establish identity and provenance under their assumptions; they cannot prove that the observation captured the right fact.
Recorded v7 policy replay
Both deterministic policies made 252 tool calls and 60 evaluations across the same 36 authored scenarios. These are JSON byte counts. They cannot be substituted for tokenizer counts, API prices or human review time.
The v7 representation retains required failures, unknowns, coverage and identity. The report records 13,116 finite combinations with no mismatch between full and compact acceptance predicates. That establishes equivalence for the specified predicate, not sufficiency for every repair or human judgment.
When two designs produce the same observations but need different verdicts, a judge restricted to those observations cannot always distinguish them. The earlier lab demonstrated a missing-observation problem with an invisible action. More agreement on the same incomplete observation does not supply the missing fact.
Jev-style interfaces and frontier-agent policies were mocked in this research. Real-model capability, calibration, adversarial robustness and human outcomes remain unmeasured. The original code, raw results and limitations are preserved, not rerun as new evidence for this article.
06 / The payoff
Reuse can repay setup. A one-off page or noisy checker can make the work slower.
Under these assumptions. Change them; the answer can reverse.
Illustrative time value: A$105. Not a forecast or realised cash saving.
Hypothetical, editable inputs. Assumes comparable work quality. Missed defects, conversion, provider pricing and customer losses are not modelled here; lower review time alone is not a reason to ship.
N is the number of changes, r₀ and r₁ the review minutes before and after, S the one-off setup, u upkeep, h the value assigned to an hour and c incremental cash per change. Setup is in minutes, not money. Only convert it to money once.
With these assumptions, positive net value begins at change 7.
The strict break-even count is floor((Sh/60) / per-change benefit) + 1 when per-change benefit is positive. Otherwise there is no positive break-even within this model. A subscription may make an extra call cost no additional cash while still consuming time or allowance. Metered use, quotas and review labour are different resources.
Reuse only holds if the checks remain relevant. Changing the task, content, design system or dependency graph can require more coverage. The deeper notebook includes calibration, correlated error and review-risk models; their probabilities are synthetic, not plug-in estimates for your agent.
07 / Use it
Start with the supplied failing and passing examples. Then map the checks to your page.
Read the generated brief below before using it.
08 / What is known
The code can test the contract. People still have to validate the contract—and the experience.
Browser failures and finite-state counterexamples. Targeted scripted repairs. Compact evidence retaining a specified decision. Raw records are in the research archive.
The checkout story, style controls, layout budget and receipt. These explain mechanisms; they are not new frontier-model evaluations.
Less inspection, lower accepted-delivery cost and more freedom to explore. Real-agent and human outcomes still need to be measured.
The Codex release described browser inspection and attached screenshots. That makes browser evidence a practical handoff ingredient. [C]
The Claude Design announcement described reusable design systems and a coding handoff. It did not validate this kit. [D]
Dated first-party capabilities, not screenshots, endorsements, pricing claims or an expiring offer. The reason to try this is your next repeated task—not a manufactured countdown.
Layout constraints, browser testing and agent tool design have substantial prior art. The project’s narrower research question is whether we can give agents useful evidence while detecting failures and preserving legitimate design variation. Global novelty and publishable superiority are not established.
Explore typography, Fitts’ law, judge correlation, stale evidence and deeper maths. The original v7 code, qualitative/quantitative models, tests, source ledgers and subsequent article history remain in the ZIP.
The reason to bother
Give the next agent one requirement it can test. Ask what passed, what changed and what it could not check.
In this lab, a layout could fit, the mock judges could agree, and the checks could pass.
None of them proved how well a person could use the result.
That is the next test. Until then: keep the useful checks, keep the gaps visible, and keep your judgment.
One last distinction: the price, layout and decline tests pass. What can the agent honestly claim?
Research informs the mechanisms and editorial choices. It does not certify this article’s comprehension or aesthetics. Source notes distinguish evidence, design advice and practitioner opinion.
First-party release, 15 September 2025. Browser inspection and screenshot evidence. Historical capability, not performance or pricing evidence for this kit.
First-party announcement, 17 April 2026. Design systems and a coding handoff. Not independent validation of this research.
A mechanism for semantic token-backed styling. No Tailwind benchmark was run for this article.
NN/G research interpretation: descriptive, visually distinct headings help scanning. The application here is unvalidated; no reading-speed multiplier is claimed.
NN/G guidance: primary content stays accessible; advanced detail is available on demand. Critical limitations remain outside the drawers. This is a principle, not a timed comprehension law.
Hohman, Conlen, Heer and Chau, 2020. Synthesises representations, interaction and learning. Animation is not uniformly better than static figures; no universal layout solution.
Bret Victor: let a reader manipulate assumptions and see their consequences. Design precedent, not empirical proof for this article.
NN/G guidance on explanatory motion and continuity. Our timings and transitions are authored choices, not validated optimums.
W3C guidance on disabling nonessential interaction-triggered motion. Implementing a motion control is not complete WCAG conformance.
Editorial review for vague significance, promotion and unsupported attributions. Contextual guide, not a detector or proof of human authorship.
Jeff Humble’s practitioner critique of competing emphasis and decorative patterns. Opinion, not an authorship classifier. Claims about training causes are not adopted.
Existing design instructions emphasise deliberate, context-specific visual direction. A prompt or skill is not a measured outcome.
Official guidance on verification criteria and runnable feedback. Instructions are not a security boundary.
Defines visibility and actionability checks; opacity-zero can satisfy its visibility definition. That definition is not a complete model of human-visible content.
Writing review: use concrete claims and attributable evidence, not decorative significance or invented urgency. Wikipedia’s AI-writing guide is contextual advice, not an authorship detector. Design review: explanatory objects earn their space; style stereotypes are not proof of provenance.