← All writingCalvin Kennedy Writing

Design agents / the work after “done”

The AI finished.
You’re still
working.

Give it the checks. Get back the page, the evidence, and the exceptions.

A reusable brief and a runnable example. Not another instruction to “make it better”.

ONE CHECKOUT. SAME BRIEF.Illustration
A$29

You sell an invoice template.
One purchase. A phone buyer. No subscription.

“Done”

Freelancer
invoice template

A$290Pay A$29

Payment confirmed.
Even when declined.

You still have to find the mistakes.

Checked against the brief

Freelancer
invoice template

A$29Pay A$29

Payment declined.
Try another card.

The agent returns the result and the checks.

Same design. Fewer unchecked assumptions.
Authored example, not a live AI run. The kit checks selected requirements; it does not prove usability or payment security.

Start with the figures. Nothing important waits for an animation.

01 / The problem

Three mistakes.
One “finished” page.

The price is wrong. The pay button runs off the phone. A declined card gets a success message.

I do not want to spend the evening checking whether a generated checkout obeys the brief. I want the agent to bring back the answer to those questions. Start with the page a customer would actually use.

THE INDEPENDENT WORKSHOP01

Freelancer invoice template

A reusable template for your next invoice.
A$29. One purchase.

Pay todayA$290
Pay A$29

Your download follows a confirmed payment.

The total disagrees with the brief.

Compare the amount. No model judgment is needed.

A / Preserve the fact

It says A$290.
You asked for A$29.

Check the displayed total against an independent requirement.

B / Fit the screen

A button outside the phone is not a checkout.

Open the narrow layout. Check the page and the action.

C / Try the failure

Declined is not confirmed.

Exercise the failure state. Keep the next step usable.

D / Return the evidence

“Here is what passed. Here is what I could not check.”

Keep the approved visual direction. Attach the current result.

The useful change is in the handoff. The customer gets the page. You get evidence about its requirements.

Rule-driven teaching illustration. No card is charged. The separate local fixture in the kit can actually be run.

02 / The steps

Give the agent a job
with a finish line.

Write the requirements once. Make the agent return current evidence, not another confident paragraph.

YOU DEFINE

What must stay true

A$29 once.
Fits the phone.
Decline means retry.

TOOLS OBSERVE

Run the awkward cases

Exact amount.
Narrow viewport.
Failed payment.

AGENT REPAIRS

Fix the cause

Keep the style.
Bound the attempts.
Run checks again.

YOU REVIEW

Page + evidence + gaps

What changed.
What passed.
What is still unknown.

Keep the tests outside the agent’s proposed answer. Otherwise it can accidentally change the meaning of “correct”.

This is not automatic publishing.
Local review is the end of this loop. Production access and release approval remain separate.

price = A$29viewport = 320px → action fitspayment = declined → no success

The project began with mathematical design models, then tested browser evidence, qualitative findings, mock judges, failure cases and bounded repair tools. The checkout is one way into that work, not the whole research programme. The broader question is whether checks can catch important failures without erasing legitimate design freedom.

What “ready” means in the model
Ready = complete ∧ current ∧ well-formed ∧ all required checks pass

Each result belongs to an artifact, requirement set, observer and context. Missing or unrun checks are not passing checks. A later unsuccessful recheck blocks reuse of an older success until a new valid evaluation is available.

The v7 prototype exposes nine local JSON-lines tools, not an MCP server. Deterministic policies exercised its interface; no real model integration or human study was run. The small checkout kit is a separate fixture, not the entire broker.

Find a design D ∈ F, then compare preferences.
F = {D : gj(render(D,e)) ≤ 0 for required j,e}

Hard requirements define feasibility. Typography, hierarchy, brand suitability and preference still need appropriate models and evidence. A high aesthetic score cannot compensate for a failed required payment-state check.

03 / Use the maths

Don’t pay a model
to be a ruler.

Calculate what fits. Use the model for the decisions the calculation does not settle.

320 px
A price and payment button must share the available widthAt 320 pixels the 328-pixel row cannot fit. Stacking the elements meets the selected minimum widths.A$29Pay

8 px short. Stack the row.

24 + 132 + 12 + 136 + 24= 328 px
CODE
Is the amount correct?

Compare A$29 with the requirement.

BROWSER
Can the customer reach Pay?

Exercise the actual layout and state.

MODEL
What repair preserves the design?

Propose a change. Recheck its consequences.

PEOPLE
Does the page make sense?

Observe understanding, preference and use.

The breakpoint follows the content. The minimum sizes are design choices. The arithmetic is not.

Live numerical model; the diagram is scaled to fit this article. It is not a browser measurement of your site. There is no universal “good design” equation.

Open the geometry, design-system and typography models
One row fits when W ≥ 2m + p + g + b
W ≥ 2(24) + 132 + 12 + 136 = 328 px

W is the viewport width, m the side margin, p the price block’s minimum width, g the gap and b the button’s minimum width. At 327px a rigid row is one pixel too wide. At 328px it fits under these assumptions. Real text, focus treatments and localisation still need rendering.

n = min(nmax, max(0, ⌊(W − 2m + g)/(cmin + g)⌋))

For equal-width cards, n selects the greatest permitted column count that fits. A zero means infeasible under the chosen assumptions; it is not permission to shrink text until the page passes.

Shared decisions, not a particular framework

CSS variables or Tailwind theme variables can represent the same semantic tokens. Repeated components then share spacing, colours and control dimensions. A shared mistake can also spread everywhere. Updating a token should therefore trigger checks on its consumers. Tailwind supplies a mechanism, not proof that the choices are right. [T]

Keep the other models available

The embedded notebook retains interactive line-breaking optimisation, Fitts’ pointing index, correlated-judge errors and review economics. These are different models with different certainty. Exact geometry is not measured usability, and mock judgment is not evidence of human preference.

04 / Keep the freedom

Fix the mistake.
Not the personality.

A strict checker can catch faults by rejecting everything unfamiliar. That is not the result I want.

Invoice template

A$29
Pay A$29

Approved style. Faint action.

Invoice template

A$29
Pay A$29

The action is visible. The style stays.

1 targeted change.The selected visual direction is preserved.
Detection and unnecessary rejection belong in the same evaluation. Passing more tests is not useful if it quietly removes valid choices.

Authored styling illustration. In the earlier v7 replay, targeted policies preserved 34/34 permitted focus styles; a reset control preserved 32/34. Those are project cases, not general model performance.

Open the repair objective and the counterexample
D* ∈ arg min distance(D, Dapproved)
subject to required checks passing

The distance must represent meaningful changes, not reward the smallest edit at any cost. Necessary behaviour changes can justify a larger patch. A narrow colour change that hides a symptom without fixing its cause is not a good repair.

The research found that an invisible parent could make a valid focus treatment invisible too. Repairing the shared cause before replacing the focus style preserved the approved design. The rule was authored; it was not a generally learned causal model.

v6 also found implementation-specific checks that blocked legitimate alternatives. The failure corpus and counterexamples remain in the archive. The objective is to detect consequential failures without erasing legitimate design freedom.

05 / What comes back

“Done” needs
a receipt.

Which page? Which tests? What failed? A later edit cannot borrow an earlier pass.

LOCAL REVIEWRevision 7

Selected checks passed.

Displayed total
A$29 ✓
320px layout
Fits ✓
Declined payment
Retry ✓
Human usability
NOT TESTED

This is evidence for revision 7. Not authorisation to publish.

Current page: revision 7.
The recorded checks apply to this version.

Ready for local review.

Keep passing evidence small, but never trim away a failed or missing check. Ask for the measurements only when you need them.

Illustrative state model. Signatures and hashes establish identity and provenance under their assumptions; they cannot prove that the observation captured the right fact.

Same decision. Less report to inspect.

Recorded v7 policy replay

Full
897,984 B
Compact
186,448 B
79.24% fewer response bytes. Identical final designs in 36 authored scenarios. No measured token saving, human speedup or live-model benefit.
Open the evidence model, denominator and mock-judge limits
Reduction = 1 − 186,448 / 897,984 = 0.79237…

Both deterministic policies made 252 tool calls and 60 evaluations across the same 36 authored scenarios. These are JSON byte counts. They cannot be substituted for tokenizer counts, API prices or human review time.

The v7 representation retains required failures, unknowns, coverage and identity. The report records 13,116 finite combinations with no mismatch between full and compact acceptance predicates. That establishes equivalence for the specified predicate, not sufficiency for every repair or human judgment.

Φ(D₁) = Φ(D₂), but Y(D₁) ≠ Y(D₂)

When two designs produce the same observations but need different verdicts, a judge restricted to those observations cannot always distinguish them. The earlier lab demonstrated a missing-observation problem with an invisible action. More agreement on the same incomplete observation does not supply the missing fact.

Jev-style interfaces and frontier-agent policies were mocked in this research. Real-model capability, calibration, adversarial robustness and human outcomes remain unmeasured. The original code, raw results and limitations are preserved, not rerun as new evidence for this article.

06 / The payoff

Keep the checks
that give time back.

Reuse can repay setup. A one-off page or noisy checker can make the work slower.

Manual review200 min
Setup + review + upkeep95 min
SetupReviewUpkeep
105 minutes less

Under these assumptions. Change them; the answer can reverse.

Illustrative time value: A$105. Not a forecast or realised cash saving.

This is the bet, not the result. The goal is less repeated inspection while preserving outcomes—not merely fewer minutes of review.

Hypothetical, editable inputs. Assumes comparable work quality. Missed defects, conversion, provider pricing and customer losses are not modelled here; lower review time alone is not a reason to ship.

Open the assumptions, cash costs and break-even maths

T₀ = N × r₀
T₁ = S + N(r₁ + u)
Value = (T₀ − T₁) × h/60 − N × c

N is the number of changes, r₀ and r₁ the review minutes before and after, S the one-off setup, u upkeep, h the value assigned to an hour and c incremental cash per change. Setup is in minutes, not money. Only convert it to money once.

Per-change benefit = (r₀ − r₁ − u)h/60 − c

With these assumptions, positive net value begins at change 7.

The strict break-even count is floor((Sh/60) / per-change benefit) + 1 when per-change benefit is positive. Otherwise there is no positive break-even within this model. A subscription may make an extra call cost no additional cash while still consuming time or allowance. Metered use, quotas and review labour are different resources.

Reuse only holds if the checks remain relevant. Changing the task, content, design system or dependency graph can require more coverage. The deeper notebook includes calibration, correlated error and review-risk models; their probabilities are synthetic, not plug-in estimates for your agent.

07 / Use it

Give your AI
one checked task.

Start with the supplied failing and passing examples. Then map the checks to your page.

Get the small kit ↓

Read the generated brief below before using it.

Read and select the complete brief

08 / What is known

A useful tool.
An unfinished claim.

The code can test the contract. People still have to validate the contract—and the experience.

Observed
in this project

Browser failures and finite-state counterexamples. Targeted scripted repairs. Compact evidence retaining a specified decision. Raw records are in the research archive.

Illustrated
in this article

The checkout story, style controls, layout budget and receipt. These explain mechanisms; they are not new frontier-model evaluations.

Hypothetical
benefit

Less inspection, lower accepted-delivery cost and more freedom to explore. Real-agent and human outcomes still need to be measured.

A coding agent can inspect its page.

The Codex release described browser inspection and attached screenshots. That makes browser evidence a practical handoff ingredient. [C]

Visual iteration gets a design-system handoff.

The Claude Design announcement described reusable design systems and a coding handoff. It did not validate this kit. [D]

Dated first-party capabilities, not screenshots, endorsements, pricing claims or an expiring offer. The reason to try this is your next repeated task—not a manufactured countdown.

Layout constraints, browser testing and agent tool design have substantial prior art. The project’s narrower research question is whether we can give agents useful evidence while detecting failures and preserving legitimate design variation. Global novelty and publishable superiority are not established.

The checkout is the entry point. The design research stays.

Explore typography, Fitts’ law, judge correlation, stale evidence and deeper maths. The original v7 code, qualitative/quantitative models, tests, source ledgers and subsequent article history remain in the ZIP.

Save the notebook as HTML ↓

The reason to bother

I want the evening back.
Not another score.

Give the next agent one requirement it can test. Ask what passed, what changed and what it could not check.

In this lab, a layout could fit, the mock judges could agree, and the checks could pass.

None of them proved how well a person could use the result.

That is the next test. Until then: keep the useful checks, keep the gaps visible, and keep your judgment.

One last distinction: the price, layout and decline tests pass. What can the agent honestly claim?

Sources and boundaries

Research informs the mechanisms and editorial choices. It does not certify this article’s comprehension or aesthetics. Source notes distinguish evidence, design advice and practitioner opinion.

Open the source ledger and communication research
  1. Introducing upgrades to Codex ↗

    First-party release, 15 September 2025. Browser inspection and screenshot evidence. Historical capability, not performance or pricing evidence for this kit.

  2. Introducing Claude Design ↗

    First-party announcement, 17 April 2026. Design systems and a coding handoff. Not independent validation of this research.

  3. Tailwind theme variables ↗

    A mechanism for semantic token-backed styling. No Tailwind benchmark was run for this article.

  4. The layer-cake pattern of scanning ↗

    NN/G research interpretation: descriptive, visually distinct headings help scanning. The application here is unvalidated; no reading-speed multiplier is claimed.

  5. Progressive disclosure ↗

    NN/G guidance: primary content stays accessible; advanced detail is available on demand. Critical limitations remain outside the drawers. This is a principle, not a timed comprehension law.

  6. Communicating with Interactive Articles ↗

    Hohman, Conlen, Heer and Chau, 2020. Synthesises representations, interaction and learning. Animation is not uniformly better than static figures; no universal layout solution.

  7. Explorable Explanations ↗

    Bret Victor: let a reader manipulate assumptions and see their consequences. Design precedent, not empirical proof for this article.

  8. Animation for Attention and Comprehension ↗

    NN/G guidance on explanatory motion and continuity. Our timings and transitions are authored choices, not validated optimums.

  9. Animation from Interactions ↗

    W3C guidance on disabling nonessential interaction-triggered motion. Implementing a motion control is not complete WCAG conformance.

  10. Wikipedia: Signs of AI writing ↗

    Editorial review for vague significance, promotion and unsupported attributions. Contextual guide, not a detector or proof of human authorship.

  11. 7 Signs a UI Has Been Vibe Coded ↗

    Jeff Humble’s practitioner critique of competing emphasis and decorative patterns. Opinion, not an authorship classifier. Claims about training causes are not adopted.

  12. Anthropic frontend-design skill ↗

    Existing design instructions emphasise deliberate, context-specific visual direction. A prompt or skill is not a measured outcome.

  13. Claude Code best practices ↗

    Official guidance on verification criteria and runnable feedback. Instructions are not a security boundary.

  14. Playwright actionability ↗

    Defines visibility and actionability checks; opacity-zero can satisfy its visibility definition. That definition is not a complete model of human-visible content.

Writing review: use concrete claims and attributable evidence, not decorative significance or invented urgency. Wikipedia’s AI-writing guide is contextual advice, not an authorship detector. Design review: explanatory objects earn their space; style stereotypes are not proof of provenance.

Published 2026-08-28 · Source edition v5 · About 16 minutes for the complete article, including optional sections. Return to Writing