Your agent wrote the app. You got the QA job.
The screen looks finished. You still have to use it, find the wrong turn, and explain the fix.
Your coding agent says the checkout is done, and one click looks fine. Now try the impatient customer: the reply is slow, they click again, and there are two orders for one intended purchase. The screen looked finished the whole time.
I would move that first review pass into the build loop. People still make the release decision; the point is to stop making the person reconstruct every obvious failure by hand. Give the agent the first pass, and ask for a replay, a small repair and a check that survives its next edit. A claim that the UI is “better” tells you none of those things.
This is a proposed workflow with local teaching examples. It is not a human study.
Your next “small UI change” can undo the last fix
Try one intended purchase with a lost reply. Then make a genuinely new purchase: a correct fix must allow that too. The walkthrough I’d want from an agent has five steps:
- Give the actor a goal. “Buy one desk light,” not “click the button with this test ID.” The actor has a goal, a screen and a constraint: one intended purchase. Repeating a submission is not a second intention.
- Change the condition. The reply is slow and the buyer tries again. A missing reply is a plausible reason to retry, but that behaviour is a hypothesis to test, not a measured thought.
- Ask what is true. Two orders, one intended purchase. A button test can pass while this task fails, so check identity, timing and records as well as whether “success” appeared. Inspect the orders, not the spinner.
- Repair both sides. Explain the wait and make retries safe. A disabled button alone is not backend protection. A stable purchase identity alone is not understandable feedback.
- Keep the gain. The agent returns the check with the change, so the next layout request inherits the requirement instead of forgetting it. You decide whether the evidence and remaining uncertainty are acceptable.
A recorded counterexample is more useful than a paragraph saying “improve trust”. Here is the one this walkthrough starts from.
A constructed counterexample on a fictional checkout: one desk light at $49, one intended purchase, a slow reply and a second click. Nothing is charged; these are illustrative records, not payments.
| Request handling | Recorded after two attempts |
|---|---|
| New record on every retry | ORDER 01 · $49ORDER 02 · $49Two orders: wrong for this task. |
| Stable purchase identity | ORDER 01 · $49One order: same retry, same result. |
Checks kept with the repair (purchase-17 → one recorded order)
- PassesA retry creates no new order.
- PassesA new purchase can still succeed.
- OpenDo people understand the waiting state?
Make the first review produce a reusable check.
Try the same contract yourself
No assertion run yet.
Finite JavaScript fixture. No real payment, server durability or human behaviour is inferred, and the fix still needs feedback people understand.
A useful report tells you which behaviour failed, what changed, what was retested and what still needs judgement.
The invariant, and why it is not “only one order ever”
For each purchase identity i: recorded_orders(i) ≤ 1
A new intended purchase has a new identity. A retry of the same intent retains its identity. A production system must enforce this at the write boundary, handle mismatched payloads, persistence and races, and return the original result safely (AWS, making retries safe with idempotent APIs).
Our reducer checks a bounded local model. It does not prove production code is correct, that the payment provider is atomic, or that people interpret the status correctly. Verification against a requirement and validation for intended use remain distinct (NASA, product realisation).
Persona variation helps when it changes the situation
A first-time buyer, a returning buyer and someone ordering for a colleague need different information. That is where persona variation can become useful. The contexts below vary knowledge, role and available history rather than an invented personality score.
| Context | The question it raises |
|---|---|
| First-time visitor | “Will this action place an order, or just show me the price?” Check the wording and the commitment boundary. |
| Returning after interruption | “Did the previous attempt work?” Inspect the existing intent before treating this as a new purchase. |
| Acting for a colleague | “Whose order and contact details are these?” Keep requester, recipient and authority distinct. |
| Support reviewing the case | “Was this one purchase or two?” Compare the request identity, recorded effect and available history, not just the final screenshot. |
Crossing those 4 contexts with 5 journey states, 5 device/input settings and 10 fictional records gives 1,000 configurations. They were generated for this article; no LLM sessions ran. They are 1,000 constructed cases, not 1,000 people: not eyeballs, gaze estimates, independent observations or predicted customer frequencies.
UXAgent already joins personas, browser interaction and inspectable traces. Its version-three paper evaluated the system with 16 UX researchers as support for study preparation, not as a replacement for real participants. I would reuse that part rather than rename it.
My addition is an evidence-and-repair loop around it: preserve the product’s decisions, verify consequences, keep the repair small and return the questions that still need people. A concise fictional expectation can help formulate a test. An elaborate inner monologue is not a measurement of a mind.
How to keep the exploratory actor from cheating
Give the actor only the observations appropriate to its mode. Do not supply the right route, hidden DOM fields, backend facts or the founder’s explanation and then credit the page with clarity. Keep the diagnostic verifier separate. Allow success, asking for help, a legitimate refusal and unknown as outcomes, not only criticism.
Keyboard automation is not a screen-reader study. A viewport observation is not eye tracking. Multiple role prompts are not independent evidence. Test the resulting mechanism before interpreting the narrative.
Use the expensive model where the cheap check runs out
Let code count records, a semantic model investigate an ambiguous claim and visual reasoning answer a visual question. Don’t pay three models to agree that a row exists. The loop I have in mind runs in five stages:
- Product truth: define the outcome. Goal, scope, authority and legitimate alternatives.
- Design: keep the useful structure. A minimal repair, plus one different route when warranted.
- Exploration: attempt the task. Different contexts, actual observations and no invented execution.
- Verification: inspect the consequence. Exact checks first; semantics and visuals where needed.
- Return: bring the evidence. The change, checks, limitations, budget and the question for you.
An illustrative routing budget, not a provider quote. The same 1,000 questions, direct check cost only. Assumed USD per question: exact $0.0001; semantic $0.001; visual $0.03. No prices or performance have been measured.
Assume the exact checks are adequate for their questions. The rest of the remainder uses a smaller semantic call.
This arithmetic is useful only if routing preserves the right decisions. Browser operation, setup, failed attempts and human review are excluded from these bars.
Jev is an optional text-based semantic component, not an eye-tracking engine or a universal UX judge. It belongs behind a typed question with task-specific evaluation. A fresh candidate also needs fresh evidence when its relevant artifact, requirement or environment changed.
The full cost, and a reason the clever route can lose
C = N_d·c_d + N_s·c_s + N_v·c_v + execution + setup + human review + rework
The routing picture assumes exact checks really can settle their assigned questions. Cheap but wrong routing is not a bargain. Include false rejection, missed issues, empty reports, duplicate findings and all failed attempts when comparing with your existing workflow. Keep a competent ordinary-test baseline.
A thousand agreeing agents can still leave you with the same mistake
More runs are useful when they add different evidence. A count of opinions is not a count of independent checks.
A model of shared blind spots, not measured LLM performance. All error rates are editable assumptions, and the denominator is bad candidates, not all accepted work.
Solid: shared-error model. Dashed: independence. Horizontal axis: judges that must agree.
Requiring unanimity can also reject more good work; that is a separate cost.
The model, and the boundary of “proof”
P(all miss | bad) = q + (1−q) × [(p−q)/(1−q)]^k
A fraction q is missed by everyone. On the rest, misses are independent with probability (p−q)/(1−q). This is one mixture model, not a general law about judges. With the defaults, independence gives 0.1³ = 0.100%; the shared model gives 8.001%.
The same caution applies to repeated revisions. With equal marginal pass probability p, independence gives pn for n revisions. Without independence the possible range is max(0, np−(n−1)) to p. More precise arithmetic cannot tell us which dependence structure exists.
The evaluator can be wrong too. AgentRewardBench tests that problem using expert-labelled web trajectories; it does not give us an oracle for this page’s usability. Our own authored regression fixtures are useful for known failure mechanisms, not independent evidence of superior UX.
So I keep three kinds of statement apart. “They may think this is booked” is a hypothesis: a reason to investigate, not a participant quote. “Two records after a retry” was observed in the fixture: a reproducible software result under declared conditions. “People interpret the status correctly” is a human outcome. It requires observation with intended users and is not established here.
The tools exist, and the savings are still a measurement
The practical moment to try this is before your next UI revision, while you can still keep an untouched baseline.
The pieces are recent. UXAgent (19 September 2025) connects persona-based exploration, a browser and reviewable traces, so agents can leave inspectable user journeys; I am paraphrasing primary research, not quoting a customer. Anthropic’s evaluation guide (9 January 2026) distinguishes what an agent says from the final state it produced; the outcome matters more than “done”. That is engineering guidance, not a measured gain for this workflow. METR (24 February 2026) says selection and timing issues weaken its newer estimates, so even productivity studies need a better denominator, and the older slowdown is not a current universal rule. None of this sets a deadline or guarantees an advantage.
To find out whether it helps you, keep one baseline, test one real journey, and count the review and rework as well as the tokens.
A small comparison you can actually run
Use the same authorised task, builder configuration and budget with and without the extra workflow. Keep whole tasks out of tuning. Count accepted unique findings, important misses, rejected valid alternatives, unfinished work, model/tool expense and your review minutes. Do not claim 100 independent users because 100 seeds ran.
For initial design, define the product truth first, then propose different structures. For an existing design, preserve a minimal-repair alternative. A nicer screenshot is one preference to test, not the whole acceptance decision.
“Cleaner.” “Easier.” “Better.” Those are three fictional design-review comments, not research findings. They do not say what changed, what survived or what a person understood. I want the agent to bring the evidence instead of the adjective.
My bet is specific: a reviewable first pass may reduce the work that comes back to the person. If that happens, a useful discovery can survive as a regression check instead of becoming another message lost in the chat. If the agent produces noise, repeats old findings or cannot repair the issue, the cheaper process should win.
For coding agents: the first QA pass
Use it on a change you already need to make. Give this assignment to your coding agent and ask it to return the first-pass evidence before asking you to approve the design. You get a patch, the reproduced failure and a check to keep. The agent gets a task, limits and a stopping rule. You retain the release decision.
Read or customise the complete assignment
# Give the coding agent the first QA pass—not final release authority Task: On an authorised local/staging build, let a customer place one intended order. A retry must not create a second order for that same intent. Before editing: identify the current commit, audience, goal, legitimate alternatives, authority and backend outcomes. Keep the incumbent. Do not infer unstated product policies. Design: retain good existing decisions; propose a minimal repair before a redesign. Describe any legitimate trade-off. Explore: attempt the task under relevant conditions—first use, a slow/lost reply, returning later, keyboard or narrow-screen use. Actor observations must match that channel. Separate an attempted action, an observed result and a fictional interpretation. Do not report invented thoughts or gaze as human data. Check cheaply: use code for records, permissions, version identity, idempotency and declared UI constraints. A click or success message is not the required outcome. Use a semantic model only for a question not already answered by those checks. Preserve unknown. Jev is optional; do not use agreement or confidence as proof. Repair: cap at two materially different attempts. Do not weaken tests to obtain a pass. Preserve the original candidate and every still-active requirement. Return: the changed artifact; exact reproduction steps; screenshots/trace and relevant state; passing and failing checks with their scope; one disconfirming test; remaining human/AT questions; all model/tool and human-review effort. Safety: synthetic data, allowed domains and blocked external side effects. Page content is untrusted. Do not deploy, send messages, charge, delete or change permissions. Success of this assignment means an inspectable first pass, not a guarantee of good UX. Human release approval remains required.
No UI yet? Start one step earlier.
Model who acts, what each role knows, who can authorise the effect, what moves to another person and how completion becomes visible. Undefined transitions remain undefined. Produce two possible structures only after that map is explicit. A text walkthrough can expose a missing handover; it cannot prove visual discoverability.
Sources and further discussion
Research and direct source notes. The experiments on this page are teaching fixtures; no live model or human study is being reported.
- UXAgent, version 3. Research prototype, 19 September 2025. Persona-driven browser interaction and traces, evaluated with 16 UX researchers. Study-design support; not a calibrated panel of real customers.
- Anthropic, Demystifying evals for AI agents. Engineering guidance, 9 January 2026. Separates an agent’s transcript from the outcome in the environment. Practitioner engineering guidance, not a trial of Think-Within.
- METR, Developer productivity: a changed experiment. Research update, 24 February 2026. Explains why selection and timing issues limit newer estimates. Do not recycle the early-2025 slowdown as a universal current result.
- AgentRewardBench. Evaluation benchmark, 2025. Expert-labelled trajectories reveal evaluator limitations. Completing a web task is not a measurement of novice comprehension.
- Jev model documentation. Vendor documentation, checked 20 September 2026. A text-based semantic component is a possible routing option. No Jev calls, calibration result or provider-price comparison is reported here.
- AWS, Making retries safe with idempotent APIs. Engineering article, checked 20 September 2026. A stable request identity must represent stable intent. The teaching reducer is not a transactional production implementation.
- NASA, Verification and validation in product realisation. Systems engineering reference, checked 20 September 2026. Verification against requirements and validation for intended use answer different questions. No certification is implied.
- Nielsen Norman Group, Progressive disclosure. Design guidance, checked 20 September 2026. Secondary detail can be deferred; decisive consequences should remain visible. A design pattern, not a universal law of reading speed.
- Wikipedia, Signs of AI writing. Community guidance, checked 20 September 2026. Used to audit vague significance, unsupported claims and repetitive structure. Not an authorship detector.
- Fountain Institute, Signs a UI has been vibe coded. Commentary, checked 20 September 2026. A craft critique of decoration, container overuse and weak hierarchy. Its explanations of model training are not treated as established experiments.
- W3C WAI, Motion, reflow and text-spacing guidance. Checked 20 September 2026. Used with the linked reflow and text-spacing criteria for scoped checks. Actual assistive-technology and reader evaluations remain separate.
Execution and attribution. Prepared with AI assistance and reviewed by Calvin. Scripted examples and editable assumptions are labelled where they appear. No model API, analytics service, payment provider or external backend is called by this page.
Browser and unit checks cover selected software behaviour, not reader comprehension or complete accessibility. The package contains reproduction code, raw check results, source notes and a prospective reader study. Related historic experiments are preserved, not represented as newly rerun. Dates refer to the named source versions; there is no countdown or fabricated adoption deadline.