← Back to writing

How to let an AI use your browser without watching every click

The model is not the control system. A production run needs bounded authority, durable state, a narrow decision layer, an executor, independent readback and a receipt that survives inspection.

A demo on my phone showed a browser agent finding a flight from Zürich to London in 7.1 seconds for less than half a cent. My first thought was obvious: use this for the browser work I keep putting off, such as university portals, finance dashboards, job sites and services with no useful API.

Social post showing a fast Jev-powered browser agent
The social post with the flight-search demo.

The demo answered the clicking question. It did not answer the trust question. I started with a fast browser demo and ended up redesigning the whole control loop.

I am testing the design on one read-only university workflow before I let it spread to anything more consequential. The companion article, The browser is becoming the API for every app, explains why the browser may become the API for every app, and why verified workflow evidence is the durable advantage. This one is the implementation.

Split the run into six responsibilities

The system became easier to reason about once I stopped treating “the agent” as one thing. Give each uncertainty to the least powerful controller that can handle it, then require independent evidence before success.

  1. Outcome contract. What state must exist? Which sources count? What is prohibited? The objective is frozen before the run starts.
  2. Frontier planner. Handles novelty and chooses a typed macro-option.
  3. Capability gate. Compiles the legal action set, and deterministic policy checks authority, scope and budget. The model cannot widen it.
  4. Jev or specialist. Makes bounded local choices from observed options.
  5. Deterministic executor. Acts idempotently and preserves unknown outcomes.
  6. Independent verifier. Reads back reality and issues a receipt, so the run can finish without the model grading itself.

When a step fails, the loop goes to recovery, which means re-observing or stopping, instead of declaring success. Nothing moves until authority and evidence exist.

In my setup, OpenClaw schedules and coordinates the run. Jev chooses among typed local actions. Software owns permission, state and side effects. Evidence decides whether the job actually finished.

The example: check UON at 6:00 and never submit anything

This is the bounded outcome used throughout the article. It reads Canvas, MyHub and UON Outlook, returns changed obligations plus source evidence, and never submits, messages, enrols, accepts or pays. It is specific enough to test and consequential enough to expose weak assumptions.

A click is not a result

My first design had a familiar agent loop: observe the page, choose an action, act, repeat. It looked sensible in a diagram. It was too weak in practice.

The problem was the unit of success. I was measuring whether the agent completed a sequence of actions. I needed to measure whether the external state I cared about now existed.

Missing source data looked like an assessment removal

The previous snapshot said Assessment A exists. The first version then read an incomplete source; the repaired version handles the same fetch in three further steps.

StepCurrent fetchSystem conclusionWhat it shows
1 · first versionSource timed outAssessment removedThe first version compared records without proving source coverage. Absence was mistaken for evidence.
2 · repairedCoverage flag: incompleteState unknownThe source adapter now reports completeness separately from the records it returns.
3 · repairedRetry still incompleteUNKNOWN_RECONCILINGThe workflow enters a durable uncertainty state. It does not issue a removal alert or a blind retry.
4 · repairedAuthoritative page readAssessment still existsA complete read resolves the uncertainty and updates the snapshot with current evidence.

The repair was simple and strict: absence can become evidence only after the system proves that every expected source was read completely. A timeout now produces UNKNOWN_RECONCILING, not a removal alert.

The four steps replay the defect and its repair. They are not a live portal run.

I also reused an idempotency key too broadly. A new scheduled observation could return an old receipt. The key now identifies one logical run, not one snapshot forever.

These were contract failures. A larger model would have made the same weak assumptions more fluently.

Use the least powerful controller that can reliably do the job

Different kinds of uncertainty need different controllers. Four tasks from the same workflow show where each one belongs.

TaskMinimum sufficient controllerWhy, and what bounds it
Read a balance through a supported APIAPI or deterministic software
Precision and repeatability
The operation is structured, supported and directly verifiable. Adding a model would increase cost and uncertainty without adding useful capability. Bounded by a typed contract and tests.
Choose which observed button means “Current courses”Jev
Fast local ambiguity resolution
The page state is observed, the operation is bounded, and the choice is semantic rather than strategic. Jev can select from compatible targets, then a verifier checks the state change. Bounded by a calibrated typed action set.
Recover after the portal changes its layoutFrontier model
Novel planning and diagnosis
The state is unfamiliar and the plan must change. A frontier model can diagnose the failure and propose a recovery option, but it still cannot widen authority. Bounded by typed macro-options and policy compilation.
Grant permission to submit an assignmentCalvin
Judgement and authority
Submitting an assignment changes an external record and represents me. That is a new consequential authority class, so the decision stays human-controlled. Bounded by explicit current approval.
The hierarchy as a formal control model

The frontier model selects a typed macro-option. A deterministic policy compiler turns that option into a legal action set. Jev may select a local action inside that set. The executor performs it. A separate verifier updates the system's belief about the world.

frontier plan → typed option → policy mask → Jev choice → tool → evidence

The information can move between layers. Authority does not move with it.

Jev is useful when the choice is fuzzy, local and bounded

TypeSafe's Jev evaluates typed questions against structured state and returns typed answers with probability distributions. In a browser loop, that can mean choosing an operation and one compatible element from the page that was actually observed.

That is a smaller and safer problem than asking a model to invent JavaScript, selectors, destinations and recovery logic. In the reference package's demo, the observed page has four controls: #12 Past courses, #17 Current courses, #24 Help and #31 Profile. The target is #17.

The typed decision, from the demo distribution
Typed optionProbability
CLICK #17 Current courses82%
CLICK #12 Past courses11%
WAIT5%
DONE2%

The decision still escalates. The top choice is 82%, the local pair-error upper bound is 1.8% and the action budget is 1.0%. The typed choice is useful, but it has not earned autonomous execution.

The Jev Ultrafast repository treats DONE as loop termination, not proof. It also warns against blind mutation retries. Jev cannot grant authority, retrieve an unrestricted vault, or establish whether the real-world outcome happened. Those jobs remain outside the model.

Small per-step error becomes a large system problem

If a workflow has independent steps with the same success rate, the chance that every step succeeds is pⁿ. That is a useful warning even though it is not a production model.

A small per-step error compounds across a run

A toy model: every probabilistic step is independent and succeeds with the same probability. Change either input.

End-to-end result66.8%

About 67 of 100 runs would complete every step.

A teaching model, not a production model. Real failures are often correlated.

Asking the model to be more careful does not fix this. Reduce probabilistic steps, checkpoint meaningful stages, store durable state, and verify postconditions.

A trustworthy action leaves a chain you can reconstruct

Operational logs are not enough. I need to know which source created a datum, which transform created a feature, which model produced an inference, which policy allowed the effect, and which evidence verified the outcome.

Link in the chainWhat it isThe question it must answer
SourceAuthority and freshnessThe system or observation that claims to know something about the world.Is this source authorised, current, complete and correctly identified?
DatumProvenance and timeA source-bound record with event, ingest and effective time.Can I reconstruct where it came from and when it was true?
FeaturePoint-in-time transformA derived value built only from data available at the decision time.Were all inputs available then, and which transform version produced it?
InferenceProbability, not permissionA model estimate with uncertainty, calibration and an abstention path.Is the estimate locally calibrated and inside its validity region?
DecisionAlternatives and policyA selected course of action with its rationale and rejected alternatives.Which policy and risk budget made this choice acceptable?
PermissionNarrow capabilityA task-scoped lease that limits subject, operation, origin, object, quota and expiry.Is this lease current, genuine and no broader than its parent authority?
EffectIntent and reconciliationThe attempted change, including full intent, idempotency and state.Did the effect commit, fail or remain unknown?
EvidenceIndependent proofCurrent authoritative evidence about the resulting state.Does the evidence come from a genuinely separate failure domain?
OutcomeValue and burdenThe useful result, cost, baseline and human intervention.Did this reduce work or improve the real result?

Existing standards cover parts of this. OpenTelemetry is useful for model, agent and tool spans. OpenLineage models jobs, runs and datasets. W3C PROV describes entities, activities and agents. They overlap, but they answer different questions. None of them grants permission or proves causal value.

The reference system is well tested. The live system has not earned that claim.

By the time I froze the architecture, the reference package had 683 automated tests, 220 defect records and 13 seeded unsafe mutants detected. Those numbers helped find real bugs. They do not prove the university portal, OpenClaw runtime, browser session or deletion path works in production.

Reference evidence compared with live evidence
EvidenceReferenceLive
Automated tests683 passedUseful, but indirect
Scheduled UON runsFixtures only0 of 28 required
Jev labelsDemo distribution0 of 500 required
Observed valueModelledNot established

A large test suite can make a weak claim look stronger than one complete live trace. I would rather keep the release state at HOLD than blur the difference.

One read-only workflow gets to earn trust first

The first production-shaped Autonomy Cell is a UON material-change monitor. It checks every current course source, detects changed obligations, keeps the evidence, and interrupts me only when something material changes or the state cannot be established. The review is pre-registered, and each gate has an exit condition.

  1. Freeze the real source contract. List every current course, portal, page, pagination rule, freshness rule and authentication state. A source is either checked or explicitly unresolved. Exit condition: every expected source is accounted for.
  2. Build the deterministic baseline. Use APIs or Playwright first. Capture authoritative snapshots and compare them without model interpretation wherever possible. Exit condition: missing data can never produce “no change”.
  3. Operate in shadow. Schedule real checks, survive session expiry and crashes, and record model, browser, verifier, latency and cost receipts. Exit condition: 28 complete scheduled runs.
  4. Calibrate the AI layers. Collect adversarial verifier negatives and labelled Jev operation-target decisions against the deterministic source of truth. Exit condition: 300 negative cases and 500 Jev labels.
  5. Make the promotion decision. Review each action family using the pre-registered threshold and the exact version tuple that produced the evidence, and promote only the action families whose error bounds fit the risk budget. Exit condition: each action family gets one result, PROMOTE, HOLD or DEMOTE.

If it passes, the platform expands by outcome, not by data source

The likely sequence is GitHub and CI outcome tracking, calendar reconciliation, then PocketSmith read-only monitoring. OwnTracks and screen activity come later, after the system can connect context to a verified result.

  1. GitHub and CI outcome graph. Connect tasks to commits, review, deployment health and product outcomes.
  2. Calendar reconciliation. Prove reversible writes, conflict ownership and authoritative readback.
  3. PocketSmith monitor. Keep financial context behind its own sensitive-data boundary.
  4. Later, personal context. Add location and activity only for a named workflow with a measured benefit.

That order avoids building a detailed surveillance system that cannot tell whether anything useful happened.

My current rule is plain: build one bounded outcome, prove it in the world and keep the parts that earned trust. The model can say “done”, but the system must prove what happened. That is the difference between a demo and something you can leave running.

For coding agents: the contract, schemas and evaluator

The article ships with the contract, schemas and evaluator, so an AI can use the package directly without inferring the operating rules from prose.

The blueprint below summarises the six responsibilities as rules an agent can follow.

OUTCOME CONTRACT
- Freeze the requested external state.
- Enumerate required sources and forbidden effects.

AUTHORITY
- Issue a short-lived capability for exact operations, origins, objects, quota and expiry.
- Never infer authority from model confidence.

EXECUTION
- Prefer API or deterministic software.
- Use Playwright for stable GUI paths.
- Use Jev for bounded local ambiguity.
- Use a frontier model only for novelty or recovery.

VERIFICATION
- Read back authoritative state.
- Preserve UNKNOWN after ambiguous writes.
- Return a machine-readable receipt.

PROMOTION
- Promote only action families whose local error bounds fit their risk budget.
Sources and further discussion

The technical argument comes from the linked primary documentation and papers, plus the defects found while building the Calvin Ops reference package.

  1. TypeSafe, Jev introduction. Typed questions, values and distributions.
  2. TypeSafe, confidence. Risk-sensitive thresholds and local calibration.
  3. Jev Ultrafast, agent rules. Observed targets, mutation rules and independent completion checks.
  4. Jev Ultrafast, performance notes. A narrow speed experiment, not a broad reliability benchmark.
  5. ReAct. Interleaved reasoning and acting.
  6. τ-bench. Repeated consistency and passk.
  7. Temporal workflow execution. Event history and deterministic replay.
  8. METR time horizons. Task duration as one view of capability.
  9. OpenTelemetry GenAI conventions. Operational model, agent and tool telemetry.
  10. OpenLineage object model. Jobs, runs and datasets.
  11. W3C PROV-O. Semantic provenance for entities, activities and agents.
  12. OWASP GenAI Security Project. Prompt injection, data disclosure and excessive agency.

I rewrote the interactive edition of this page after auditing it against Wikipedia's field guide to common AI-writing patterns and recent work on design homogenisation in vibe-coded websites. The guide is descriptive, not a detector, so I used it as an editing prompt rather than a score. The design paper on web vibe coding argues that frictionless generation can increase homogenisation and proposes productive friction as a countermeasure.

The first version used the usual dark AI-product palette, gradient text, decorative glow, orbit animation, nested cards and repeated slogan headings. The interactive edition replaced them with one accent colour, flat surfaces and typographic grouping. This edition uses the site's own type and colours.

AI assisted with research, code and editing. I chose the argument, reviewed the sources, rewrote the article in my own voice, and tested the interactions.

  1. Wikipedia: Signs of AI writing
  2. Interrogating Design Homogenization in Web Vibe Coding
  3. Good Vibrations? A qualitative study of vibe coding
  4. Fountain Institute: 7 signs a UI has been vibe coded
  5. ONS Service Manual: structuring content
  6. W3C: pause, stop, hide
Published 2026-09-01 · Updated 2026-09-28 · Source edition series