How to let an AI use your browser without watching every click
The model is not the control system. A production run needs bounded authority, durable state, a narrow decision layer, an executor, independent readback and a receipt that survives inspection.
A demo on my phone showed a browser agent finding a flight from Zürich to London in 7.1 seconds for less than half a cent. My first thought was obvious: use this for the browser work I keep putting off, such as university portals, finance dashboards, job sites and services with no useful API.
The demo answered the clicking question. It did not answer the trust question. I started with a fast browser demo and ended up redesigning the whole control loop.
I am testing the design on one read-only university workflow before I let it spread to anything more consequential. The companion article, The browser is becoming the API for every app, explains why the browser may become the API for every app, and why verified workflow evidence is the durable advantage. This one is the implementation.
Split the run into six responsibilities
The system became easier to reason about once I stopped treating “the agent” as one thing. Give each uncertainty to the least powerful controller that can handle it, then require independent evidence before success.
- Outcome contract. What state must exist? Which sources count? What is prohibited? The objective is frozen before the run starts.
- Frontier planner. Handles novelty and chooses a typed macro-option.
- Capability gate. Compiles the legal action set, and deterministic policy checks authority, scope and budget. The model cannot widen it.
- Jev or specialist. Makes bounded local choices from observed options.
- Deterministic executor. Acts idempotently and preserves unknown outcomes.
- Independent verifier. Reads back reality and issues a receipt, so the run can finish without the model grading itself.
When a step fails, the loop goes to recovery, which means re-observing or stopping, instead of declaring success. Nothing moves until authority and evidence exist.
In my setup, OpenClaw schedules and coordinates the run. Jev chooses among typed local actions. Software owns permission, state and side effects. Evidence decides whether the job actually finished.
The example: check UON at 6:00 and never submit anything
This is the bounded outcome used throughout the article. It reads Canvas, MyHub and UON Outlook, returns changed obligations plus source evidence, and never submits, messages, enrols, accepts or pays. It is specific enough to test and consequential enough to expose weak assumptions.
A click is not a result
My first design had a familiar agent loop: observe the page, choose an action, act, repeat. It looked sensible in a diagram. It was too weak in practice.
The problem was the unit of success. I was measuring whether the agent completed a sequence of actions. I needed to measure whether the external state I cared about now existed.
The previous snapshot said Assessment A exists. The first version then read an incomplete source; the repaired version handles the same fetch in three further steps.
| Step | Current fetch | System conclusion | What it shows |
|---|---|---|---|
| 1 · first version | Source timed out | Assessment removed | The first version compared records without proving source coverage. Absence was mistaken for evidence. |
| 2 · repaired | Coverage flag: incomplete | State unknown | The source adapter now reports completeness separately from the records it returns. |
| 3 · repaired | Retry still incomplete | UNKNOWN_RECONCILING | The workflow enters a durable uncertainty state. It does not issue a removal alert or a blind retry. |
| 4 · repaired | Authoritative page read | Assessment still exists | A complete read resolves the uncertainty and updates the snapshot with current evidence. |
The repair was simple and strict: absence can become evidence only after the system proves that every expected source was read completely. A timeout now produces UNKNOWN_RECONCILING, not a removal alert.
The four steps replay the defect and its repair. They are not a live portal run.
I also reused an idempotency key too broadly. A new scheduled observation could return an old receipt. The key now identifies one logical run, not one snapshot forever.
These were contract failures. A larger model would have made the same weak assumptions more fluently.
Use the least powerful controller that can reliably do the job
Different kinds of uncertainty need different controllers. Four tasks from the same workflow show where each one belongs.
| Task | Minimum sufficient controller | Why, and what bounds it |
|---|---|---|
| Read a balance through a supported API | API or deterministic software Precision and repeatability | The operation is structured, supported and directly verifiable. Adding a model would increase cost and uncertainty without adding useful capability. Bounded by a typed contract and tests. |
| Choose which observed button means “Current courses” | Jev Fast local ambiguity resolution | The page state is observed, the operation is bounded, and the choice is semantic rather than strategic. Jev can select from compatible targets, then a verifier checks the state change. Bounded by a calibrated typed action set. |
| Recover after the portal changes its layout | Frontier model Novel planning and diagnosis | The state is unfamiliar and the plan must change. A frontier model can diagnose the failure and propose a recovery option, but it still cannot widen authority. Bounded by typed macro-options and policy compilation. |
| Grant permission to submit an assignment | Calvin Judgement and authority | Submitting an assignment changes an external record and represents me. That is a new consequential authority class, so the decision stays human-controlled. Bounded by explicit current approval. |
The hierarchy as a formal control model
The frontier model selects a typed macro-option. A deterministic policy compiler turns that option into a legal action set. Jev may select a local action inside that set. The executor performs it. A separate verifier updates the system's belief about the world.
frontier plan → typed option → policy mask → Jev choice → tool → evidence
The information can move between layers. Authority does not move with it.
Jev is useful when the choice is fuzzy, local and bounded
TypeSafe's Jev evaluates typed questions against structured state and returns typed answers with probability distributions. In a browser loop, that can mean choosing an operation and one compatible element from the page that was actually observed.
That is a smaller and safer problem than asking a model to invent JavaScript, selectors, destinations and recovery logic. In the reference package's demo, the observed page has four controls: #12 Past courses, #17 Current courses, #24 Help and #31 Profile. The target is #17.
| Typed option | Probability |
|---|---|
CLICK #17 Current courses | 82% |
CLICK #12 Past courses | 11% |
WAIT | 5% |
DONE | 2% |
The decision still escalates. The top choice is 82%, the local pair-error upper bound is 1.8% and the action budget is 1.0%. The typed choice is useful, but it has not earned autonomous execution.
The Jev Ultrafast repository treats DONE as loop termination, not proof. It also warns against blind mutation retries. Jev cannot grant authority, retrieve an unrestricted vault, or establish whether the real-world outcome happened. Those jobs remain outside the model.
Small per-step error becomes a large system problem
If a workflow has independent steps with the same success rate, the chance that every step succeeds is pⁿ. That is a useful warning even though it is not a production model.
A toy model: every probabilistic step is independent and succeeds with the same probability. Change either input.
About 67 of 100 runs would complete every step.
A teaching model, not a production model. Real failures are often correlated.
Asking the model to be more careful does not fix this. Reduce probabilistic steps, checkpoint meaningful stages, store durable state, and verify postconditions.
A trustworthy action leaves a chain you can reconstruct
Operational logs are not enough. I need to know which source created a datum, which transform created a feature, which model produced an inference, which policy allowed the effect, and which evidence verified the outcome.
| Link in the chain | What it is | The question it must answer |
|---|---|---|
| SourceAuthority and freshness | The system or observation that claims to know something about the world. | Is this source authorised, current, complete and correctly identified? |
| DatumProvenance and time | A source-bound record with event, ingest and effective time. | Can I reconstruct where it came from and when it was true? |
| FeaturePoint-in-time transform | A derived value built only from data available at the decision time. | Were all inputs available then, and which transform version produced it? |
| InferenceProbability, not permission | A model estimate with uncertainty, calibration and an abstention path. | Is the estimate locally calibrated and inside its validity region? |
| DecisionAlternatives and policy | A selected course of action with its rationale and rejected alternatives. | Which policy and risk budget made this choice acceptable? |
| PermissionNarrow capability | A task-scoped lease that limits subject, operation, origin, object, quota and expiry. | Is this lease current, genuine and no broader than its parent authority? |
| EffectIntent and reconciliation | The attempted change, including full intent, idempotency and state. | Did the effect commit, fail or remain unknown? |
| EvidenceIndependent proof | Current authoritative evidence about the resulting state. | Does the evidence come from a genuinely separate failure domain? |
| OutcomeValue and burden | The useful result, cost, baseline and human intervention. | Did this reduce work or improve the real result? |
Existing standards cover parts of this. OpenTelemetry is useful for model, agent and tool spans. OpenLineage models jobs, runs and datasets. W3C PROV describes entities, activities and agents. They overlap, but they answer different questions. None of them grants permission or proves causal value.
The reference system is well tested. The live system has not earned that claim.
By the time I froze the architecture, the reference package had 683 automated tests, 220 defect records and 13 seeded unsafe mutants detected. Those numbers helped find real bugs. They do not prove the university portal, OpenClaw runtime, browser session or deletion path works in production.
| Evidence | Reference | Live |
|---|---|---|
| Automated tests | 683 passed | Useful, but indirect |
| Scheduled UON runs | Fixtures only | 0 of 28 required |
| Jev labels | Demo distribution | 0 of 500 required |
| Observed value | Modelled | Not established |
A large test suite can make a weak claim look stronger than one complete live trace. I would rather keep the release state at HOLD than blur the difference.
One read-only workflow gets to earn trust first
The first production-shaped Autonomy Cell is a UON material-change monitor. It checks every current course source, detects changed obligations, keeps the evidence, and interrupts me only when something material changes or the state cannot be established. The review is pre-registered, and each gate has an exit condition.
- Freeze the real source contract. List every current course, portal, page, pagination rule, freshness rule and authentication state. A source is either checked or explicitly unresolved. Exit condition: every expected source is accounted for.
- Build the deterministic baseline. Use APIs or Playwright first. Capture authoritative snapshots and compare them without model interpretation wherever possible. Exit condition: missing data can never produce “no change”.
- Operate in shadow. Schedule real checks, survive session expiry and crashes, and record model, browser, verifier, latency and cost receipts. Exit condition: 28 complete scheduled runs.
- Calibrate the AI layers. Collect adversarial verifier negatives and labelled Jev operation-target decisions against the deterministic source of truth. Exit condition: 300 negative cases and 500 Jev labels.
- Make the promotion decision. Review each action family using the pre-registered threshold and the exact version tuple that produced the evidence, and promote only the action families whose error bounds fit the risk budget. Exit condition: each action family gets one result,
PROMOTE,HOLDorDEMOTE.
If it passes, the platform expands by outcome, not by data source
The likely sequence is GitHub and CI outcome tracking, calendar reconciliation, then PocketSmith read-only monitoring. OwnTracks and screen activity come later, after the system can connect context to a verified result.
- GitHub and CI outcome graph. Connect tasks to commits, review, deployment health and product outcomes.
- Calendar reconciliation. Prove reversible writes, conflict ownership and authoritative readback.
- PocketSmith monitor. Keep financial context behind its own sensitive-data boundary.
- Later, personal context. Add location and activity only for a named workflow with a measured benefit.
That order avoids building a detailed surveillance system that cannot tell whether anything useful happened.
My current rule is plain: build one bounded outcome, prove it in the world and keep the parts that earned trust. The model can say “done”, but the system must prove what happened. That is the difference between a demo and something you can leave running.
For coding agents: the contract, schemas and evaluator
The article ships with the contract, schemas and evaluator, so an AI can use the package directly without inferring the operating rules from prose.
- Operating contract (YAML). Allowed actions, unknown-state behaviour and receipt requirements.
- Trust receipt schema. Machine-checkable evidence returned by every run.
- Capability lease schema. Narrow operations, domains, objects, quotas and expiry.
- Jev decision log schema. Local probability, entropy, OOD, freshness and correctness labels.
- Evaluation checklist. Verifier negatives, drift, crash, replay and promotion gates.
- AI reader prompt. Tells an engineering agent exactly how to apply the field guide.
The blueprint below summarises the six responsibilities as rules an agent can follow.
OUTCOME CONTRACT - Freeze the requested external state. - Enumerate required sources and forbidden effects. AUTHORITY - Issue a short-lived capability for exact operations, origins, objects, quota and expiry. - Never infer authority from model confidence. EXECUTION - Prefer API or deterministic software. - Use Playwright for stable GUI paths. - Use Jev for bounded local ambiguity. - Use a frontier model only for novelty or recovery. VERIFICATION - Read back authoritative state. - Preserve UNKNOWN after ambiguous writes. - Return a machine-readable receipt. PROMOTION - Promote only action families whose local error bounds fit their risk budget.
Sources and further discussion
The technical argument comes from the linked primary documentation and papers, plus the defects found while building the Calvin Ops reference package.
- TypeSafe, Jev introduction. Typed questions, values and distributions.
- TypeSafe, confidence. Risk-sensitive thresholds and local calibration.
- Jev Ultrafast, agent rules. Observed targets, mutation rules and independent completion checks.
- Jev Ultrafast, performance notes. A narrow speed experiment, not a broad reliability benchmark.
- ReAct. Interleaved reasoning and acting.
- τ-bench. Repeated consistency and passk.
- Temporal workflow execution. Event history and deterministic replay.
- METR time horizons. Task duration as one view of capability.
- OpenTelemetry GenAI conventions. Operational model, agent and tool telemetry.
- OpenLineage object model. Jobs, runs and datasets.
- W3C PROV-O. Semantic provenance for entities, activities and agents.
- OWASP GenAI Security Project. Prompt injection, data disclosure and excessive agency.
I rewrote the interactive edition of this page after auditing it against Wikipedia's field guide to common AI-writing patterns and recent work on design homogenisation in vibe-coded websites. The guide is descriptive, not a detector, so I used it as an editing prompt rather than a score. The design paper on web vibe coding argues that frictionless generation can increase homogenisation and proposes productive friction as a countermeasure.
The first version used the usual dark AI-product palette, gradient text, decorative glow, orbit animation, nested cards and repeated slogan headings. The interactive edition replaced them with one accent colour, flat surfaces and typographic grouping. This edition uses the site's own type and colours.
AI assisted with research, code and editing. I chose the argument, reviewed the sources, rewrote the article in my own voice, and tested the interactions.