← Back to writing

The browser is becoming the API for every app

If this generalises, one AI can operate the university portal, finance dashboard, job sites and internal tools you already use. That could erase whole categories of admin. The durable advantage is the private proof layer that lets you stop checking.

Almost every app already exposes a human interface. Browser agents turn that interface into a possible automation surface. The social post that started this question showed a Jev-powered browser agent finding a flight: 7.1 seconds of action and a cost claim of $0.0039.

Screenshot of a social post demonstrating a fast Jev-powered browser agent finding a flight
The post that triggered the question. It proves speed. It does not prove the final state.

That collapses the integration barrier. Old portals, dashboards and internal systems become reachable without waiting for an API. But action is not the outcome: a stale page, missing source or timed-out write can still produce a polished false success.

My bet is simple: the model will commoditise, and your verified workflow history can compound. The companion technical field guide covers the implementation. This is the argument for building it now.

Your AI could operate the software you already pay for

That changes the unit of integration. Instead of waiting for every vendor to expose the right API, you can define one outcome, one authority boundary and one proof receipt, and use the interface the software already exposes. The next integration may be a workflow, not a six-week software project.

  • Reach. Legacy portals and internal tools become addressable.
  • Time. Repeated tab-checking can collapse into one verified result. Three hours a week is roughly 156 hours a year before verification and exceptions.
  • Cost. Frontier models handle novelty, and cheap reflexes handle routine choices, so frontier intelligence is not spent on every click.
  • Moat. Your source maps, exceptions and verifier cases improve with use. Each verified run can make the next run cheaper and safer.

If browser agents generalise, the browser stops being a place you visit and becomes a universal compatibility layer your software can operate.

Several companies are converging on the same capability

This is no longer one clever demo. Different companies are building models that can see software, choose actions and operate it.

Action is moving from bespoke demos into products and APIs. The operational advantage is still proof. The dates and claims come from official announcements or the linked project repository. Product announcements show direction. They are not production-reliability certificates.

Disconnected software could become one system

If the browser is the integration, your disconnected software becomes one system. It becomes reachable, though not magically and not safely by default.

WorkflowWhat the agent does, and where it stopsThe result
University admin becomes one outcomeCheck every source. Report only what changed. Never submit or message.One verified morning digest. Canvas, MyHub and Outlook checked. Missing sources remain unknown.
Finance becomes an exception feedRead balances and transactions. Surface anomalies. Never move money.One read-only financial exception feed. Balances and anomalies surfaced. No transaction authority exists.
Job search becomes a pipelineCollect, deduplicate and score roles. Draft applications. Never submit silently.One deduplicated opportunity pipeline. Roles collected and scored. Applications remain drafts until authorised.
Software delivery becomes observable workIssue to branch to test to PR to deploy to verified product outcome.One source-to-production result chain. Issue, branch, tests, PR, deploy and product evidence remain linked.

The demos stop one frame too early

They show the agent moving. Delegation needs the next frame: did the requested state actually exist?

A visible crash is annoying. A polished false success is dangerous because the human stops checking.

A clean run can still end in a quiet false success

One plain task: at 6:00, check my university and do not make me check it again. Read Canvas, MyHub and UON Outlook. Report only changed obligations. Never submit, message, enrol or pay. If a source is missing, return UNKNOWN and do not guess.

Demo loop

Observe → choose → click → say “done”. It stops at motion.

  1. Canvas checked.
  2. MyHub checked.
  3. Outlook timed out. Empty result accepted.
  4. No change reported.

Result: a false “no change”.

Delegation loop

Observe → act inside authority → read back reality → issue receipt. It stops at evidence.

  1. Canvas current and complete.
  2. MyHub current and complete.
  3. Outlook source missing. Preserving UNKNOWN.
  4. Source recovered. Deadline change found.
  5. Private calendar updated and read back.

Result: a verified change.

Both runs are scripted to show the difference between the loops. They are not live results.

Put the proof layer outside the model

The model can propose. Software controls what happens, and evidence decides whether it worked.

  1. Freeze the outcome. Name the sources, permitted actions, prohibited effects and stop condition before execution.
  2. Issue narrow authority. The agent receives a short-lived capability. It cannot invent a new permission.
  3. Preserve uncertainty. Missing, stale or contradictory evidence blocks a success claim.
  4. Read back the real state. “Done” becomes verified only after an authoritative postcondition check.
  5. Return one receipt. Sources, authority, effects, evidence, cost and version travel together.

For the university task, the receipt lists four expected sources, read-only authority, preserved unknowns, one reconciled effect and authoritative evidence. Its result reads HOLD while a source is unknown, and VERIFIED only after the readback.

The cost collapse is real. The supervision cost is the trap.

A frontier model does not need to deliberate over every click. Plan once, use cheap bounded decisions, and compile stable paths into software.

Human checking and failure recovery often cost more than tokens

Two editable models. The first prices the browser admin you could safely delegate. The second compares the model cost of one run of browser decisions made three ways.

Price the tabs you keep reopening

Potentially returned each year110 hours
Illustrative annual valueA$5,475

Buy cognition where it changes the decision

Frontier every step$0.60
Plan + Jev + verify$0.064
Compiled stable path$0.008

Illustrative 9.4× lower model cost

Show the cost model
C_hybrid = C_plan + nC_jev + C_verify + C_browser

The hybrid route counts two frontier calls (plan and verify), n Jev decisions priced at $0.042 per million input tokens, and $0.00034 of browser cost. The compiled path is a fixed $0.006 plus $0.00008 a decision, with the per-decision part capped at $0.004. The time model is minutes a day × 365 × the delegated share.

Illustrative and editable inputs. Neither model includes human review or failure recovery.

Long chains turn “pretty reliable” into fragile

If each probabilistic step succeeds with probability p, a simple independent chain succeeds with pn. Checking and recovering at stage boundaries changes that picture.

A raw chain loses runs that checked stages can recover

P(all succeed) = pn for the raw chain. Change the inputs.

Raw chain66.8%
Illustrative checked stages85.5%

Checked stages group the chain into ⌈n ÷ 5⌉ stages, at least two. A failed stage is recovered with probability 0.75 × the checkpoint recovery rate.

The reliability equation is a teaching model. Steps are assumed independent, and the recovery factor is illustrative.

One confidence threshold is meaningless

The same confidence score cannot license every action. The error budget tightens as the consequence grows.

ConsequenceError budgetWho may act
Read-only observationA wider error budget may be acceptable because no external state changes.Jev may execute
Reversible internal writeThe decision must clear a tighter budget and support authoritative readback.Verify before commit
External communication or commitmentThe action leaves the private system. Confidence alone cannot create authority.Calvin or exact policy

The model advantage will shrink. The workflow evidence can compound.

Everyone can rent the same frontier model. Fewer people will have your source map, failure history, authority rules, recovery paths and verified receipts.

Model access, browser libraries, generic agent loops and basic tool calling are commoditising, and their price is falling. Source contracts, exception cases, verifier negatives and outcome receipts are private, and the evidence they hold grows with use.

Every successful run can expose a stable pattern. The system can move that pattern down the stack until the general model is needed only for exceptions.

How a workflow moves down the stack as verified successful traces accumulate
StageWhat has changedCostReliabilityHuman burden
Frontier explorationfrom 0 tracesUnknown states still need general reasoning and human review.highestunprovenfrequent
Typed macro-optionfrom 30 tracesThe goal, inputs, outputs, stop rules and authority are now explicit.highmeasurableregular
Jev reflexfrom 100 tracesBounded ambiguous choices are locally calibrated and cheap.lowstate-specificexceptions
Playwright routinefrom 250 tracesStable interaction has become deterministic browser software.very lowtestabledrift only
API or ordinary softwarefrom 500 tracesThe browser disappears where a structured operation now exists.lowestcontractedrare

Generic agent access is becoming ordinary. Workflow-specific evidence is still scarce. The window is the chance to build source maps, exception history, authority rules and trusted receipts before the generic interface absorbs the obvious tasks. The time-sensitive move is to start earning private, verified operational knowledge while the interface layer is still changing.

Start with one real, read-only workflow

Choose a repetitive, read-only outcome. By Friday, have a source map, a deterministic baseline and a shadow receipt.

  1. Today: name one outcome. One result, not a universal agent.
  2. Day 1: freeze sources and authority. What must be checked? What is prohibited?
  3. Day 2: build the boring baseline. API or Playwright first.
  4. Week 1: run shadow. Compare the agent against truth.
  5. After evidence: promote, hold or demote. Never on “looks good”.

The browser is becoming an automation layer for the whole internet. That could return weeks of repetitive software use each year. It could also produce polished, silent errors at machine speed.

The useful system is the one that can act inside a narrow contract, preserve uncertainty and return evidence. The technical field guide turns this thesis into a concrete architecture: capability leases, Jev gates, durable state, verifier independence and promotion evidence.

The demos proved the agent could act. None of them proved how to leave it alone.

For coding agents: an operating contract

Instead of asking an agent to “automate this”, give it a contract. Pick one workflow, copy the contract, and make the agent return the missing fields before it builds anything. The contract separates outcome, authority, evidence and unknown handling.


      

The same package as plain files:

Sources and further discussion

Vendor claims are labelled. Cost assumptions are editable. The reliability equation is a teaching model. This article does not claim the described system has earned production autonomy.

  1. Anthropic, computer use announcement. Official product announcement, October 2024.
  2. OpenAI, Introducing ChatGPT agent. Official announcement, July 2025.
  3. Browser Use, Jev Ultrafast. Project repository and narrow flight-search demonstration.
  4. OpenAI, Previewing Ultrafast mode. Official speed announcement, August 2026.
  5. TypeSafe Jev documentation. Typed questions, structured outputs and probabilities.
  6. τ-bench. Repeated consistency in tool-agent interaction.
  7. Temporal workflow execution. Durable event-history model.
  8. OpenAI API, computer use guide. Developer documentation for browser and desktop control.

These informed the original interactive edition’s format rather than the argument:

  1. Nielsen Norman Group, Progressive Disclosure.
  2. Sweller, Cognitive Load During Problem Solving.
  3. Mayer, Coherence Principle.
  4. W3C, Pause, Stop, Hide.
  5. Wikipedia’s signs-of-AI-writing field guide. Used as an editing checklist, not a detector.
  6. The Fountain Institute, signs of vibe-coded UI. Used as a design-audit prompt.
Published 2026-09-02 · Updated 2026-09-28 · Source edition series