← Back to writing

I tried to give an AI agent real authority. The model wasn’t the hard part.

Intelligence is cheap. Trusted authority is scarce. Capable AI agents still need bounded actions, commit-time authority, external verification and evidence-gated autonomy.

The models are already useful. The dangerous leap is treating a plausible next action as permission to change the world. I wanted a system that could answer five questions before any effect: who authorised this, what exact action is legal, is the evidence still fresh, did it happen once, and what is true now?

Agents can act before businesses verify what actually happened outside. Bounded authority converts model capability into safer operational leverage today. If that control layer works, better models also let the same business safely delegate more work, reduce review, and compound a private corpus of verified operating evidence.

Model access is common. Verified authority is scarce. My bounded reference system passed 106 automated checks across 10,097,601 approximate mock and model exposures, and its evidence still stops at A2: prepare, never commit. Production autonomy is blocked. The tests are bounded and synthetic. The strategic thesis is mine. Neither is presented as production proof.

The opening is the gap between agent capability and trusted deployment

2026 has two simultaneous stories. Agents are being pushed into real workflows. Almost everyone still lacks a scalable way to decide when those agents may act. Six current signals point the same way, and each has its own limit.

SignalWhat it showsIts limit
Cisco Security, March 202685% → 5%85% of surveyed organisations were experimenting with, piloting or deploying agentic AI. Only 5% reported broad production. Agents are everywhere in pilots and almost nowhere in broad production.Cisco surveyed its own customer base; the results are not a census of all organisations.
Dynatrace, The pulse of Agentic AI in 202669%69% of organisations manually verify agentic-AI decisions. Most still put a person behind every important decision, which is an approval queue waiting to become a bottleneck.A vendor survey of 919 global leaders; methodology and wording should be inspected in the full report before generalising.
Agentic AI in Industry, arXiv 2605.1467512 companiesA “capability-deployment verification gap”: experimental capability was ahead of production integration because credible verification was absent. The missing layer is qualification.Interviews with 16 practitioners from 12 companies. A small qualitative sample and preprint; useful for mechanisms, not population estimates.
ABC News, 19 September 2026Real reviewTasmania’s Justice Department will review AI use in parole board decisions after decision materials cited fictitious case law likely produced by AI.One reported incident under review; it does not establish prevalence across the justice system.
Business Insider, September 2026$2,000 capJPMorgan reportedly set a $2,000 monthly Claude cap for selected employees, plus an isolated Devspace environment.Paywalled reporting based on company sources; implementation details are incomplete.
Reuters, 16 September 2026900B agentsHuawei forecast up to 900 billion active agents by 2035 and more than 90% of AI-token traffic coming from agents.A corporate forecast, not an observed deployment count or neutral prediction.

The parole review shows this stopped being a hypothetical governance problem. Once generated material enters consequential decisions, “the model usually gets it right” is useless. The grown-up response is the one JPMorgan’s reported caps point to: budgets, isolation, identity, monitoring and an effect boundary the model cannot rewrite. The 900 billion figure is speculative, but the infrastructure race is already real. If agents become a major share of machine activity, authority infrastructure becomes basic business infrastructure.

My strategic bet is that this is a temporary build window. Model quality will keep spreading through APIs. What will not arrive automatically is your private map of legal actions, your failure corpus, your process integrations, your verified trajectories, and your evidence for safely reducing review. Model intelligence, general tools and commodity orchestration are rented. Operational truth, action contracts, release evidence and learned failure boundaries are owned, and every verified run can improve policy, evaluation, routing and economics. This is an inference from current adoption and verification gaps, not a forecast with a known expiry date.

The payment succeeded. The agent thought it failed.

The model acts on an incomplete picture of reality: timeouts, stale evidence, retries, drift and hidden external effects. It can be locally reasonable and globally wrong. A payment may already exist even when the response never arrived, and retrying from the agent’s local state can create a second real effect.

No response is not evidence that no effect occurred

A worked example: one invoice, INV-2048 for AUD 18,400, followed from proposal to timeout. The business system and the payment provider end up knowing different things.

TimeEventBusiness systemProviderKnown state
08:41:12.018The invoice passes every business rule. Purchase order found, vendor approved, amount matched. The model proposes PAY.Validated: one payment proposedReady: no effect yetREADY
08:41:12.441The payment request leaves the system with one intent, one idempotency identity and one permitted effect.Request sent: awaiting responseProcessing: intent receivedEXECUTING
08:41:12.827The provider commits the payment. A real external effect now exists. The agent has not seen the receipt.Awaiting receipt: effect status not yet observedCommitted: AUD 18,400 movedCOMMITTED
08:41:42.441The response never arrives. The only known fact is that no response arrived.Timed out: no response observedCommitted: effect exists externallyUNKNOWN
08:41:42.442What should happen next? A retry may duplicate the payment. Doing nothing may leave work unresolved.Uncertain: retry forbidden until reconciliationCommitted: authoritative state availableDECISION

Retry now

Fast and plausible, and potentially catastrophic.

The invoice is paid twice. Two payment attempts create two effects, and AUD 36,800 moves in total. The second request creates a second physical effect because the retry path did not reconcile authoritative state first. The agent followed a reasonable story and still caused a business loss.

Reconcile first

Slower, bounded and evidence-driven.

The existing payment is found and verified. The system reads authoritative external state, matches the immutable intent, records the receipt, and closes without creating another effect. There is one effect, and the receipt and state match.

The safe result also depends on the provider preserving idempotency semantics, the reconciliation read being authoritative, and restored state not losing prior effect history.

This is why “the model called the right tool” is not a safety claim. The business system must represent uncertainty honestly. It needs a durable UNKNOWN state, stable intent identity, authoritative reconciliation, and a prohibition against creating a second effect until the first is known to be absent. AWS’s guide to making retries safe describes the same pattern for idempotent APIs.

The test for this part of the system is simple: can it distinguish “failed” from “outcome unknown”?

Three useful agents have the same missing proof

Browsing, coding, planning, emailing and paying are visible. Permission, uniqueness, reconciliation, evidence and liability usually are not. Demos reveal capability. Businesses inherit consequences.

DemoWhat it showsThe proof it still needsBounded implementation
Refund $120A support agent sees an eligible order and proposes a $120 refund.Was this customer, order, amount, reason, and payment target authorised at commit time?Bind customer + order + amount; reject prior refunds; commit once; verify the payment ledger.
Deploy a patchA coding agent fixes a production bug and every unit test passes.Did it use the approved branch, pass integration checks, survive canary traffic, and remain reversible?Enforce branch, CI, canary, rollback, and health checks outside the model.
Send 2,000 renewalsA sales agent drafts and queues personalised renewal emails.Were recipients, claims, consent, timing, rate limits, and unsubscribe rules authorised?Bind audience + template; rate-limit; sample; verify delivery and CRM state.

A business needs a system that remains correct when certainty is unavailable, whether or not its agent sounds certain.

My opinion is that most “autonomous business” claims are capability demonstrations wearing an operations costume. They show that a model can produce a plausible next action. They rarely show that the action remained authorised at commit time, occurred once, produced the intended external state, survived restart, or created risk-adjusted value.

The numbers become uncomfortable before the model becomes stupid

Long trajectories turn tiny per-step errors into ordinary failures. Each equation below answers a different question and exposes a different way to fool yourself.

Even high per-step reliability compounds downward across a long workflow

P(all steps succeed) = pn. Toy assumption: steps are independent and equally reliable. Change either input; the curve is calculated in your browser.

Whole-trajectory success90.5%

About one trajectory in 11 contains at least one failed step.

Trajectory reliability chart A curve showing whole-trajectory success declining as the number of steps rises from 1 to 500.

A toy model. Real trajectories often share correlated failure modes.

A small failure probability can dominate when the downside is large

EV = P(S)B − P(H)L − C, per run. Toy assumption: probabilities and losses are stable, estimable, and expressed over the same decision horizon.

Expected benefit$95.00
Expected harm$100.00
Operating cost$2.00
Expected value−$7.00

This agent completes useful work and still destroys expected value because downside dominates.

Illustrative inputs. The conclusion changes if tail loss, review cost, or failure dependence was underestimated.

Zero observed failures is evidence. It is not the same statement as zero risk. The rule of three gives a rough upper 95% failure bound after zero failures in n independent trials: about 3 ÷ n. After 10,000 zero-failure trials, the upper failure rate is about 0.0300%, so the data remain compatible with roughly one failure per 3,333 runs. Independence, representativeness and stable conditions still matter.

Then there is the question of what an agent programme should optimise. Task volume is not the numerator. Verified economic value is.

Leverage = verified value / (attention + compute + risk)

Verified business value is incremental outcome × confidence × durability. The denominator is human attention, compute and infrastructure, and expected and tail-risk cost. This is a decision ratio, not an accounting identity; units and time horizons must be defined before comparison. The useful question is how much verified value the agent produced per scarce unit of oversight and risk, rather than how many tasks it completed.

The model belongs inside the control system, not above it

This is the architecture that survived longer than the diagram. Probabilistic intelligence proposes, deterministic software constrains, independent evidence verifies, and authority contracts when evidence weakens.

Reasoning can remain probabilistic. Authority cannot. A model may choose among candidates; software must define the legal action set, enforce the boundary, preserve evidence, and contract authority when assumptions fail. The probabilistic core interprets, plans, classifies, generates and abstains. The deterministic shell around it authorises, constrains, commits, reconciles, verifies, records and stops.

ComponentOwnsMust neverResidual failure
Frontier modelInterprets ambiguity and proposes intent.Reasoning, synthesis, candidate actions.Grant itself authority or declare external effects successful.Confidently proposes an illegal, stale, or underspecified action.
JEV-style selectorChooses inside a legal action set constructed by software.Bounded classification and routing.Invent arbitrary tools, selectors, fields, or side effects.Selects a legal action that is still wrong for this state.
Policy and authorityChecks current identity, scope, limits, evidence, and expiry.Permission and deterministic invariants.Accept model confidence as evidence of authorisation.Time-of-check and time-of-use drift invalidates a prior decision.
Effect boundaryCommits one authorised effect at the authoritative boundary.Idempotency, fencing, transactionality, reconciliation hooks.Assume a lost response means no effect occurred.Split brain or an external connector creates duplicate effects.
Independent verifierChecks what actually happened through a separate evidence path.Outcome verification and contradiction detection.Rely only on the executor’s own story about success.Shared credentials, provider, or database create common-mode error.
Evidence ledgerPreserves causal, versioned, tamper-evident decision evidence.Traceability, provenance, expiry, replay support.Turn logs into proof without validating their source.Gaps, stale evidence, or rollback hide the real causal sequence.
What happens when a fault is injected at each layer
  • Frontier model. The model proposes an action outside the legal set. The selector rejects it before authority is considered.
  • JEV-style selector. The selector chooses a legal but state-inappropriate action. Preconditions fail, and the workflow escalates instead of committing.
  • Policy and authority. The policy version changes after planning. Commit-time fencing rejects stale authority and freezes the effect.
  • Effect boundary. The connector times out after commit. The state becomes UNKNOWN and reconciliation replaces blind retry.
  • Independent verifier. Executor and verifier disagree. Contradictory evidence contracts authority and routes the case to investigation.
  • Evidence ledger. A causal sequence gap appears. The evidence chain is invalidated, preventing it from supporting release or recovery.

The test for the control layer: can any prompt, model, or tool bypass the effect boundary?

Do not spend frontier tokens on deterministic work

Use the frontier model for ambiguity. Use a small decision model for bounded choice. Compile stable invariants into software. The correct objective is total cost, not the lowest token bill.

C_run = n_F·c_F + n_D·c_D + C_review + E[L]

That is frontier calls, plus decision calls, plus human review, plus expected loss. Each layer of the routing stack earns its place differently:

  1. Frontier model. Novel ambiguity, synthesis and hard planning. Expensive, flexible and probabilistic.
  2. JEV-style decision model. Chooses one legal action from a typed set. Cheaper, bounded and testable.
  3. Deterministic software. Policy, budgets, identity, idempotency and invariants. Fast, auditable and enforceable.

The failures I would rather have found before publication

A serious assurance programme has to attack its own evidence instead of only generating more green checks. The most valuable tests were the ones that defeated my own claims.

DefectBefore and attackRepairStill unresolved
Evidence expired after validationCriticalFreshness was checked when work began. The request waited, then executed after the evidence expired.Recheck trusted time and evidence at the commit boundary.Trusted time and restart behaviour still require production evidence.
Policy changed between check and effectCriticalPolicy validation happened in application code. The policy version changed before the external commit.Move policy fencing into the same authoritative transaction as the effect.Third-party systems may not expose an atomic policy boundary.
The executor graded its own workCriticalThe adapter executed and reported the outcome. A colluding adapter fabricated both execution and verification evidence.Separate executor, verifier, credentials, and evidence paths.Common providers and trust roots can still correlate failures.
Poisoned memory became future policyHighThe agent reused prior success memories. A false memory encouraged unsafe repetition in later cases.Add provenance, expiry, scope, quarantine, and independent outcome confirmation.A compromised authoritative source remains difficult to detect.
The API schema stayed stable while meaning changedHighCompatibility checks compared fields and types. Units and business semantics changed without breaking the schema.Bind semantic digest, units, mode, version, and idempotency contract.Semantic contracts still require owners and requalification discipline.
The evaluator learned to be impressedHighA visible model grader scored the agent’s performance. The agent optimised style and rubric cues instead of real outcomes.Use sealed holdouts, deterministic outcome checks, counterbalancing, and human calibration.Open-ended quality judgments remain contestable and expensive.
Human review became a denial-of-service attackHighUncertain work was escalated to people. The queue exceeded reviewer capacity and encouraged rubber-stamping.Use risk-ranked admission, reserved capacity, and safe no-effect freezing.Safe freezing can still destroy useful throughput and customer experience.
Two correct systems both committedCriticalEach effect boundary enforced local uniqueness. A network partition created two independent authoritative views.Use consensus, fencing, exclusive partition ownership, or one serialising authority.This remains an open production blocker across independent targets.

The purpose of verification is to find the strongest reason the system should not ship.

“It passed the eval” is an unfinished sentence

Passed which evaluation, against which oracle, under which assumptions, covering which trajectory, for which authority decision? A green eval cannot prove what its evaluator never measured. The evidence has three separate jobs:

  1. Prove the shell. Can software violate the encoded boundaries? Use contracts, model checking, tests, and fault injection.
  2. Evaluate the agent. How does probabilistic behaviour change across real trajectories? Score tools, paths, abstention, outcomes, and repeated trials.
  3. Validate reality. Did people and the business receive measurable net value? Use independent state, field trials, harms, costs, and counterfactuals.
Six layers of evidence, and what each one can and cannot establish
LayerCan establishCannot establishExample property
1. RequirementsWhat must always be true?Define losses, authority boundaries, invariants, operating conditions, and prohibited states.Show that implementation or deployment actually satisfies them.No consequential effect may occur without current authority and fresh evidence.
2. Formal methodsCan the encoded state machine violate its invariants?Explore finite states, ordering, retries, authority transitions, and deadlocks.Prove open-world model behaviour, unmodelled connectors, people, or business value.Unknown outcomes cannot transition directly into a second effect.
3. Software testsDoes the implementation enforce the intended mechanics?Exercise code, contracts, fault handling, mutation sensitivity, and regressions.Establish performance on unrepresented real-world distributions.A policy version changing before commit must reject the operation.
4. Agentic evaluationsHow does the model-plus-harness behave across trajectories?Score tools, paths, outcomes, abstention, long horizons, and repeated stochastic trials.Grant authority when graders, datasets, or assumptions are weak.A correct final answer reached through a prohibited tool path still fails.
5. Runtime assuranceAre assumptions and evidence still valid now?Detect drift, freeze effects, reconcile uncertainty, and contract authority.Recover facts that no independent source can observe.Critical evidence expiry switches the workflow into no-effect mode.
6. Business validationDid the system create incremental risk-adjusted value?Compare outcomes, costs, harms, review burden, and opportunity cost.Be replaced by task counts, benchmark scores, or polished demonstrations.Shadow and controlled field trials estimate value against a counterfactual.

Formal methods do not prove business value. Agentic evals do not prove commit-time authority. Runtime monitors do not repair an unknowable external state. Human approval does not prove meaningful review. The stack is deliberately non-substitutable: strong evidence at one layer cannot compensate for a critical failure at another, and a utility score cannot average away unauthorised money movement, privacy leakage, or duplicate irreversible effects.

The numbers looked impressive until I asked what they authorised

The same four numbers from my execution record read two ways.

The bounded system passed its tests. The evidence still stopped at A2.

The promotional reading and the defensible reading of the same evidence, from the Bounded Agentic Assurance V10 execution record.

EvidencePromotional readingDefensible reading
106 automated checks passedA broad programme of tests passed without a reported failure.The tests validate encoded behaviours inside the reference package. They do not represent production traffic, credentials, connectors, or people.
10,097,601 approximate mock/model exposuresMillions of virtual cases exercised the bounded system.The exposure count aggregates synthetic trials and model obligations. It is useful for stress, not a substitute for representative field evidence.
0 safe-model violation classesNo encoded safety invariant was violated in the explored safe model.Zero applies only to represented states, transitions, invariants, depth, and assumptions. Unmodelled reality remains outside the claim.
10/10 control ablations detectedEvery removed control caused a detectable degradation.This is useful mutation sensitivity, but the ablation catalogue cannot prove that every important control or failure mode was included.

Authority earned: A2, meaning observe, recommend, draft, prepare, and simulate. Authority withheld: A3 and above, meaning no autonomous consequential effect.

Synthetic and bounded reference evidence. No production autonomy claim.

Intelligence is continuous; permission should be discrete

Authority should expand slower than capability. The ladder marks the highest level the current evidence supports.

LevelPermissionRequired evidenceExampleCurrent evidence
A0ObserveRead state and summarise. No proposed change.Identity, data scope, provenance, privacy boundaries.Summarise an inbox without modifying messages.Supported
A1RecommendPropose an action and explain evidence. No external effect.Grounding, calibration, escalation, user comprehension.Recommend which invoice needs human review.Supported
A2PrepareCreate drafts or sandbox artifacts for independent verification.Deterministic validation, bounded tools, trace capture, rollback.Prepare a payment instruction that cannot execute.Supported. Current ceiling.
A3Approved effectExecute one consequential action after meaningful approval.Independent approval, commit-time authority, idempotency, reconciliation.Submit an approved low-value payment once.Blocked
A4Bounded autonomyExecute within a narrow, reversible, monitored envelope.Shadow and canary history, rare-event bounds, runtime contraction.Resolve validated low-risk tickets below fixed limits.Blocked
A5Continuous operatorRun workflows continuously and escalate exceptions.Sustained field reliability, incident drills, human-capacity proof.Operate a business process with humans handling anomalies.Blocked
A6Broad autonomyChoose and coordinate consequential goals across functions.No credible evidence currently supports this general claim.Independently manage finance, sales, engineering, and compliance.Blocked

The durable advantage is verified operating evidence

Everyone will gain access to stronger models. Fewer organisations will accumulate the evidence required to delegate real authority without adding an equal amount of human review.

General intelligence is what gets cheaper: reasoning tokens, browser and coding agents, generic orchestration and standard tool protocols. Competitors can buy those models and tools too. What you own, and what compounds, is your operating system: action contracts and process truth (your exact operating constraints), verified trajectories, a failure and incident corpus, and domain-specific release evidence that can safely reduce review.

Every verified run can convert supervision into reusable operating knowledge. Trajectories, overrides, failures and reconciled outcomes can improve policies and evaluations. Authority expands only while fresh evidence supports it.

Qualified evidence can release review time

An illustrative planner. Verified trajectories accumulate each month; once the evidence qualifies, a smaller share of them needs review.

Verified trajectory corpus12,000
Human hours released450
Eight-hour workdays56
Authority claimNone without field validation

Hours released = trajectories a month × months × (review rate before − review rate after) × minutes per review ÷ 60.

This is a planning model, not a productivity promise. It shows why the evidence loop can become more valuable than temporary model access.

The test for the payoff: did verified value rise faster than attention, compute, and expected loss?

The unit of company leverage could change

This is conditional, not promised. It requires the controls, field evidence and economics to survive reality. If they do:

  • A small team can supervise a private digital workforce instead of manually operating every workflow.
  • Humans move from touching every task to handling novel states, exceptions, policy changes, and relationships.
  • Businesses that were uneconomic because coordination was expensive can become viable.
  • Every cheaper or stronger model improves the same controlled operating system rather than forcing a redesign.

The next experiment should replace synthetic confidence, not expand scope

The next step is one boring workflow, with OpenClaw and JEV in effect-disabled shadow mode, complete trajectories, independent outcomes and no production heroics.

  1. Choose one boring workflow. Frequent, observable, low-consequence, reversible, and measurable. Avoid broad business management.
  2. Run effect-disabled shadow mode. Let OpenClaw and JEV propose actions while software prevents real-world effects.
  3. Preserve complete trajectories. Capture prompts, tools, arguments, policy versions, states, evidence, abstentions, and outcomes.
  4. Grade outcome and process. Use authoritative state, deterministic checks, protected evaluators, and sampled expert review.
  5. Replace synthetic parameters. Estimate real error, review burden, cost, drift, and incremental value before seeking A3 authority.

During the experiment the agent may read approved inputs, propose bounded actions, explain evidence, abstain, and create sandbox artifacts. It may not commit external effects, contact third parties, move money, change identity, or redefine its own authority. The experiment succeeds when real trajectories replace synthetic assumptions and support a narrow, reversible, monitored A3 decision.

The interesting future is a company where intelligence is cheap, actions are typed, authority is explicit, effects are reconciled, and every run produces evidence that can justify the next small expansion of autonomy. If that works, a founder saves time and also owns an operating system that becomes more useful every time models improve and every time the business learns where the software fails.

That is the advantage worth acting on now. The future is not guaranteed, but the evidence, process knowledge, and control infrastructure take time to accumulate. The missing proof is the product, the moat, and the permission to act. Everyone showed what agents could do. None of them proved how.

For coding agents: the operating contract and brief

The page gives a human a decision framework: where the strategic opening is, which failures matter first, what each assurance method can prove, and when more autonomy is economically justified. It gives an agent an implementation brief: allowed and prohibited actions, required evidence and stop conditions, retry and reconciliation semantics, and evaluation and release-gate questions.

The prompt I would actually use:

Read this article and its agent-brief.json. Do not implement effects. Produce: the operating domain, legal action set, authority matrix, evidence model, top ten failure modes, evaluation plan, runtime monitors, release gates, and the smallest shadow-mode experiment. Mark every unsupported assumption.

To build a bounded operating contract for your own workflow, choose its limits. The contract stays blocked when the choices are inconsistent with the evidence.

A2 is structurally consistent. Effects remain disabled until field qualification.

{
  "schemaVersion": "1.0",
  "articleRelease": "v6",
  "operatingDomain": "operations",
  "maximumAuthority": "A2",
  "effectiveExternalEffect": "none",
  "effectEnabled": false,
  "releaseState": "STRUCTURALLY_CONSISTENT_BUT_UNQUALIFIED",
  "unknownOutcomeAction": "RECONCILE_THEN_FREEZE"
}
Sources and further discussion

Current signals are presented as signals. Forecasts stay forecasts. Synthetic evidence stays synthetic. Each source says what it was used for and where its claim stops.

  1. NIST, TEVV-Athlon Framework for Evaluating AI Systems. Evaluation framing: objectives, context, measurement concepts, events, tools, and lifecycle evidence. Boundary: the cited 2026 document is an initial public draft.
  2. Anthropic Engineering, Demystifying evals for AI agents. Evaluation vocabulary: tasks, trials, graders, transcripts, trajectories, outcomes, harnesses, and repeated trials. Boundary: vendor engineering guidance; domain-specific assurance still requires independent design.
  3. AWS Builders’ Library, Making retries safe with idempotent APIs. Stable caller intent, idempotency identity, retries, and unknown-outcome handling. Boundary: patterns must be adapted to each external system’s actual guarantees.
  4. Google Site Reliability Engineering, Addressing Cascading Failures. Retry budgets, load shedding, admission control, graceful degradation, and overload testing. Boundary: service reliability patterns do not by themselves validate agent decisions.
  5. IETF RFC 9334, Remote ATtestation procedureS Architecture. Separate evidence producer, verifier, and authority-deciding relying party. Boundary: remote attestation concepts do not prove external business outcomes without appropriate evidence.
  6. Model Context Protocol: Authorization. Protocol authentication and authorization boundaries for tool connectivity. Boundary: protocol authorization does not substitute for local business policy and effect controls.
  7. Bounded Agentic Assurance V10 execution record. My research artifact: article-specific synthetic metrics, failure simulations, and authority ceiling. Boundary: synthetic and bounded reference evidence; no production autonomy claim.
  8. Cisco Security, The Agent Trust Gap, March 2026. Current adoption gap: 85% experimenting, piloting, or deploying; 5% broad production; security and access control remain core barriers. Boundary: Cisco surveyed its own customer base; the results are not a census of all organisations.
  9. Dynatrace, The pulse of Agentic AI in 2026. LinkedIn post and survey report; the published graphic states that 69% manually verify agentic-AI decisions. Boundary: a vendor survey of 919 leaders; methodology and wording should be inspected in the full report before generalising.
  10. Apostolou, Bosch, and Holmström Olsson, Agentic AI in Industry: Adoption Level and Deployment Barriers. Industrial evidence for a capability-deployment verification gap across interviews with 16 practitioners from 12 companies. Boundary: a small qualitative sample and preprint; useful for mechanisms, not population estimates.
  11. ABC News, Tasmania’s Justice Department to review AI use in parole board decisions, 19 September 2026. A current Australian high-stakes example where fictitious case law, likely AI-generated, entered decision materials. Boundary: one reported incident under review; it does not establish prevalence across the justice system.
  12. Business Insider, JPMorgan rolls out Claude changes: spending limits and extra security, September 2026. Illustrates cost caps and isolated execution environments appearing in serious enterprise deployment practice. Boundary: paywalled reporting based on company sources; implementation details are incomplete.
  13. Reuters, Huawei forecasts billions of agents will dominate AI traffic by 2035, 16 September 2026. Illustrates the scale of current industry expectations for agent-generated traffic and infrastructure demand. Boundary: a corporate forecast, not an observed deployment count or neutral prediction.
  14. Michels et al., Vibe Coding: Practice, Performance, Productivity, and Risk, arXiv 2608.20446. Research synthesis showing that productivity claims vary substantially with task, measurement method, codebase maturity, and time horizon. Boundary: a recent preprint review; some included studies and claims may change after peer review.

These informed the original interactive edition’s editing, design and format rather than the argument:

  1. Wikipedia: Signs of AI writing. Editorial audit: generic significance claims, vague attribution, superficial analysis, formulaic structure, excessive formatting. Boundary: the page describes possible patterns, not reliable proof of AI authorship.
  2. Shin et al., Interrogating Design Homogenization in Web Vibe Coding, arXiv 2603.13036. Frictionless generation can reproduce dominant conventions; productive friction can preserve creator intent. Boundary: a 2026 preprint and sociotechnical analysis, not a universal visual-quality test.
  3. Nielsen Norman Group, Progressive Disclosure. Reveal advanced detail when needed while keeping core actions visible. Boundary: progressive disclosure can hide important information when hierarchy is designed badly.
  4. Nielsen Norman Group, 10 Usability Heuristics for User Interface Design. Visible state, user control, error prevention, recognition over recall, and recovery support. Boundary: heuristics guide review; they do not replace testing with representative users.
  5. W3C WCAG 2.2, Animation from Interactions. Reader-triggered animation must be disable-able unless essential. Boundary: passing one motion criterion does not establish complete accessibility.
  6. Nielsen Norman Group, State of UX 2026: Design Deeper to Differentiate. AI trust requires transparency, control, consistency, and support when systems fail. Boundary: expert synthesis, not a controlled validation of this article’s interface.
  7. Business Insider, There are 3 telltale signs that you used AI to make your app, July 2026. Critique of homogeneous visuals, polished surfaces masking weak function, and neglected edge cases. Boundary: journalistic synthesis and expert commentary, not a formal UI-quality instrument.
  8. Brysbaert, How many words do we read per minute? A review and meta-analysis of reading rate, 2019. Reading-time calibration; the interactive edition used a conservative 200-wpm technical baseline. Boundary: population averages do not predict one reader, one device, or comprehension of specialised material.
  9. W3C WCAG 2.2, Understanding Success Criterion 1.4.10: Reflow. Text must reflow without two-dimensional scrolling at narrow equivalent widths. Boundary: a single criterion does not establish complete accessibility.
  10. W3C WCAG 2.2, Understanding Success Criterion 2.5.8: Target Size (Minimum). Minimum target-size and spacing checks for interactive controls across mobile layouts. Boundary: minimum conformance is not the same as comfortable use.
  11. W3C WCAG 2.2, Understanding Success Criterion 2.3.3: Animation from Interactions. Motion controls and reduced-motion fallbacks for reader-triggered animation. Boundary: disabling motion does not by itself guarantee cognitive or vestibular accessibility.
  12. Nielsen Norman Group, Progressive Disclosure. The reading lenses of the interactive edition: essential claim first, narrative context on request, engineering evidence on demand. Boundary: poorly chosen hierarchy can conceal important information; the split still requires user testing.
  13. Mayer, Multimedia Learning, third edition: signaling and contiguity principles. Keep explanatory words, visual state, and causal timing close enough to be mentally integrated. Boundary: learning principles do not yield a universal scroll formula.
  14. Segel and Heer, Narrative Visualization: Telling Stories with Data. Balance authored narrative sequence with reader-controlled exploration. Boundary: a design-space analysis, not evidence that any one narrative interface improves comprehension.
  15. Morkes and Nielsen, Concise, SCANNABLE, and Objective: How to Write for the Web. Concise language, scannable structure, and objective wording as separate usability levers. Boundary: a small late-1990s study; its exact effect sizes should not be treated as universal contemporary constants.
  16. Pirolli, Rational Analyses of Information Foraging on the Web, 2005. Readers continually compare expected information value against the cost of continuing. Boundary: a cognitive modelling framework, not a direct comprehension test of this article.
  17. Chotisarn et al., VISHIEN-MAAT: Scrollytelling visualization design for explaining Siamese Neural Network concepts, 2023. Coordinate authored sequence, reader pace, and visual explanation for a complex AI concept. Boundary: a specific study and concept; it does not prove every scrollytelling interface improves understanding.
  18. Gooding et al., Predicting Text Readability from Scrolling Interactions, 2021. Scroll behaviour and reading difficulty as related signals. Boundary: interaction patterns predict readability statistically; they do not reveal whether one reader understood a specific claim.
  19. Nielsen Norman Group, Information Foraging: A Theory of How People Navigate on the Web. Information-scent guidance for labels, routes, and deciding whether continuing appears worth the effort. Boundary: a practitioner synthesis; representative user testing remains necessary.
  20. Kong et al., Getting the Most from Eye-Tracking: User-Interaction Based Reading Region Estimation, 2023. Distinguishing skip, skim, and detailed reading at the content-region level. Boundary: the interactive edition used heuristic simulations, not eye-tracking or validated per-user reading inference.

The editorial audit used Wikipedia’s field guide as a prompt to remove generic significance claims, vague attribution, superficial analysis, formulaic conclusions, excessive formatting, and language that smooths away specific facts. The guide itself warns that these patterns are descriptive rather than proof of AI authorship.

The interactive edition used three reading lenses and viewport choreography that estimated exposure time from word count, animation duration and decision time, with an appendix of those heuristics (exposure time, scroll runway, viewport occupancy, time to core, attention survival and comprehension rate). They were explicit design heuristics, not cognitive laws, and they governed that edition’s layout only. This edition uses the site’s reading layout.

Published 2026-09-03 · Updated 2026-09-28 · Source edition v6