← All writingCalvin Kennedy Writing
BESF / 01 · The work after the code

Your agent finished.
You’re still working.

The patch is ready. The decisions are still yours. I want the next run to bring back a checked result and the one thing it couldn’t settle.

Cheap, broad inspection. Code for facts. Frontier reasoning for the part that still needs it.

See what changes ↓
YOUR NEXT REVIEW SESSIONFICTIONAL BACKLOG

Six tasks. All back on your desk.

01Keep access after cancellationUnresolved
02Stop duplicate bookingsUnresolved
03Correct the refund totalUnresolved
04Will buyers pay for export?Unresolved
05Can we delete old accounts?Unresolved
06Does the signup work on iOS?Unresolved

See the proposed division of labour, not a live agent run.

Scripted outcomes. “Reviewable” is not “shipped”; the two evidence gaps remain unresolved.

FOR YOULess work deciding what “done” means.

Get evidence and the unresolved choice together.

FOR THE AGENTA next action it can justify.

Build, check, gather evidence, repair or stop.

FOR THE PROJECTThe next run starts further ahead.

Keep the rules, failures and checks—not only the chat.

01 / A problem worth caring about

They cancelled renewal.
You cancelled their access.

One paid month. One button. A plausible patch can still break the promise the customer bought. This is a constructed example, not a reported incident.

01

A customer paid for October. Turning off renewal must not remove the access already bought. That rule has to come from the product owner, not from a guess.

02

Date boundaries, retries, tenant isolation, interface copy, missing evidence. Code handles exact facts. A Jev-like layer flags interpretation and gaps.

03

Once the outcome and failure are explicit, the generator has a target. Another long opinion is unnecessary when a small executable check can answer.

04

A passing check belongs to a specific revision. It earns a reviewable packet, not permission to deploy. Missing rules and withdrawn authority stop the path.

FIG. 01 · THE SAME CUSTOMER THROUGH THE SYSTEMLOCAL DEMO
1 OctPays $29 for one month
12 OctTurns renewal off
1 NovPaid access ends
Still has access
Locked out
The obvious-looking implementation treats “not renewing” as “not entitled.”
ExampleExpectedObservedCheck
6 Oct: paid accountAccessAccessPass
12 Oct: renewal offAccessNo accessFail
1 Nov: period endedNo accessNo accessPass
REPAIR REQUIRED

The missing entitlement is observable. A confident approval does not change it.

October is an illustrative billing period. The browser executes two small candidate functions against the stated examples. There is no account, model call or deployment.

Open the actual contract, candidate and failure boundary
promise: cancel renewal, keep access until paid_through
wrong:   active = !cancelled && now < paid_through
fixed:   active = now < paid_through

ready = tests_pass && rules_present && authority_current

The end-to-end product would also need billing events, time zones, identity, concurrency, recovery and actual UI tests. The included offline Python fixture checks a declared finite day domain. Missing product meaning is not solved by asserting the wrong rule more thoroughly.

02 / The original BESF idea

Give it more ways to think.
Fewer reasons to interrupt.

The system is broader than regression tests: inspect the task from many useful angles, then spend effort on whichever next action could change the decision.

FIG. 02 · Wide observation, bounded action
One authorised task Relevant state + accepted outcomesources · revision · budget · permissions Code: exact checksfacts / tests / arithmetic Jev: broad inspectiontyped observations Missing evidenceretrieve / measure / ask Choose the next permitted actionfrontier builds or repairs only where needed Verify → reviewable packet, or explicit stop new evidence ↺One authorised taskRelevant state + outcomesources · revision · budgetInspect in parallelCode: exact facts and testsJev: narrow judgmentsGaps: retrieve, measure, askChoose a permitted actionfrontier builds or repairs as neededVerify → reviewable resultor return a specific reason to stop

Proposed architecture. The complete orchestration has not been validated on live BESF work. Existing research and vendor examples already cover important parts; the combination is not claimed as a new invention.[1][6][14]

A giant tree is the wrong thing to store. Keep the distinctions that matter; explore the relevant branches.

For the cancellation change, the system should examine the paid-through date, repeat clicks, tenant scope, the interface promise and the authority to act. It should not ask a language model to subtract dates or invent a refund policy.

Paid-through field
CODE
Repeated cancellation
CODE
Tenant boundary
CODE
UI promise matches rule
SEMANTIC
Support explanation
SEMANTIC
Refund rule absent
UNKNOWN
Patch needed
FRONTIER
Runtime result
VERIFIER
Permission current
AUTHORITY
Earlier failure
RETRIEVAL
Buyer demand
OBSERVATION
Retained cost
CODE

The source research contains 48 named questions across product, demand, marketing, software, AI, harness, economics and evidence. Their coverage and Jev accuracy remain unvalidated on real tasks.

What the decision graph stores—and what it must never claim

Store task types, evidence requirements, action preconditions, measured outcomes, cost and prior counterexamples. Let a generator propose missing distinctions; test them on held-out tasks before promotion. Search can retain several candidate actions while code enforces budgets and permissions. A typed probability is an observation, not a fact or a permission. Never infer purchasing from a coding benchmark, or silently convert a simulated result into field evidence.

03 / The stake is your capacity

More code isn’t the prize.
More work you can take on is.

If the routine decisions stop returning to you, the same person might handle more useful work. If exceptions and upkeep grow instead, the system has simply moved the workload.

Hypothetical workload. Baseline: 12 owner minutes per task. Prepared result: 2 minutes to review. An exception needs the original 12 minutes plus 3 minutes of overhead. No quality uplift is assumed.

FIG. 03 · Does the work actually leave your desk?
4.6 hoursof owner time freed in this assumed week
Current workflow600 min
Prepared results + exceptions + upkeep322.5 min

A gain only counts if accepted quality and business outcomes hold up. Deferrals, missed requirements, support and fatigue are not measured here.

Inspect the equation and the reversal
T₀ = N h₀
T₁ = N[(1−d)r + d(h₀+e)] + u
Time freed = T₀ − T₁

N is task count; d the exception share; r the residual review; e exception overhead; u upkeep. These are editable planning inputs, not outputs of a validated capability model. Holding other defaults fixed, 80% exceptions makes the system take more owner time. A queue can also become unstable even when its average looks acceptable.

04 / Why experiment now

The ingredients are here.
The bottleneck is worth testing.

The opportunity is to turn current model access into reusable capability before the next repeated task. There is no sourced expiry date. You do not need a new platform or a larger model bill to begin.

OPENAI · 11 FEB 2026

Human QA became the constraint.

OpenAI’s internal engineering account identifies QA capacity as a bottleneck and describes giving agents direct access to runnable environments and feedback.

Read the engineering account ↗
TYPESAFE · CHECKED 20 SEP 2026

Wide inspection is cheap enough to try.

Current Jev pricing is $0.042 per million input tokens. Shared context makes a broad question set economical; its usefulness still needs local evaluation.

Read the current model documentation ↗
ANTHROPIC · 26 NOV 2025

The next session needs an inheritance.

Anthropic’s harness work describes incremental progress and persistent artifacts between sessions. The result should leave tomorrow’s agent a better starting point.

Read the harness account ↗

Attributed source notes, not social screenshots or independent replications. This is the current rationale for a bounded experiment, not evidence of a time-limited commercial advantage.

Before your next repeated task: record the owner time, run the plain baseline, then change one part of the workflow. Keep the change only if the outcome improves.
05 / What the research earns

A working mechanism.
A larger hypothesis.

The laboratory supports some implementation claims. It does not establish autonomous delivery, safe production use or paying demand. Passing more tests does not fill those evidence gaps.

EXECUTABLE HERE

A confident mock judgment cannot repair the wrong result.

The teaching controller checks the outcome and stops on missing evidence or withdrawn permission. The runnable starter has no network or production adapter.

RETAINED RESEARCH

Correlation, selection and missing requirements can defeat reassuring numbers.

The complete prior laboratory is preserved in the combined ZIP. Its historical test counts are not presented as new runs of this article.

NOT YET ESTABLISHED

Lower owner workload on representative BESF tasks.

Run code-only where applicable, verified frontier, and the proposed router on the same frozen task families and permissions. Measure final state, deferrals, rework and complete cost.

NOT YET ESTABLISHED

Real buyer, user and deployment outcomes.

Correct code does not show that anyone needs it. Simulated personas do not show that people understand this page. Real task and reader evaluation remain separate.

06 / Take one useful thing

Give it a task.
Require a decision packet.

Use a task already authorised in your workspace. Ask for the outcome, evidence and smallest remaining choice together. The actual harness or review process must enforce the boundary.

I don’t need my agent to be more convincing. I need the next decision to cost me less attention.

Start with a known failure or repeated piece of work. Keep the simple baseline. Add broad interpretation only where it changes what happens next. This is how I would test the BESF idea before building the giant version.

No signup. No installation required to read it. The ZIP adds a Python demonstration, checks, methods and the original research.

FIG. 04 · What should come back
OUTCOME Cancel renewal. Keep paid access. EVIDENCE Rule + account state + revision FAILURE Day 12: customer was locked out RESULT Corrected candidate passes checks BOUNDARY No live release or customer action UNKNOWN Refund policy still needs its owner NEXT Review the patch and that one choice

Illustrative packet. A structured format makes missing evidence visible; it does not make the content true.

They all said it worked.
None of them proved how.

“They” are the fictional approving reviewers in this demonstration. The missing paid access is the counterexample. The check shows its rule and result; it is not a universal mathematical proof.

Evidence trail

Open the source.
Keep the boundary.

Web sources checked 20 September 2026. Source notes are paraphrases unless explicitly marked as a short quote. Simulations are not live-model or customer results.

[S1] TypeSafe · Models ↗

Jev 1.13.0: $0.042/M input tokens, free output; text only; 64k request and 32k state-plus-longest-question limits. Provider documentation, checked 20 September 2026.

[S2] TypeSafe · Parallel questions ↗

Provider cookbook reports $0.000497 for one 13-question batch versus $0.006090 for separate calls. Its speed comparison sums sequential call latencies, not concurrent wall time. The example uses Jev 1.12 and five repeats; it is not our benchmark.

[S3] OpenAI · Harness engineering ↗

11 February 2026. An internal engineering account describes human QA becoming a bottleneck and making the environment, repository knowledge and feedback legible to agents. Not a controlled BESF trial.

[S4] OpenAI · GPT-6 Astra pricing ↗

Standard short-context uncached input/output: $10/$50 per million tokens, checked 20 September 2026. Cache, batch, tool and long-context pricing differ. Token quantities in our examples are assumptions.

[S5] Anthropic · Claude Fable 5.1 ↗

Official model ID claude-fable-5-1; standard input/output $10/$50 per million tokens, checked 20 September 2026. Equal pricing does not establish equal capability.

[S6] Anthropic · Effective harnesses for long-running agents ↗

26 November 2025. Engineering account about incremental work, persistent progress and end-to-end feedback across sessions. Scope: their environment, not our measured outcomes.

[S7] OpenAI · Separate ChatGPT and API billing ↗

Subscriptions and API billing are separate. No account allowance or permission was inspected. Included-call and 99%-subsidy cases are hypothetical.

[S8] Bret Victor · Explorable Explanations ↗

Design precedent for an argument readers can manipulate and challenge. Does not prove that these articles improve comprehension.

[S9] Nielsen Norman Group · Progressive Disclosure ↗

3 December 2006. Defer secondary complexity; retain the immediate task and important limitations. Practitioner HCI guidance, not a universal reading-time law.

[S10] Nielsen Norman Group · Information Scent ↗

Design guidance for making destinations and likely value apparent. Used for specific action labels and figure captions; no conversion claim.

[S11] W3C · Animation from Interactions ↗

Nonessential interaction animation should be disableable. Implemented motion controls do not by themselves certify WCAG conformance.

[S12] Wikipedia · Signs of AI writing ↗

WikiProject advice page, not Wikipedia policy or an authorship detector. Used to challenge inflated claims, boilerplate and unsupported significance—not to disguise AI assistance.

[S13] TypeSafe · Current-model limitations ↗

Context relevance, indirection and judgment limitations need domain-specific evaluation. Probabilities are neither business outcomes nor authority.

[S15] Jeff Humble · Signs of vibe-coded UI ↗

24 April 2026. Practitioner critique of decorative glow, arbitrary status, repeated cards and weak hierarchy. Useful design challenges, not an authorship detector or causal study of this article.

[S14] RouteLLM ↗

Prior work on model routing. Our cost equations and illustrative policy do not reproduce its experiments or establish a new routing algorithm.

One already-authorised task

A bounded brief for your agent.

This does not grant tools, spending, deployment or data permissions. The real workflow must enforce the boundary.

Retained research · Unvalidated in the field

Which distinctions matter?

48 named questions

Published 2026-08-31 · Source edition two-article complete · About 12 minutes for the complete article, including optional sections. Return to Writing