← All writingCalvin Kennedy Writing
BESF / 02 · The economic experiment

What should
48 AI checks
buy you?

At today’s listed Jev price, this example costs about a tenth of a cent. Interesting. Now ask whether it saves a frontier call, catches a mistake, or just gives you 48 more answers to inspect.

Try your workload ↓
FIG. 01 · ONE STATE, MANY QUESTIONSTOKEN MODEL
Relevant task state · 20,000 tokens
48 short questions · 3,840 tokens
Request overhead · 1,000 tokens
$0.00104324,840 input tokens × $0.042 / million

Each tile is one budgeted question, not a measured answer. Schema overhead and tokenisation must be measured. Jev is text-only; image or audio extraction is additional.[1]

THE MOVESeparate inspection from invention.

Code exact facts; use a cheap model for suitable judgments.

THE TESTCompare complete outcomes.

The same task, permission and independent verifier.

THE PAYOFFKeep only the useful complexity.

Cheaper tokens are not enough if rework or owner time grows.

01 / A real change in the cost structure

Stop resending the library
for every question.

The arithmetic is simple: context can dominate the bill. Pack independent questions against the same relevant state, then use code to ignore irrelevant answers. Statistical errors can still be dependent.

TypeSafe’s own 13-question cookbook reports $0.000497 for a batch versus $0.006090 for separate requests. It uses one document, five repeats and Jev 1.12. That is a provider demonstration—not our measured BESF speed or capability.[2]

Spend the savings on evidence that can change the decision.

Request limits, baseline fairness and the missing measurement

This comparison sends the same state once per individual question versus a packed request; it is not the only possible baseline. A frontier model can also answer several questions in one call, and caching can lower its effective cost. Jev has a 64k total request limit and a 32k state-plus-longest-question limit. At 1,000 assumed 80-token questions, the model repeats context across two requests. More context is not automatically better. This is an all-semantic packing scenario, not a recommendation to send every archived question to a model: 31 of the 48 registry entries prefer code, 14 Jev and 3 generation. We have not measured query relevance, packing overhead or marginal decision quality.

FIG. 02 · The duplicated part of the bill
One question per request1,011,840 input tokens
Packed questions + repeated state only when required24,840 input tokens
$0.04250Separate requests
$0.00104Packed requests

48 questions fit in one modeled request.

20,000 state tokens, 80 per question, 1,000 overhead per request. Fixed linear axes; no claim of a matching quality or latency advantage.

02 / The clever bit should survive the invoice

Buy frontier reasoning
where it changes the result.

The default uses Jev-style inspection on every task and a frontier review on 10%. Change the assumptions. Near-free frontier use can reverse the recommendation.

Default frontier review: 10k input + 1.5k billed output at $10/$50 per million = $0.175. Jev pack: $0.00104328. Both paths include $0.002 per task for the same check. USD throughout.[4][5]

FIG. 03 · The full comparison, including extra work
Frontier review each task + checks$177.00
Screen + sparse frontier + checks + owner time$95.54
$81.46less total resource cost in this assumed period
Screening inference
$1.04
Routed cash, before extra owner time
$20.54
Additional owner time
$75.00

Equal useful outcomes are a condition, not a demonstrated result. This comparison excludes differential rework, errors, latency and customer impact. The point is to measure those—not assume them away.

Open the equation, included-capacity boundary and caching caveat
F(c) = max(0,c−Q) × C_F × (1−s)
C₀ = F(N) + NC_test + fee
C₁ = NC_J + F(Nr) + NC_test + fee + hw

Q is a finite allowance; s is a hypothetical subsidy; h is extra time; w its opportunity value. A subscription does not automatically provide API credit. API billing and ChatGPT billing are separate. The finite allowance is applied separately to each counterfactual workload, not spent twice in one actual account. Defaults use uncached standard tokens; real cache hits, batch rates and agent traces can change the result. No production routing decision should be made on this calculator alone.

[7]
03 / Cheap agreement can be expensive

Three approving reviewers.
One shared blind spot.

Keep each check’s error rate unchanged. Only change how their mistakes overlap. The combined risk moves dramatically, even when pairs of checks appear independent.

Which independence assumption did the calculation actually earn?

Among deliberately incorrect proposals, each check misses 10%. All three therefore miss 0.1% only under mutual conditional independence. Pairwise independence alone can allow 1%; a shared blind spot can allow 10%.

This is a constructed probability counterexample, not an actual Astra, Fable or Jev benchmark. A reviewer’s confidence is not permission to act.

FIG. 04 · Joint misses among 1,000 incorrect proposals
1incorrect proposals approved by every check
Stopped by at least one checkPassed every check

Same individual miss rates. The independent case needs a stronger assumption than pairwise independence.

Follow the denominator—and see why agreement is not enough
Risk among approvals = (1−p)q / [p t + (1−p)q]

Let p=.579 be a historical mock prior, not a current local model accuracy. Let correct proposals pass all three checks with t=.99³. Changing q from .001 to .01 to .1 gives 0.749, 7.438 and 69.713 incorrect outputs per 1,000 approved. The two denominators are different: the visible grid contains only incorrect inputs; this formula concerns all approvals. More questions cannot identify an omitted fact they never receive.

04 / Which check comes next?

Don’t ask whether it’s uncertain.
Ask whether the answer matters.

One extra observation is worth buying when it is likely to improve the eventual action enough to cover its cost. A cheap test can be worthwhile; an expensive one can lose.

Illustrative loss: $1,000 for an incorrect release; $20 for holding correct work. Test sensitivity 95%; specificity 98%. These are scenario assumptions, not measured BESF consequences.

FIG. 05 · Compare the best action with and without a test
Minimum expected loss without test$18.40
Expected loss after result + test cost$6.37
$12.03expected net value of the test in this model

A hard safety or permission rule is not for sale. These expected-value calculations apply only among actions already permitted.

Exact two-outcome value-of-information calculation
R₀ = min(pL, (1−p)H)
R₁ = min(pdL, (1−p)(1−t)H)
  + min(p(1−d)L, (1−p)tH) + c
Net information value = R₀ − R₁

L and H are the two losses; d and t are sensitivity and specificity; c the check cost. The terms already include each test result’s probability. This is a small decision model, not a validated solution to every business choice.

06 / The experiment to give your agent

Make the fancy version
beat the boring one.

Test on one bounded task family before building the giant graph. Compare the same final outcome with the same verifier, permissions and budget. Keep the result that costs you less complete work.

BASELINES

Code where possible. Verified frontier everywhere else.

Astra or Fable-grade generation is appropriate to evaluate, but the named model is not an outcome guarantee. Preserve actual token and tool traces. Do not manufacture a local success probability from a leaderboard score.

PROPOSED ADDITION

Jev-like screening plus selective frontier reasoning.

Record the joint mistakes, not just individual accuracies. Allow “unknown.” Freeze the policy before the final evaluation and split by underlying task or customer.

MEASURE

Accepted result, missed rules, rework and owner minutes.

Report deferrals as deferrals. Count setup, maintenance, human review and false decisions. A pipeline that refuses more work can look cheap per incoming task.

STOP CONDITION

Quality falls, permission changes, or the saving disappears.

Do not promote a simulated success into operational authority. This page executes teaching mathematics; the accompanying starter executes bounded code. Neither validates live model quality.

The useful advantage is
less work left for you.

If the router costs more attention than it saves, delete the router. Keep the evidence, the explicit outcome and the checks.

Evidence trail

Open the source.
Keep the boundary.

Web sources checked 20 September 2026. Source notes are paraphrases unless explicitly marked as a short quote. Simulations are not live-model or customer results.

[S1] TypeSafe · Models ↗

Jev 1.13.0: $0.042/M input tokens, free output; text only; 64k request and 32k state-plus-longest-question limits. Provider documentation, checked 20 September 2026.

[S2] TypeSafe · Parallel questions ↗

Provider cookbook reports $0.000497 for one 13-question batch versus $0.006090 for separate calls. Its speed comparison sums sequential call latencies, not concurrent wall time. The example uses Jev 1.12 and five repeats; it is not our benchmark.

[S3] OpenAI · Harness engineering ↗

11 February 2026. An internal engineering account describes human QA becoming a bottleneck and making the environment, repository knowledge and feedback legible to agents. Not a controlled BESF trial.

[S4] OpenAI · GPT-6 Astra pricing ↗

Standard short-context uncached input/output: $10/$50 per million tokens, checked 20 September 2026. Cache, batch, tool and long-context pricing differ. Token quantities in our examples are assumptions.

[S5] Anthropic · Claude Fable 5.1 ↗

Official model ID claude-fable-5-1; standard input/output $10/$50 per million tokens, checked 20 September 2026. Equal pricing does not establish equal capability.

[S6] Anthropic · Effective harnesses for long-running agents ↗

26 November 2025. Engineering account about incremental work, persistent progress and end-to-end feedback across sessions. Scope: their environment, not our measured outcomes.

[S7] OpenAI · Separate ChatGPT and API billing ↗

Subscriptions and API billing are separate. No account allowance or permission was inspected. Included-call and 99%-subsidy cases are hypothetical.

[S8] Bret Victor · Explorable Explanations ↗

Design precedent for an argument readers can manipulate and challenge. Does not prove that these articles improve comprehension.

[S9] Nielsen Norman Group · Progressive Disclosure ↗

3 December 2006. Defer secondary complexity; retain the immediate task and important limitations. Practitioner HCI guidance, not a universal reading-time law.

[S10] Nielsen Norman Group · Information Scent ↗

Design guidance for making destinations and likely value apparent. Used for specific action labels and figure captions; no conversion claim.

[S11] W3C · Animation from Interactions ↗

Nonessential interaction animation should be disableable. Implemented motion controls do not by themselves certify WCAG conformance.

[S12] Wikipedia · Signs of AI writing ↗

WikiProject advice page, not Wikipedia policy or an authorship detector. Used to challenge inflated claims, boilerplate and unsupported significance—not to disguise AI assistance.

[S13] TypeSafe · Current-model limitations ↗

Context relevance, indirection and judgment limitations need domain-specific evaluation. Probabilities are neither business outcomes nor authority.

[S15] Jeff Humble · Signs of vibe-coded UI ↗

24 April 2026. Practitioner critique of decorative glow, arbitrary status, repeated cards and weak hierarchy. Useful design challenges, not an authorship detector or causal study of this article.

[S14] RouteLLM ↗

Prior work on model routing. Our cost equations and illustrative policy do not reproduce its experiments or establish a new routing algorithm.

One already-authorised task

A bounded brief for your agent.

This does not grant tools, spending, deployment or data permissions. The real workflow must enforce the boundary.

Retained research · Unvalidated in the field

Which distinctions matter?

48 named questions

Published 2026-08-30 · Source edition two-article complete · About 11 minutes for the complete article, including optional sections. Return to Writing