What should
48 AI checks
buy you?
At today’s listed Jev price, this example costs about a tenth of a cent. Interesting. Now ask whether it saves a frontier call, catches a mistake, or just gives you 48 more answers to inspect.
48 short questions · 3,840 tokens
Request overhead · 1,000 tokens
Each tile is one budgeted question, not a measured answer. Schema overhead and tokenisation must be measured. Jev is text-only; image or audio extraction is additional.[1]
Code exact facts; use a cheap model for suitable judgments.
The same task, permission and independent verifier.
Cheaper tokens are not enough if rework or owner time grows.
Stop resending the library
for every question.
The arithmetic is simple: context can dominate the bill. Pack independent questions against the same relevant state, then use code to ignore irrelevant answers. Statistical errors can still be dependent.
TypeSafe’s own 13-question cookbook reports $0.000497 for a batch versus $0.006090 for separate requests. It uses one document, five repeats and Jev 1.12. That is a provider demonstration—not our measured BESF speed or capability.[2]
Spend the savings on evidence that can change the decision.
Request limits, baseline fairness and the missing measurement
This comparison sends the same state once per individual question versus a packed request; it is not the only possible baseline. A frontier model can also answer several questions in one call, and caching can lower its effective cost. Jev has a 64k total request limit and a 32k state-plus-longest-question limit. At 1,000 assumed 80-token questions, the model repeats context across two requests. More context is not automatically better. This is an all-semantic packing scenario, not a recommendation to send every archived question to a model: 31 of the 48 registry entries prefer code, 14 Jev and 3 generation. We have not measured query relevance, packing overhead or marginal decision quality.
48 questions fit in one modeled request.
20,000 state tokens, 80 per question, 1,000 overhead per request. Fixed linear axes; no claim of a matching quality or latency advantage.
Buy frontier reasoning
where it changes the result.
The default uses Jev-style inspection on every task and a frontier review on 10%. Change the assumptions. Near-free frontier use can reverse the recommendation.
Default frontier review: 10k input + 1.5k billed output at $10/$50 per million = $0.175. Jev pack: $0.00104328. Both paths include $0.002 per task for the same check. USD throughout.[4][5]
- Screening inference
- $1.04
- Routed cash, before extra owner time
- $20.54
- Additional owner time
- $75.00
Equal useful outcomes are a condition, not a demonstrated result. This comparison excludes differential rework, errors, latency and customer impact. The point is to measure those—not assume them away.
Open the equation, included-capacity boundary and caching caveat
C₀ = F(N) + NC_test + fee
C₁ = NC_J + F(Nr) + NC_test + fee + hw
Q is a finite allowance; s is a hypothetical subsidy; h is extra time; w its opportunity value. A subscription does not automatically provide API credit. API billing and ChatGPT billing are separate. The finite allowance is applied separately to each counterfactual workload, not spent twice in one actual account. Defaults use uncached standard tokens; real cache hits, batch rates and agent traces can change the result. No production routing decision should be made on this calculator alone.
[7]Three approving reviewers.
One shared blind spot.
Keep each check’s error rate unchanged. Only change how their mistakes overlap. The combined risk moves dramatically, even when pairs of checks appear independent.
Which independence assumption did the calculation actually earn?
Among deliberately incorrect proposals, each check misses 10%. All three therefore miss 0.1% only under mutual conditional independence. Pairwise independence alone can allow 1%; a shared blind spot can allow 10%.
This is a constructed probability counterexample, not an actual Astra, Fable or Jev benchmark. A reviewer’s confidence is not permission to act.
Same individual miss rates. The independent case needs a stronger assumption than pairwise independence.
Follow the denominator—and see why agreement is not enough
Let p=.579 be a historical mock prior, not a current local model accuracy. Let correct proposals pass all three checks with t=.99³. Changing q from .001 to .01 to .1 gives 0.749, 7.438 and 69.713 incorrect outputs per 1,000 approved. The two denominators are different: the visible grid contains only incorrect inputs; this formula concerns all approvals. More questions cannot identify an omitted fact they never receive.
Don’t ask whether it’s uncertain.
Ask whether the answer matters.
One extra observation is worth buying when it is likely to improve the eventual action enough to cover its cost. A cheap test can be worthwhile; an expensive one can lose.
Illustrative loss: $1,000 for an incorrect release; $20 for holding correct work. Test sensitivity 95%; specificity 98%. These are scenario assumptions, not measured BESF consequences.
A hard safety or permission rule is not for sale. These expected-value calculations apply only among actions already permitted.
Exact two-outcome value-of-information calculation
R₁ = min(pdL, (1−p)(1−t)H)
+ min(p(1−d)L, (1−p)tH) + c
Net information value = R₀ − R₁
L and H are the two losses; d and t are sensitivity and specificity; c the check cost. The terms already include each test result’s probability. This is a small decision model, not a validated solution to every business choice.
The giant tree is a budget.
You don’t have to spend it all.
Keep several promising paths, gather discriminating evidence, and stop when the next expansion costs more than it is worth. Pruning can also discard the best path.
Counting model only. No tree solver or language model predicts these futures in the browser. Beam quality depends on its scoring evidence.
Inspect the search count and original research boundary
At depth i, full search scores bⁱ nodes. A beam scores b times the number retained at the preceding depth, retaining at most K after each expansion. At b=5, depth=8 and K=5, this is 5 + 7×25 = 180. The saved work is exact arithmetic; preserving the best answer is not guaranteed. The earlier full laboratory, including its search, dependence, statistical-assurance and controller counterexamples, is retained unchanged in the combined package.
Make the fancy version
beat the boring one.
Test on one bounded task family before building the giant graph. Compare the same final outcome with the same verifier, permissions and budget. Keep the result that costs you less complete work.
Code where possible. Verified frontier everywhere else.
Astra or Fable-grade generation is appropriate to evaluate, but the named model is not an outcome guarantee. Preserve actual token and tool traces. Do not manufacture a local success probability from a leaderboard score.
Jev-like screening plus selective frontier reasoning.
Record the joint mistakes, not just individual accuracies. Allow “unknown.” Freeze the policy before the final evaluation and split by underlying task or customer.
Accepted result, missed rules, rework and owner minutes.
Report deferrals as deferrals. Count setup, maintenance, human review and false decisions. A pipeline that refuses more work can look cheap per incoming task.
Quality falls, permission changes, or the saving disappears.
Do not promote a simulated success into operational authority. This page executes teaching mathematics; the accompanying starter executes bounded code. Neither validates live model quality.
The useful advantage is
less work left for you.
If the router costs more attention than it saves, delete the router. Keep the evidence, the explicit outcome and the checks.