← Back to writing

Which job gets the expensive AI?

A small model can spot a concern, a source can settle a fact, a strong agent can write the fix and code can keep the rule. The useful skill is knowing which of those you need next.

Repeated model calls can repeat the same blind spot. The proposal here is to buy the evidence that can change the decision. The possible gain is less repeated reasoning and less work handed back to you.

The output I want is a completed job with evidence.

This is for full-stack and applied-AI builders. It has a concrete route, a cost comparison and an executable companion. The figures use local code and synthetic assumptions, not live model calls. It is the second of two articles; Ten apps. One you. asks what the next product costs to keep.

“Add paid credits” contains several different jobs

The feature should grant a paid purchase once, preserve the next legitimate purchase, and keep accounts separate. Refund policy is a different decision.

Each concern goes to the evidence that can settle it

Adding paid credits to an image app, as one complete job. Cheap typed concerns say where to look; they are fallible signals. Concern labels and route steps are authored examples, not model outputs.

Concern → useful evidence → action
RouteTyped concernUseful next step
Code: reproduce the failureRepeat delivery? Check the business key. Wrong account? Check the tenant boundary.A replay can settle the duplicate-grant question. Another opinion cannot replace it.
Source: check the semanticsNew event, same purchase? Look up external semantics.An external fact needs a current source. Do not infer the provider contract from agreement.
Agent: implement the changeThe evidence calls for a revised business key and implementation.Use strong reasoning to implement the justified change. Keep its regression and its limits.
Owner: decide the policyRefund after credits spent? An unmade product decision.No refund policy was authorised. Preserve the work and ask for the missing decision.

This is the proposed arrangement. The simpler single-agent baseline stays in the comparison. A typed answer can still be wrong.

  1. Cheap signals suggest where to look. Ask short questions about a concrete change: Can a purchase repeat? Is the tenant checked? What is the refund rule? Accept “unknown.” A Jev-like typed answer is a clue, not a permission or a proof. The useful output is the next investigation.
  2. Use a replay for code and a source for facts. If two deliveries might grant twice, execute the sequence. If the uncertainty concerns what an external provider sends, inspect the provider’s current contract. More judges reading the same incomplete brief do not supply that missing fact. A different question needs a different evidence path.
  3. Repair the cause and keep the new case. The stronger agent can interpret the evidence, revise the business key and update the implementation. Code can rerun the approved rule next time. Count the review and failed attempts in the bill. A useful retained result can be code, a source note or a boundary; it is not always another prompt.
  4. “Should we refund spent credits?” is not a test failure. That policy was never specified. The agent can explain options, but it must not silently decide what you promised customers. Keep the rest of the work; ask for the missing authorised decision. Holding this part is honest. Calling it complete is not.
Inspect the local example (not a payment integration)

The fixtures below are already authenticated, paid and normalised. This illustrates a business-key distinction; no payment-provider integration or refund operation is implemented. Each balance is recomputed from the fixture for each grant rule when the page loads.

Balance in credits for each grant rule
Input sequenceExpectedGrant every messageDeduplicate eventDeduplicate tenant + purchase
evt_01 / order_17, twice100200100100
evt_01 + evt_02 / order_17100200200100
order_17, then order_18200200200200

Highlighted balances do not match the expected value. The tiny in-memory demo omits concurrent storage, spending, refunds and cross-service side effects. The prior v5 SQLite example and its fuller limits remain in the private history. Verification of a specified rule does not establish the right intended use (NASA, Verification and validation).

Buy the next check only if it can change the decision

A check is worth more when the consequence is large, and less when it tells you nothing new.

Suppose a result has an assumed 8% chance of being wrong. Releasing an error costs A$100; holding the work costs A$8. An A$0.40 review catches 80% of wrong results and flags 8% of correct ones.

The calculator asks whether either possible review result would change release versus hold. It includes the price of the review. Move the loss down: sometimes the next check should not be purchased at all.

Here the A$0.40 review cuts expected loss from A$8.00 to A$3.10

One hypothetical decision: expected loss plus the review price, for deciding now or reviewing first.

Decide now, without another checkA$8.00
Review first, including its priceA$3.10

Buy the review in this example: A$4.90 less expected loss.

The equation: expected value of another observation

V(s) = min{ loss(release | s), loss(hold | s),
cost(check) + E[ V(s + new observation) ] }

The full research model uses a finite decision graph. This one-check illustration uses Bayes’ rule with known hypothetical sensitivity and false-positive rate. It computes the chance of each flag, the posterior error probability and the cheaper authorised action after each possible observation.

When sensitivity equals the false-positive rate, the review is independent of correctness: posterior risk does not change. It cannot improve the decision and adds its cost. Real estimated probabilities can be miscalibrated, dependent or stale. Fit, calibration and evaluation require separate evidence; more decimal places do not fix that.

The question is which costly uncertainty to reduce, not how to fit every concern into a weighted score. Record qualitative reasons (an unmade policy, an unknown external contract, a mismatch with the user’s purpose) alongside the numbers.

Loss values are invented. The state is a hypothetical authorised decision; real authority and hard safety constraints cannot be traded for a favourable score. This is not a deployment controller.

The four-cent answer can be the expensive one

It is the expensive one only when its review, repair and failure bill make it so. The “No quality gain” setting below is the counterexample.

Compare two ways to complete the same job. The cheaper call costs A$0.04, needs eight minutes of review and repair, and is correct 80% of the time. The stronger call costs A$0.80, needs one minute and is correct 95% of the time.

Counting review time, the four-cent call costs more per correct result

AUD per correctly completed result, on the same absolute scale. Every number is assumed, not a Jev or frontier-model benchmark. Time is valued at A$75/hour initially; it is not necessarily cash saved.

Cheaper callA$12.55
Stronger callA$2.16

Model callsReview + repair

The stronger-call route is cheaper under these assumptions.

“Included” sets marginal call price to zero. It does not make plan fees, limits, queueing, fallback or human time free.

Reproduce the arithmetic; batch and cache only eligible work

Cost per correct result = (model spend + review minutes × hourly value / 60) / correct fraction.

At the defaults: (0.04 + 8 × 75 / 60) / 0.80 = A$12.55. For the stronger case: (0.80 + 1 × 75 / 60) / 0.95 ≈ A$2.16. With equal time and quality, the cheaper call wins. This is a population-average ratio under a fixed task mix, not a guarantee that repeated retries are independent.

For non-urgent eligible API work, OpenAI documents a 50% Batch discount and a 24-hour completion window (OpenAI, Batch API). Batch expiry and latency still matter. Prefix caching can also alter the bill; measure actual eligible prefixes and the current pricing rather than assuming every request is cached (OpenAI, Prompt caching). Neither discount validates the result.

The formula counts variable effort across attempts. It does not include demand, liability, integration setup or fixed overhead. Being “accepted by a judge” is not the definition of correct.

Routing saved money in the mock. It also failed under change

The comparison has to include buying the same evidence everywhere, so the baseline is not unfairly weak.

Routing cost about 11% less on familiar cases but let errors above 1% on an unfamiliar family

A saved v3 mock-agent experiment: 60,000 assigned cases per world, with the same evidence choices available to each policy. No live model performance. Four of the constructed worlds are shown.

Each cell: variable cost per correct release (AUD); wrong among released outputs (percent, not of all assigned cases); requests receiving an output (percent of 60,000 assigned)
Synthetic settingBuy all evidenceSelective routingIndependent-check-only controlReading
Familiar casesA$7.240.045% wrong77.37% outputA$6.440.045% wrong77.37% outputA$6.750.178% wrong77.75% outputThey released the same 46,419 cases. Routing spent about 11% less per correct result.
Unfamiliar familyA$8.451.191% wrong65.05% outputA$7.391.191% wrong65.05% outputA$7.882.290% wrong66.00% outputThe familiar-world release rule did not keep error below 1% in this new family.
Checker deterioratesA$7.251.517% wrong78.47% outputA$6.461.517% wrong78.47% outputA$6.765.010% wrong81.60% outputThe checking channel deteriorated. Old calibration did not make it trustworthy.
Semantic clues uselessA$7.570.194% wrong73.96% outputA$6.980.194% wrong73.96% outputA$6.740.193% wrong77.72% outputThe cheap semantic clues no longer help. Keep a simpler control in the experiment.

In familiar cases, a simpler independent-check-only policy cost A$6.75 per correct result, with 0.178% wrong among releases and 77.75% coverage. The routed system is not automatically the best trade-off.

What is mocked, what is learned, and what could fool it

The lab generated hidden implementation/specification errors and fallible, dependent observations. It fitted and calibrated policies separately, then evaluated six constructed worlds. Call prices, error mechanisms, label availability and human checking costs were inputs. “Jev-like” describes typed observations, not a reproduction of Jev’s architecture or measured accuracy.

The core question survives the abstraction: does selective evidence acquisition earn its overhead? It can fail when clues carry no information, modelled dependencies are wrong, a source is contaminated, or the checker changes. Preserve simple controls, severe subgroups, abstentions and the full evidence bill.

Agent-evaluation guidance emphasises inspectable outcomes and realistic task design (Anthropic, Demystifying evals for AI agents). Public model claims cannot supply a quality coefficient for this workload without a relevant test. Primary literature and the earlier code are retained in the research history.

Synthetic policies from a saved mock, not product ratings or live-model performance.

Eight answers can be less than three independent signals

Changing the model name does not remove a shared missing fact. This approximation illustrates why an eight-way vote deserves a dependence check: at the defaults, eight judgments with a 30% shared error correlation carry about 2.58 effective observations.

Effective observations2.58

n / [1 + (n − 1)ρ], an exchangeable-correlation approximation for variance. Not a measured effective sample size for real models, and not a reliability guarantee.

The part I want to keep is the learning

A useful result should make the next relevant job easier, without making every job depend on this whole essay.

Keep a rule when it is justified. Keep the source and counterexample that explain it. Keep the boundary where it stops applying. Retire it when the environment changes or a simpler method becomes better.

That is the link back to ownership: useful knowledge, code and recovery could compound across work. A growing pile of prompts can just as easily create maintenance. I have no interest in buying that by the gigabyte.

For a one-off script or a competent agent already meeting the requirement cheaply, do less. For repeated valuable work, compare the complete arrangement before feeding it more tasks.

Try it on one consequential job

The cheap judgments are not the advantage by themselves. Neither is the stronger model. The advantage would be choosing a better next step and retaining something that genuinely reduces the next bill.

Try it on one consequential job. If the simple approach wins, keep it. If the method transfers to unfamiliar work and another operator, there may be something worth owning.

Part 1, Ten apps. One you., starts from the owner’s question: what the next product costs to keep.

For coding agents: take one job into your existing repo

The brief separates finding a concern, checking a fact, writing a fix and deciding what you actually authorised. It asks for evidence, costs and a clear stop.

This text requests a bounded experiment. It does not enforce a sandbox, spending cap, identity check or production permission.

No external API, new model account or deployment permission is required to read or interact with this article.

Sources and further discussion

Saved experiments, live calculations and proposed benefits have different meanings here. The replay and calculators run locally. The policy table is sourced from the v3 saved CSV; it is not rerunning millions of agents in the browser. The brief requests an experiment rather than promising any particular productivity gain.

  1. TypeSafe, Introducing System One Models & Jev. 15 September 2026. Vendor announcement of early-access typed decision outputs. Not independent validation; no vendor speed, price or accuracy claim is imported into these calculations.
  2. Anthropic, Demystifying evals for AI agents. Engineering guidance for outcome-based agent evaluation. The diagrams here are a proposal, not a reproduction of an Anthropic benchmark.
  3. NASA, Verification and validation. Conformance to requirements and intended-use suitability are distinct. No certification or real system safety claim is implied.
  4. OpenAI, Batch API. Checked 20 September 2026. Documents 50% lower eligible API costs and a 24-hour completion window. Eligibility, expiry and workload latency still matter.
  5. OpenAI, Prompt caching. Checked 20 September 2026. Prefix caching is a latency/billing mechanism. Measure actual eligible usage; it does not establish answer quality.

Written and implemented with AI assistance from Calvin’s research and direction. No invented customer incidents, testimonials or live-model measurements. Reviewed and published by Calvin Kennedy. All interactive calculations run locally; this page sends no prompts or telemetry.

Published 2026-09-09 · Updated 2026-09-28 · Source edition v6