← All writingCalvin Kennedy Writing
ck.Calvin Kennedy
Agent brief ↗
Field notes / AI engineeringEvidence checked 20 Sep 2026

Keep the
big model.
Change its job.

Your agent builds the system.
It doesn’t have to be every model inside it.

For builders using Claude Code or Codex. More useful outputs. A potentially smaller bill. A test before you commit.

SAME FRONTIER CODING AGENTresearch → build → compare
THE FEATURE / PHOTO → LISTINGS
“A lamp, a mug and headphones.”
Editable imagemask, not description
Short draftwriter worth testing
$4.00
Exact fee$40 × assumed 10%

One model is a baseline.
It is not the only design.

Illustrated outputs, not model results. The baseline may win.
Explore: follow the pictures; open detail when useful.No essential point requires animation.
AThe problem / wrong output, wrong repeated cost

“It’s a lamp”
isn’t a cut-out.

You’re building a resale app.
The seller wants three editable listings from one photo—not an essay about their desk.

Describe it.

Useful for the title.
Not the image editor.

Return its pixels.

A mask gives the app
something it can edit.

Different output, different candidate. SAM 3 supplies segmentation masks; that does not prove condition or ownership. A tool-using frontier baseline may already solve the job. Compare against that. [SAM]

Why this is not “replace the frontier model”

Keep the model that helps you understand the requirement, research options, write code and diagnose failures. Distinguish that build-time agent from the components called repeatedly inside the shipped application. They can be the same model, but need not be.

This is not a claim that frontier models cannot use tools, return structured data or do arithmetic. A capable baseline retains its normal tools. The question is whether a different output representation, a tested specialist or deterministic implementation improves the entire feature.

For this example, segmentation produces editable pixels; a writer produces a draft; supplied seller information establishes condition; fee code applies a declared rule. A 3D reconstruction is an optional preview. These are proposed boundaries, not a tested marketplace architecture.

BThe steps / keep the product goal fixed

Same photo.
Better questions.

Ask the agent to build a comparison, not a six-model stack. One useful change is enough.

01 / Baseline

Start with what works.

Let the capable vision-and-language baseline draft the listings. Record what the seller still has to fix.

Keep the baseline strong.

02 / Decompose

Match each output.

A mask for the crop. A short draft. Fees from code. Ask the seller what the photo cannot establish.

Use the fewest parts that do the job.

03 / Compare

Change one part.

Try a smaller writer on the same listing tasks. Count escalation, bad claims, missed outputs and edits.

Same tasks. Same acceptance. Full cost.

04 / Retain

Keep what helps.

Does the seller get usable listings with less work? Keep the result, scope and regression tests—or the baseline.

A failed candidate is a useful result.

DESKDROP / SELL THESE THREE THINGSConcept
Claude Code / Codexresearch → implement → test
same listingFrontier baselineordinary tools retainedSmaller writerescalation countedSame checksclaims · editstime · full costPROPOSED COMPARISON · NO MODEL OUTPUTS SHOWN
Green desk lampCeramic mugHeadphonesConfirm conditionCheck for chipsConfirm it worksLess seller work? Measure it. Don’t assume it.
Same items throughout. No live model calls.
Inputphoto
Methodfrontier baseline
Outputeditable drafts

One photo. Three draft listings.

The obvious version might already be the right one.

Desktop scroll follows the steps. Buttons give direct control.

Pick an output, not a leaderboard.

What does the app actually need? Each choice shows a candidate, a test and a reason to leave it out.

PHOTOOBJECT MASKAn image your crop tool can use.
SAM 3 / segmentation

Give the crop tool a mask.

Select the lamp’s pixels instead of asking the app to interpret a description.

Test: missed items, damaged edges, glare and occlusion.

Limit: pixels do not tell you whether the lamp works. [SAM]

SAM 3 / promptable segmentation

Locate each item and give the seller an editable cut-out—not another description.

Test: Missed objects, edges, glare and overlap.

Limit: A mask does not identify hidden damage or prove a brand.

Jev-like decision model / bounded output

Choose between declared actions, such as requesting seller confirmation. A typed choice is easier to handle in code.

Test: False acceptance, calibration and refusal.

Limit: A structured answer can still be wrong. Confidence needs task-specific checks.

Small language model / a candidate writer

Try a smaller writer for a listing draft. Send uncertain or failed cases to the capable baseline.

Test: Unsupported claims, seller edits and full cost.

Limit: Cheap tokens do not establish equal quality. Escalation and extra review count.

SAM 3D Objects / reconstruction

A 3D preview is a different output. Test whether it helps the buyer understand the item.

Test: View consistency, shape and useful interaction.

Limit: Reconstruction estimates unseen surfaces. It does not prove dimensions or condition.

Cosmos / world-model family

A robot might compare possible outcomes of a grasp. This resale feature does not need that capability.

Test: Action-conditioned outcomes and failure under shift.

Limit: A plausible future is not a safety certificate. Leave the model out unless the task needs it.

Ordinary code / explicit rules

Obtain the current rule from a reliable source. Calculate it and test the boundary cases.

Test: Rounding, currency, limits and changed rules.

Limit: Correct arithmetic cannot repair an incorrect fee rule or untrusted price.

The parts are real.
The gains aren’t settled.

The useful deadline is your next dependency. Compare an alternative before the rest of the app depends on it. There is no evidenced expiry date on this opportunity. [notes]

More model families, selection criteria and why some do not belong here

A product search may need SQL, BM25, embeddings or a reranker. A stock forecast needs a time-series baseline. A delivery schedule may need a constraint solver. Choose the mathematical object first: a category, score, vector, mask, mesh, action, forecast or text.

The Nemotron family offers open-weight language-model options; the 120B-total/12B-active Super release is not a tiny laptop model. Sparse compute changes a tradeoff but does not remove memory, hosting or idle-capacity costs. Check hardware, licences, latency and maintenance before choosing a deployment. [NVIDIA]

World models predict environment transitions, potentially conditioned on actions. That may matter for a robot trying to grasp a lamp. It does not automatically help a seller write a listing. Generated futures must not be treated as a safety guarantee. [Cosmos]

For each candidate, record the output contract, official model version, data requirements, access, licence, complete cost and a falsifiable test. Vendor announcements establish candidates to investigate—not an independent ranking for this feature.

CThe payoff / keep the capability, question repeated spend

A cheaper call.
A better system?

Keep expensive reasoning where it earns its place. Test a smaller writer for repeated drafts. Then add your time back in.

Calculated scenario / not measured savingsUSD · text only · equal quality untested
100200,000
noneevery draft
no extra work8 seconds
Possible monthly headroom
$2,200

less modelled cost. Not an accepted-quality result.

Frontier each time$3,000
Small + escalation + overhead$800
small callsescalationyour timeoverhead
1.32 extra seconds per draft eats the saving.

Default assumptions: 2,000 input + 200 output tokens; Fable 5.1 versus Haiku 4.5; 10% escalation; $60/hour review; $200/month added overhead. Prices are dated inputs, not evidence of equal performance. [pricing]

Two extra review seconds adds $3,333.33 per month. The candidate becomes $1,133.33 more expensive. At 100 drafts and no extra review, fixed overhead also makes it lose.

Inspect the equation, edit overhead, and challenge the assumptions
ΔC = N[cF − cS − f cF − tw / 3600] − I

N: monthly drafts. f: escalation fraction.

cF, cS: frontier and smaller-model cost per call.

t: additional review seconds averaged over every draft. w: review cost per hour.

I: added monthly infrastructure, maintenance and amortised setup.

20 Sep 2026 standard API snapshot · USD per million tokens
ModelInputOutputCache read
Fable 5.1$10$50$0.25
Haiku 4.5$1$5$0.10

Both receive the same assumed token count, not necessarily the same amount of text: tokenisers differ. Small calls are charged even when they escalate. The cache option omits cache-write and warm-up costs, so it is a steady-state sensitivity case, not a complete cache bill. [pricing]

Images, geometry generation, taxes, tool calls and errors are not separately priced. Add incremental costs to overhead or replace the model with measured itemised costs. Your coding subscription is a different expense from the runtime API bill. A subsidised subscription is not unlimited free deployment compute.

Do not optimise this number alone. Compare cost per accepted outcome only when the task population, acceptance criteria and complete costs are available. Keep severe failures separate from an average score. Try caching, batching and simpler code before introducing another service.

Saved synthetic result / E11

A higher score.
Fewer delivered answers.

The lab’s candidate accepted 7,054 tasks. Its flattering score excluded 2,561 timeouts.

Candidate acceptancecompleted only
94.82%

7,054 ÷ 7,439

7,054 accepted385 rejected2,561 timed out

The missing work changes the conclusion.

Baseline: 86.61% of all tasks. Candidate: 70.54% when everything counts. [record]
Does the guide itself help? Test repairs against regressions.

A helpful guide can also disrupt work the agent already did correctly. Let p be baseline success, r the share of failures repaired, and h the share of successes broken. These are conditional rates.

Net gain = (1 − p)r − ph
Before
80%
Repairs
+6 pp
Breaks
−4 pp
After
82%

Net: +2 percentage points

Shared 0–100% scale. Assumptions, not observed agents.

The guide, the helpers and their combination are separate interventions. Compare them with the existing capable agent. Use pilot tasks to choose one package, then freeze it and confirm on fresh task families. Preserve baseline successes and near-duplicate lineage; repeated attempts are not new independent tasks.

SkillsBench reports improvements on selected tasks; its construction rejects tasks without measurable separation. An AGENTS.md study found increased cost without general success improvement in its settings. Neither establishes the effect of this article or kit. [research]

Verification, validation, uncertainty and a stopping rule

Test the obligation, not the confidence.

Verification: do code, constraints and calculations satisfy the declared contract? Use unit, property, integration and numerical checks, with independent oracles where feasible. Validation: does the finished feature meet the seller’s actual needs? Inspect corrections, missing items, usefulness and the work displaced to people.

For a mask, inspect ground-truth coverage and edge errors. For a decision model, examine false acceptance, calibration and abstention under realistic shift. For a small writer, count invented claims and edits. For geometry, inspect topology, views and dimensions against suitable reference data. No single score covers all these contracts.

Numbers plus failure explanations.

Record all planned tasks, latency, provider charges, operations and human review. Qualitatively code wrong assumptions, unnecessary steps, hidden edits, requirement mismatch, severe failures and unknowns. A second model’s agreement is not independent truth.

Declare the primary comparison, useful-effect threshold, risk boundaries, budget, independent unit and stopping rule before acceptance data. Use grouped inference for repeated task families. Report uncertainty; an inconclusive pilot is not proof of equivalence. More peeking can manufacture apparent gains.

The archived lab’s prospective calculation found 27.2% power for one assumed five-point improvement with 120 independent pairs and 20% discordance under a fixed one-sided test. That is a chance of detecting an assumed effect, not the chance this kit works. Its model, assumptions and raw results are preserved in the original archive. [lab]

Stop when the next test is not worth its cost.

Test the smallest consequential uncertainty first. Expand after a result survives its checks. Quarantine stale lessons, remove redundant modules, and revalidate when models, tools or requirements change. The intended result is an accumulating evidence base—not an ever-growing testing ritual.

If this works

One feature request.
A wider set of options.

Give your coding agent model knowledge, executable tools and a bounded test. It should return a working comparison you can decide on—not just a longer explanation.

Your app gets different capabilities.

Masks and geometry where text alone is the wrong output.

You stop buying every answer the same way.

Keep expensive reasoning. Test cheaper repeated work.

The next agent inherits evidence.

Keep runnable checks and scoped lessons, not folklore.

Put the idea on one real task

Give it your
next feature.

You set the outcome. The agent proposes the smallest credible comparison. You decide what can run.

1.Current baseline + relevant alternatives.
2.One bounded test, with costs and failure criteria.
3.Keep / reject / unknown, with evidence.
What a guide can request—and cannot enforce

Skills provide procedures and tool-use information. They do not independently protect hidden labels, authorise paid calls or prevent deployment. Host permissions, an external evaluator, spend limits and CI must enforce those boundaries.

The starter is a proposal generator and a pair of local calculation/scoring helpers, not an autonomous service. Review the files before installing anything. Do not replace your existing AGENTS.md or CLAUDE.md wholesale. Both hosts have documented skill mechanisms; check the current version and effective loaded instructions. [host docs]

Local export. No uploads, paid calls, permissions or deployment are authorised.

Read exactly what the agent gets
Define the outcome and keep a capable baseline. Research only relevant model families. Propose one bounded comparison. Count all planned outcomes and full cost. Return keep, reject or unknown with scope and tests. No execution is authorised by this brief.

With JavaScript off, copy the brief in the disclosure above. The publication ZIP also contains the full agent kit.

Sixteen synthetic examples in the archived lab

None of them proved how much better your next build would be.

They exposed ways to fool ourselves. The opportunity is to give an already capable agent better options—and a way to find out which ones matter.

Keep the big model.
Make the next decision less of a guess.

Sources, dates and evidence boundaries

Linked source notes, not screenshots or endorsements. Capability and price references were checked on 20 September 2026. The diagrams are authored illustrations. The calculator uses declared assumptions. No real model or human-comprehension trial was run for this article.

Primary implementation

Meta / SAM 3 ↗

Promptable object detection, segmentation and tracking. Availability does not establish accuracy on the example app.

Primary implementation

Meta / SAM 3D Objects ↗

Single-image object reconstruction. A generated shape is not a certified scan. Checkpoints released 19 November 2025; not described here as a new September release.

Vendor release · 15 September 2026

TypeSafe / introducing Jev ↗

Bounded choices and calibrated decision outputs. Vendor positioning is not an independent benchmark or our measured result.

Primary model release · 10 March 2026

NVIDIA / Nemotron 3 Super ↗

Open-weight hybrid model; 120B total and 12B active parameters. Another deployment tradeoff—not a tiny laptop model, a free service or a claim to be the newest model.

Primary technical description

NVIDIA / Cosmos 3 ↗

Physical reasoning, world generation and action modeling. Illustrative futures here are authored, not generated by Cosmos.

First-party USD pricing snapshot

Anthropic / model pricing ↗

Fable 5.1 input/output $10/$50 per million tokens; Haiku 4.5 $1/$5. Cache reads $0.25/$0.10. Equal quality, token counts and eligible cache fractions are assumptions. Recheck before spending.

Arithmetic identity · not a marketplace price

Article’s fee calculation

40 × 0.10 = 4. The fee rule is illustrative; the real rule and its effective date must be acquired separately.

Preserved local evidence · synthetic only

AI Engineering Evidence Lab v0.3

The full original lab is nested unchanged in archive/prior-publication.zip. E11: 7,054 accepted, 7,439 completed, 10,000 planned. Sixteen archived synthetic examples; no real model, agent or reader-productivity trial.

Research synthesis · 11 September 2020

Distill / Communicating with Interactive Articles ↗

Informs connected representations, segmentation and reader control. Animation is not automatically superior to a static explanation. This page’s comprehension effects remain unmeasured.

Published 2026-08-29 · Source edition v2 · About 15 minutes for the complete article, including optional sections. Return to Writing