Ten apps. One you.
Each launch can leave work that keeps coming back. A monthly time budget, a saved 180-day simulation and a payback calculation test whether a shared way of working makes the next product cheaper to keep.
I want to build the next product without the last one taking another evening. The next build can be cheap while the work after launch keeps growing.
The relevant limit is the work you can keep doing, not the number of repos your agents can create.
This is for people building things they intend to keep. It has a live time budget, saved simulations and a usable next experiment. It is not a forecast or proof of an autonomous business.
Every launch leaves a time bill
Here is one invented month for the same person, in four moves.
- The first product still fits. A paid image app needs occasional fixes, billing checks and support. In this invented month, the ongoing work takes four hours. There is room to build something else. A small obligation can be worth owning.
- Add five more. The builds finish, and the work accumulates. Give each product just two routine hours a month. With shared upkeep, that is fourteen hours. Of the twenty-four you set aside, ten remain for the next thing. That calls for a whole-work comparison, not a faster coding demo.
- One shared dependency fails, and they can all need you at once. Now add the work of diagnosing and fixing a common failure. The month fills faster than the product count suggests. If you stop at your limit, the figure keeps the unfinished hours. A time cap does not finish the job.
- Change the production method to remove repeat work before adding more. A stable rule becomes code. A known failure gets a recovery procedure. Another operator can use the evidence. Shared upkeep rises, but routine work per product might fall. That “might” is the experiment. It is not a promised 65% saving.
One invented month for one person. The six hypothetical products are an image studio, a booking tool, a report builder, an export service, a member portal and an asset library. Assume 2 routine hours per product, 2 shared hours and 24 available hours per month. No measured workload is implied. Change every assumption below.
Routine work servedUnfilled track: time left for new work
10.0 hours remain for the next build. The older products keep their claim.
The accounting behind the month
Work due = shared upkeep + products × routine hours + common-failure work
New-work time = max(0, available hours − work due)
Backlog = max(0, work due − available hours)
Baseline upkeep is 2 h. The changed-process scenario uses 4 h plus 35% of the per-product work. A shared failure adds 2 + N hours in the baseline, or 2 + 0.25N after recovery investment. These are editable conceptual assumptions, not fitted coefficients. This simplified month excludes revenue, task severity, stochastic queues, operator cost and the value of your chosen work.
The earlier operating models cover some of those omitted mechanisms. Do not add this illustration to their results as if the two were independent evidence.
Illustrative workload, not observed productivity. “Changed process” assumes 65% less per-product work but more shared upkeep. Hidden work is still counted.
What would make the second product cheaper?
The reusable part might be how the work gets done, rather than the same app sold twice.
I do not need every product to share a customer. I need to know whether some part of the reasoning, delivery and recovery can be used again without bringing the same amount of work with it.
Picture the paid-access feature in an image app. A small model flags a possible duplicate grant. Code reproduces it. A source lookup resolves delivery semantics. A stronger agent changes the implementation. The next build reruns the recorded case.
That is one candidate production arrangement. A single capable agent may already do it more cheaply. The comparison must let that answer win.
In the proposed method, what an image app or a booking app passes to a new build is a justified rule, the check that enforces it and a recovery method. Share only what stays valid. New semantics, tenants and operators can invalidate an old check. Customer demand does not travel with it.
From a technical advantage to something I retain
Owner value over a horizon can be represented as distributions received + your share of the realisable remaining business value − capital contributed. Do not count retained cash twice. Compare the cash needed, obligations and routine workload alongside that objective.
Routine-replacement surplus = operating cash before founder pay − incremental cost of replacing residual routine founder work.
This is not company valuation, passive income or a prescription to delegate everything. Shared tooling can be a poor investment. A product can be useful without paying. Paid services can be worthwhile without becoming a product. The decision is whether this particular arrangement serves the company you want to own.
The cheaper arrangement shipped more errors
This is a result I would lose by showing only “hours saved”.
A saved v4 simulation: 120 synthetic paths per setting over 180 days, with the same constructed demand for every arrangement. Not live agents. Founder rescue is the comparison.
| Arrangement | Adjusted surplus (AUD) | Routine founder work (hours) | Incorrect releases (mean count, not percent) | Unfinished work at the end (mean hours still owed) |
|---|---|---|---|---|
| Founder rescue | A$12,002 | 67.16 h | 1.12 | 0 h |
| Cap my hours | −A$5,321 | 26.97 h | 0.86 | 45.36 h |
| Pay for capacity | A$10,202 | 1.57 h | 1.12 | ≈0 h |
| Executable checks + capacity | A$13,153 | 0.22 h | 2.94 | 0 h |
- Cap my hours: fewer founder hours, but work remains unfinished. The cap is not the solution.
- Pay for capacity: routine work moves to paid capacity. It is not made free.
Executable checks + capacity: less founder work and more wrong outputs. That is not an acceptable win if the quality boundary fails.
Adjusted surplus prices residual founder work and ending backlog. It excludes taxes, financing, chosen development/strategy work and exit value. Transferability of routine work and all commercial inputs were assumed. Do not interpret fractional mean errors as fractions of actual incidents.
What these numbers establish, and what they do not
The saved operating_summary.csv is reproduced without changing its values. It is a model of capacity, errors and costs, not a current model benchmark or a prospective customer study. The “cap” policy can show fewer wrong outputs because fewer requests finish. The table does not erase the queue.
Operational overload is a real design concern (Google SRE, Handling overload). The particular coefficients here are not borrowed from Google or measured at Calvin Kennedy.
A real test needs independently assessed outcomes, appropriate severity thresholds, all attempts, paid and unpaid work, and a holdout task/operator. “No observed errors” on repeated known examples does not establish general reliability.
A model of capacity, errors and costs from saved synthetic paths, not a benchmark of current models or a study of customers.
A reusable advantage still has to repay its own bill
This is where “build the platform first” can fail.
In an earlier mock experiment, making some checks executable reduced variable spending. Suppose the applicable saving were A$2.769 per assigned job, setup cost A$7,000 and upkeep A$200 a month. How often would you really use it?
Cumulative savings after the investment, in AUD over 36 months. Savings came from a synthetic scenario assuming equal checker quality at lower execution cost. That assumption is not established in practice. Revenue and demand are not being predicted.
The dashed line is zero. Move usage down: the line is allowed to stay below zero. An impressive shared method with too little use is still a cost.
Break-even and the value of an hour
Net savings at T months = T × (jobs/month × savings/job − upkeep/month) − setup
Break-even time = setup ÷ monthly net savings, only when that net is positive.
The source value is A$2.7689950833333326 per assigned case, not per correct release. The equation does not double-count avoided spend as new revenue. It is undiscounted and excludes financing, taxes and omitted adoption costs. At 200 cases/month, the unrounded recovery time is 19.785 months. At 50, upkeep exceeds savings, so there is no positive payback.
Do not apply a labour price twice. Reducing paid hours can save cash; freeing your own time creates capacity that may or may not earn revenue. Your strategy work still has a cost even when you choose to keep doing it.
An undiscounted illustration built on a synthetic saving, not a forecast of revenue or demand.
Why test this now? The building blocks changed
This is an opportunity to compare methods, not a countdown to easy money.
Strong agents can work inside an engineered environment. OpenAI’s harness-engineering account (11 February 2026) describes repository knowledge, mechanical constraints and feedback around agent-produced software. It is a reported engineering experience, not your expected speedup.
Typed decisions are a separate model product. TypeSafe’s early-access announcement of Jev (15 September 2026) makes cheap structured observations worth investigating. It does not establish that a committee of them is accurate on your task.
My practical timing rule is simple: record the current method before changing it. A new model, price or check can change the comparison. There is no evidence here of a closing window in which an unproven business must be launched.
The next product would not inherit the whole workload
That is what I would like to be true. If the method transfers, I could spend more of the same week building and deciding what to own. A lower token bill is only useful if the rest of the work does not come back through another door.
The simulations found conditions where this could work, and conditions where it became a worse business. They did not establish a profitable company that operates without my routine involvement. The next evidence has to come from actual intended use, complete costs and a real commercial arrangement.
Part 2, Which job gets the expensive AI?, looks at the production method itself.
For coding agents: test one recurring job
Give your agent the experiment, not the whole archive. Ask it to find whether a recurring job can cost less next month, and to keep the simpler approach if the advantage fails to transfer.
This text requests a bounded experiment. It does not enforce a sandbox, spending cap, identity check or production permission.
Sources and further discussion
Saved experiments, live calculations and proposed benefits have different meanings here. The time-budget and payback figures are local calculations. The operating comparison uses saved v4 results; the savings coefficient comes from the v3 mock lab. The two models answer different questions and are not pooled into a forecast. Exact tables, parameter files and lineage are in the ZIP.
- OpenAI, Harness engineering. 11 February 2026. First-party engineering account. Demonstrates a reported way of organising agent work, not a controlled estimate for this reader.
- TypeSafe, Introducing System One Models & Jev. 15 September 2026. Vendor announcement of early-access typed decision outputs. Not independent validation; no vendor speed, price or accuracy claim is imported into these calculations.
- Google SRE, Handling overload. Operational treatment of load and capacity. Does not supply the numerical workload or customer response assumptions in this essay.
- NASA, Verification and validation. Conformance to requirements and intended-use suitability are distinct. No certification or real system safety claim is implied.
Written and implemented with AI assistance from Calvin’s research and direction. No invented customer incidents, testimonials or live-model measurements. Reviewed and published by Calvin Kennedy. All interactive calculations run locally; this page sends no prompts or telemetry.