A new kind of company just became possible
AI made implementation cheap before companies learned how to use it. For a brief window, a small technical team can sell the finished work. The hard part is proving the accepted result.
Imagine a workflow already costing a business A$30,000 a month. A small technical team delivers the accepted outcome for A$13,500 and prices the spread so both sides win. That opening exists because implementation got cheap first. It narrows as experts learn, vertical vendors package the work, and incumbents adapt.
Code got cheaper. Organisational truth did not. In one dated benchmark a frontier-model task cost US$3.26. Finding the real workflow, making the data usable, defining authority and recovery, and winning adoption and supporting it are all still scarce. That benchmark cost is only the hook; the opportunity is the remaining work.
This is a field note for people who already know the stack. Public benchmarks and papers are sourced. The timing model, example economics and strategy are mine, and illustrative inputs are not forecasts.
The feed saw the opportunity before anyone proved the operating model
Three signals now point in the same direction. None is sufficient alone.
People can feel that the unit of leverage changed.
The screenshots are cultural signals, not proof. They show that investors and operators are already looking beyond software licences.
The models are the second signal. Delegated implementation is now real enough to alter company design: capability moved from suggestion to delegation. Astra and Fable can perform long-running tool work, and Astra scores 57.9% on Terminal-Bench 4.0. Vendor benchmarks need context; the direction still matters.
The third is money. The budget around software may be much larger than the software budget. Sequoia’s services thesis puts it at $1 of software for every $6 of services. That 6× claim is a thesis, not an accounting identity. The useful question is which work dollars can become AI-native delivery.
They all point at the same gap: more code still attenuates before it becomes a release. A 2026 working paper estimates much larger gains at commits than releases (+240 at commits, +30 at releases, estimated). That points at the unautomated work between output and production.
The missing work sits between model output and a business outcome
Generating code, reading documents, using tools and drafting decisions are all cheapening fast. Five layers between that output and a business outcome are not:
- Truth. What is the real workflow?
- Authority. Who may decide?
- Evidence. What proves completion?
- Adoption. Will anyone change behaviour?
- Ownership. Who carries failure and support?
My bet is to engineer that gap before vendors package it and experts learn it themselves.
The accepted state is the product
Monday, 8:42 am. A casual teacher starts at 9. This composite school workflow is deliberately ordinary, and that is why the economics matter.
The request arrives late as a PDF, an email, a form and a CSV. The name differs across two forms, working-with-children evidence is missing, and the requested role includes student-data access. The input is not clean enough to automate blindly: four documents disagree, the start time is close, and one permission is consequential.
- Request, by email and forms. It arrives incomplete and inconsistent, 18 minutes before the start.
- The agent reconciles and prepares. It builds the case, not the final authority. It matches identity where possible, checks and links sources, drafts accounts, flags missing evidence and prepares the exact approval question.
- An approver decides the exception. A named person approves the risky part. The system does not infer privileged access from urgency; a responsible approver accepts, narrows or rejects the request, and the system records who decided what and why.
- Ready, with correct access and evidence. Ready means more than “account created”: correct identity, least privilege, device, MFA, training, payroll state, evidence, and a tested offboarding path. Reversal is tested.
The interface is a component. The accepted state is the product.
No one wins by pretending the model replaces all three roles
The domain operator brings truth, the customer and authority. They know the exceptions, the consequences and what an acceptable outcome looks like. You bring the boundary, the system and the evidence, turning tacit workflow into interfaces, controls, recovery and an inspectable result. Frontier AI brings execution, iteration and scale: it produces, tests and operates more of the implementation loop at marginal cost.
Together they sell an accepted outcome, a result with justified reliance. Your job is to reduce total effort to an accepted result, compared with the buyer’s next-best alternative.
Price the workflow, not the software
The deal works only when the customer keeps value and delivery funds itself. Start with observed time and cost, then expose who keeps the difference.
A worked model with illustrative monthly inputs in Australian dollars. Choose a workflow, then edit the assumptions.
Both sides retain value before sales, switching, support, tax and overhead.
Edit assumptions
Show the equations
S = C₀ − C₁ Vc = C₀ − P − K Mp = P − C₁ − O
S is the surplus, C₀ the old cost, C₁ the AI delivery cost and P the price. Vc is what the customer keeps and Mp the provider’s margin. K and O are the costs this calculator leaves out: switching for the customer; sales, support, tax and overhead for the provider.
The test is ΔVyou = V(SME + you + AI) − V(best alternative). If the operator using the same model alone gets the same accepted result at lower total cost, your advantage is not real.
The cheapest call can produce the most expensive result
Route by accepted-result cost, not token price or benchmark rank. In the Artificial Analysis snapshot of 9 September 2026, an Astra task cost US$3.26 and a Fable task US$7.63 on the same aggregate index. These numbers expire, so re-run them before deployment. The cost that matters also counts review, retries and escaped failure.
Expected cost per accepted result: E[Ca] = (cm + w·hr) / pa + pf·L. Run cost plus paid review time, divided by the acceptance rate, plus the chance a failure escapes times the loss it causes. US dollars.
Fable is cheaper by $9 per accepted result. Its higher run cost is outweighed by the entered acceptance, review and failure assumptions.
Edit local-eval assumptions
The run costs come from the dated snapshot. Acceptance, review time and failure rates are assumptions to replace with your own local evaluation.
The routing rule follows: use the cheapest model that reaches acceptance after human review and failure cost.
“Build the app” is the wrong contract
A strong model will efficiently satisfy the wrong definition of done. The same staff-access job looks very different depending on where the agent is told to stop.
The composite school staff-access workflow, given to an agent under two definitions of done.
Build the app
- Prompt Build the staff access app.
- Generate Schema, routes, UI and tests.
- Run The happy path passes.
- Declare victory Authority, recovery, adoption and economics remain unknown.
“Implemented successfully.” 14 screens. 86 tests. No proof the school can grant the right access safely. Fast and impressive, but still not a business.
Build for acceptance
- Map Input, actors, baseline and exceptions.
- Contract Observable accepted state and named authority.
- Build Smallest complete path through the real workflow.
- Verify Resulting state, evidence, permissions and recovery.
- Price Customer value and provider burden.
- Decide Continue, reprice, narrow, hand over or kill.
“Three blockers remain.” Missing authority, unresolved identity conflict, and no recovery after a partial directory update. Slower to start, and much harder to fool yourself with at the end.
A composite example of two agent contracts, not a measured deployment.
More conversations, fewer obligations
Parallelise discovery. Serialise responsibility. A suggested discovery funnel runs from forty operator conversations to ten evidenced workflow maps, three paid bounded pilots and one continuing venture.
Cofounder or another idea?
The partner filter scores a potential partner from 0 to 5 on domain depth, customer access, ownership of commercial work, evidence access and delivery reliability. Customer access and commercial ownership each carry 25% of the weighted fit, domain depth 20%, and evidence access and reliability 15% each. The relationship follows from the first rule that applies:
| When | Relationship | Why |
|---|---|---|
| Customer access or commercial ownership is 1 or less | Useful expert. Not yet a venture partner. | Domain knowledge without customer access or commercial ownership leaves you holding product and go-to-market. |
| Evidence access is 1 or less, or reliability 2 or less | Paid diagnostic before a pilot. | The opportunity may be real, but access or behaviour is too uncertain for an equity commitment. |
| Fit of 4 or more, with customer access, commercial ownership and reliability all 4 or more | Founder discussion earned. | The partner appears to own a complementary function. Now define continuing duties, vesting, governance and departure terms. |
| Fit of 3 or more | Paid pilot before founder talks. | Useful access, but continuing commercial ownership and reliability still need evidence. |
| Anything else | Discovery only. | Keep the conversation. Do not let enthusiasm become an operating obligation. |
A partner rated 4, 4, 3, 4 and 3 scores 3.6: a paid pilot before founder talks.
About 50% can make sense for half the continuing founder responsibility. It is not the automatic price of version one. Define duties, vesting, governance and departure before the percentage, and move from trial to paid pilot to continuing duties, vesting and governance in that order. Discovery can run in parallel; responsibility should not, until capacity is proven.
The naive formula assumes your attention is free: the chance of at least one success is 1 − (1 − p)n. This capacity model cuts each venture’s chance once weekly load exceeds capacity, by e−γ·x² where x is the overload above capacity, before combining them.
Naive Capacity-adjusted Selected number of ventures
3 ventures fit the entered capacity. The overload penalty has not activated.
Illustrative inputs. The overload penalty is an assumption that makes attention visible, not an estimate of anyone’s odds.
Your evidence must compound faster than the market catches up
Models improve. Experts learn. Vertical tools package the work. Incumbents adapt.
I use a scenario model to make that race explicit; it is not a forecast. The advantage A(t) = G₀·e−kt·D(t)·E(t) starts at today’s gap, decays with market catch-up, grows with your learning and reuse, and is scaled by your delivery evidence. With a gap of 55 today, 40% catch-up a year, 20% learning and reuse a year and 60% delivery evidence, the gap has an illustrative half-life of 1.7 years and a capture index today of 33.0.
Under those inputs your delivery capability roughly tracks market catch-up, so the reading is to run a paid test: make a small, reversible investment and measure the result. If catch-up outruns your learning by 22 points or more, move and stay narrow; run a paid test, not a portfolio announcement. Delivery evidence below 40% means earning evidence before equity. If learning and reuse outpace catch-up with at least 70% evidence, you may be building an accumulating asset, and the next question is whether a second delivery is actually cheaper.
The temporary advantage is only useful if you convert it into customers, evidence, reusable systems and ownership.
One workflow, one buyer, one bill
The next evidence should come from a user, a payment or a failed acceptance test. A 30-day version:
- Days 1–5: find the ugly workflow. Baseline time, error, delay and consequence. Output: a signed problem boundary.
- Days 6–10: sell the diagnostic. Charge to map acceptance before software. Output: a paid decision contract.
- Days 11–24: ship one complete slice, from real input to accepted state. Output: accepted or rejected evidence.
- Days 25–30: count the delta. Compare SME + AI against SME + you + AI. Output: a fee, a founder talk, a revision or a stop.
Keep score on time (human hours per accepted result), quality (representative cases and critical failures), burden (review, rework and support), value (customer gain after switching) and reuse (whether the second build is actually cheaper).
I wouldn’t start ten companies. I’d start forty conversations, find ten real workflows, sell three bounded pilots and let one earn the right to become a company.
Five ways this dies
- The expert learns. Experts using frontier models alone reach the same accepted outcome at lower total cost. Then my intermediary role is weak. I should narrow into specialist engineering, own distribution, or stop pretending orchestration is scarce.
- The model absorbs it. The base model absorbs my orchestration, evaluation and recovery patterns. Then prompts and scaffolding were never the moat. I need domain data, distribution, reusable operations or a different role.
- Support wins. Delivery looks profitable until exceptions and support consume the margin. Then the service is underpriced or badly chosen. Count total burden, reprice, narrow the accepted state or stop.
- Nobody pays. Operators praise the demo but nobody pays to change the workflow. Then the pain, buyer, timing or offer is wrong. Do not convert compliments into market evidence.
- Equity traps me. Paper ownership accumulates while cash and attention disappear. Then the portfolio is an obligation portfolio. Require funding, vesting, continuing duties and explicit stop conditions.
The model bill collapsed; the company bill did not. That mismatch could let one engineer and one operator build what once needed a team. It could also disappear as experts learn and vendors package the workflow. So I’d move now, stay narrow and measure everything. Everyone has a thesis. None of them proved how.
The companion operating note is Your AI can build the app. Make it prove the outcome.
For coding agents: the accepted-outcome contract
The contract should change what an agent builds next. You get a decision system: workflow, economics, partner gate and kill condition. The agent gets a definition of done that a polished demo cannot satisfy: authority, evidence, recovery, alternatives and a fixed return format.
Agent stopping rule. Do not claim completion until the resulting environment matches the accepted state, authority is explicit, recovery is tested, and both sides of the economics are visible.
To write a contract for your own workflow, fill in the fields and paste the result into Codex, Claude Code or another agent.
Sources and further discussion
Public benchmarks and papers are sourced. The timing model, example economics and strategy are mine. Illustrative inputs are not forecasts. The supplied social screenshots are cultural signals, not empirical proof.
- OpenAI, GPT-6 Astra launch, September 2026. Vendor capability, price and benchmark claims.
- Anthropic, Claude Fable 5.1 launch, September 2026. Vendor capability and long-running work claims.
- Artificial Analysis, 9 September 2026. Independent cost and aggregate benchmark snapshot.
- Sequoia Capital, Services: The New Software, March 2026. The $1 software to $6 services thesis.
- Demirer, Musolff and Yang, Writing Code vs. Shipping Code, 2026. Working paper; observational caveats apply.
- Anthropic, Demystifying evals for AI agents. Resulting-state verification guidance.
These informed the original interactive edition’s format rather than the argument:
- Segel and Heer, Narrative Visualization, 2010. Author-driven and reader-driven narrative structures.
- Méndez and Such, Scrollytelling as an Alternative Format, 2026. Scrollytelling evidence in one context.
- Nielsen Norman Group, How People Read Online. Scanning patterns and information hierarchy.
- Nielsen Norman Group, Progressive Disclosure. Layering advanced detail behind clear cues.
- Wikipedia, Signs of AI writing. Editorial lint, not authorship detection.
- W3C, Animation from Interactions. Motion control and accessibility.