AI is bigger than the apps selling it
I think AI can be underestimated while AI businesses are overpriced. Why agents could make today's tools disappear from view, and why implementation matters.
The interface is the visible edge of a much deeper capability.AI-generated conceptual setting.
Someone tries an AI app, gets a disappointing result and concludes that AI is overhyped. What did they test: the model, the product around it, or their way of using it?
I think AI can be underestimated while AI businesses are overpriced. I expect an expectations correction, but I could be wrong about whether, when or how deeply it happens. A falling valuation would not settle the question of capability.
The app is one way of buying the capability
An app, the implementation behind it and the underlying models answer different questions. Paying for one should be a considered choice.
-
01
What does the interface add?
Convenience, creative control and support can be worth paying for. An app is one route to the capability.
01The interfaceControls, convenience and support02The implementationContext, connections, checks and recovery03The capabilityThe models doing part of the workThe value of the business is a separate question.
-
02
What completes the whole job?
Context, tools, checks and recovery determine whether the output becomes useful work.
01The interfaceControls, convenience and support02The implementationContext, connections, checks and recovery03The capabilityThe models doing part of the workThe value of the business is a separate question.
-
03
What could the model do elsewhere?
Direct access or a narrow custom workflow may be another option. Include time, usage, failed runs, maintenance and support in the comparison.
01The interfaceControls, convenience and support02The implementationContext, connections, checks and recovery03The capabilityThe models doing part of the workThe value of the business is a separate question.
A conceptual stack, not a vendor architecture. Direct access is not automatically cheaper, and capability does not establish a business valuation.
Higgsfield’s video workspace packages multiple models and creative controls. Lovable’s documentation explains how to sync code to GitHub, edit elsewhere and deploy outside Lovable. Those are legitimate products; they are also routes into capabilities that can exist elsewhere.
The visible interface, the software joining the work together and the underlying capability are different parts of the purchase.
AI-generated conceptual cutaway, not a vendor architecture. Another route still carries integration, maintenance and support costs.
A coding agent might help you build the narrow part you need. I don’t know these companies’ margins, and direct access is not automatically cheaper. Convenience and someone else fixing a broken integration have value.
I don’t think the future is a bigger collection of AI tools
I expect more work to begin with the intended outcome, with an agent coordinating the tools underneath.
One hypothetical job: prepare a product launch. Both routes still need the same underlying work.
- Explain the product
- Move work between tools
- Check the pieces together
- Decide what can go live
- You set the brief
- Agent joins the tools
- Agent brings checked work
- You retain the release decision
The test is the finished launch. Count coordination, review and repair as work, even when the interface makes them less visible.
A proposed division of work, not a working autonomous launch service. If supervising the agent costs more than coordinating the tools, this example has not delivered its promised benefit.
This is the direction I want to build towards. Specialised interfaces could remain valuable for direct creative control; providers could supply capabilities an agent selects. My prediction is about less tool management, not the disappearance of the software.
It connects personal AI that carries context forward with JustSwipe’s prepared decisions. I want less of my day spent explaining my life to one piece of software after another.
Implementation is where this argument has to earn its keep
The surrounding files, instructions, tools, tests and feedback—often called a harness—determine whether a model’s output becomes completed work. Research gives us reasons to test that carefully.
Task count, indexed to a baseline of 100. Estimates varied; individual experiments were noisy. Code completion, not autonomous businesses.
Company study · February 2026Completion time, indexed to a baseline of 100. Experienced contributors in familiar repositories, using the early-2025 tools studied.
METR study · July 2025Separate within-study comparisons, not a combined estimate. More tasks and more time are different outcomes. Neither chart establishes the effect of current autonomous agents.
The METR developers felt faster even while taking longer. That gap should bother anyone building with AI, including me. Its February 2026 follow-up says newer participation, task selection and parallel-work timing prevented a reliable current-effect estimate. Neither an old slowdown nor new enthusiasm settles today’s result.
Anthropic’s November 2025 engineering account describes agents losing unfinished work and declaring completion too early. Progress records, smaller steps and actual tests helped. That is vendor experience, not proof that every business needs an elaborate system.
I think dependable implementation is less common than talk about AI. I have no credible headcount. The useful test is a finished job at an acceptable total cost.
A disappointing website generator doesn’t define the physical frontier
NVIDIA’s EgoScale research explores transfer into physical tasks, including unscrewing bottle caps. Its February 2026 result needs to be read with its training conditions.
- 20,000+ hoursHuman video pretraining
The starting point already contains extensive learning.
- 100 human demonstrationsAligned examples per object
These accompany the new robot task.
- 1 robot demonstrationAdaptation in the one-shot setup
Evidence of bounded transfer, not general understanding.
NVIDIA EgoScale, February 2026. This is a training recipe, not a proportional chart. The result does not demonstrate an autonomous factory.
Prior learning can make adaptation more efficient. That direction interests me far more than whether a consumer app gives a convincing demo.
Bezos’s June 2026 account of Prometheus describes an ambition to shorten engineering cycles. It does not demonstrate an autonomous factory. My interest is a loop that proposes, predicts, tests and revises a design; neither the robot result nor the business ambition establishes that complete loop.
What I’m betting my work on
I’ve felt concern about losing work. Someone’s anxiety isn’t explained away by saying they haven’t tried the right app.
I think business judgement, applied AI and software architecture are moving closer together: the problem shapes the build, model capability shapes what is feasible, and architecture determines whether it works in a business.
A correction could punish useful experiments alongside weak ones. Being right about a technology’s importance does not guarantee surviving its business cycle.
I expect agents to take over more of today’s tool coordination. If reliability, economics or people’s preference for direct control keep defeating that shift, I’ll have to revise the prediction. For now, I’d rather learn to build the system that completes the job.
Sources and further discussion
- Suproteem K. Sarkar: AI Agents and Higher-Order Work (working paper, May 2026). Research using Cursor activity examines the shift from implementation toward supervision, with delegation varying by how easily outputs can be checked and by worker experience. It concerns one coding platform; it does not establish that agents will replace every interface or business model.
Related Posts
A new kind of company just became possible
AI made implementation cheap before companies learned how to use it. For a brief window, a small technical team can sell the finished work. The hard part is proving the accepted result.
The check you didn’t know to ask for
Your agent can pass every test you wrote. Who notices the test you forgot? A working local example of keeping the check; discovery and human benefit are unmeasured.
The reel sold dinner. Your AI built homework.
Before you hand an idea to your agent, find the work its users won’t do. A social-media reply becomes a smaller product test.