Skip to content
← Back to writing

AI is bigger than the apps selling it

I think AI can be underestimated while AI businesses are overpriced. Why agents could make today's tools disappear from view, and why implementation matters.

The interface is the visible edge of a much deeper capability.AI-generated conceptual setting.

Someone tries an AI app, gets a disappointing result and concludes that AI is overhyped. What did they test: the model, the product around it, or their way of using it?

I think AI can be underestimated while AI businesses are overpriced. I expect an expectations correction, but I could be wrong about whether, when or how deeply it happens. A falling valuation would not settle the question of capability.

The app is one way of buying the capability

An app, the implementation behind it and the underlying models answer different questions. Paying for one should be a considered choice.

Separate three questions
  1. 01

    What does the interface add?

    Convenience, creative control and support can be worth paying for. An app is one route to the capability.

    01
    The interfaceControls, convenience and support
    02
    The implementationContext, connections, checks and recovery
    03
    The capabilityThe models doing part of the work

    The value of the business is a separate question.

  2. 02

    What completes the whole job?

    Context, tools, checks and recovery determine whether the output becomes useful work.

    01
    The interfaceControls, convenience and support
    02
    The implementationContext, connections, checks and recovery
    03
    The capabilityThe models doing part of the work

    The value of the business is a separate question.

  3. 03

    What could the model do elsewhere?

    Direct access or a narrow custom workflow may be another option. Include time, usage, failed runs, maintenance and support in the comparison.

    01
    The interfaceControls, convenience and support
    02
    The implementationContext, connections, checks and recovery
    03
    The capabilityThe models doing part of the work

    The value of the business is a separate question.

A conceptual stack, not a vendor architecture. Direct access is not automatically cheaper, and capability does not establish a business valuation.

Higgsfield’s video workspace packages multiple models and creative controls. Lovable’s documentation explains how to sync code to GitHub, edit elsewhere and deploy outside Lovable. Those are legitimate products; they are also routes into capabilities that can exist elsewhere.

The interface is one way into the capability
An architectural cutaway separates a thin glass interface, an off-white layer of connections, and a substantial graphite computational lattice underneath.

The visible interface, the software joining the work together and the underlying capability are different parts of the purchase.

AI-generated conceptual cutaway, not a vendor architecture. Another route still carries integration, maintenance and support costs.

A coding agent might help you build the narrow part you need. I don’t know these companies’ margins, and direct access is not automatically cheaper. Convenience and someone else fixing a broken integration have value.

I don’t think the future is a bigger collection of AI tools

I expect more work to begin with the intended outcome, with an agent coordinating the tools underneath.

The work stays. Who carries it changes.

One hypothetical job: prepare a product launch. Both routes still need the same underlying work.

You coordinate
  1. Explain the product
  2. Move work between tools
  3. Check the pieces together
  4. Decide what can go live
An agent coordinates
  1. You set the brief
  2. Agent joins the tools
  3. Agent brings checked work
  4. You retain the release decision
Still underneath either routeModels · software · product facts · integrations · checks · repair

The test is the finished launch. Count coordination, review and repair as work, even when the interface makes them less visible.

A proposed division of work, not a working autonomous launch service. If supervising the agent costs more than coordinating the tools, this example has not delivered its promised benefit.

This is the direction I want to build towards. Specialised interfaces could remain valuable for direct creative control; providers could supply capabilities an agent selects. My prediction is about less tool management, not the disappearance of the software.

It connects personal AI that carries context forward with JustSwipe’s prepared decisions. I want less of my day spent explaining my life to one piece of software after another.

Implementation is where this argument has to earn its keep

The surrounding files, instructions, tools, tests and feedback—often called a harness—determine whether a model’s output becomes completed work. Research gives us reasons to test that carefully.

Two studies. Two different measures. Opposite directions.
Company experiments · 4,867 developers≈26% more completed tasks
Without assistant
100
With assistant
≈126

Task count, indexed to a baseline of 100. Estimates varied; individual experiments were noisy. Code completion, not autonomous businesses.

Company study · February 2026
METR · 16 experienced developers19% longer to complete tasks
Without AI tools
100
With AI tools
119

Completion time, indexed to a baseline of 100. Experienced contributors in familiar repositories, using the early-2025 tools studied.

METR study · July 2025

Separate within-study comparisons, not a combined estimate. More tasks and more time are different outcomes. Neither chart establishes the effect of current autonomous agents.

The METR developers felt faster even while taking longer. That gap should bother anyone building with AI, including me. Its February 2026 follow-up says newer participation, task selection and parallel-work timing prevented a reliable current-effect estimate. Neither an old slowdown nor new enthusiasm settles today’s result.

Anthropic’s November 2025 engineering account describes agents losing unfinished work and declaring completion too early. Progress records, smaller steps and actual tests helped. That is vendor experience, not proof that every business needs an elaborate system.

I think dependable implementation is less common than talk about AI. I have no credible headcount. The useful test is a finished job at an acceptable total cost.

A disappointing website generator doesn’t define the physical frontier

NVIDIA’s EgoScale research explores transfer into physical tasks, including unscrewing bottle caps. Its February 2026 result needs to be read with its training conditions.

A one-shot result carries a great deal of prior learning
  1. 20,000+ hoursHuman video pretraining

    The starting point already contains extensive learning.

  2. 100 human demonstrationsAligned examples per object

    These accompany the new robot task.

  3. 1 robot demonstrationAdaptation in the one-shot setup

    Evidence of bounded transfer, not general understanding.

NVIDIA EgoScale, February 2026. This is a training recipe, not a proportional chart. The result does not demonstrate an autonomous factory.

Prior learning can make adaptation more efficient. That direction interests me far more than whether a consumer app gives a convincing demo.

Bezos’s June 2026 account of Prometheus describes an ambition to shorten engineering cycles. It does not demonstrate an autonomous factory. My interest is a loop that proposes, predicts, tests and revises a design; neither the robot result nor the business ambition establishes that complete loop.

What I’m betting my work on

I’ve felt concern about losing work. Someone’s anxiety isn’t explained away by saying they haven’t tried the right app.

I think business judgement, applied AI and software architecture are moving closer together: the problem shapes the build, model capability shapes what is feasible, and architecture determines whether it works in a business.

A correction could punish useful experiments alongside weak ones. Being right about a technology’s importance does not guarantee surviving its business cycle.

I expect agents to take over more of today’s tool coordination. If reliability, economics or people’s preference for direct control keep defeating that shift, I’ll have to revise the prediction. For now, I’d rather learn to build the system that completes the job.

Sources and further discussion