← Back to writing

Stop asking AI for posts. Give it a world that can surprise you.

A content system should run experiments, capture proof and return with a story you could not know in advance.

An agent can fill a calendar without learning anything. Give it a topic and it can turn that topic into twenty finished assets without producing one new event.

A better system changes one bounded world, records what actually follows, and lets the next episode inherit the result. What that system produces is a renewable source of events that neither you nor the agent can fully script in advance.

My position: the future will not have a shortage of content. It will have a shortage of things that actually happened.

I had automated the least scarce part

The pipeline could research a subject, draft a script, choose footage, generate a voice, add captions and render a short. It never had to wait for a real result. That convenience was the defect.

The machine had no reason to fail. If a claim was awkward, it softened the claim. If evidence was missing, it replaced evidence with explanation. If the topic was dull, it inflated the topic until it sounded historically important.

The output was usually competent, which made the problem harder to see. Bad content is easy to reject. Smooth, plausible, unnecessary content can fill a calendar for months. I had built a factory that could manufacture the appearance of having something to say.

The failure I observed was that throughput rose faster than specificity. The system measured completed assets. It did not measure whether a new fact, conflict or consequence had entered the world. A system can complete every production step without creating one new fact.

Six polished, interchangeable outputs the machine can make without an event
Authored examples. Outputs: 6. New evidence: 0. World state changed: no.
OutputWhat gives it away
5 ways AI is transforming modern work“AI is reshaping how teams collaborate, innovate and unlock efficiency across an evolving digital landscape.” Importance inflation, generic attribution, no event.
The future of autonomous agents“From productivity to creativity, intelligent agents are poised to redefine the way we live and work.” Future fog, broad trend, no proof.
Why verification matters“Robust verification plays a pivotal role in ensuring trustworthy, reliable and responsible AI systems.” Pivotal role, rule of three, no concrete failure.
Building smarter systems“By combining automation with human insight, organisations can foster innovation and drive meaningful outcomes.” Superficial synthesis, business fog, no stakes.
What comes after software?“The convergence of AI and automation marks a significant shift in the broader technology landscape.” Broader landscape, significance claim, interchangeable.
A practical guide to AI workflows“This guide explores key strategies, common challenges and actionable steps for navigating AI-powered work.” Outline voice, canned triad, no lived detail.

The production layer is becoming infrastructure

Production is becoming cheap, and events are not. The three images below are exact visual assets committed to real GitHub repositories. Editing, rendering and evaluation are becoming agent-addressable. The remaining strategic question is what the machine is being asked to discover.

The actual video-use timeline view showing frames, waveform, transcript labels and cut candidates
Editing becomes an agent operation.browser-use/video-use The repository exposes transcript, waveform and cut evidence instead of asking a model to guess from raw frames.
The actual Remotion logo from its brand repository
Rendering becomes code.remotion-dev/brand Programmatic video is increasingly a callable production primitive rather than a specialist bottleneck.
The actual video-evaluator benchmark snapshot showing operational success and semantic failure
Running is not understanding.45ck/video-evaluator The benchmark is valuable precisely because it records semantic failure instead of converting pipeline completion into proof.

Exact upstream files: the packaged bytes reproduce the Git blob identifiers recorded in the asset register.

AI writing and vibe-coded interfaces fail in the same way

They regress toward a competent middle. The result is rarely unusable. It is simply hard to believe that a particular person needed to make it. Writing and interface sameness are symptoms of the same missing thing: a particular decision.

The problem goes deeper than the word “delve”

Wikipedia’s field guide to signs of AI writing is useful because it notices deeper patterns: generic importance claims, superficial analysis, promotional tone, vague attribution and prose that could attach to almost any subject. The page also warns that no single pattern proves AI authorship.

Generic draft (authored specificity score: 18%)

In today’s rapidly evolving digital landscape, autonomous content systems are playing a pivotal role in helping creators unlock meaningful opportunities. By combining innovation, automation and human insight, these systems highlight the importance of building engaging experiences that resonate with modern audiences.

Edited draft (authored specificity score: 91%)

I asked the system for ten episode ideas about autonomous agents. Eight opened with the same promise: productivity. None required a real agent to attempt a task. I deleted them and kept the failed calendar benchmark because it produced a trace I could inspect.

Editing rule: replace importance with consequence, category words with objects, and unsupported significance with an event the reader can inspect.

AI UI has an accent too

Current critiques of vibe-coded products point to visual sameness, rounded-card repetition, muted gradients, polished default states and neglected error or edge states. Beige and rounded corners are not forbidden. The issue is that the interface has no evidence of a decision.

The same mock product page, left to defaults and art-directed
Default generated UIArt-directed system
HeadlineBuild better stories with AIA world should be able to prove you wrong.
Supporting lineTransform your creative workflow with intelligent automation.One change. One open result. One permanent consequence.
Call to actionGet startedInspect the evidence
Error stateEverything looks ready.ERROR 04 · proof missing · publication blocked
What it showsEvery block uses the same visual grammar. The copy could describe hundreds of products. The error state has not been designed.The headline carries the argument, not a category label. Form follows the editorial hierarchy instead of a component kit. The broken state is designed as carefully as the happy state.

Distinctive design is evidence that somebody chose what matters, what can wait and what should happen when the happy path breaks.

A post is an output, and a world is a renewable source of events

The world does not need to be fictional. It can be a benchmark, a physical prototype, a simulated town, a drawer, a competition or a museum. It needs state, rules and a way to answer back.

A post generator starts with a subject and fabricates an output. A world engine starts with persistent state, changes one thing, then waits for reality, a simulation or an audience to answer back.

The difference is memory. A good episode leaves the world changed, and the next episode inherits that change. Viewers are no longer asked to remember a creator because of branding. They remember what happened.

world + intervention + uncertainty + proof + consequence + memory = episode

Persistent state, one intervention and permission to be wrong create renewable story pressure.

Each episode starts from the state the last one left behind

Six stages of a world engine feed an episode, an event with receipts, and return to persistent state.

Six connected stages of a world engine feeding an episode and returning to persistent state Worldstate Changeintervention Openuncertainty Proofcapture Changeconsequence Keepmemory Episodean event with receipts
  1. Start with state. Records, debts, laws, beliefs, objects, rivalries and unanswered questions already exist before the episode begins. Persistent state makes the first episode a beginning instead of an isolated asset.
  2. Change one thing on purpose. Add a rule. Remove a dependency. Change the lighting. Introduce a new event. The intervention should be clear enough to argue about later, and it gives the episode a question precise enough to test.
  3. Leave room to be wrong. If the planned conclusion cannot lose, the episode is an illustrated opinion. Genuine uncertainty gives the evidence a job: the result must be able to contradict the creator’s preferred conclusion.
  4. Decide what would count as proof. A trace, a clock, a sensor, a ledger, a replay, a state diff or a human judgement must be able to contradict the script. The evidence contract prevents the narration from replacing the event.
  5. Make the consequence visible. The record moves. The monster grows. An exhibit changes label. A faction gains power. Something is different because the event occurred, and visible consequence makes abstract change legible in a glance.
  6. Do not reset. The next episode inherits the record, the debt, the law or the unresolved contradiction. History becomes part of the format: it compounds, and the next episode begins where the previous one actually ended.

A design model for planning episodes, not a measured production process.

The Maintenance Monster turns invisible upkeep into plot

An automation claims to save six hours each week. Every hidden intervention, patch and exception feeds a physical creature. The episode ends when the ledger and the monster agree on what the automation is worth.

Every silent repair is charged against the six hours the automation claims to save

A live model of one automation world. Each action changes the weekly repair time by the amount shown; net weekly value is the six-hour claim minus recorded repair time.

Claimed saving6.0 h
Repair time0.0 h
Exceptions0
Net weekly value+6.0 h

Maintenance debtThe mark is the six-hour claim

  1. The automation enters the world with a six-hour claim and no recorded maintenance.

Authored increments for illustration. The episode’s real ledger would come from time logs, exception counts and patch history.

When every silent repair changes the creature, maintenance becomes plot instead of hidden labour. Different worlds should fail differently.

Six worlds, and what each one can prove
01 · Agent Olympics
Work becomes sport when the rules, clocks and disqualifications are visible. Plot engine: competition and prediction. Intervention: introduce one new discipline or change one rule. Proof: fixtures, traces, intervention logs, costs and a reproducible scoreboard. Memory: future events inherit records, rivalries and rule changes. Episode premise: “Three agents enter the Calendar Collision Relay. The fastest result is disqualified because a hidden human correction changed the schedule.”
02 · The Failure Museum
A failure becomes useful when it stays available for inspection. Plot engine: evidence mystery and revision. Intervention: reconstruct one incident or challenge an accepted explanation. Proof: logs, screenshots, timestamps, conflicting accounts and a controlled replay. Memory: new incidents can connect to older exhibits and overturn the museum catalogue. Episode premise: “Exhibit 014 was labelled ‘model hallucination’. A replay shows the model followed a corrupted dashboard exactly.”
03 · Lost Files from the Autonomous Office
The audience reconstructs a world from evidence the world accidentally left behind. Plot engine: artifact-first investigation. Intervention: recover one file, room, message or system log. Proof: clearly labelled fictional records, maps, tapes and incident files. Memory: every file becomes part of a persistent timeline with unresolved contradictions. Episode premise: “A training tape warns employees never to empty the approval queue. The office has had no employees for fourteen years.”
04 · The Maintenance Monster
Every hidden repair makes the creature larger. Plot engine: physical metaphor backed by cost accounting. Intervention: add, repair, simplify or retire an automation dependency. Proof: time logs, exception counts, patch history and a maintenance ledger. Memory: old shortcuts return as future maintenance debt. Episode premise: “The workflow claims to save six hours. After three undocumented repairs, the monster is larger than the task it replaced.”
05 · Agent Civilisation Lab
One constitutional change can create consequences several episodes later. Plot engine: institutional simulation. Intervention: change one law, resource rule or communication constraint. Proof: seeded runs, state snapshots, rule history and replayable events. Memory: the world never resets; later episodes inherit earlier decisions. Episode premise: “The town creates a court for hallucinations. By the next episode, the court is manufacturing cases to justify its existence.”
06 · Drawer Ecology Lab
A domestic object becomes interesting when it can hold a wrong belief. Plot engine: physical sensing and classification failure. Intervention: change lighting, move objects or introduce an unknown item. Proof: frames, sensor readings, confidence histories and manual corrections. Memory: the drawer carries a history of belief and correction. Episode premise: “A black sock disappears under warm light, returns as a phone charger and becomes a new species after two sensors disagree.”

The equations cannot predict a hit, but they can stop me combining different failures into one fake score

Opening acceptance, retention, diffusion, audience memory and portfolio diversity are different processes. Each needs its own evidence and its own failure state. Do not compress attention, spread, return and creative diversity into one flattering score.

M1: A hook is an expectation contract

The opening needs to be legible, interesting and credible. It also needs to avoid promising a payoff the episode cannot deliver.

â = (ℓ c e)1/3 (1 − m), ℓ, c, e, m ∈ [0, 1]

M1, opening acceptance. Legibility, curiosity and evidence all matter; promise mismatch reduces the opening contract.

Modelled acceptance0.62

A workable opening contract, but at least one component is limiting acceptance.

Authored hypothesis. The shape is deliberate, but its coefficients are not estimated from an audience.

M2: Attention is a survival curve

The opening earns the next moment. Every later beat either preserves that decision or adds an abandonment hazard. YouTube itself reports time-indexed retention rather than a universal “attention span”.

S(t) = exp(−∫t0 h(u) du), h(u) ≥ 0
𝔼[min(T, D)] = ∫D0 S(t) dt

M2, viewer survival. Retention is time-indexed. Expected watch is the area under the survival curve.

Move the authored hazards and watch the curve

100%0%60 sec
Opening accepted84%
Completion52%
Expected watch40.6 s

Research-aligned form. The survival structure is established mathematics. The creative hazard inputs are authored and shadow-only; at their defaults the illustrative 60-second curve opens at 84%, completes at 52% and gives an expected watch of 40.6 seconds.

M3: A large audience can arrive through different shapes

One platform broadcast and a deep person-to-person cascade may create similar view counts. Structural virality measures topology rather than size.

N = Next + ΣGg=1 Ng
𝔼[N | R < 1] = μ / (1 − R)

M3, exposure and diffusion. A broadcast and a person-to-person cascade can reach similar totals through different structures.

Three illustrative diffusion patterns
PatternEstimated reachEndogenous shareMean depthQualified return
Broadcast-heavy540k8%1.73.2%
Cascade-heavy506k73%7.95.8%
Smaller, durable188k61%4.812.6%

Illustrative. The network figures are explanatory. They are not fitted to a platform or account.

M4: Distribution becomes valuable when exposure becomes memory

A person can sample once, recognise the series, return, become regular or disappear. Those transitions should not be compressed into follower count. YouTube similarly distinguishes new, casual and regular viewers.

πt+1 = πtP, Pij ≥ 0, Σj Pij = 1

M4, audience memory. Exposure becomes an audience only when people move into remembered and returning states.

Episode 0, starting from 1,000 sampled people

  • Sampled1,000
  • Recognised0
  • Returning0
  • Regular0
  • Dormant0

The first release creates exposure. It has not created an audience yet.

The transition matrix behind each release
Share moving from each state (row) to each state (column) per episode
FromSampledRecognisedReturningRegularDormant
Sampled0.400.350.100.000.15
Recognised0.050.350.350.050.20
Returning0.020.050.480.250.20
Regular0.010.020.120.680.17
Dormant0.050.080.070.030.77

State model. Useful for accounting and mock tests. Not validated as a literal model of real people.

M5: One winner can make the whole system worse

AI-assisted creators may produce stronger individual outputs while the collection becomes more similar. A quality-diversity archive keeps strong candidates across different behavioural niches instead of allowing one early winner to consume every production slot.

maxx ∈ {0,1}n Σni=1 qixi
subject to Σni=1 cixi ≤ B, Σi∈𝒞m xi ≥ 1 ∀m, Σi∈𝒲r xi ≤ ρ Σni=1 xi ∀r

M5, creative portfolio. The constraints stop one short-term winner from erasing the creative search space.

Two illustrative allocation policies for 30 production slots across six worlds
PolicySlots per worldWorlds preservedLargest shareNovel mechanism slots
Greedy score17, 0, 0, 13, 0, 02 / 657%1
Diverse portfolio6, 5, 4, 5, 4, 66 / 620%8

Research analogy. MAP-Elites is not a validated content strategy. The transfer is the preservation of multiple strong niches.

Notation and LaTeX for the five models
M1  \widehat a=(\ell c e)^{1/3}(1-m),\qquad \ell,c,e,m\in[0,1]
M2  \begin{aligned}S(t)&=\exp\!\left(-\int_0^t h(u)\,\mathrm du\right),\qquad h(u)\ge 0,\\[2pt]\mathbb E[\min(T,D)]&=\int_0^D S(t)\,\mathrm dt\end{aligned}
M3  \begin{aligned}N&=N_{\mathrm{ext}}+\sum_{g=1}^{G}N_g,\\[2pt]\mathbb E[N\mid R<1]&=\frac{\mu}{1-R}\end{aligned}
M4  \boldsymbol{\pi}_{t+1}=\boldsymbol{\pi}_tP,\qquad P_{ij}\ge0,\quad \sum_jP_{ij}=1
M5  \begin{aligned}\max_{\boldsymbol{x}\in\{0,1\}^{n}}\quad&\sum_{i=1}^{n}q_ix_i\\[2pt]\text{subject to}\quad&\sum_{i=1}^{n}c_ix_i\le B,\\&\sum_{i\in\mathcal C_m}x_i\ge1\quad\forall m,\\&\sum_{i\in\mathcal W_r}x_i\le\rho\sum_{i=1}^{n}x_i\quad\forall r\end{aligned}

Test which world makes relevant people return

Do not search for the viral hook. The first pilot should compare different plot engines. Cosmetic hook variants can wait until there is evidence that the underlying world deserves another episode. Compare content worlds over repeated episodes, not isolated hooks over one weekend.

Candidate worlds for the pilot. Choose four and run four episodes of each, sixteen runs in an interleaved order.
WorldMechanismProduction burden
Agent OlympicsCompetition3 / 5
Failure MuseumEvidence mystery2 / 5
Maintenance MonsterPhysical metaphor4 / 5
Lost FilesFictional investigation5 / 5
Civilisation LabPersistent simulation5 / 5
Drawer EcologyPhysical observation4 / 5

The primary outcome is qualified subsequent-episode consumption: did a relevant viewer voluntarily consume another episode from the same world within the declared window? Four guardrails sit beside it: evidence completeness, trust incidents, production burden and portfolio concentration.

One hit does not crown a format. The system should earn confidence through repeated episodes, comparable cohorts and a reasoned account of what changed.

Automate the labour, and keep public authority bounded

An agent can research, propose experiments, write, render, test and package a release. That does not entitle it to decide that a claim is true, a person is safe to depict or a piece is worth publishing.

Authority levels for a content agent. The selected boundary is A3.
LevelWhat the system may do
A0 · ObserveThe system can inspect public or approved material and report findings. It cannot create or change artifacts.
A1 · DraftThe system can propose ideas, scripts and plans. Every artifact remains a suggestion awaiting human review.
A2 · BuildThe system can produce local assets, simulations and renders. Nothing leaves the controlled environment.
A3 · Prepare (selected)The system can complete the production package and evidence checks. Calvin reviews the finished piece and explicitly authorises public release.
A4 · Bounded publishOnly calibrated, low-risk, allowlisted formats may publish through official APIs, with kill switches, incident logging and strict scope.
A5 · UnboundedProhibited. The system does not receive unrestricted authority to publish claims, depict people or alter public channels.

I still want automation. I want agents to do the dull research, assemble the evidence pack, render the variants, run the checks and remember what happened last time.

I no longer want the agent to begin with “make content”. I want it to begin with a world, a rule and a question whose answer is not yet known. Distribution, for me, is accumulated memory rather than accumulated output.

For coding agents: give the agent a world, not a topic

The useful instruction is a bounded operating contract that forces state, uncertainty, proof and memory before production, rather than “make content about AI”.

Before creating content:
1. Name the persistent world state.
2. Change one thing on purpose.
3. State what remains genuinely unknown.
4. Lock the evidence needed before the run.
5. Keep an unflattering result.
6. Write the state change back into the world.
7. Ask for human approval before public release.
Open the World Engine protocol

Sources and further discussion

What is sourced, what is modelled and what still needs a real audience. Research-backed means a source directly supports the narrow claim in the studied or documented context. An authored hypothesis is a useful model shape or design judgement that still needs prospective testing. Illustrative means a visual explanation using arbitrary units, not an estimated platform model. My position marks an explicit argument, not a disguised empirical finding.

  1. Wikipedia: Signs of AI writing, Wikipedia WikiProject AI Cleanup. Descriptive field guide, not a detector or policy. Used to audit generic prose patterns.
  2. Progressive Disclosure, Jakob Nielsen. Supports keeping the primary reading path simple and making technical depth available on request.
  3. Understanding Animation from Interactions, W3C Web Accessibility Initiative. Motion is optional, user-controlled and removed under reduced-motion preferences.
  4. Narrative Visualization: Telling Stories with Data, Edward Segel and Jeffrey Heer. Informs the balance between a guided narrative and reader-controlled exploration.
  5. Generative AI enhances individual creativity but reduces the collective diversity of novel content, Anil R. Doshi and Oliver P. Hauser. Supports treating output quality and portfolio diversity as separate concerns.
  6. The Structural Virality of Online Diffusion, Sharad Goel, Ashton Anderson, Jake Hofman and Duncan Watts. Supports separating diffusion size from diffusion topology.
  7. Measure key moments for audience retention, YouTube Help. Supports time-indexed retention and matched comparisons, not a universal attention-span claim.
  8. New, casual and regular viewers, YouTube Help. Supports distinguishing one-time exposure from recurring audience behaviour.
  9. Illuminating search spaces by mapping elites, Jean-Baptiste Mouret and Jeff Clune. Used as an analogy for preserving strong candidates across different creative niches.
  10. Telltale signs of AI-coded products, Paul Bakaus and other product-design practitioners, reported by Business Insider. Used only as a current practitioner critique of visual sameness and neglected edge states.

The interactive models are designed to make assumptions visible. They do not predict future views, virality, revenue or audience return. The honest next step is a manually approved pilot.

Published 2026-09-13 · Updated 2026-09-28 · Source edition v3