← Back to writing

AI can make a thousand posts. Make it prove one thing happened.

Editing, animation and publishing are becoming agent tools. The scarce advantage is an event, a measurement or a failure that only your system can supply.

The production stack can research, script, edit, render and publish. It can finish an asset while nobody has supplied a reason to care. So ask the stack one question first: what happened that only you can prove? If the answer is nothing, the stack should stop before it renders.

The rule I propose is a proof gate. Your agent names the event and the evidence before production, or it stops. Below are one ordinary worked example and a machine-readable contract. The advantage is a lower cost per original fact.

For you, that means killing polished nothing before you review a finished asset: approve the question, evidence and failure condition, and reject the empty idea while it is still cheap. For your agent, it is a hard stop before scripting, editing or publishing. The machine must name what will happen, what remains unknown and what evidence could prove its preferred answer wrong. The economic edge is that one real result can support many useful explanations. The experiment is the fixed cost, and agentic production makes each honest derivative cheaper without inventing another conclusion.

𝒫 = W ∧ I ∧ U ∧ E ∧ F ∧ M

P1, the content-proof gate and acceptance rule. All six conditions must exist: world state, intervention, uncertainty, evidence, failure permission and memory. Polish cannot compensate for a missing event.

My position: the best content agent knows when no publishable event exists. Finishing fastest is the wrong test.

The price of polish is falling, and the price of being believed is rising

That combination changes the strategy. When everyone can produce a competent-looking video, competence stops proving that anything happened. The scarce object becomes the trace, test, failure, measurement or changed world underneath it. Editing and rendering are becoming callable. Evidence is still a separate job.

The logic runs in four steps:

  1. Agents can now make the artefact.
  2. Audiences are learning to discount generic polish.
  3. Evidence remains expensive, specific and difficult to copy.
  4. So make the agent operate the evidence, rather than merely describe it.

agentic production ↑ + generic-output discount ↑ → proof advantage

Production is becoming infrastructure, and trust is becoming a filter. Polish is becoming abundant while generic AI output is being discounted. My inference, as of September 2026: visible evidence and accumulated history become more valuable as production itself stops being scarce. It is a reason to test the strategy before proof protocols become ordinary, not proof that evidence-backed content will win.

The three images below are the exact SVG files committed to the cited repositories, not mock screenshots. The article only places them inside its own editorial frame.

The actual video-use timeline view asset showing a filmstrip, speaker tracks, waveform, word timestamps, silence gaps and suggested cut points.
The editing agent exposes its evidence.browser-use/video-use · static/timeline-view.svg Git blob bb3f11476546, unredrawn. The editing agent does more than promise automation: its repository exposes the timeline evidence it reasons over before cutting.
The official Remotion logo from the Remotion brand repository.
“Video tools for the agent era.”remotion-dev/brand · logo.svg Git blob 6c6b4ed7956c, unredrawn. Remotion’s own repository now frames video as an agentic, interactive and programmatic medium rather than a specialised finishing step.
The actual benchmark snapshot from Calvin’s video-evaluator repository, reporting three operational successes and zero semantic passes.
Running is not understanding.45ck/video-evaluator · assets/benchmark-snapshot.svg Git blob 7ca0bb3c9d8c, unredrawn. My own evaluator already showed the more uncomfortable distinction: a pipeline can run perfectly while still failing to understand what mattered.

The response has started outside the repositories too:

The old advantage was making polished media faster than the next person, and tools are compressing it. The advantage I propose is owning a result the next person cannot prompt into existence: they would have to reproduce the experiment, history and evidence. That window lasts until proof becomes the normal interface, so the useful move is to build the evidence habit now instead of publishing more generic volume.

I built three systems to make and inspect content, then the production layer stopped being the interesting part

Content Machine could assemble shorts. Demo Machine could turn executable product flows into videos. Video Evaluator could inspect the result. Meanwhile, public tools began solving the same production work faster. That was useful. It also made the missing layer impossible to ignore.

The stack could research, script, edit, render and review. It could manufacture a convincing object before reality had supplied a convincing event. When production was expensive, the finished artefact itself looked like evidence of effort. Agentic production breaks that shortcut. A finished-looking video may now mean very little.

The mistake was giving the machine permission to finish before it had something specific to lose. The failure I observed was that the state of the world changed faster than my architecture: editing and rendering became replaceable infrastructure, and the unanswered question moved upstream. What must be true before the agent is allowed to produce anything? Pipeline completion is not proof that anything new was observed.

Give it a topic and the calendar fills. The outputs are plausible, and none require a new event, a falsifiable claim or a changed world.

Six polished, interchangeable outputs the machine can make without an event
Authored examples. Outputs: 6. New evidence: 0. World state changed: no.
OutputWhat gives it away
5 ways AI is transforming modern work“AI is reshaping how teams collaborate, innovate and unlock efficiency across an evolving digital landscape.” Importance inflation, generic attribution, no event.
The future of autonomous agents“From productivity to creativity, intelligent agents are poised to redefine the way we live and work.” Future fog, broad trend, no proof.
Why verification matters“Robust verification plays a pivotal role in ensuring trustworthy, reliable and responsible AI systems.” Pivotal role, rule of three, no concrete failure.
Building smarter systems“By combining automation with human insight, organisations can foster innovation and drive meaningful outcomes.” Superficial synthesis, business fog, no stakes.
What comes after software?“The convergence of AI and automation marks a significant shift in the broader technology landscape.” Broader landscape, significance claim, interchangeable.
A practical guide to AI workflows“This guide explores key strategies, common challenges and actionable steps for navigating AI-powered work.” Outline voice, canned triad, no lived detail.

The stack still has no owner for “why care?”

The new stack has a tool for almost every production step. It can research (collect material), script (shape the claim), edit (assemble footage), render (make the artefact), publish (reach the platform) and measure (observe response). The reason to care, what changed in the world, has no owner. None of those capabilities proves that the source event was original, uncertain or real.

This is not a complaint about the tools. video-use can turn raw footage into an edited video. Remotion can make programmatic video the source of truth. Postiz can schedule and measure distribution. My own Content Machine tried to connect even more of that pipeline.

The tooling stack can improve every step after an idea exists. It does not establish that the idea deserves to become an asset. A system can be excellent at production while remaining indifferent to whether the source event is generic, invented or already known. Making the artefact and justifying the artefact are different jobs, and I had assigned the first job to the agent and quietly kept the second in my own head. “Why care?” must become a system responsibility, not a line added after rendering.

Finished-looking output is the trap

Generic prose and generic interfaces are symptoms of the same shortcut: the system completes the surface before anybody has chosen what the surface is supposed to prove. Finished-looking language and interface polish can conceal a missing event, and the article should survive its own test.

The problem goes deeper than the word “delve”

Wikipedia’s field guide to signs of AI writing is useful because it notices deeper patterns: generic importance claims, superficial analysis, promotional tone, vague attribution and prose that could attach to almost any subject. The page also warns that no single pattern proves AI authorship.

Generic draft (authored specificity score: 18%)

In today’s rapidly evolving digital landscape, autonomous content systems are playing a pivotal role in helping creators unlock meaningful opportunities. By combining innovation, automation and human insight, these systems highlight the importance of building engaging experiences that resonate with modern audiences.

Edited draft (authored specificity score: 91%)

I asked the system for ten episode ideas about autonomous agents. Eight opened with the same promise: productivity. None required a real agent to attempt a task. I deleted them and kept the failed calendar benchmark because it produced a trace I could inspect.

Editing rule: replace importance with consequence, category words with objects, and unsupported significance with an event the reader can inspect.

AI UI has an accent too

An empirical study of AI-generated interface prototypes found them usable but conventional: pragmatic ratings were stronger than hedonic originality and innovation. Practitioner critiques add the familiar symptoms: rounded-card repetition, muted gradients, polished default states and neglected edge cases. A forbidden colour or corner radius is not the problem. The problem is the absence of a visible decision.

The same mock product page, left to defaults and art-directed
Default generated UIArt-directed system
HeadlineBuild better stories with AIA world should be able to prove you wrong.
Supporting lineTransform your creative workflow with intelligent automation.One change. One open result. One permanent consequence.
Call to actionGet startedInspect the evidence
Error stateEverything looks ready.ERROR 04 · proof missing · publication blocked
What it showsEvery block uses the same visual grammar. The copy could describe hundreds of products. The error state has not been designed.The headline carries the argument, not a category label. Form follows the editorial hierarchy instead of a component kit. The broken state is designed as carefully as the happy state.

Distinctive design is evidence that somebody chose what matters, what can wait and what should happen when the happy path breaks.

Ask an agent to find out whether AI saved your time

Do not ask an agent to explain how AI saves time. The difference is one week, one ledger and permission for the answer to be disappointing. Count setup, review, repair and mistakes before claiming that automation saved time.

The easy request is “Make a video about how AI can save time on email.” The agent writes the conclusion first, finds supporting claims and produces a polished asset. No one knows whether your inbox improved. Content exists, and knowledge did not change.

The proof-bound request is “For seven workdays, measure whether an agent reduces the time I spend processing email. Count review, repairs, missed messages and setup. Publish the result even if it loses.” Measure a manual baseline. Predeclare what counts as intervention and failure. Run the week before writing the ending, and preserve the worst day, not just the average. The result becomes the story.

A week of full-cost accounting shrinks the headline to a few minutes a day

A seven-day inbox trial with illustrative numbers. The structure is the useful part: baseline, full cost, failure and a result that can survive disappointment. The manual baseline is 42 minutes a day; setup is 70 minutes, spread across the days run.

Minutes per day for agent time, review and repair
DayAgentReviewRepairTotalWhat happened
1814426Clean day. The agent triaged 63 messages and required two corrections.
2916631A newsletter was marked urgent. Six minutes disappeared into repair.
3812323Best day. The queue was simple and the review rules held.
410201545Failure day. A message from a real person was filed as routine. The automation lost three minutes and created risk. The week produces a failure worth keeping.
5915529The rule was repaired. Review remained the largest cost.
6711220Low-volume day. The agent was useful but the comparison is not like-for-like.
7814426Final day. The week ends with a modest saving, not the promised transformation.
Manual baseline42 min/day
Agent + review + repair28.6 min/day
First-week net benefit+3.4 min/day

Bnet = Tmanual − (Tagent + Treview + Trepair + Tsetup / n)

E1, full-cost automation benefit. Count every minute the automation creates, including the minutes spent checking and fixing it. Here: 42 − 28.6 − 70 / 7 ≈ +3.4 minutes a day.

The headline gets smaller. The evidence gets stronger.

Illustrative accounting. Run the week before writing the title: the system is not allowed to replace the missing run with a general explanation of email automation.

The publishable result might be: “I let an AI run my inbox for a week. It saved three minutes a day and nearly hid the message that mattered.” That title is a compressed record of something that happened, not generated cleverness.

The agent should be allowed to fail before it is allowed to create

A weak agent answers the prompt. A useful one can stop, point to the missing event and tell me exactly what has to be observed before production begins. No world, intervention, unknown, evidence, failure path and memory means no production.

Same request, different authority. Prompt: “Make a video proving whether AI saved me time on email.” A scripted comparison of two agent behaviours.
Output-first agentProof-bound agent
What it did00:01 Chose conclusion: AI saves time.
00:03 Wrote five benefits and a 40-second script.
00:08 Added inbox footage, captions and a confident title.
00:09 Your measured time saving: not requested.
00:01 Blocked: the claimed saving has not been measured.
00:02 Proposed seven-day inbox trial against a manual baseline.
00:03 Unknown: does saved sorting time survive review and repair?
00:04 Required: time ledger, intervention log, missed-message audit and negative-result rule.
StateOutput readyExperiment ready
Assets proposed111 plan
Evidence requested04 logs
World changedNoAfter run

The output-first agent appears faster because it quietly assumes the answer. The proof-bound agent spends its speed creating a result you can inspect. The stronger behaviour is the refusal to manufacture a conclusion before the event exists.

A hard gate, not another weighted content score

Each condition asks a different question. A high score on polish cannot compensate for a missing event or fabricated proof.

𝒫 = W ∧ I ∧ U ∧ E ∧ F ∧ M

P2, the agent acceptance condition. The gate is conjunctive: one missing term blocks the episode.

The six conditions, applied to three example requests
ConditionThe question it asksGeneric explainerSeven-day inbox trialAutonomous Office file
W · Persistent state existsWhat is already true before the episode?MissingDeclaredDeclared
I · One intervention is declaredWhat changes on purpose?MissingDeclaredDeclared
U · The result is genuinely unknownWhat answer could disappoint us?MissingDeclaredDeclared
E · Evidence is named in advanceWhat would let another person inspect it?MissingDeclaredDeclared
F · Failure is publishableCan the episode survive an unflattering result?MissingDeclaredDeclared
M · The world changes afterwardsWhat will the next episode inherit?MissingDeclaredDeclared
VerdictReject, 0 / 6Allow planning, 6 / 6Allow planning, 6 / 6
  • Generic explainer: the agent has a topic and a preferred conclusion. It has no event.
  • Seven-day inbox trial: the baseline, seven-day run, intervention log, error policy and publish-even-if-negative rule are declared.
  • Autonomous Office file: the fiction is labelled, the canon can change and the audience can test competing theories.

A complete content contract allows planning, not publication. Production still needs rights, safety and human release checks. The same six fields form the machine-readable contract in the section for coding agents.

A working gate turns the agent from content generator into experiment operator

That is a more valuable job. It moves the machine from paraphrasing the world to creating controlled situations that can teach the human something. Before, the agent receives a topic and returns an artefact to fill a calendar. After, it runs a world: it proposes an intervention, captures the result and updates state. The compounding asset is history, because the next episode begins with evidence, consequences and unanswered questions.

What happens when the gate gets stricter? In this authored scenario, twelve candidate ideas have different evidence, persistence, failure tolerance and production costs. Raise the threshold and watch the trade-off.

Episodes allowed6
Evidence-backed100%
Worlds updated5
Future episode seeds6

Fewer outputs survive, but most remaining episodes carry evidence or create a follow-up.

The twelve candidate ideas and their authored scores
IdeaScoreEvidenceUpdates a worldSeeds a follow-upCostAt this threshold
5 ways agents improve productivity18NoNoNo1Blocked
Three agents enter the Calendar Collision Relay92YesYesYes4Allowed
Why verification matters24NoNoNo1Blocked
Exhibit 014: the model followed the broken dashboard88YesYesYes3Allowed
The future of autonomous work12NoNoNo1Blocked
The Maintenance Monster audits a six-hour saving84YesYesYes5Allowed
A practical guide to AI agents31NoNoNo2Blocked
A village court starts manufacturing hallucination cases79YesYesYes5Allowed
What comes after software?27NoNoNo1Blocked
The drawer loses a sock when the light turns warm73YesYesYes4Allowed
Can an agent recover after a hidden human correction?67YesNoYes3Allowed
AI is changing creativity16NoNoNo1Blocked

Illustrative. The candidate values are authored to expose failure modes. They are not estimates of audience performance.

The work changes for everyone involved:

  • For the human, review shifts from blank-page invention to judgement. The agent brings back a proposed intervention, evidence plan and unresolved decision instead of six finished scripts.
  • For the agent, the acceptance test becomes part of the build. A failed contract is a useful output. It tells the system to gather reality before generating polish.
  • For distribution, every release can create its own follow-up. A changed scoreboard, larger monster or new contradiction gives the next episode a cause, not merely another theme.
  • For product work, content can double as an experiment. The same run can expose a software failure, a user misunderstanding or a research question worth fixing.

AI makes one honest result reusable

AI does make the post cheaper, but that is not the interesting part. Evidence is still costly. You have to run the week, build the fixture, capture the trace or tolerate the failed prototype. Once that object exists, agents can explain it at different depths without inventing a second event. Amortise one honest result across useful explanations; do not amortise invented truth.

Assume a real evidence-producing run costs A$240 in time and materials, and each additional explanation costs A$18 to produce and review.

Average cost per useful outputA$7870% below a one-off asset

C̄(k) = Ce / k + Cp, k ≥ 1

E2, evidence-cost amortisation. The fixed evidence cost is shared; the claim itself is not duplicated or exaggerated. With four outputs: 240 / 4 + 18 = A$78, against A$258 for a one-off asset.

Illustrative economics. These are not market estimates. They expose the fixed-cost logic.

Evidence leaves assets behind. A result can leave a dataset, footage, failure case, public method, audience question and the next unresolved test, so the following episode begins with more state than the previous one.

At+1 = (1 − δ) At + et, 0 ≤ δ ≤ 1

E3, evidence accumulation with decay, an authored state model. The decay term prevents pretending that old evidence stays relevant forever.

That stock is harder to imitate than a style and cheaper to explain again than to rediscover. It is more useful to an agent, because the agent can retrieve real state, and more useful to a human, because it changes decisions.

This is a time-limited bet. For a short period, the field may overinvest in production and underinvest in proof. If that is true, the edge is building a history of inspectable results before everyone else realises the interface has changed, rather than publishing first.

Once the contract exists, the unit changes from post to world

A one-off proof gate prevents an empty post. A persistent world does more: it remembers the result, carries the consequence forward and gives the agent a cheaper, more specific next question. Store the result so the next episode starts from history instead of a blank prompt.

Each episode starts from the state the last one left behind

Six stages of a world engine feed an episode, an event with receipts, and return to persistent state.

Six connected stages of a world engine feeding an episode and returning to persistent state Worldstate Changeintervention Openuncertainty Proofcapture Changeconsequence Keepmemory Episodean event with receipts
  1. Start with state. Records, debts, laws, beliefs, objects, rivalries and unanswered questions already exist before the episode begins. Persistent state makes the first episode a beginning instead of an isolated asset.
  2. Change one thing on purpose. Add a rule. Remove a dependency. Change the lighting. Introduce a new event. The intervention should be clear enough to argue about later, and it gives the episode a question precise enough to test.
  3. Leave room to be wrong. If the planned conclusion cannot lose, the episode is an illustrated opinion. Genuine uncertainty gives the evidence a job: the result must be able to contradict the creator’s preferred conclusion.
  4. Decide what would count as proof. A trace, a clock, a sensor, a ledger, a replay, a state diff or a human judgement must be able to contradict the script. The evidence contract prevents the narration from replacing the event.
  5. Make the consequence visible. The record moves. The monster grows. An exhibit changes label. A faction gains power. Something is different because the event occurred, and visible consequence makes abstract change legible in a glance.
  6. Do not reset. The next episode inherits the record, the debt, the law or the unresolved contradiction. History becomes part of the format: it compounds, and the next episode begins where the previous one actually ended.

A design model for planning episodes, not a measured production process.

The Maintenance Monster turns invisible upkeep into plot

An automation claims to save six hours each week. Every hidden intervention, patch and exception feeds a physical creature. The episode ends when the ledger and the monster agree on what the automation is worth.

Every silent repair is charged against the six hours the automation claims to save

A live model of one automation world. Each action changes the weekly repair time by the amount shown; net weekly value is the six-hour claim minus recorded repair time.

Claimed saving6.0 h
Repair time0.0 h
Exceptions0
Net weekly value+6.0 h

Maintenance debtThe mark is the six-hour claim

  1. The automation enters the world with a six-hour claim and no recorded maintenance.

Authored increments for illustration. The episode’s real ledger would come from time logs, exception counts and patch history.

Visible maintenance creates story pressure and honest accounting. Different worlds should fail differently.

Six worlds, and what each one can prove
01 · Agent Olympics
Work becomes sport when the rules, clocks and disqualifications are visible. Plot engine: competition and prediction. Intervention: introduce one new discipline or change one rule. Proof: fixtures, traces, intervention logs, costs and a reproducible scoreboard. Memory: future events inherit records, rivalries and rule changes. Episode premise: “Three agents enter the Calendar Collision Relay. The fastest result is disqualified because a hidden human correction changed the schedule.”
02 · The Failure Museum
A failure becomes useful when it stays available for inspection. Plot engine: evidence mystery and revision. Intervention: reconstruct one incident or challenge an accepted explanation. Proof: logs, screenshots, timestamps, conflicting accounts and a controlled replay. Memory: new incidents can connect to older exhibits and overturn the museum catalogue. Episode premise: “Exhibit 014 was labelled ‘model hallucination’. A replay shows the model followed a corrupted dashboard exactly.”
03 · Lost Files from the Autonomous Office
The audience reconstructs a world from evidence the world accidentally left behind. Plot engine: artifact-first investigation. Intervention: recover one file, room, message or system log. Proof: clearly labelled fictional records, maps, tapes and incident files. Memory: every file becomes part of a persistent timeline with unresolved contradictions. Episode premise: “A training tape warns employees never to empty the approval queue. The office has had no employees for fourteen years.”
04 · The Maintenance Monster
Every hidden repair makes the creature larger. Plot engine: physical metaphor backed by cost accounting. Intervention: add, repair, simplify or retire an automation dependency. Proof: time logs, exception counts, patch history and a maintenance ledger. Memory: old shortcuts return as future maintenance debt. Episode premise: “The workflow claims to save six hours. After three undocumented repairs, the monster is larger than the task it replaced.”
05 · Agent Civilisation Lab
One constitutional change can create consequences several episodes later. Plot engine: institutional simulation. Intervention: change one law, resource rule or communication constraint. Proof: seeded runs, state snapshots, rule history and replayable events. Memory: the world never resets; later episodes inherit earlier decisions. Episode premise: “The town creates a court for hallucinations. By the next episode, the court is manufacturing cases to justify its existence.”
06 · Drawer Ecology Lab
A domestic object becomes interesting when it can hold a wrong belief. Plot engine: physical sensing and classification failure. Intervention: change lighting, move objects or introduce an unknown item. Proof: frames, sensor readings, confidence histories and manual corrections. Memory: the drawer carries a history of belief and correction. Episode premise: “A black sock disappears under warm light, returns as a phone charger and becomes a new species after two sensors disagree.”

The equations cannot predict a hit, but they can stop me combining different failures into one fake score

Opening acceptance, retention, diffusion, audience memory and portfolio diversity are different processes. Each needs its own evidence and its own failure state. Do not compress attention, spread, return and creative diversity into one score.

M1: A hook is an expectation contract

The opening makes a promise under uncertainty. Legibility, curiosity and an evidence signal may help; promise/payoff mismatch should hurt. The form below is an authored model, not a measured platform law.

Pr(open) = σ(β0 + βLL + βCC + βEE − βMM), σ(z) = 1 / (1 + e−z)

M1, opening acceptance model. A compact expectation contract. Coefficients remain uncalibrated.

Authored hypothesis. The shape is deliberate, but its coefficients are not estimated from an audience.

M2 and M3: Attention is a survival curve

The opening earns the next moment. Every later beat either preserves that decision or adds an abandonment hazard. YouTube itself reports time-indexed retention rather than a universal “attention span”.

S(t) = exp(−∫t0 h(u) du), h(u) ≥ 0

M2, viewer survival. The curve can only stay level or fall.

𝔼[min(T, D)] = ∫D0 S(t) dt

M3, expected watch within duration. Two videos can have the same completion but different useful watch paths.

Move the authored hazards and watch the curve

100%0%60 sec
Opening accepted84%
Completion52%
Expected watch40.6 s

Research-aligned form. The survival structure is established mathematics. The creative hazard inputs are authored and shadow-only; at their defaults the illustrative 60-second curve opens at 84%, completes at 52% and gives an expected watch of 40.6 seconds.

M4: A large audience can arrive through different shapes

One platform broadcast and a deep person-to-person cascade may create similar view counts. Structural virality measures topology rather than size.

N = Next + ΣGg=1 Ng
𝔼[N | R < 1] = μ / (1 − R)

M4, exposure and endogenous diffusion. Popularity and propagation shape are different measurements.

Three illustrative diffusion patterns
PatternEstimated reachEndogenous shareMean depthQualified return
Broadcast-heavy540k8%1.73.2%
Cascade-heavy506k73%7.95.8%
Smaller, durable188k61%4.812.6%

Illustrative. The network figures are explanatory. They are not fitted to a platform or account.

M5: Distribution becomes valuable when exposure becomes memory

A person can sample once, recognise the series, return, become regular or disappear. Those transitions should not be compressed into follower count. YouTube similarly distinguishes new, casual and regular viewers.

πt+1 = πtP, Pij ≥ 0, Σj Pij = 1

M5, audience-state transition, an illustrative state model. Every row must conserve probability mass.

Episode 0, starting from 1,000 sampled people

  • Sampled1,000
  • Recognised0
  • Returning0
  • Regular0
  • Dormant0

The first release creates exposure. It has not created an audience yet.

The transition matrix behind each release
Share moving from each state (row) to each state (column) per episode
FromSampledRecognisedReturningRegularDormant
Sampled0.400.350.100.000.15
Recognised0.050.350.350.050.20
Returning0.020.050.480.250.20
Regular0.010.020.120.680.17
Dormant0.050.080.070.030.77

State model. Useful for accounting and mock tests. Not validated as a literal model of real people.

M6 and M7: One winner can make the whole system worse

AI-assisted creators may produce stronger individual outputs while the collection becomes more similar. A quality-diversity archive keeps strong candidates across different behavioural niches instead of allowing one early winner to consume every production slot.

maxx ∈ {0,1}n Σni=1 qixi

M6, portfolio objective, an illustrative optimisation. This is only the objective.

Σi cixi ≤ B,
Σi dijxi ≥ mj ∀j,
Σi∈𝒲w xi ≤ ρ Σi xi ∀w

M7, portfolio constraints. Production cost stays within budget, every required creative dimension receives minimum coverage, and no content world exceeds the maximum portfolio share. The constraints stop a short-term winner erasing the search space.

Two illustrative allocation policies for 30 production slots across six worlds
PolicySlots per worldWorlds preservedLargest shareNovel mechanism slots
Greedy score17, 0, 0, 13, 0, 02 / 657%1
Diverse portfolio6, 5, 4, 5, 4, 66 / 620%8

Research analogy. MAP-Elites is not a validated content strategy. The transfer is the preservation of multiple strong niches.

Notation and LaTeX for all twelve equations
P1  \mathcal{P}=W\land I\land U\land E\land F\land M
E1  \begin{aligned}
B_{\mathrm{net}}
&=T_{\mathrm{manual}}\\
&\quad-\left(T_{\mathrm{agent}}+T_{\mathrm{review}}+T_{\mathrm{repair}}+\frac{T_{\mathrm{setup}}}{n}\right)
\end{aligned}
P2  \mathcal{P}=W\land I\land U\land E\land F\land M
E2  \bar C(k)=\frac{C_e}{k}+C_p,\qquad k\ge 1
E3  A_{t+1}=(1-\delta)A_t+e_t,\qquad 0\le\delta\le 1
M1  \Pr(\mathrm{open})=\sigma\!\left(\beta_0+\beta_L L+\beta_C C+\beta_E E-\beta_M M\right),\quad \sigma(z)=\frac{1}{1+e^{-z}}
M2  S(t)=\exp\!\left(-\int_0^t h(u)\,\mathrm du\right),\qquad h(u)\ge 0
M3  \mathbb E\!\left[\min(T,D)\right]=\int_0^D S(t)\,\mathrm dt
M4  \begin{aligned}
N&=N_{\mathrm{ext}}+\sum_{g=1}^{G}N_g,\\[2pt]
\mathbb E[N\mid R<1]&=\frac{\mu}{1-R}
\end{aligned}
M5  \boldsymbol{\pi}_{t+1}=\boldsymbol{\pi}_t P,\qquad P_{ij}\ge 0,\quad \sum_j P_{ij}=1
M6  \max_{\mathbf x\in\{0,1\}^n}\;\sum_{i=1}^{n}q_i x_i
M7  \begin{aligned}
\sum_i c_i x_i&\le B,\\
\sum_i d_{ij}x_i&\ge m_j &&\forall j,\\
\sum_{i\in\mathcal W_w}x_i&\le \rho\sum_i x_i &&\forall w
\end{aligned}

Automate the labour, and keep public authority bounded

An agent can research, propose experiments, write, render, test and package a release. That does not entitle it to decide that a claim is true, a person is safe to depict or a piece is worth publishing.

Authority levels for a content agent. The selected boundary is A3.
LevelWhat the system may do
A0 · ObserveThe system can inspect public or approved material and report findings. It cannot create or change artifacts.
A1 · DraftThe system can propose ideas, scripts and plans. Every artifact remains a suggestion awaiting human review.
A2 · BuildThe system can produce local assets, simulations and renders. Nothing leaves the controlled environment.
A3 · Prepare (selected)The system can complete the production package and evidence checks. Calvin reviews the finished piece and explicitly authorises public release.
A4 · Bounded publishOnly calibrated, low-risk, allowlisted formats may publish through official APIs, with kill switches, incident logging and strict scope.
A5 · UnboundedProhibited. The system does not receive unrestricted authority to publish claims, depict people or alter public channels.

Test which world makes relevant people return

Do not search for the viral hook. The first pilot should compare different plot engines. Cosmetic hook variants can wait until there is evidence that the underlying world deserves another episode. The real test is whether relevant people recognise the series and return.

Candidate worlds for the pilot. Choose four and run four episodes of each, sixteen runs in an interleaved order.
WorldMechanismProduction burden
Agent OlympicsCompetition3 / 5
Failure MuseumEvidence mystery2 / 5
Maintenance MonsterPhysical metaphor4 / 5
Lost FilesFictional investigation5 / 5
Civilisation LabPersistent simulation5 / 5
Drawer EcologyPhysical observation4 / 5

The primary outcome is qualified subsequent-episode consumption: did a relevant viewer voluntarily consume another episode from the same world within the declared window? Four guardrails sit beside it: evidence completeness, trust incidents, production burden and portfolio concentration.

One hit does not crown a format. The system should earn confidence through repeated episodes, comparable cohorts and a reasoned account of what changed.

What I am betting on

If the gate works, AI content stops being a faster way to state an opinion and becomes a cheaper way to operate and explain reality.

I still want the automation. I want the agent to search, set up the run, record the evidence, cut the video, visualise the maths and remember what changed.

But the machine should spend its new speed on the expensive part: asking questions that reality can reject. The human advantage is moving judgement earlier. The agent advantage is receiving a world with state instead of another empty prompt.

If it fails, I have a stricter planning system that produces fewer empty assets. If it succeeds, an agent can build a body of work whose originality comes from accumulated reality, not a more elaborate prompt.

Every system I found could make more content. The production layer is already arriving, and the useful question is whether we can force it to build knowledge rather than merely simulate confidence.

For coding agents: give your agent the contract

Give the agent an acceptance test, not a vague request to “make better content”. A human should be able to skim the rule in under a minute. An agent should be able to validate it before spending tokens, render time or public trust.

The operating loop has five steps:

  1. Inspect the actual work. Find the inbox, failed build, customer question, prototype, decision or measurement that already contains tension.
  2. Name one intervention. Change a rule, process, object or input. Do not ask for “content about” a broad topic.
  3. Lock the evidence contract. Declare the metric, trace, footage, ledger or human judgement before seeing the result.
  4. Run before writing the title. Keep the boring, negative or embarrassing outcome. That is where trust enters.
  5. Store the state change. Give the next agent the result, unresolved problem and authority boundary instead of resetting to a blank prompt.

At the A3 boundary, the agent can prepare the complete release. The human still authorises publication.

The contract an agent can actually use
content_proof_contract:
  world_state_before: ""
  intervention: ""
  unknown_outcome: ""
  evidence_required: []
  failure_story: ""
  world_state_after: ""
reject_if:
  - the outcome is already written
  - the proof can be replaced by narration
  - an unflattering result would be discarded
  - the next episode resets the world
Download the agent brief

The article explains the judgement. The brief gives the agent the rule.

A copyable brief: do not make content until the event is defined

Use this before research, scripting, rendering or publishing. A failed gate is a valid result.

When asked to create content:

1. Inspect the real work for a stateful world.
2. Propose one controlled intervention.
3. Name the outcome that remains genuinely unknown.
4. Declare the evidence required before any claim can ship.
5. Preserve an unflattering or uneventful result.
6. State what changes permanently after the run.

Reject the idea when:
- the conclusion is already written;
- narration could replace proof;
- failure would be hidden or rerun until flattering;
- the next episode would reset the world;
- public release exceeds the authorised boundary.

Return the completed content_proof_contract before producing assets.

Sources and further discussion

What is sourced, what is authored and what the pilot still has to prove. Research-backed means a source directly supports the narrow claim in the studied or documented context. An authored hypothesis is a useful model shape or design judgement that still needs prospective testing. Illustrative means a visual explanation using arbitrary units, not an estimated platform model. My position marks an explicit argument, not a disguised empirical finding.

  1. Wikipedia: Signs of AI writing, Wikipedia WikiProject AI Cleanup. Descriptive field guide, not a detector or policy. Used to audit generic prose patterns.
  2. Progressive Disclosure, Jakob Nielsen. Supports keeping the primary reading path simple and making technical depth available on request.
  3. Understanding Animation from Interactions, W3C Web Accessibility Initiative. Motion is optional, user-controlled and removed under reduced-motion preferences.
  4. Narrative Visualization: Telling Stories with Data, Edward Segel and Jeffrey Heer. Informs the balance between a guided narrative and reader-controlled exploration.
  5. Generative AI enhances individual creativity but reduces the collective diversity of novel content, Anil R. Doshi and Oliver P. Hauser. Supports treating output quality and portfolio diversity as separate concerns.
  6. The Structural Virality of Online Diffusion, Sharad Goel, Ashton Anderson, Jake Hofman and Duncan Watts. Supports separating diffusion size from diffusion topology.
  7. Measure key moments for audience retention, YouTube Help. Supports time-indexed retention and matched comparisons, not a universal attention-span claim.
  8. New, casual and regular viewers, YouTube Help. Supports distinguishing one-time exposure from recurring audience behaviour.
  9. Illuminating search spaces by mapping elites, Jean-Baptiste Mouret and Jeff Clune. Used as an analogy for preserving strong candidates across different creative niches.
  10. Telltale signs of AI-coded products, Paul Bakaus and other product-design practitioners, reported by Business Insider. Used only as a current practitioner critique of visual sameness and neglected edge states.
  11. Information Scent: How Users Decide Where to Go Next, Raluca Budiu, Nielsen Norman Group. Supports making the value of the article and each interaction legible before the reader commits attention.
  12. Website Reading: It (Sometimes) Does Happen, Jakob Nielsen, Nielsen Norman Group. Supports a scan-first structure that lets interested readers move from headings and figures into deeper prose.
  13. Usable but Conventional: An Empirical Study on the UX of AI-Generated Interface Prototypes, Karoline Romero, Igor Wiese, Renato Balancieiri, Gislaine Camila Leal and Guilherme Guerino. Reports positive pragmatic evaluations alongside weaker hedonic originality and innovation for AI-generated prototypes. (The source edition listed this paper twice, as sources 13 and 20.)
  14. video-use, Browser Use. Official project documentation for agent-led raw-footage editing, transcription, rendering and self-evaluation.
  15. Video tools for the agent era, Remotion. Official project documentation for agentic, interactive and programmatic video creation.
  16. Open-source social scheduling and analytics, Postiz. Official project documentation for scheduling, APIs and social analytics.
  17. YouTube’s top AI slop channels are disappearing, The Verge. Current reporting used to establish that platforms are already reacting to low-quality synthetic volume.
  18. Tampa Bay businesses are getting called out for AI slop, Axios Tampa Bay. Current reporting used as evidence of audience and community backlash against repetitive AI promotional output.
  19. A coffee shop owner used AI to make a menu poster. Then came the angry DMs., Business Insider. A recent small-business example showing that production savings can create a trust cost.
  20. ScrollyVis: Interactive visual authoring of guided dynamic narratives for scientific scrollytelling, Eric Mörth, Stefan Bruckner and Noeska N. Smit. Supports scroll-controlled narrative as a way to connect explanation with changing visual state while retaining reader control.

The interactive models are designed to make assumptions visible. They do not predict future views, virality, revenue or audience return. The honest next step is a manually approved pilot.

Published 2026-09-14 · Updated 2026-09-28 · Source edition v4.2