AI can make a thousand posts. Make it prove one thing happened.
Editing, animation and publishing are becoming agent tools. The scarce advantage is an event, a measurement or a failure that only your system can supply.
The production stack can research, script, edit, render and publish. It can finish an asset while nobody has supplied a reason to care. So ask the stack one question first: what happened that only you can prove? If the answer is nothing, the stack should stop before it renders.
The rule I propose is a proof gate. Your agent names the event and the evidence before production, or it stops. Below are one ordinary worked example and a machine-readable contract. The advantage is a lower cost per original fact.
For you, that means killing polished nothing before you review a finished asset: approve the question, evidence and failure condition, and reject the empty idea while it is still cheap. For your agent, it is a hard stop before scripting, editing or publishing. The machine must name what will happen, what remains unknown and what evidence could prove its preferred answer wrong. The economic edge is that one real result can support many useful explanations. The experiment is the fixed cost, and agentic production makes each honest derivative cheaper without inventing another conclusion.
𝒫 = W ∧ I ∧ U ∧ E ∧ F ∧ M
P1, the content-proof gate and acceptance rule. All six conditions must exist: world state, intervention, uncertainty, evidence, failure permission and memory. Polish cannot compensate for a missing event.
My position: the best content agent knows when no publishable event exists. Finishing fastest is the wrong test.
The price of polish is falling, and the price of being believed is rising
That combination changes the strategy. When everyone can produce a competent-looking video, competence stops proving that anything happened. The scarce object becomes the trace, test, failure, measurement or changed world underneath it. Editing and rendering are becoming callable. Evidence is still a separate job.
The logic runs in four steps:
- Agents can now make the artefact.
- Audiences are learning to discount generic polish.
- Evidence remains expensive, specific and difficult to copy.
- So make the agent operate the evidence, rather than merely describe it.
agentic production ↑ + generic-output discount ↑ → proof advantage
Production is becoming infrastructure, and trust is becoming a filter. Polish is becoming abundant while generic AI output is being discounted. My inference, as of September 2026: visible evidence and accumulated history become more valuable as production itself stops being scarce. It is a reason to test the strategy before proof protocols become ordinary, not proof that evidence-backed content will win.
The three images below are the exact SVG files committed to the cited repositories, not mock screenshots. The article only places them inside its own editorial frame.
browser-use/video-use · static/timeline-view.svg Git blob bb3f11476546, unredrawn. The editing agent does more than promise automation: its repository exposes the timeline evidence it reasons over before cutting.remotion-dev/brand · logo.svg Git blob 6c6b4ed7956c, unredrawn. Remotion’s own repository now frames video as an agentic, interactive and programmatic medium rather than a specialised finishing step.45ck/video-evaluator · assets/benchmark-snapshot.svg Git blob 7ca0bb3c9d8c, unredrawn. My own evaluator already showed the more uncomfortable distinction: a pipeline can run perfectly while still failing to understand what mattered.The response has started outside the repositories too:
- Platforms. YouTube channels built on low-quality AI output are already being removed or disappearing (The Verge, 28 January 2026).
- Audiences. Local businesses are being publicly called out for repetitive AI-made promotional work (Axios Tampa Bay, 17 September 2026).
- Trust. A small business lost followers and received angry messages over one AI-made menu poster (Business Insider, 14 September 2026).
The old advantage was making polished media faster than the next person, and tools are compressing it. The advantage I propose is owning a result the next person cannot prompt into existence: they would have to reproduce the experiment, history and evidence. That window lasts until proof becomes the normal interface, so the useful move is to build the evidence habit now instead of publishing more generic volume.
I built three systems to make and inspect content, then the production layer stopped being the interesting part
Content Machine could assemble shorts. Demo Machine could turn executable product flows into videos. Video Evaluator could inspect the result. Meanwhile, public tools began solving the same production work faster. That was useful. It also made the missing layer impossible to ignore.
The stack could research, script, edit, render and review. It could manufacture a convincing object before reality had supplied a convincing event. When production was expensive, the finished artefact itself looked like evidence of effort. Agentic production breaks that shortcut. A finished-looking video may now mean very little.
The mistake was giving the machine permission to finish before it had something specific to lose. The failure I observed was that the state of the world changed faster than my architecture: editing and rendering became replaceable infrastructure, and the unanswered question moved upstream. What must be true before the agent is allowed to produce anything? Pipeline completion is not proof that anything new was observed.
Give it a topic and the calendar fills. The outputs are plausible, and none require a new event, a falsifiable claim or a changed world.
Six polished, interchangeable outputs the machine can make without an event
| Output | What gives it away |
|---|---|
| 5 ways AI is transforming modern work | “AI is reshaping how teams collaborate, innovate and unlock efficiency across an evolving digital landscape.” Importance inflation, generic attribution, no event. |
| The future of autonomous agents | “From productivity to creativity, intelligent agents are poised to redefine the way we live and work.” Future fog, broad trend, no proof. |
| Why verification matters | “Robust verification plays a pivotal role in ensuring trustworthy, reliable and responsible AI systems.” Pivotal role, rule of three, no concrete failure. |
| Building smarter systems | “By combining automation with human insight, organisations can foster innovation and drive meaningful outcomes.” Superficial synthesis, business fog, no stakes. |
| What comes after software? | “The convergence of AI and automation marks a significant shift in the broader technology landscape.” Broader landscape, significance claim, interchangeable. |
| A practical guide to AI workflows | “This guide explores key strategies, common challenges and actionable steps for navigating AI-powered work.” Outline voice, canned triad, no lived detail. |
The stack still has no owner for “why care?”
The new stack has a tool for almost every production step. It can research (collect material), script (shape the claim), edit (assemble footage), render (make the artefact), publish (reach the platform) and measure (observe response). The reason to care, what changed in the world, has no owner. None of those capabilities proves that the source event was original, uncertain or real.
This is not a complaint about the tools. video-use can turn raw footage into an edited video. Remotion can make programmatic video the source of truth. Postiz can schedule and measure distribution. My own Content Machine tried to connect even more of that pipeline.
The tooling stack can improve every step after an idea exists. It does not establish that the idea deserves to become an asset. A system can be excellent at production while remaining indifferent to whether the source event is generic, invented or already known. Making the artefact and justifying the artefact are different jobs, and I had assigned the first job to the agent and quietly kept the second in my own head. “Why care?” must become a system responsibility, not a line added after rendering.
Finished-looking output is the trap
Generic prose and generic interfaces are symptoms of the same shortcut: the system completes the surface before anybody has chosen what the surface is supposed to prove. Finished-looking language and interface polish can conceal a missing event, and the article should survive its own test.
The problem goes deeper than the word “delve”
Wikipedia’s field guide to signs of AI writing is useful because it notices deeper patterns: generic importance claims, superficial analysis, promotional tone, vague attribution and prose that could attach to almost any subject. The page also warns that no single pattern proves AI authorship.
In today’s rapidly evolving digital landscape, autonomous content systems are playing a pivotal role in helping creators unlock meaningful opportunities. By combining innovation, automation and human insight, these systems highlight the importance of building engaging experiences that resonate with modern audiences.
I asked the system for ten episode ideas about autonomous agents. Eight opened with the same promise: productivity. None required a real agent to attempt a task. I deleted them and kept the failed calendar benchmark because it produced a trace I could inspect.
Editing rule: replace importance with consequence, category words with objects, and unsupported significance with an event the reader can inspect.
AI UI has an accent too
An empirical study of AI-generated interface prototypes found them usable but conventional: pragmatic ratings were stronger than hedonic originality and innovation. Practitioner critiques add the familiar symptoms: rounded-card repetition, muted gradients, polished default states and neglected edge cases. A forbidden colour or corner radius is not the problem. The problem is the absence of a visible decision.
| Default generated UI | Art-directed system | |
|---|---|---|
| Headline | Build better stories with AI | A world should be able to prove you wrong. |
| Supporting line | Transform your creative workflow with intelligent automation. | One change. One open result. One permanent consequence. |
| Call to action | Get started | Inspect the evidence |
| Error state | Everything looks ready. | ERROR 04 · proof missing · publication blocked |
| What it shows | Every block uses the same visual grammar. The copy could describe hundreds of products. The error state has not been designed. | The headline carries the argument, not a category label. Form follows the editorial hierarchy instead of a component kit. The broken state is designed as carefully as the happy state. |
Distinctive design is evidence that somebody chose what matters, what can wait and what should happen when the happy path breaks.
Ask an agent to find out whether AI saved your time
Do not ask an agent to explain how AI saves time. The difference is one week, one ledger and permission for the answer to be disappointing. Count setup, review, repair and mistakes before claiming that automation saved time.
The easy request is “Make a video about how AI can save time on email.” The agent writes the conclusion first, finds supporting claims and produces a polished asset. No one knows whether your inbox improved. Content exists, and knowledge did not change.
The proof-bound request is “For seven workdays, measure whether an agent reduces the time I spend processing email. Count review, repairs, missed messages and setup. Publish the result even if it loses.” Measure a manual baseline. Predeclare what counts as intervention and failure. Run the week before writing the ending, and preserve the worst day, not just the average. The result becomes the story.
A seven-day inbox trial with illustrative numbers. The structure is the useful part: baseline, full cost, failure and a result that can survive disappointment. The manual baseline is 42 minutes a day; setup is 70 minutes, spread across the days run.
| Day | Agent | Review | Repair | Total | What happened |
|---|---|---|---|---|---|
| 1 | 8 | 14 | 4 | 26 | Clean day. The agent triaged 63 messages and required two corrections. |
| 2 | 9 | 16 | 6 | 31 | A newsletter was marked urgent. Six minutes disappeared into repair. |
| 3 | 8 | 12 | 3 | 23 | Best day. The queue was simple and the review rules held. |
| 4 | 10 | 20 | 15 | 45 | Failure day. A message from a real person was filed as routine. The automation lost three minutes and created risk. The week produces a failure worth keeping. |
| 5 | 9 | 15 | 5 | 29 | The rule was repaired. Review remained the largest cost. |
| 6 | 7 | 11 | 2 | 20 | Low-volume day. The agent was useful but the comparison is not like-for-like. |
| 7 | 8 | 14 | 4 | 26 | Final day. The week ends with a modest saving, not the promised transformation. |
Bnet = Tmanual − (Tagent + Treview + Trepair + Tsetup / n)
E1, full-cost automation benefit. Count every minute the automation creates, including the minutes spent checking and fixing it. Here: 42 − 28.6 − 70 / 7 ≈ +3.4 minutes a day.
The headline gets smaller. The evidence gets stronger.
Illustrative accounting. Run the week before writing the title: the system is not allowed to replace the missing run with a general explanation of email automation.
The publishable result might be: “I let an AI run my inbox for a week. It saved three minutes a day and nearly hid the message that mattered.” That title is a compressed record of something that happened, not generated cleverness.
The agent should be allowed to fail before it is allowed to create
A weak agent answers the prompt. A useful one can stop, point to the missing event and tell me exactly what has to be observed before production begins. No world, intervention, unknown, evidence, failure path and memory means no production.
| Output-first agent | Proof-bound agent | |
|---|---|---|
| What it did | 00:01 Chose conclusion: AI saves time.00:03 Wrote five benefits and a 40-second script.00:08 Added inbox footage, captions and a confident title.00:09 Your measured time saving: not requested. | 00:01 Blocked: the claimed saving has not been measured.00:02 Proposed seven-day inbox trial against a manual baseline.00:03 Unknown: does saved sorting time survive review and repair?00:04 Required: time ledger, intervention log, missed-message audit and negative-result rule. |
| State | Output ready | Experiment ready |
| Assets proposed | 11 | 1 plan |
| Evidence requested | 0 | 4 logs |
| World changed | No | After run |
The output-first agent appears faster because it quietly assumes the answer. The proof-bound agent spends its speed creating a result you can inspect. The stronger behaviour is the refusal to manufacture a conclusion before the event exists.
A hard gate, not another weighted content score
Each condition asks a different question. A high score on polish cannot compensate for a missing event or fabricated proof.
𝒫 = W ∧ I ∧ U ∧ E ∧ F ∧ M
P2, the agent acceptance condition. The gate is conjunctive: one missing term blocks the episode.
| Condition | The question it asks | Generic explainer | Seven-day inbox trial | Autonomous Office file |
|---|---|---|---|---|
| W · Persistent state exists | What is already true before the episode? | Missing | Declared | Declared |
| I · One intervention is declared | What changes on purpose? | Missing | Declared | Declared |
| U · The result is genuinely unknown | What answer could disappoint us? | Missing | Declared | Declared |
| E · Evidence is named in advance | What would let another person inspect it? | Missing | Declared | Declared |
| F · Failure is publishable | Can the episode survive an unflattering result? | Missing | Declared | Declared |
| M · The world changes afterwards | What will the next episode inherit? | Missing | Declared | Declared |
| Verdict | Reject, 0 / 6 | Allow planning, 6 / 6 | Allow planning, 6 / 6 |
- Generic explainer: the agent has a topic and a preferred conclusion. It has no event.
- Seven-day inbox trial: the baseline, seven-day run, intervention log, error policy and publish-even-if-negative rule are declared.
- Autonomous Office file: the fiction is labelled, the canon can change and the audience can test competing theories.
A complete content contract allows planning, not publication. Production still needs rights, safety and human release checks. The same six fields form the machine-readable contract in the section for coding agents.
A working gate turns the agent from content generator into experiment operator
That is a more valuable job. It moves the machine from paraphrasing the world to creating controlled situations that can teach the human something. Before, the agent receives a topic and returns an artefact to fill a calendar. After, it runs a world: it proposes an intervention, captures the result and updates state. The compounding asset is history, because the next episode begins with evidence, consequences and unanswered questions.
What happens when the gate gets stricter? In this authored scenario, twelve candidate ideas have different evidence, persistence, failure tolerance and production costs. Raise the threshold and watch the trade-off.
Fewer outputs survive, but most remaining episodes carry evidence or create a follow-up.
The twelve candidate ideas and their authored scores
| Idea | Score | Evidence | Updates a world | Seeds a follow-up | Cost | At this threshold |
|---|---|---|---|---|---|---|
| 5 ways agents improve productivity | 18 | No | No | No | 1 | Blocked |
| Three agents enter the Calendar Collision Relay | 92 | Yes | Yes | Yes | 4 | Allowed |
| Why verification matters | 24 | No | No | No | 1 | Blocked |
| Exhibit 014: the model followed the broken dashboard | 88 | Yes | Yes | Yes | 3 | Allowed |
| The future of autonomous work | 12 | No | No | No | 1 | Blocked |
| The Maintenance Monster audits a six-hour saving | 84 | Yes | Yes | Yes | 5 | Allowed |
| A practical guide to AI agents | 31 | No | No | No | 2 | Blocked |
| A village court starts manufacturing hallucination cases | 79 | Yes | Yes | Yes | 5 | Allowed |
| What comes after software? | 27 | No | No | No | 1 | Blocked |
| The drawer loses a sock when the light turns warm | 73 | Yes | Yes | Yes | 4 | Allowed |
| Can an agent recover after a hidden human correction? | 67 | Yes | No | Yes | 3 | Allowed |
| AI is changing creativity | 16 | No | No | No | 1 | Blocked |
Illustrative. The candidate values are authored to expose failure modes. They are not estimates of audience performance.
The work changes for everyone involved:
- For the human, review shifts from blank-page invention to judgement. The agent brings back a proposed intervention, evidence plan and unresolved decision instead of six finished scripts.
- For the agent, the acceptance test becomes part of the build. A failed contract is a useful output. It tells the system to gather reality before generating polish.
- For distribution, every release can create its own follow-up. A changed scoreboard, larger monster or new contradiction gives the next episode a cause, not merely another theme.
- For product work, content can double as an experiment. The same run can expose a software failure, a user misunderstanding or a research question worth fixing.
AI makes one honest result reusable
AI does make the post cheaper, but that is not the interesting part. Evidence is still costly. You have to run the week, build the fixture, capture the trace or tolerate the failed prototype. Once that object exists, agents can explain it at different depths without inventing a second event. Amortise one honest result across useful explanations; do not amortise invented truth.
Assume a real evidence-producing run costs A$240 in time and materials, and each additional explanation costs A$18 to produce and review.
C̄(k) = Ce / k + Cp, k ≥ 1
E2, evidence-cost amortisation. The fixed evidence cost is shared; the claim itself is not duplicated or exaggerated. With four outputs: 240 / 4 + 18 = A$78, against A$258 for a one-off asset.
Illustrative economics. These are not market estimates. They expose the fixed-cost logic.
Evidence leaves assets behind. A result can leave a dataset, footage, failure case, public method, audience question and the next unresolved test, so the following episode begins with more state than the previous one.
At+1 = (1 − δ) At + et, 0 ≤ δ ≤ 1
E3, evidence accumulation with decay, an authored state model. The decay term prevents pretending that old evidence stays relevant forever.
That stock is harder to imitate than a style and cheaper to explain again than to rediscover. It is more useful to an agent, because the agent can retrieve real state, and more useful to a human, because it changes decisions.
This is a time-limited bet. For a short period, the field may overinvest in production and underinvest in proof. If that is true, the edge is building a history of inspectable results before everyone else realises the interface has changed, rather than publishing first.
Once the contract exists, the unit changes from post to world
A one-off proof gate prevents an empty post. A persistent world does more: it remembers the result, carries the consequence forward and gives the agent a cheaper, more specific next question. Store the result so the next episode starts from history instead of a blank prompt.
Six stages of a world engine feed an episode, an event with receipts, and return to persistent state.
- Start with state. Records, debts, laws, beliefs, objects, rivalries and unanswered questions already exist before the episode begins. Persistent state makes the first episode a beginning instead of an isolated asset.
- Change one thing on purpose. Add a rule. Remove a dependency. Change the lighting. Introduce a new event. The intervention should be clear enough to argue about later, and it gives the episode a question precise enough to test.
- Leave room to be wrong. If the planned conclusion cannot lose, the episode is an illustrated opinion. Genuine uncertainty gives the evidence a job: the result must be able to contradict the creator’s preferred conclusion.
- Decide what would count as proof. A trace, a clock, a sensor, a ledger, a replay, a state diff or a human judgement must be able to contradict the script. The evidence contract prevents the narration from replacing the event.
- Make the consequence visible. The record moves. The monster grows. An exhibit changes label. A faction gains power. Something is different because the event occurred, and visible consequence makes abstract change legible in a glance.
- Do not reset. The next episode inherits the record, the debt, the law or the unresolved contradiction. History becomes part of the format: it compounds, and the next episode begins where the previous one actually ended.
A design model for planning episodes, not a measured production process.
The Maintenance Monster turns invisible upkeep into plot
An automation claims to save six hours each week. Every hidden intervention, patch and exception feeds a physical creature. The episode ends when the ledger and the monster agree on what the automation is worth.
A live model of one automation world. Each action changes the weekly repair time by the amount shown; net weekly value is the six-hour claim minus recorded repair time.
- The automation enters the world with a six-hour claim and no recorded maintenance.
Authored increments for illustration. The episode’s real ledger would come from time logs, exception counts and patch history.
Visible maintenance creates story pressure and honest accounting. Different worlds should fail differently.
Six worlds, and what each one can prove
- 01 · Agent Olympics
- Work becomes sport when the rules, clocks and disqualifications are visible. Plot engine: competition and prediction. Intervention: introduce one new discipline or change one rule. Proof: fixtures, traces, intervention logs, costs and a reproducible scoreboard. Memory: future events inherit records, rivalries and rule changes. Episode premise: “Three agents enter the Calendar Collision Relay. The fastest result is disqualified because a hidden human correction changed the schedule.”
- 02 · The Failure Museum
- A failure becomes useful when it stays available for inspection. Plot engine: evidence mystery and revision. Intervention: reconstruct one incident or challenge an accepted explanation. Proof: logs, screenshots, timestamps, conflicting accounts and a controlled replay. Memory: new incidents can connect to older exhibits and overturn the museum catalogue. Episode premise: “Exhibit 014 was labelled ‘model hallucination’. A replay shows the model followed a corrupted dashboard exactly.”
- 03 · Lost Files from the Autonomous Office
- The audience reconstructs a world from evidence the world accidentally left behind. Plot engine: artifact-first investigation. Intervention: recover one file, room, message or system log. Proof: clearly labelled fictional records, maps, tapes and incident files. Memory: every file becomes part of a persistent timeline with unresolved contradictions. Episode premise: “A training tape warns employees never to empty the approval queue. The office has had no employees for fourteen years.”
- 04 · The Maintenance Monster
- Every hidden repair makes the creature larger. Plot engine: physical metaphor backed by cost accounting. Intervention: add, repair, simplify or retire an automation dependency. Proof: time logs, exception counts, patch history and a maintenance ledger. Memory: old shortcuts return as future maintenance debt. Episode premise: “The workflow claims to save six hours. After three undocumented repairs, the monster is larger than the task it replaced.”
- 05 · Agent Civilisation Lab
- One constitutional change can create consequences several episodes later. Plot engine: institutional simulation. Intervention: change one law, resource rule or communication constraint. Proof: seeded runs, state snapshots, rule history and replayable events. Memory: the world never resets; later episodes inherit earlier decisions. Episode premise: “The town creates a court for hallucinations. By the next episode, the court is manufacturing cases to justify its existence.”
- 06 · Drawer Ecology Lab
- A domestic object becomes interesting when it can hold a wrong belief. Plot engine: physical sensing and classification failure. Intervention: change lighting, move objects or introduce an unknown item. Proof: frames, sensor readings, confidence histories and manual corrections. Memory: the drawer carries a history of belief and correction. Episode premise: “A black sock disappears under warm light, returns as a phone charger and becomes a new species after two sensors disagree.”
The equations cannot predict a hit, but they can stop me combining different failures into one fake score
Opening acceptance, retention, diffusion, audience memory and portfolio diversity are different processes. Each needs its own evidence and its own failure state. Do not compress attention, spread, return and creative diversity into one score.
M1: A hook is an expectation contract
The opening makes a promise under uncertainty. Legibility, curiosity and an evidence signal may help; promise/payoff mismatch should hurt. The form below is an authored model, not a measured platform law.
Pr(open) = σ(β0 + βLL + βCC + βEE − βMM), σ(z) = 1 / (1 + e−z)
M1, opening acceptance model. A compact expectation contract. Coefficients remain uncalibrated.
Authored hypothesis. The shape is deliberate, but its coefficients are not estimated from an audience.
M2 and M3: Attention is a survival curve
The opening earns the next moment. Every later beat either preserves that decision or adds an abandonment hazard. YouTube itself reports time-indexed retention rather than a universal “attention span”.
S(t) = exp(−∫t0 h(u) du), h(u) ≥ 0
M2, viewer survival. The curve can only stay level or fall.
𝔼[min(T, D)] = ∫D0 S(t) dt
M3, expected watch within duration. Two videos can have the same completion but different useful watch paths.
Move the authored hazards and watch the curve
Research-aligned form. The survival structure is established mathematics. The creative hazard inputs are authored and shadow-only; at their defaults the illustrative 60-second curve opens at 84%, completes at 52% and gives an expected watch of 40.6 seconds.
M4: A large audience can arrive through different shapes
One platform broadcast and a deep person-to-person cascade may create similar view counts. Structural virality measures topology rather than size.
N = Next + ΣGg=1 Ng
𝔼[N | R < 1] = μ / (1 − R)
M4, exposure and endogenous diffusion. Popularity and propagation shape are different measurements.
| Pattern | Estimated reach | Endogenous share | Mean depth | Qualified return |
|---|---|---|---|---|
| Broadcast-heavy | 540k | 8% | 1.7 | 3.2% |
| Cascade-heavy | 506k | 73% | 7.9 | 5.8% |
| Smaller, durable | 188k | 61% | 4.8 | 12.6% |
Illustrative. The network figures are explanatory. They are not fitted to a platform or account.
M5: Distribution becomes valuable when exposure becomes memory
A person can sample once, recognise the series, return, become regular or disappear. Those transitions should not be compressed into follower count. YouTube similarly distinguishes new, casual and regular viewers.
πt+1 = πtP, Pij ≥ 0, Σj Pij = 1
M5, audience-state transition, an illustrative state model. Every row must conserve probability mass.
Episode , starting from 1,000 sampled people
The first release creates exposure. It has not created an audience yet.
The transition matrix behind each release
| From | Sampled | Recognised | Returning | Regular | Dormant |
|---|---|---|---|---|---|
| Sampled | 0.40 | 0.35 | 0.10 | 0.00 | 0.15 |
| Recognised | 0.05 | 0.35 | 0.35 | 0.05 | 0.20 |
| Returning | 0.02 | 0.05 | 0.48 | 0.25 | 0.20 |
| Regular | 0.01 | 0.02 | 0.12 | 0.68 | 0.17 |
| Dormant | 0.05 | 0.08 | 0.07 | 0.03 | 0.77 |
State model. Useful for accounting and mock tests. Not validated as a literal model of real people.
M6 and M7: One winner can make the whole system worse
AI-assisted creators may produce stronger individual outputs while the collection becomes more similar. A quality-diversity archive keeps strong candidates across different behavioural niches instead of allowing one early winner to consume every production slot.
maxx ∈ {0,1}n Σni=1 qixi
M6, portfolio objective, an illustrative optimisation. This is only the objective.
Σi cixi ≤ B,
Σi dijxi ≥ mj ∀j,
Σi∈𝒲w xi ≤ ρ Σi xi ∀w
M7, portfolio constraints. Production cost stays within budget, every required creative dimension receives minimum coverage, and no content world exceeds the maximum portfolio share. The constraints stop a short-term winner erasing the search space.
| Policy | Slots per world | Worlds preserved | Largest share | Novel mechanism slots |
|---|---|---|---|---|
| Greedy score | 17, 0, 0, 13, 0, 0 | 2 / 6 | 57% | 1 |
| Diverse portfolio | 6, 5, 4, 5, 4, 6 | 6 / 6 | 20% | 8 |
Research analogy. MAP-Elites is not a validated content strategy. The transfer is the preservation of multiple strong niches.
Notation and LaTeX for all twelve equations
P1 \mathcal{P}=W\land I\land U\land E\land F\land M
E1 \begin{aligned}
B_{\mathrm{net}}
&=T_{\mathrm{manual}}\\
&\quad-\left(T_{\mathrm{agent}}+T_{\mathrm{review}}+T_{\mathrm{repair}}+\frac{T_{\mathrm{setup}}}{n}\right)
\end{aligned}
P2 \mathcal{P}=W\land I\land U\land E\land F\land M
E2 \bar C(k)=\frac{C_e}{k}+C_p,\qquad k\ge 1
E3 A_{t+1}=(1-\delta)A_t+e_t,\qquad 0\le\delta\le 1
M1 \Pr(\mathrm{open})=\sigma\!\left(\beta_0+\beta_L L+\beta_C C+\beta_E E-\beta_M M\right),\quad \sigma(z)=\frac{1}{1+e^{-z}}
M2 S(t)=\exp\!\left(-\int_0^t h(u)\,\mathrm du\right),\qquad h(u)\ge 0
M3 \mathbb E\!\left[\min(T,D)\right]=\int_0^D S(t)\,\mathrm dt
M4 \begin{aligned}
N&=N_{\mathrm{ext}}+\sum_{g=1}^{G}N_g,\\[2pt]
\mathbb E[N\mid R<1]&=\frac{\mu}{1-R}
\end{aligned}
M5 \boldsymbol{\pi}_{t+1}=\boldsymbol{\pi}_t P,\qquad P_{ij}\ge 0,\quad \sum_j P_{ij}=1
M6 \max_{\mathbf x\in\{0,1\}^n}\;\sum_{i=1}^{n}q_i x_i
M7 \begin{aligned}
\sum_i c_i x_i&\le B,\\
\sum_i d_{ij}x_i&\ge m_j &&\forall j,\\
\sum_{i\in\mathcal W_w}x_i&\le \rho\sum_i x_i &&\forall w
\end{aligned}
Automate the labour, and keep public authority bounded
An agent can research, propose experiments, write, render, test and package a release. That does not entitle it to decide that a claim is true, a person is safe to depict or a piece is worth publishing.
| Level | What the system may do |
|---|---|
| A0 · Observe | The system can inspect public or approved material and report findings. It cannot create or change artifacts. |
| A1 · Draft | The system can propose ideas, scripts and plans. Every artifact remains a suggestion awaiting human review. |
| A2 · Build | The system can produce local assets, simulations and renders. Nothing leaves the controlled environment. |
| A3 · Prepare (selected) | The system can complete the production package and evidence checks. Calvin reviews the finished piece and explicitly authorises public release. |
| A4 · Bounded publish | Only calibrated, low-risk, allowlisted formats may publish through official APIs, with kill switches, incident logging and strict scope. |
| A5 · Unbounded | Prohibited. The system does not receive unrestricted authority to publish claims, depict people or alter public channels. |
Test which world makes relevant people return
Do not search for the viral hook. The first pilot should compare different plot engines. Cosmetic hook variants can wait until there is evidence that the underlying world deserves another episode. The real test is whether relevant people recognise the series and return.
| World | Mechanism | Production burden |
|---|---|---|
| Agent Olympics | Competition | 3 / 5 |
| Failure Museum | Evidence mystery | 2 / 5 |
| Maintenance Monster | Physical metaphor | 4 / 5 |
| Lost Files | Fictional investigation | 5 / 5 |
| Civilisation Lab | Persistent simulation | 5 / 5 |
| Drawer Ecology | Physical observation | 4 / 5 |
The primary outcome is qualified subsequent-episode consumption: did a relevant viewer voluntarily consume another episode from the same world within the declared window? Four guardrails sit beside it: evidence completeness, trust incidents, production burden and portfolio concentration.
One hit does not crown a format. The system should earn confidence through repeated episodes, comparable cohorts and a reasoned account of what changed.
What I am betting on
If the gate works, AI content stops being a faster way to state an opinion and becomes a cheaper way to operate and explain reality.
I still want the automation. I want the agent to search, set up the run, record the evidence, cut the video, visualise the maths and remember what changed.
But the machine should spend its new speed on the expensive part: asking questions that reality can reject. The human advantage is moving judgement earlier. The agent advantage is receiving a world with state instead of another empty prompt.
If it fails, I have a stricter planning system that produces fewer empty assets. If it succeeds, an agent can build a body of work whose originality comes from accumulated reality, not a more elaborate prompt.
Every system I found could make more content. The production layer is already arriving, and the useful question is whether we can force it to build knowledge rather than merely simulate confidence.
For coding agents: give your agent the contract
Give the agent an acceptance test, not a vague request to “make better content”. A human should be able to skim the rule in under a minute. An agent should be able to validate it before spending tokens, render time or public trust.
- CONTENT-PROOF-PROTOCOL.md, the Markdown protocol: readable operating instructions, sequence, failure modes and required artefacts.
- content-proof-contract.schema.json, the machine contract: a JSON Schema for validating the six content fields before production begins.
- pass.json / fail.json, the worked cases: one stateful experiment that passes and one polished explainer that is rejected.
The operating loop has five steps:
- Inspect the actual work. Find the inbox, failed build, customer question, prototype, decision or measurement that already contains tension.
- Name one intervention. Change a rule, process, object or input. Do not ask for “content about” a broad topic.
- Lock the evidence contract. Declare the metric, trace, footage, ledger or human judgement before seeing the result.
- Run before writing the title. Keep the boring, negative or embarrassing outcome. That is where trust enters.
- Store the state change. Give the next agent the result, unresolved problem and authority boundary instead of resetting to a blank prompt.
At the A3 boundary, the agent can prepare the complete release. The human still authorises publication.
The contract an agent can actually use
content_proof_contract: world_state_before: "" intervention: "" unknown_outcome: "" evidence_required: [] failure_story: "" world_state_after: "" reject_if: - the outcome is already written - the proof can be replaced by narration - an unflattering result would be discarded - the next episode resets the world
The article explains the judgement. The brief gives the agent the rule.
A copyable brief: do not make content until the event is defined
Use this before research, scripting, rendering or publishing. A failed gate is a valid result.
When asked to create content: 1. Inspect the real work for a stateful world. 2. Propose one controlled intervention. 3. Name the outcome that remains genuinely unknown. 4. Declare the evidence required before any claim can ship. 5. Preserve an unflattering or uneventful result. 6. State what changes permanently after the run. Reject the idea when: - the conclusion is already written; - narration could replace proof; - failure would be hidden or rerun until flattering; - the next episode would reset the world; - public release exceeds the authorised boundary. Return the completed content_proof_contract before producing assets.
Sources and further discussion
What is sourced, what is authored and what the pilot still has to prove. Research-backed means a source directly supports the narrow claim in the studied or documented context. An authored hypothesis is a useful model shape or design judgement that still needs prospective testing. Illustrative means a visual explanation using arbitrary units, not an estimated platform model. My position marks an explicit argument, not a disguised empirical finding.
- Wikipedia: Signs of AI writing, Wikipedia WikiProject AI Cleanup. Descriptive field guide, not a detector or policy. Used to audit generic prose patterns.
- Progressive Disclosure, Jakob Nielsen. Supports keeping the primary reading path simple and making technical depth available on request.
- Understanding Animation from Interactions, W3C Web Accessibility Initiative. Motion is optional, user-controlled and removed under reduced-motion preferences.
- Narrative Visualization: Telling Stories with Data, Edward Segel and Jeffrey Heer. Informs the balance between a guided narrative and reader-controlled exploration.
- Generative AI enhances individual creativity but reduces the collective diversity of novel content, Anil R. Doshi and Oliver P. Hauser. Supports treating output quality and portfolio diversity as separate concerns.
- The Structural Virality of Online Diffusion, Sharad Goel, Ashton Anderson, Jake Hofman and Duncan Watts. Supports separating diffusion size from diffusion topology.
- Measure key moments for audience retention, YouTube Help. Supports time-indexed retention and matched comparisons, not a universal attention-span claim.
- New, casual and regular viewers, YouTube Help. Supports distinguishing one-time exposure from recurring audience behaviour.
- Illuminating search spaces by mapping elites, Jean-Baptiste Mouret and Jeff Clune. Used as an analogy for preserving strong candidates across different creative niches.
- Telltale signs of AI-coded products, Paul Bakaus and other product-design practitioners, reported by Business Insider. Used only as a current practitioner critique of visual sameness and neglected edge states.
- Information Scent: How Users Decide Where to Go Next, Raluca Budiu, Nielsen Norman Group. Supports making the value of the article and each interaction legible before the reader commits attention.
- Website Reading: It (Sometimes) Does Happen, Jakob Nielsen, Nielsen Norman Group. Supports a scan-first structure that lets interested readers move from headings and figures into deeper prose.
- Usable but Conventional: An Empirical Study on the UX of AI-Generated Interface Prototypes, Karoline Romero, Igor Wiese, Renato Balancieiri, Gislaine Camila Leal and Guilherme Guerino. Reports positive pragmatic evaluations alongside weaker hedonic originality and innovation for AI-generated prototypes. (The source edition listed this paper twice, as sources 13 and 20.)
- video-use, Browser Use. Official project documentation for agent-led raw-footage editing, transcription, rendering and self-evaluation.
- Video tools for the agent era, Remotion. Official project documentation for agentic, interactive and programmatic video creation.
- Open-source social scheduling and analytics, Postiz. Official project documentation for scheduling, APIs and social analytics.
- YouTube’s top AI slop channels are disappearing, The Verge. Current reporting used to establish that platforms are already reacting to low-quality synthetic volume.
- Tampa Bay businesses are getting called out for AI slop, Axios Tampa Bay. Current reporting used as evidence of audience and community backlash against repetitive AI promotional output.
- A coffee shop owner used AI to make a menu poster. Then came the angry DMs., Business Insider. A recent small-business example showing that production savings can create a trust cost.
- ScrollyVis: Interactive visual authoring of guided dynamic narratives for scientific scrollytelling, Eric Mörth, Stefan Bruckner and Noeska N. Smit. Supports scroll-controlled narrative as a way to connect explanation with changing visual state while retaining reader control.
The interactive models are designed to make assumptions visible. They do not predict future views, virality, revenue or audience return. The honest next step is a manually approved pilot.