← Back to writing

Your AI can make 500 videos. Which 495 should never exist?

I built 19,200 synthetic viewing sessions to design a content preflight. They exposed brittle moments and better experiments. None of them proved how people watch.

Imagine asking an AI for hooks for one short clip and getting 500 candidates. One person still has to choose. Generation moved from scarce to abundant, and the cost didn’t disappear: it moved into review, experimentation and the risk of teaching an agent the wrong lesson.

I built 19,200 synthetic viewing sessions to design a filter for that pile. The system is useful only if it cuts review without silently killing the winner. My bet is that the scarce asset will be a private history of what survived a fair test.

One missing premise can turn proof into generic AI theatre

Take a hypothetical seven-second applied-AI clip. A support agent reproduces a bug, fails the right test and opens a reviewed pull request. The safety gate is the product. In the original edit:

  1. 0.0 s, a generic claim. “I built an AI support engineer.” The claim is broad.
  2. 1.7 s, an audio-only gate. “It reproduces the ticket before opening a PR.” The critical premise is speech-only.
  3. 3.4 s, a code diff. After > reproduce ticket #1842, the terminal shows running tests…; the viewer sees action, not the safety cause.
  4. 5.0 s, the pull request. The terminal reads opening pull request, and the payoff arrives without proof.
The reveal needs its premise on screen as well as in speech

A preflight of the hypothetical clip. A sound-off viewer sees code and a pull request, but misses that the agent reproduced the bug and failed a test first. The “proof” collapses into generic coding automation. The paired edit makes the gate visible as well as spoken.

Original editPreflight edit
On screen“I built an AI support engineer.” The safety gate exists only in speech.“It reproduced the bug before touching the code.” The gate is visible and spoken as three states: reproduced, test failed, PR opened.
Preflight readingThe reveal depends on a premise many viewers may never receive.The edit repairs one dependency without adding noise. The viewer receives the premise in two channels, sees the failed test, then sees the pull request. The reveal now has a cause.
Single point of failure1.7–2.4 s audio-only safety premiseNo single audio-only dependency
Competing explanationThe audience may simply not care about support automation.The product promise may still be irrelevant to this audience.
Paired editAdd a three-state visual proof before the pull request appears.Original versus dual-channel premise; hold every later scene constant.
Primary outcomeProtected recognition: “What happened before the PR opened?”Primary: proposition recognition. Secondary: six-second completion.
Predicted premise registration38%76%
Evidence levelE0 hypothesisStill E0 until tested

No real participant result is implied. The point is to replace “punchier” with a mechanism and a test: the example is specific enough for an engineer to implement and specific enough for an experiment to disprove.

A retention dashboard would show a drop somewhere in a clip like this. It rarely shows the cause, and the same visible symptom, 31% leaving within 0.8 seconds, can call for three opposite edits.

Three diagnoses for one symptom. The drop times are marks on illustrative curves.
Likely mechanismTempting editBetter testMeasurement
Missed premise (drop at 2.7 s). The viewer saw the reveal but missed the setup. A caption carried the premise while a face and object competed for selection. The next scene is unintelligible rather than boring.Add faster cuts and another visual surprise.Repeat the premise through speech or a compact recovery cue.Proposition recognition after the exact drop point.
Low information gain (drop at 4.1 s). The viewer understood the setup and stopped learning. The video repeats the same claim without adding evidence, stakes or surprise. Comprehension is intact; expected information gain collapses.Repeat the premise more loudly and add another recap.Move the reveal forward or introduce meaningful new evidence.Moment-to-moment curiosity and expected-value ratings.
Wrong audience (drop at 0.9 s). The content reached people who never wanted the promise. The edit may be clear and well-paced; the initial audience or platform exploration pool is simply a poor match for the topic.Re-edit the entire video around a more frantic hook.Change seeding, packaging or audience targeting before changing the core explanation.Retention conditional on acquisition source and audience cluster.

In my opinion, the first useful product is a failure-diagnosis instrument rather than a virality oracle.

The production surplus is here, and the judgement system is not

The platforms are becoming generation surfaces. They do not hand you a defensible evaluator for your audience.

Distribution is already enormous: on 18 June 2025, YouTube said Shorts averaged more than 200 billion daily views and announced Veo 3 for Shorts. Generation moved into the feed on 25 September 2025, when Meta launched Vibes, a feed for creating, remixing and sharing short-form AI-generated video, including cross-posting to Reels and Stories. Evaluation is becoming build infrastructure: on 29 January 2026, OpenAI published a repeatable image-evaluation harness for turning subjective media judgements into structured comparisons. And WCAG guidance (5 October 2023) recommends providing a way to disable non-essential motion triggered by interaction and respecting reduced-motion preferences: motion should explain, not merely decorate.

Each of those claims links to the original page. That they add up to a short strategic window is my inference, not something the publishers proved.

My bet is that model access compresses and clean outcome history does not. Generation capacity arrives first, with more variants than humans can inspect. Evaluation becomes the bottleneck next, and teams optimise whatever proxy is easiest. Instrumented experiments are what compound, because each real test improves the next selection. The old constraint was producing: making one credible variant took meaningful time. The new one is judging, because generation multiplied faster than human review capacity. The durable asset is learning, where paired outcomes improve the next rejection decision.

Safe rejection is the hard part

Generation scales almost for free, and bad judgement scales with it.

For the person choosing, a preflight is worth having if it spends attention where it can change the decision. It should show the exact premise, caption or transition the model thinks is fragile, let you compare edits before paying with real audience attention, and say whether a recommendation is evidence, a hypothesis or a bluff.

For the agent, the useful thing is a contract rather than another inspirational essay. It separates perception, attention, comprehension, memory, action and exposure. It states uncertainty and evidence level beside every recommendation, and it ends with the cheapest real experiment that could disprove the analysis.

One input, two agents. Input: “Retention drops at four seconds. Improve the edit.” The difference is operating discipline rather than prose.
AgentOutputWhat it carries
UnconstrainedMake the hook stronger. Add faster cuts, a more dramatic caption and more movement in the opening.No mechanism, no competing explanation and no falsification plan.
Under contractThe drop may follow a missing premise, not a weak hook. At 2.1 seconds the critical premise is audio-only. The reveal at 4.0 seconds depends on it. Compare a dual-channel premise against the original and measure proposition recognition before changing pace.E0 mechanism hypothesis · two competing explanations · E3 test required

I’m not after “viral every time”. The upside is an agent that can tell you where it thinks a video breaks, how sure it is and which real test deserves your money. Generating 500 hooks is already common; an AI that could safely reject 495 of them would be economically valuable. Content generation would then work like search: generate widely, reject carefully, test narrowly and keep what was learned.

Five levels of content agent, from generator to learner

A speculative operating model, not measured productivity gains.

LevelWhat it doesHuman burdenMain failure
0 · GeneratorIt makes variants. That is all. The system produces more options than a human can inspect. It has no causal model, no evidence boundary and no reason to prefer one version beyond surface plausibility. Output: hundreds of polished candidates.You still review everythingVolume masquerades as intelligence
1 · ScorerIt ranks the pile, but cannot explain the ranking. A score compresses work. It can also compress different failures into one number and encourage the agent to optimise whatever proxy happens to move. Output: a leaderboard of candidates.You inspect the top of an opaque listProxy gaming with confidence theatre
2 · DiagnosticianIt names where the video may break and why. The agent separates missed selection, lost context, weak relevance, memory failure and poor audience match. Every diagnosis carries a competing explanation. Output: timestamped failure hypotheses.You review mechanisms, not every variantA plausible diagnosis can still be wrong
3 · ExperimenterIt turns each diagnosis into a paired test. The agent generates treatment and control edits that differ on one mechanism, names a primary outcome and refuses to treat synthetic lift as causal evidence. Output: counterfactual edit pairs and protocols.Humans spend attention on high-value uncertaintyBad randomisation or weak measurement
4 · LearnerIt updates from reality without rewriting history. Point-in-time telemetry recalibrates the simulator. Drift, creator effects and platform changes can force abstention. Synthetic results remain subordinate to observed outcomes. Output: a closed evidence loop.Humans govern objectives, rights and promotion gatesOptimising the wrong objective at scale
Cheap generation is only an advantage if evaluation is cheaper than uncertainty

An illustrative scenario, not a forecast. It shows the best-case operating leverage if the rejection model eventually earns external validity. Review time is candidates × minutes each. The model assumes microtests before the filter are 4% of candidates (at most 20) and after it 10% of survivors. Australian dollars.

Change the cost assumptions
Survive the model50
Review before16.7 h
Review after1.67 h
Potentially recovered15.0 h
Review savedA$1,275
Microtests avoidedA$1,200
Simulation spend−A$20
Expected false-reject loss−A$100
Illustrative net valueA$2,355

Net value = review saved + tests avoided − simulation − (1 − top-k recall) × winner value

Under these assumptions, the preflight has positive expected operating value. Its top-k recall remains the dangerous assumption.

The calculator assumes the synthetic rejection is safe. That assumption is the entire research problem: the rejection model must earn high top-k recall on real outcomes before it can safely kill anything.

“Engagement” hides eight different systems

A single score makes a clean dashboard and a poor diagnosis. A frame can be visible, selected, attended, understood and remembered, and those are different events. One synthetic second shows the chain of losses.

One synthetic second. The frame is available throughout; the other columns are the model’s illustrative state estimates, not measurements.
TimeWhat happensSelectedRegisteredContext
0.0 sThe premise appears. Everything is present. Nothing is yet registered.22%7%12%
0.2 sMotion wins selection. The face captures likely gaze before the caption receives enough processing.74%18%19%
0.4 sThe premise enters working memory. Attention and relevance make registration plausible. Prior knowledge still matters.82%67%58%
0.6 sA notification breaks the bridge. The next proposition arrives while context is weak.42%38%43%
0.8 sThe reveal lands without its cause. Visual attention returns. Comprehension does not automatically return with it.62%36%21%
1.0 sA recovery cue repairs context. A short recap can restore the missing dependency. That is a hypothesis to test.79%71%76%

The chain runs present → selected → registered → understood → retrievable, and the model treats the whole path as eight stages: stimulus, selection, attention, meaning, memory, choice, sharing and exposure. Each asks a different question and fails differently.

The eight stages: question, output, failure mode and evidence needed
  1. Stimulus. What changed in the video at this moment? Frames, motion, speech, captions, music, cuts and propositions form a time-indexed stimulus stream. Output: an event timeline and semantic proposition graph. Failure mode: a visually dramatic event may carry no useful information. Evidence needed: frame, audio and transcript extraction accuracy.
  2. Selection. Which element is likely to win perceptual selection? Dynamic saliency, gaze goals, audio onsets and visual competition shape likely fixation targets. Output: gaze distribution and target-notice latency. Failure mode: the intended caption remains visible but loses to motion or a face. Evidence needed: mobile eye tracking on unseen viewers and videos.
  3. Attention. How much processing is allocated after selection? Persistence, relevance, novelty, fatigue, distraction and load drive a continuous state with intermittent lapses. Output: attention state and lapse trajectory. Failure mode: a viewer looks at the screen while shallowly processing it. Evidence needed: independent lapse labels plus passive multimodal signals.
  4. Meaning. Did the viewer integrate the proposition and its prerequisites? A proposition graph separates noticing words from understanding how a claim connects to what came before. Output: proposition registration and comprehension probabilities. Failure mode: the reveal is seen but its cause is missing. Evidence needed: protected comprehension items and free-response scoring.
  5. Memory. What remains retrievable later? Encoding strength, interference and delay produce a distribution over immediate and delayed recall. Output: immediate recognition and delayed recall probability. Failure mode: the emotional gist survives while the factual claim disappears. Evidence needed: immediate, 24-hour and 7-day follow-up data.
  6. Choice. What does the viewer do next? Continue, swipe, replay, like, save and comment are distinct actions with different costs and motives. Output: action-specific hazard and probability. Failure mode: a person understands the video and leaves because the expected value is complete. Evidence needed: timestamped session actions with acquisition context.
  7. Sharing. Why would this viewer transmit the video? Perceived usefulness, identity, emotion, memory and social friction combine into a separate decision. Output: conditional share propensity. Failure mode: high completion is mistaken for high transmission. Evidence needed: observed sharing and recipient context, not stated intent alone.
  8. Exposure. Who receives the video next? Creator history, platform exploration, network structure, timing and prior shares determine downstream reach. Output: conditional exposure and cascade distribution. Failure mode: content response is confused with recommender allocation. Evidence needed: versioned impression logs, propensities and real cascade trees.
The equations, term by term

The equations should expose assumptions, not hide them. This is an explorable structure with synthetic coefficients: a proposed state structure, not a fitted universal law of human attention.

Attention. Attention persists when the content remains relevant and informative. Load, fatigue and distraction pull it down.

A(t+Δ) = σ( ρ·A(t) + β_r·R + β_n·N − β_l·L − β_f·F + ε )

ρAt: attention has inertia; a focused viewer is more likely to remain focused over the next short interval. βrR: personal relevance raises the value of continuing to process the video. βnN: meaningful prediction error can restore attention, while decorative movement may instead add load. βlL: competing speech, captions and motion consume limited processing capacity. βfF: fatigue lowers persistence and slows recovery after interruption. ε: unobserved state and stochastic variation prevent identical viewers from following identical paths. With the default inputs (relevance 68, novel information 54, cognitive load 46, fatigue 28, confusion 24 and friction 32, each out of 100), the attention model gives 0.61 at 3.0 seconds.

Exit. Exit risk rises with confusion, boredom and lapses. Curiosity, relevance and unresolved value delay the swipe.

h(t) = σ( γ_c·C + γ_l·L + γ_d·D − γ_r·R − γ_q·Q )
S(t) = ∏_(k≤t) (1 − h_k)

γcC: confusion increases exit risk when the viewer cannot reconstruct what the current scene means. γlL: high load can make continuing feel expensive even when the topic is relevant. γdD: a lapse is not automatically an exit, but it can remove prerequisites and raise later risk. γrR: relevance reduces exit risk by increasing expected value. γqQ: an unresolved, credible question creates a reason to remain.

Share. Sharing is a separate decision. A viewer may finish, understand and still have no reason to transmit the video.

P(share) = σ( δ_v·V + δ_i·I + δ_e·E + δ_m·M − δ_f·F )

δvV: the viewer expects the recipient to gain something useful, entertaining or socially relevant. δiI: sharing can signal taste, belonging or a position to other people. δeE: arousal can increase transmission, but valence and context shape direction. δmM: a retrievable message is easier to describe and justify to somebody else. δfF: recipient choice, privacy, uncertainty and interface cost can suppress an otherwise positive intention.

Cascade. Downstream reach combines outside injection with the decaying influence of prior exposures and shares.

λ(t) = μ(t) + Σ_j α_j · e^(−β(t−t_j))

μ(t): platform exploration, creator history, paid distribution, timing and trends create exposure that does not come from prior shares. αj: different sharers create different numbers and qualities of subsequent exposures. β(t−tj): the influence of a prior event usually declines as it becomes older.

Rare actions need large samples, and a 1% share event is unstable in a population of 100. With 100 simulated viewers and a true share probability of 1.0%, the expected number of shares is 1.0, the chance of observing none is 36.6% and an approximate 95% interval is 0–3, so a single run can reverse the apparent ranking. With 10,000 viewers the expectation is 100 shares, with an interval of about 80–120. One hundred personas are useful theatre; ten thousand particles estimate distributions. The particles are stochastic, not biographies, and the five viewer profiles (focused, typical feed, attention-variable, multitasking and sound-off) describe functional conditions, not diagnoses. Simulation error is only one source of uncertainty; model misspecification remains.

Virality is not a property of the file. The same video can die, travel or explode under different seeds, creator priors and recommendation pressure. In the cascade equation above, μ(t) is outside exposure from the platform, creator and timing, and each prior share contributes an additional decaying term.

Every synthetic gate passed, and reality remained untouched

The mock study generated 2,400 participants, 120 videos and 19,200 sessions. It recovered every mechanism deliberately placed inside its artificial world, and all 13 synthetic gap gates, G01 to G13, passed. Those passes establish pipeline sensitivity to known synthetic truth. They do not establish prediction of real people or platforms.

The preserved release receipt, an excerpt from the v0.4 release verification, reads overall_pass: false:

{
  "automated_tests": { "passed": false, "count": 82 },
  "component_campaigns": { "passed": false },
  "seed_sweep": { "passed": false },
  "wheel_clean_install": { "passed": false, "summary": "no wheel" },
  "external_gap_guard": { "passed": true, "open": 13 }
}

The receipt itself contained inconsistencies: some summaries said “passed” while their boolean remained false. I’d rather show it than hide it. It is evidence that the assurance machinery also needs verification.

A claim should carry the evidence level it has earned. I use six: E0 theory (mechanism and sources), E1 synthetic (code and mock execution), E2 benchmark (public data conformance), E3 controlled human (a protected participant study), E4 field (target platform and context) and E5 replication (independent confirmation). This work has reached E0 and E1.

The minimum evidence before I should publish a claim as fact
ClaimMinimum defensible levelWhy
The simulator preserves its encoded state invariants.E1 / synthetic executionThe claim is about the software’s own encoded behaviour. Protected synthetic tests can support it.
The engine predicts where real viewers look.E3 / controlled human studyReal gaze requires mobile eye-tracking ground truth on unseen viewers and videos, with center and motion baselines.
The recap edit improves platform retention.E4 / target-platform field evidenceA controlled edit can establish a human effect. Platform retention also requires versioned field telemetry and drift controls.
The system predicts whether a video will go viral.E5 / independent replicationA general virality claim spans content response, exposure allocation and diffusion. It needs replicated platform-specific evidence.
The 13 open validation questions

Every question below is still open. Each names what would close it, the data it needs, its primary metrics and the rule for promoting a claim.

  1. G01 · Real gaze and perceptual-selection validity (human). Close with mobile eye tracking on unseen viewers and videos. Data: fixations, scanpaths, device geometry and target-notice latency. Metrics: NSS, SIM, scanpath distance and calibration. Rule: beat center and motion baselines without subgroup collapse.
  2. G02 · Real attention-lapse detection (human). Close with independent lapse labels with grouped participant holdout. Data: experience sampling, gaze, pupil, head pose and interaction traces. Metrics: event F1, latency, false alarms, Brier score and ECE. Rule: preregistered threshold plus independent replication.
  3. G03 · Diagnosed-participant validity (inclusive). Close with an approved clinician-confirmed cohort with functional measures. Data: diagnosis, medication, context and matched comparison groups. Metrics: subgroup calibration, interaction effects and heterogeneity. Rule: no diagnosis claim from a synthetic persona label.
  4. G04 · Proposition comprehension in actual viewers (human). Close with a protected item bank with free-response and recognition tasks. Data: timestamped items, prerequisite graph, confidence and rater labels. Metrics: Brier score, item discrimination, reliability and DIF. Rule: hold out propositions, videos, languages and viewers.
  5. G05 · Immediate and delayed human recall (human). Close with immediate, 24-hour and 7-day follow-up. Data: free recall, recognition, delay and attrition metadata. Metrics: forgetting-curve error and recall calibration. Rule: beat delay-only and memorability baselines.
  6. G06 · Creator and account-history effects (platform). Close with a rolling-origin creator model with point-in-time features. Data: authorised creator history and future content outcomes. Metrics: unseen-creator error and incremental predictive value. Rule: no future leakage; shrinkage and uncertainty reported.
  7. G07 · Recommender-system exposure dynamics (platform). Close with impression logs with propensities or randomised exploration. Data: exposure, policy version, logging propensity and response. Metrics: IPS, SNIPS, doubly robust agreement, overlap and ESS. Rule: no counterfactual claim without support diagnostics.
  8. G08 · Real sharing cascades (platform). Close with complete-enough chronological diffusion trees. Data: reshare edges, exposures and exogenous injections. Metrics: cascade size, depth, timing, structure and tail calibration. Rule: beat branching, Hawkes and simple-history baselines.
  9. G09 · Platform-specific retention calibration (platform). Close with a versioned event ontology and raw watch-session logs. Data: sessions, creator, acquisition source and platform version. Metrics: integrated Brier, curve error and subgroup drift. Rule: abstain when metric definitions or policy versions drift.
  10. G10 · Cross-language transfer (inclusive). Close with native-speaker parallel-stimulus studies. Data: adapted stimuli, locale and comprehension responses. Metrics: worst-language error, leave-one-language-out transfer and DIF. Rule: no pooled claim when measurement invariance fails.
  11. G11 · Habituation across repeated exposures (human). Close with longitudinal within-person repeated-exposure data. Data: exposure count, interval, familiarity and novelty measures. Metrics: within-person slope and nonlinear exposure response. Rule: separate habituation, sensitisation and mere exposure.
  12. G12 · Accessibility validation (inclusive). Close with co-designed testing with disabled participants and normal assistive technology. Data: task performance, fatigue, errors, preferences and issue severity. Metrics: success, comprehension, burden and worst-group outcomes. Rule: automated conformance never substitutes for user validation.
  13. G13 · Causal effectiveness of proposed edits (causal). Close with preregistered randomised, creator-blocked edit experiments. Data: versioned edits, assignment integrity and outcome telemetry. Metrics: ATE, CATE, guardrails and sequential-valid inference. Rule: replicate direction and magnitude across creators and cohorts.

The short window is for building the feedback loop

Everyone will get cheaper generators. Fewer teams will collect clean, paired, point-in-time evidence about why their own audience responds.

The arithmetic of that history is simple. Eight paired tests a month for 12 months is 96 executed tests. If 70% give a reusable outcome, that is 67 reusable cases, and a rough uncertainty proxy of 1/√67 is 12.2%. That is not a promise of model accuracy. It shows why clean experimental history compounds while random posting mostly creates anecdotes.

The first real test I’d run is small. Pick one fifteen-second clip. Name one proposition the viewer must understand. Create one paired edit that changes only how that proposition is communicated, collect 500 real exposures, and measure recognition before celebrating views.

Turning a model claim into a falsifiable human study means a bounded next step rather than another synthetic victory lap. For a recovery cue, the draft protocol asks: does a recovery cue restore comprehension after a brief lapse?

  1. Recruit 240 participants and block randomisation by video, creator and the functional attention measures collected before viewing.
  2. Treatment: insert a compact visual-and-spoken recap after the interruption window.
  3. Control: use the same edit without the recap.
  4. Primary outcome: protected proposition items immediately after viewing.
  5. Include a prespecified muted-viewing stratum and report interaction uncertainty.
  6. Do not make delayed-memory claims from this protocol.
  7. Freeze the analysis, exclusions and stopping rule before inspecting treatment outcomes.

That needs E3, a controlled human study, analysed as a participant- and video-blocked treatment effect.

Other edit pairs and outcomes for the same protocol
  • Dual-channel premise vs audio-only premise. Does dual-channel presentation protect the premise when sound or gaze is unavailable? Treatment: present the critical premise in aligned speech and on-screen text. Control: present the premise through audio alone.
  • Simplified overlay vs dense overlay. Does simplifying the overlay improve registration without reducing interest? Treatment: use one primary caption and remove competing decorative elements. Control: use the original dense overlay.
  • Aligned hook vs misleading hook. Does an aligned hook improve qualified retention relative to a misleading hook? Treatment: open with the real promise and prerequisite. Control: open with a high-arousal claim that the video does not directly satisfy.
  • Time-to-exit as the outcome. Timestamped time-to-exit and completion, analysed as survival outcomes with a stratified survival model with creator and video effects. Needs E4, a platform field study.
  • 24-hour recall as the outcome. Immediate recognition plus blinded 24-hour free recall, analysed with a delay-aware hierarchical recall model. Needs E3, a longitudinal human study. With delayed recall included, collect it and model attrition rather than deleting incomplete follow-up.
  • Sharing as the outcome. Observed sharing where possible, separated from stated intention, analysed as an action-specific treatment effect with recipient-context guardrails. Needs E4, a behavioural field study.
  • Without a sound-off condition. Keep sound conditions fixed and record actual device volume. Participant numbers can run from 80 to 800.

What I’d build instead of a content slot machine does three jobs. It diagnoses brittleness, finding the half-second where missing one proposition destroys everything that follows. It ranks alternatives, asking whether edit B is more robust than edit A for a defined audience distribution. And it buys better evidence, using uncertainty to decide which real experiment is worth running next.

“This video will get 2.4 million views” is theatre. “This premise is a single point of failure, here are two competing explanations, and this experiment separates them” is a system.

My bet is that every real view can make the next decision cheaper and harder to fool. I built 19,200 synthetic views, and none of them proved how people watch. They showed how to make the next 500 real views teach an agent more than the last 50,000. If a rejection model can’t earn high top-k recall on real outcomes, it shouldn’t be allowed to kill anything.

For coding agents: the Audience Dynamics contract

The article explains the idea. The contract turns it into inputs, outputs, abstention rules and evidence gates an agent can follow. Four rules sit under it:

  1. Diagnose, do not prophesy. Locate likely failure mechanisms without claiming exact private mental states.
  2. Separate the layers. Do not collapse gaze, attention, comprehension, memory, choice and exposure into one engagement score.
  3. Carry uncertainty. Every recommendation names its assumptions, competing explanations and evidence level.
  4. End with a test. The output is incomplete until it says what real observation would prove it wrong.

Read the contract, machine JSON and output schema

Audience Dynamics Contract v1 asks the agent to evaluate content as a falsifiable system:

  1. Map the video timeline, propositions and prerequisite dependencies.
  2. Model availability, selection, attention, comprehension, memory, choice, sharing and exposure separately.
  3. Run functional viewer profiles and report distributions, not fictional biographies.
  4. For every predicted failure, provide at least one competing explanation.
  5. Generate paired counterfactual edits with one primary measurable outcome.
  6. Label every claim E0–E5 and abstain beyond the evidence earned.

Required output: timestamped failure map, affected cohorts, competing causes, counterfactual edits, uncertainty, evidence level and next real experiment.

Example input
Fifteen-second explainer. Critical premise appears through audio at 1.6 seconds. The reveal arrives at 4.2 seconds. A large sound-off cohort is plausible.
Contract output
The reveal depends on an audio-only premise. Primary hypothesis: sound-off viewers can see the payoff but cannot integrate why it matters. Competing hypothesis: the promise itself is irrelevant. Test a dual-channel premise against the original and measure protected proposition recognition before retention. E0 mechanism hypothesis · requires E3 human evidence.
AUDIENCE DYNAMICS CONTRACT v1

OBJECTIVE
Diagnose plausible content failure mechanisms and propose falsifiable tests. Do not claim exact private mental states or platform outcomes beyond earned evidence.

REQUIRED ANALYSIS
1. Map the time-indexed video events, propositions and prerequisite dependencies.
2. Separate availability, selection, attention, comprehension, memory, choice, sharing and exposure.
3. Simulate functional viewer distributions, not diagnosis stereotypes or fictional biographies.
4. Provide at least one competing explanation for every predicted failure.
5. Generate paired counterfactual edits with one primary measurable outcome.
6. Label every claim E0-E5 and abstain beyond the evidence earned.

REQUIRED OUTPUT
Timestamped failure map; affected viewer conditions; competing causes; counterfactual edit pairs; uncertainty and abstentions; evidence level; next real experiment.

FORBIDDEN CLAIMS
Exact mind-state inference; diagnosis from behaviour; absolute virality forecasts without platform calibration; causal edit lift from synthetic evidence alone; hidden assumptions.
Audience Dynamics first test

1. Select one 15-second clip.
2. Name one proposition viewers must understand.
3. Produce control and treatment edits that differ only in how that proposition is communicated.
4. Pre-register proposition recognition as the primary outcome.
5. Collect 500 real exposures using point-in-time platform definitions.
6. Compare recognition before interpreting retention.
7. Feed the observed result back into the evaluator; do not promote synthetic claims automatically.
{
  "name": "Audience Dynamics Contract",
  "version": "1.1.0",
  "objective": "Diagnose plausible content failure mechanisms, price the decision value of further evidence, and propose falsifiable tests without claiming exact private mental states or platform outcomes beyond earned evidence.",
  "inputs": [
    "time-indexed video events",
    "transcript and on-screen text",
    "proposition dependency graph",
    "target audience distribution",
    "viewing context and sound state",
    "creator history available at prediction time",
    "platform metric definitions and version"
  ],
  "required_analysis": [
    "separate availability, selection, attention, comprehension, memory, choice, sharing and exposure",
    "simulate functional viewer distributions rather than diagnosis stereotypes",
    "provide competing explanations for every predicted failure",
    "compare paired counterfactual edits under common assumptions",
    "state uncertainty and evidence level for every claim",
    "end with the cheapest real observation that could falsify the recommendation",
    "estimate review and experiment costs under explicit user-supplied assumptions",
    "protect top-k recall so filtering does not silently remove the strongest candidate",
    "treat each real paired test as a point-in-time calibration event",
    "abstain when the available evidence cannot separate the leading explanations"
  ],
  "required_outputs": [
    "timestamped_failure_map",
    "affected_viewer_conditions",
    "competing_explanations",
    "counterfactual_edit_pairs",
    "uncertainty_and_abstentions",
    "evidence_level_E0_to_E5",
    "next_real_experiment",
    "decision_economics",
    "top_k_recall_gate",
    "calibration_update_plan"
  ],
  "forbidden_claims": [
    "exact individual mental-state inference",
    "diagnosis from viewing behaviour",
    "absolute virality forecast without platform calibration",
    "causal edit lift from synthetic evidence alone",
    "unlabelled assumptions or hidden metric definitions"
  ]
}
{
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "title": "Audience Dynamics Evaluation Output",
  "type": "object",
  "additionalProperties": false,
  "required": ["summary", "failure_map", "competing_explanations", "edits", "uncertainty", "evidence", "next_test"],
  "properties": {
    "summary": { "type": "string", "minLength": 1 },
    "failure_map": {
      "type": "array",
      "items": {
        "type": "object",
        "additionalProperties": false,
        "required": ["start_s", "end_s", "mechanism", "affected_conditions"],
        "properties": {
          "start_s": { "type": "number", "minimum": 0 },
          "end_s": { "type": "number", "minimum": 0 },
          "mechanism": { "type": "string" },
          "affected_conditions": { "type": "array", "items": { "type": "string" } }
        }
      }
    },
    "competing_explanations": { "type": "array", "minItems": 2, "items": { "type": "string" } },
    "edits": {
      "type": "array",
      "minItems": 1,
      "items": {
        "type": "object",
        "additionalProperties": false,
        "required": ["treatment", "control", "primary_outcome"],
        "properties": {
          "treatment": { "type": "string" },
          "control": { "type": "string" },
          "primary_outcome": { "type": "string" }
        }
      }
    },
    "uncertainty": {
      "type": "object",
      "additionalProperties": false,
      "required": ["confidence", "assumptions", "abstentions"],
      "properties": {
        "confidence": { "type": "string" },
        "assumptions": { "type": "array", "items": { "type": "string" } },
        "abstentions": { "type": "array", "items": { "type": "string" } }
      }
    },
    "evidence": {
      "type": "object",
      "additionalProperties": false,
      "required": ["level", "earned_by"],
      "properties": {
        "level": { "enum": ["E0", "E1", "E2", "E3", "E4", "E5"] },
        "earned_by": { "type": "array", "items": { "type": "string" } }
      }
    },
    "next_test": {
      "type": "object",
      "additionalProperties": false,
      "required": ["question", "design", "promotion_rule"],
      "properties": {
        "question": { "type": "string" },
        "design": { "type": "string" },
        "promotion_rule": { "type": "string" }
      }
    }
  }
}
Sources and further discussion

This article combines an internal synthetic research archive with external work on simulated UX agents, video saliency, memory, short-form context switching, diffusion and interface design. External sources motivate model structure. They do not validate the simulator’s coefficients.

  1. UXAgent frames simulated users as a way to evaluate study design before real human-subject research.
  2. TASED-Net models dynamic video saliency from consecutive frames.
  3. Memento10k models video memorability and decay across viewing delays.
  4. Chiossi et al. studied prospective-memory performance after short-form video and rapid context switching.
  5. HIP separates exogenous promotion from endogenous popularity dynamics using Hawkes processes.
  6. Progressive disclosure, scannable web writing and purposeful motion informed this article’s interaction design.
  7. WCAG 2.2 animation guidance informed the decision to keep every figure still unless the reader changes it.
  8. Wikipedia’s AI-writing field guide was used as an editorial smell list, not as an authorship detector.
  9. Model Cards for Model Reporting informed the machine-readable claim, limitation and evaluation contract.
  10. Guidelines for Human-AI Interaction informed capability boundaries, correction paths and expectation setting.
  11. Recent criticism of generic AI-coded interfaces was treated as a design-risk checklist: visual sameness, polished shells over weak function and missing edge states.
  12. YouTube CEO keynote: Shorts scale, Veo 3 and AI dubbing, 18 June 2025. A primary source for the Shorts scale figure.
  13. Meta Vibes: AI-generated short-video feed, 25 September 2025. A primary source for generation moving into the feed.
  14. OpenAI Developers: image evaluation guide, 29 January 2026. A primary source for the evaluation harness.
  15. Nielsen Norman Group: users scan, concise objective pages perform better.
  16. Nielsen Norman Group: F-pattern and layer-cake scanning.

The original interactive edition also modelled what a developer was likely to skim, miss and return for, using an explicit heuristic rather than eye tracking: P(inspect section) = σ(1.2 × information scent + 0.9 × goal fit + 0.7 × novelty − 1.1 × reading cost). The coefficients were explanatory, not fitted human parameters. That model described the earlier layout, so it is not reproduced on this page.

Published 2026-09-15 · Updated 2026-09-27 · Source edition v6