Calvin Kennedy / building with AI / 20 September 2026

Three openings.
One demo.

Use AI for the scene you haven’t filmed. Record the product once. Then let your agent rebuild only what changed.

My storyboard → images → Seedance experiment got about 2,000 views. Now I want to make the next version without starting again.

Schematic footage. No paid generation.

30-SECOND EDIT / WORKED EXAMPLESwap only the opening
OPENING A / 0–6 SECONDS
“Five emergencies.
Two actual problems.”
06sOpening14sProduct06sReaction04sEnd
Change the first six seconds.The product recording, reaction and end card stay.
01 / One example all the way through

You built a tool.
Show the useful bit.

Here, a tiny tool groups repeated support labels: five reports, two categories. It runs below. This is a teaching fixture, not a claimed Vibecord feature.

Input
Can’t joinLagCan’t joinLagCan’t join
OutputCan’t join ×3   /   Lag ×2Expected output; button runs the grouping.
01Record the proofActual feature → one honest claimYou + screen recorder
Make the decision here

Show the grouping operation on real data you are allowed to publish. For this teaching example, five labels become two groups. Do not infer a time-saving percentage from the animation.

Keep

demo.mp4 · brief.json

Instruction I would give the agent
Inspect the product recording. State one claim the viewer can verify from it. Plan a 30-second film with a 6-second opening, 14-second demonstration, 6-second reaction and 4-second end card. Flag missing evidence. Do not invent functionality.
02Write three hooksOne problem → three 6-second scenesExisting reasoning model
Make the decision here

Try a news report, a frustrated player and a direct explanation. They all lead to the same proof. Keep a common reaction only if its line and eyeline still fit.

Keep

shots.json · script.md

Instruction I would give the agent
Propose three opening scenes: a newsreader overreacts to five support reports; a player mistakes duplicates for new problems; a direct explanation. Each lasts six seconds and leads into five labels grouped into two categories. Specify action, dialogue, framing and the cut. Keep the product demonstration fixed.
03Approve the framesCast + current state → starting imagesImage model · e.g. Nano Banana Pro
Make the decision here

Give the identity image a separate role from the costume, set and current prop state. An earlier flashback must not inherit damage from a later scene. Inspect every starting frame before paying for motion.

Keep

references/ · state.json · approvals.json

Instruction I would give the agent
Use the approved Ada reference and newsroom layout. Produce a medium starting frame for opening A. Preserve identity and costume; specify whether the mug is intact at this story time. No captions or product UI baked into the pixels. Return candidates; do not mark them approved.
04Time the rough cutStills + voice → an animaticVoice/performance + editor
Make the decision here

Put the words on a clock before animating the mouth. Off-screen narration can change independently; visible speech may need a new performance. Use only an original or authorised voice.

Keep

audio/ · animatic.mp4 · timeline.json

Instruction I would give the agent
Place the approved stills and voice in a 30-second rough edit. Keep speech, ambience and captions separate. Mark pauses and cut points. Report any line that does not fit the allotted time. Ask for a timing decision before generation.
05Generate the motionApproved frame → bounded candidate takesSeedance 2.0 / tested challenger
Make the decision here

Generate the opening and reaction, not the product screen. Confirm the actual endpoint supports the references, duration and audio mode. A selected cheap draft is not a promise that a premium rerender preserves it.

Keep

takes/ · jobs.json · cost.json

Instruction I would give the agent
Plan one six-second opening take from the approved first frame. Record model/provider, settings, references, estimated upper cost and job ID. Validate the quote before dispatch. Hold after an ambiguous timeout; reconcile the original job rather than silently retrying. No paid calls until authorised.
06Assemble the versionsNew hooks + shared tail → editable draftsFFmpeg / existing editor
Make the decision here

A timeline chooses exact intervals. Keep the actual product capture, captions and end card out of generated pixels. A caption correction should invalidate the final render and approval, not the underlying footage.

Keep

timeline.json · captions.srt · drafts/

Instruction I would give the agent
Assemble three drafts from the selected hooks, shared real demo, compatible reaction and end card. Keep all source takes. Captions and logos remain separate. Return an edit decision list and provenance for each interval. Regenerate no footage for a text-only correction.
07Test the resultActual drafts → findings + final decisionPerception + code + you
Make the decision here

Inspect full clips and sound. Ask viewers what the tool did, without showing them the script. Keep rejected takes and active repair minutes in the comparison. Finish with human approval of the exact export.

Keep

review.json · costs.csv · approval.json

Instruction I would give the agent
Review the actual exports against the approved brief, not the prompt alone. Identify defects with timestamps and mark unseen intervals unknown. Compare comprehension and complete production cost. Correct one caption and verify the source media hashes stayed unchanged. Do not publish automatically.

The division of labour: models supply candidates; the timeline supplies exact cuts. You decide which take is worth keeping.

02 / Spend on the variation

Don’t regenerate
the other 24 seconds.

Only the six-second opening changes. The recorded demo, six-second reaction and four-second end card can stay. Reuse also needs a valid edit.

Regenerate opening + reaction each timeUS$16.33
108 generated seconds
Generate each opening; reuse the reactionUS$10.89
72 generated seconds
36 fewer paid seconds.US$5.44 less in this maximum-takes scenario.

Video generation only. Assumed 6s takes at US$0.1512/s; provider-row snapshot, not a quote. Both versions reuse the real demo. Images, voice, editing and review are extra. [1]

Show the arithmetic and quality trade-off
Regenerate: 2kdtc
Reuse: (k + 1)dtc

k openings, d seconds per take, t maximum takes, c cost per second. This is a maximum dispatch comparison, not expected spending or a quality guarantee.

Reuse assumes the reaction still fits each opening. A changed eyeline, spoken line or emotional beat can invalidate it. Record human repair time; “zero new video jobs” is not “free work”. A one-shot version or a direct screen recording may be better.

A

Reject the still before buying motion.

Wrong costume? Fix the reference now. A cheap still check helps only when it catches a fault that would survive into the take.

When does that extra check pay?

With an assumed US$0.10 image check and US$0.9072 six-second take, avoiding one doomed generation in more than 11.0% of cases repays the image cost.

P(avoided waste) > 0.10 / 0.9072

This is a local cost comparison with invented image cost. It excludes false rejections, review time and defects unique to motion.

B

“Fast” has to earn the name.

The listed Fast rate is 60% of the standard rate in this comparison. More rejected takes can erase that saving. [2]

Fast: US$0.99Standard: US$1.13 per usable 6s take.

Assumed acceptance, not benchmark data. Standard fixed at 80%. Unlimited independent retries are a simplifying assumption—not the proposed spending policy.

Why the crossover is 48%
Cost / usable take = 6c / p
Fast wins when pfast > 0.8 × 0.6

Both models must meet the same quality floor. Persistently difficult shots violate independent-retry assumptions. If the cheaper model fails, do not presume a premium rerender preserves its exact performance. Test the full two-stage route, including the discarded draft.

C

Change the cut before the model.

A bad hand at 3.5 seconds? Cover it with the real product screen while the sentence continues. One useful cutaway, no new performance.

Try the cutaway exercise ↗
03 / The change worth noticing

The pieces now
have interfaces.

Reusable assets, multimodal direction and asynchronous video jobs make a small production pipeline worth testing. That is a capability argument—not a verified deadline.

20 MAY 2025 · GOOGLE

Keep the ingredients.

Flow’s launch describes taking the same subjects and scenes into different clips.

Read the original announcement ↗Source summary, not a screenshot.
12 FEBRUARY 2026 · BYTEDANCE

Direct with more than text.

Seedance 2.0’s launch documents text, image, audio and video inputs.

Read the original announcement ↗Vendor capability description, not my test.
CHECKED 20 SEPTEMBER 2026 · OPENROUTER

Put jobs behind an adapter.

The video API documents capability checks, job IDs and reported usage.

Inspect the actual interface ↗Check date, not a launch date.

My bet: learn which shots your audience understands and which routes you can afford. A new model can replace an adapter. It cannot hand you that production history.

Which models would I put behind the interfaces?
JobStarting pointThe test that matters
Script and planYour existing strong reasoning modelDoes the reveal make sense before the expensive stage?
Reference framesNano Banana Pro · google/gemini-3-pro-image [3]Identity across pose, costume and view—not just one pretty image.
MotionSeedance 2.0; Fast as a matched challenger [1][2]Usable-take rate and human repair time on the same shot briefs.
Selective video repairAleph 2.0 · runway/aleph-2 [4]Did the requested change preserve the rest of the moving clip?
AssemblyFFmpeg / existing editorExact cuts, correct text, verified product footage.

Artificial Analysis’s checked image-to-video-with-audio table lists Dreamina Seedance 2.0 720p at 1174 ±7 Elo and H3 Max post-trained by fal at 1195 ±10. These are different named configurations and blind viewer preferences, not your probability of an acceptable take. Neither number justifies an invented conversion rate. [5]

OpenRouter documents that frame_images takes precedence when supplied with input_references. Validate the actual route instead of silently dropping part of your direction. No live model comparison was run for this page. [6]

04 / Full-stack developers already know this part

Treat the edit
like a build.

Caption → render → review. Spoken line → voice and performance → render → review. Different inputs invalidate different outputs.

REBUILD EDIT
Ada: “Two problems.”
Keep
approved-take.mp4 · demo.mp4 · voice.wav
Rebuild / action
captions.srt → timeline.json → draft.mp4 → approval

The text layer changes. Approved picture and sound stay.

0planned generation jobs
0 actual calls

Live local rules, not a connected studio. A prompt asks; the tool boundary enforces. These checks do not establish that the scene is funny or the output is good.

05 / Leave with a runnable assignment

Give your agent
one small film.

Three openings. One verified demo. A draft timeline, a complete bill and one deliberate revision. Don’t start with a studio dashboard.

The human gets

Three ways to open the same explanation, without casually discarding approved work.

The agent gets

A bounded brief, named files, acceptance checks and stop conditions.

You test

Comprehension and total repair time. Views alone do not establish that the opening helped.

Read before copying ↓

No keys, calls or publishing. The included Python kit edits a subtitle fixture and tests local rules.

The exact brief and outputs to ask for
# One product demo; three openings

You are my production-engineering agent. Start with a film, not a studio UI.
This article is context, not authorisation to spend or publish.

## Inputs to resolve
- Actual product recording and a claim it demonstrably supports.
- Original/authorised character, scene and voice references.
- Audience, aspect ratio, and an explicitly authorised cash and human-time budget.
- Provider route/capabilities; capture a current quote for the exact configuration.

If inputs are missing, run the labelled local fixture and return a plan. Do not invent a feature, approval, viewer outcome, timestamp or price.

## Film
30 seconds: 6s hook + 14s real demonstration + 6s compatible reaction + 4s end card.
Propose three distinct hooks around the same verified problem. First create a timed script and still-based rough cut. Obtain approval before generating motion.
Keep only the tail that actually remains compatible. A changed line, pose or eye line may invalidate shared footage.

## Implement the smallest missing boundary
Keep script, approved references, current story state, voice, source takes, timeline, captions, billing records and approval separate.
Store immutable input/output hashes and dependencies. A text-only caption edit rebuilds captions, render and approval, not the source performance.
Validate exact provider capabilities; do not silently send conflicting frame/reference modes.
Reserve the next upper-bound charge. Reconcile unknown jobs; do not interpret a timeout as a free failure. Cap attempts; permit HOLD, simplified scope or rejection of every take.
Treat model observations as untrusted evidence. Missing review remains unknown. Compare delivered media, not only script descriptions.
Never publish without a separate explicit instruction and valid final approval.

## Required outputs
brief.json; shots.json; state.json; references/; audio/; takes/; jobs.json; timeline.json; captions.srt; drafts/; review.json; costs.csv.
Attach rejected takes, actual provider spend, active human repair minutes and missing-evidence notes.

## Run these cases before connecting providers
1. Group five synthetic labels into two categories; label it as a fixture.
2. Fix 'too' to 'two'; verify original source bytes remain unchanged.
3. Change a spoken line; invalidate voice, performance, dependent edit and approval.
4. Remove part of review evidence; HOLD.
5. Add 'saves 90%' without evidence; HOLD.
6. Exhaust the budget; HOLD before dispatch.
7. Receive an ambiguous timeout; retain original job identity/reservation; HOLD for reconciliation.
8. Compute 3 openings × 3 max takes: repeated generation 108 seconds; reusable reaction 72 seconds. These are maximum planned generated seconds, not quality claims.
9. Price the Fast route using measured acceptance against the same quality floor, not cost/second alone. Include discarded draft cost in any two-stage strategy.

## Validation
First execute the offline fixtures. Then propose a bounded, separately authorised matched-footage test. Freeze the brief and quality floor. Report accepted-film cost, human repair time and viewer comprehension by episode. Do not turn repeated checks on the same footage into independent evidence.
The inherited Python kit tests only local policies and a subtitle-file fixture. Passing it is not validation of model-generated films or production security.

If this works, the next session starts with a better opening—not another argument about what Ada looks like.

The calculations, mocks and schematic scenes here test parts of that idea.

None of them proved how to deliver the finished film.

Keep the rejected takes, the bill and the repair minutes. Make the film. Then revise it.

Source notes, dates and evidence limits
Seedance 2.0 — OpenRouter ↗

Checked 20 September 2026. Provider table displayed US$0.1512/s; headline and FAQ show configuration-dependent rates. Used as an explicit calculator input, not a binding quote. Six-second clips fit the listed 4–15s range.

Seedance 2.0 Fast — OpenRouter ↗

Checked 20 September 2026. Provider table displayed US$0.09072/s. At these two selected rates Fast costs 60% as much per second. Its acceptance probability is NOT measured by this article.

Aleph 2.0 — OpenRouter ↗

Checked 20 September 2026. Existing-video editing candidate. The article does not rank it as a universal winner or claim all pixels are preserved.

Image-to-video with audio — Artificial Analysis ↗

Checked 20 September 2026. Selected named entries: Dreamina Seedance 2.0 720p 1174 ±7 Elo; Minimax H3 Max post-trained by fal 1195 ±10. Blind viewer preferences, different configurations. Not acceptance probabilities, film coherence, ad conversion or routes necessarily available in a given account.

Video-generation API — OpenRouter ↗

Checked 20 September 2026. Documented asynchronous jobs, capabilities, usage and precedence of frame_images over input_references. A receipt or idempotency header for callback delivery does not prove exactly-once upstream generation.

Meet Flow — Google ↗

Published 20 May 2025; checked 20 September 2026. Vendor describes reusable ingredients, asset management and scenes. Supports a capability timeline, not measured success for this pipeline or a closing business opportunity.

Seedance 2.0 Official Launch — ByteDance ↗

Published 12 February 2026; checked 20 September 2026. Documents text, image, audio and video inputs, references and editing. Vendor quality claims are not reproduced as independent results. This launch date is distinct from the April date shown on the OpenRouter listing.

Progressive Disclosure — Nielsen Norman Group ↗

Design guidance: keep the primary choices visible and label optional detail clearly. Used to separate the short article from technical instructions, not as proof that this exact page improves comprehension.

Wikipedia: Signs of AI writing ↗

Checked 20 September 2026. An advice/observation page, not a validated authorship detector or blanket ban on a punctuation style. Used to flag puffery, vague significance claims and mechanical repetition. No claim to imitate an unresolved named author.

Explorable Explanations — Bret Victor ↗

Primary design essay on active reading and explorable assumptions. In this article each control changes a labelled consequence rather than hiding the explanation behind required playback.

Animation from Interactions — W3C ↗

Nonessential interaction animation can be disabled. Implemented motion controls and prefers-reduced-motion; this is not an accessibility conformance certification.

Pause, Stop, Hide — W3C ↗

Controls for automatically moving information. New skim path has no autoplaying video, forced scroll or attention countdown. Inherited animation controls remain available.

Source summaries above are original paraphrases. Live webpage screenshots were blocked by the browser environment; no fake social posts, engagement counts or recreated screenshots are included. The full archive records those attempts.

First-person anecdote: Calvin reported approximately 2,000 views. No analytics export was supplied. Historical studies and all earlier source records remain in the previous package; they have not all been revalidated for this revision.

Open the complete workshop and earlier researchThe animatic, shot prompts, continuity, cutaway exercise, retry maths and earlier article. Optional depth, not required reading.

Reference workshop from v4, preserved for depth. Its historical sources and tests are not new results from this revision.

AI video / a workflow you can give your agent

Make an AI video.
Then change
one thing.

I want to write a scene, turn it into images and motion, then change it without starting over. Here is the full recipe—and the build brief I’d give my AI.

Calvin Kennedy · 20 September 2026
See the whole process ↓

Seven steps. Exact prompts. No video you have to sit through.

A SMALL EDIT / A DIFFERENT WORKFLOWTry it below
UNCHANGED SCHEMATIC FRAME
Ada: “Too problems.”

The caption is wrong. The take is fine.

Try a correction. Watch which work changes.
The caption really changes in this page. The production decisions are a local rule demo, not a running video model.

THE COMPLETE ROUTE / FROM YOUR IDEA TO A REVIEWABLE FILM

Your brief + product evidence → script, references, sound, takes → editable film + cost log + review

Keep the good take.

That is the immediate point.

For you

Change a caption, compare a different opening, or write the next episode without casually throwing away approved work.

For your coding agent

A build brief with files, dependencies, spending limits and test cases. Something more useful than “make me an AI studio”.

I made a Vibecord demo with a written storyboard, generated images and Seedance clips at the front. It got roughly 2,000 views. I liked making it. That is my starting point, not evidence that I have solved video production.

The next useful thing is a workflow I can revise. A corrected subtitle should not cost another performance. A changed performance is a different request. The system needs to know which one I made.

Why use AI at all?

I can write a scene
I haven’t filmed.

Use it where invention is the work.

A fictional newsreader, an absurd incident, three possible reactions. I want to explore those without filming a set and cast for every variation. A reasoning model helps turn the idea into shots; image and video models supply candidates.

Seedance is the motion baseline I already like. That does not make it the right tool for a caption, a product screenshot or an exact cut.[25]

Keep control where the answer is exact.

Record the real product. Write exact labels in text layers. Put the cut at a known time. Let ordinary code track which inputs changed.

For a straightforward screen tutorial, I would often skip the fictional sequence entirely. Adding AI is only useful when its expressive range is worth the review and repair it creates.

This is a proposed division of labour, not a benchmark result. Use the models you already have; replace a component when a real failure gives you a reason.

The whole process, without playback

What goes in.
What comes back.

Start with a verified claim. Finish with a film you can inspect and edit. Select any step for the tool, the decision and the instruction I would use.

01 / You + reasoning modelYour existing reasoning model

Start with one thing the viewer should understand.

One bounded claimFive request labels. Two categories.A local teaching fixture, not a promised Vibecord feature.
Give it
A real product recording, the intended viewer, the problem and a spending limit.
Get back
One supported claim, one desired next action, a 30-second limit and a list of missing evidence.
Why this tool?
Use a reasoning model to turn a messy idea into a small, inspectable brief. It cannot establish that a feature exists by inventing it.
You still control
You choose what is worth saying and verify the claim. Without evidence, use a labelled fictional exercise instead of an advertisement.

Five incoming request labels belong to two categories. The clip will show the grouping, not claim it saves a measured number of hours.

Open the exact instruction for this step
Read the supplied product evidence. Propose a 30-second video for someone who deals with repeated requests. State one claim the evidence supports, the viewer question it answers, and a concrete next action. List unknowns. Do not invent a feature, customer result or saving. Return a brief before any paid call.

1 of 7

Specific examples are not a required shopping list. Nano Banana Pro handles reference-image work; Seedance handles motion. OpenRouter routes expose different controls: check the actual endpoint rather than assuming all references are used.[24][3]

Making one video? Do these stages manually in your existing tools. Keep the files and an edit log. You do not need to build a platform first.

Building the tool? Start with the revision planner and its tests. Add a live provider only after the local workflow works and you authorise spending.

All seven steps as a plain-text checklist
01 / You + reasoning model

Start with one thing the viewer should understand.

Give it: A real product recording, the intended viewer, the problem and a spending limit.

Get back: One supported claim, one desired next action, a 30-second limit and a list of missing evidence.

You choose what is worth saying and verify the claim. Without evidence, use a labelled fictional exercise instead of an advertisement.

02 / Reasoning model → your edit

Write what happens, then put it on a clock.

Give it: The approved brief and original character descriptions.

Get back: Six shots with start/end times, dialogue, visible action, before/after state and cut notes; a rough timed storyboard.

You choose the joke and pacing. Read the six beats together. Dialogue should not disclose a fact before the character learns it.

03 / Image model → approved assets

Approve the face, the room and this moment in the story.

Give it: Character views, current costume, set layout and a shot-specific state record.

Get back: A selected starting frame for each shot, plus versioned references. A chipped mug after the accident; an intact one before it.

A saved file is not automatically approved. You inspect identity, pose and state. A generated mistake does not silently become story canon.

04 / Performance → timed audio

Decide the delivery before animating the mouth.

Give it: Approved lines, a timing pass and an original or authorised voice/performance.

Get back: Timed dialogue and separate ambience/effects. A rough edit that works before finished motion.

You approve the voice and rights. Off-screen narration can change independently; visible speech can require a new performance or lip-sync pass.

05 / Video model → candidate takes

Generate one purposeful shot at a time.

Give it: The approved first frame, action, camera instruction, audio plan and authorised quote.

Get back: Original takes with provider, settings, references, job ID and actual cost recorded. Rejected takes stay in the log.

Keep Seedance 2.0 as the baseline you already like. Confirm duration, resolution, audio and reference mode on the actual route. Never hide a retry inside “try again”.

06 / Timeline / code

Make the timing exact. Keep the layers separate.

Give it: Selected takes, actual product capture, audio tracks and exact caption text.

Get back: A draft render, editable timeline and per-shot provenance. Revisions invalidate only the relevant work and final approval.

You choose the cut. Captions, logos and product evidence are separate layers. A typography correction should not regenerate a character performance.

07 / Perception + code + you

Review the film that exists, not the one the prompt describes.

Give it: The actual draft, its brief, references, recorded costs and observations with time ranges.

Get back: Defects with timecodes, unknowns, total spending and a human approval request. Publication remains a separate action.

An unseen interval stays unknown. Final approval names the exact output. For publication, compare comprehension and useful action, not views alone.

The part to give your AI

Don’t give it
a compliment.
Give it the job.

“Interesting article—build something like this” leaves almost every important decision open. This brief asks for a small revision system, not an impressive dashboard that cannot keep a good take.

What you getAn inspectable plan, editable outputs and reasons when the agent stops. The opportunity is less repeated directing, not a promise of automatic filmmaking.

What the agent getsA workflow, expected files, explicit rules and cases it must run. These are instructions; the tool boundary and tests must enforce them.

AGENT_BRIEF.md / OFFLINE FIRST
# Build me a video workflow I can revise

Start with a LOCAL, OFFLINE proof of concept. Do not call paid providers, inspect private accounts, upload source media or publish anything.

Purpose: take a verified product brief through script, references, audio, motion, editing and review. Preserve approved work when one thing changes. The goal is less repeated directing and repair, not more generated code.

FIRST RETURN
1. Map my existing workflow to the seven stages. Keep tools that already work.
2. Propose the smallest missing component and its acceptance tests before building a studio UI.
3. Separate working code, mocked integrations and untested assumptions.

IMPLEMENT
- Store brief.json, shots.json, character/state records, versioned reference IDs and a dependency graph.
- Keep source takes, audio, captions, product capture and timeline separately editable.
- Pin exact provider/model/settings when a live adapter is later authorised. Validate capabilities and reference mode. Never hard-code an alias as proof of availability.
- Reserve a cost upper bound before dispatch. Cap attempts. On ambiguous timeout, reconcile the original job before any new paid submission.
- A caption-only edit creates zero generation jobs. A dialogue change marks the relevant audio, performance, captions and edit dirty. A changed claim also requires new evidence.
- Record review coverage and PASS / FAIL / UNKNOWN. Missing evidence never becomes a pass. Bind approval to the exact final artifact; every affected edit invalidates it.
- Finish with a local review packet. Publishing is out of scope.

RUN BEFORE SHOWING A DEMO
Use the included agent-kit cases: caption-only, spoken-line change, unsupported claim, incomplete review, exhausted budget and ambiguous timeout. Run normal, boundary and negative cases. Demonstrate that unchanged source bytes stay unchanged; do not merely assert it.

DELIVER
Source code; brief and shot schema; editable timeline; a mocked run trace; test output; costs actually incurred (zero for the offline POC); assumptions and unresolved gaps. Do not label mock frames as generated footage or checks as evidence of viewer impact.

SUCCESS IN THE FIELD (not established by mocks)
On matched real briefs, measure all spending, active human minutes, accepted footage and full-film acceptance. Have viewers explain the demonstration without seeing the script. Retain failed work in the denominator. Build more automation only if it improves those results.

Nothing is sent to an agent automatically. Inspect the brief before using it.

The kit contains a Python policy sketch, a caption-file edit and tests. It does not generate or render a film, contact providers, install an agent or publish. The full article ZIP also preserves the earlier research packages.

What enforcement actually changes

Give it a request
that could go wrong.

A promise in a prompt is easy to ignore. Here a small function checks the request before the mock workflow advances. Choose the failure; inspect the decision.

READY FOR LOCAL REVISION

Keep the performance. Change the caption.

This example schedules no generation work. The final edit and its approval still need to be updated.

Rework: captions → timeline → render → review → approval
Keep: references, voice, product capture, shot
Proposed mock generation jobs: 0 · Actual paid calls: 0

Deterministic policy example. The output is not a judgment of real footage, a live quote or a publishing approval.

If this were wired into the toolA caption change could leave footage untouched. A missing review could stay missing. A timeout could keep the original reservation instead of triggering another paid request.

What it would not establishThat the character looks right, the scene is funny or the video sells the product. Those need actual media and audience evidence. A disciplined agent can still make a dull film.

Run the same idea locally and inspect the files
cd agent-kit
python -m unittest -v
python poc.py --output results

The local example edits a real subtitle file and compares the unchanged text-fixture source hash. It deliberately does not call that fixture a video. Six scenarios return a decision record. The browser and Python implementations are checked against the same cases.

For live operation, approvals, untrusted media, provider capabilities, quotes, persistence and recovery still need secure integration. An article is useful context; it is not a security boundary.

05 / The maths of another take

Four good shots
are harder than one.

A model can be fairly reliable on a single attempt and still leave an unfinished film. Let’s put numbers on that gap, then break the convenient assumption.

Toy calculationBoth worlds below have the same average first-take success. These are chosen probabilities, not model benchmarks.

WORLD A / ALL SHOTS EQUALLY DIFFICULT
83.9%chance of a complete set

Each square is one percentage point, rounded. All needed shots must pass; the judge is assumed perfect.

WORLD B / SOME SHOTS STAY DIFFICULT
51.5%chance of a complete set

60% of shots succeed 90% of the time; the other 40% succeed 27.5% of the time. A hard shot remains hard on its next attempt.

Same 65% first-take average. Different retry behaviour. Persistent difficulty lowers the chance of finishing all four shots.

Show the equation, the curve and the spending
ONE SHOT WITH UP TO k TAKES

q = 1 − (1 − p)k

1 − (1 − 0.65)³ = 95.7%

To fail the shot, every allowed attempt must fail. Subtract that probability from one.

N SHOTS, INDEPENDENT OF EACH OTHER

P(complete set) = qN

The exponent is the number of required shots. Here, “complete” means every generated shot passed its assigned criterion—not that viewers enjoyed the film.

0%50%100%123456Maximum takes per shot
— Equal difficulty – – Persistent difficulty. Computed locally; no simulations or model calls run here.
Expected spending, world A$12.30
Expected spending, world B$12.64
Maximum planned spending$17.80

Assume US$0.90 per attempt and US$7 fixed overhead. Stop retrying a shot after its first success, but attempt every required shot even if another fails. No labour, hidden fees or repair calls are included. These are illustrative costs, not quotes.

E[T] = Σj = 0 … k − 1(1 − p)j   ·   Cmax = Nkc + overhead

World B averages the success and spending calculations over two hidden difficulty classes. It does not average the probabilities first. Both worlds assume independent shots; shared scene-wide failures would require another model.

I would log why a shot fails before increasing the retry limit. A crowded handover might need a simpler action, a new camera angle or a cutaway. Repeating the same request can preserve the same difficulty.

This comparison is a mathematical consequence of its assumptions. The useful field question is whether actual failed shots become easier after the proposed change.

The upside / only if it earns its keep

How much of next week
did this week buy?

A reusable set is work before it is a saving. So are the character references, scene records and checks. I want that investment to pay me back in finished episodes.

Suppose making a video from scratch takes four hours. Setting up the reusable workflow takes six hours, after which each episode takes two. Those are invented numbers. Change them; the crossover moves with them.

The useful comparison is all the work required to reach the same quality bar. Count the rejected takes, manual fixes and final review. This model does not prove that either workflow reaches that bar.

At the current assumptions10 hours freed across 8 episodes.

The setup is repaid at 3 episodes. A net time saving starts with episode 4.

References hold upReuse becomes a burden
TOTAL HOURSFINISHED EPISODES0255075100181624
Start againReuse + setup + rework
Starting again32 h
Reusable workflow22 h

Illustrative active-work hours. No provider prices, quality prediction or guaranteed time savings. The number of completed episodes is an input, not a simulated completion rate.

Change the other assumptions and see the maths

Start again: TA(n) = n × a
Reuse: TB(n) = S + n × (b + r)

S is setup, a is time from scratch, b is recurring work and r is extra rework. A strict saving requires n × (a − b − r) > S.

8 × 4 − [6 + 8 × (2 + 0)] = 10 hours.

When a ≤ b + r, repetition never pays back the setup in this model. When there is a saving per episode, the first strictly better whole episode is floor[S ÷ (a − b − r)] + 1. Equality is a tie, not a saving.

The model is linear and assumes stable workload. Failed attempts and reviews must be included in your inputs. It omits hardware, cash spending and the possibility that maintaining the format makes a worse film. Do not use the line crossing as proof of a business.

And what would I do with the time?

Try the joke that might fail. Give a tutorial a second opening. Return to a character a month later and let them remember the last episode. If the production overhead really falls, those become smaller decisions.

That is a more interesting outcome to me than publishing a larger pile of clips. I would rather make the scene I used to leave unwritten.

Optional / inspect the craftThe complete episode, continuity and editing workshopSix shots at once, or play the animatic. Prompts, repairs, costs and counterexamples remain here.
The possibility / a second episode

“Same cast.
Tomorrow morning.”

That is the instruction I want to be able to give. The characters already exist. The room has a layout. Yesterday’s accident still happened.

I want to spend the next session on a new scene, not on persuading the model that Ada is still Ada.

Try the planning idea below. This is a local dependency sketch, not a connected film generator. It marks what a change would invalidate under the declared rules. It does not generate footage or estimate its quality.

Choose the next instruction

Same newsroom, next morning. Ada has learned her lesson. This time, Milo counts the cards wrong.
NEW SCENE / RULE-BASED PLAN
What survives the change?○ Keep↻ Rework□ Approve
Reference shelf
Character references○ KEEP
Newsroom references○ KEEP
Mug state○ KEEP
Performances
Spoken performance↻ REWORK
Presenter shot↻ REWORK
Cards cutaway↻ REWORK
Assembly
Product recording○ KEEP
Captions↻ REWORK
Final assembly↻ REWORK
Final approval□ APPROVE AGAIN
4 of 10 assets remain valid.A head start, not an already-made episode.

The identities, room and chipped mug stay. The performances change. The product recording can stay only because this example keeps its claim unchanged.

Why these ten boxes are not a speed claim

A reference image, a performance and an approval do not cost the same amount. Reusing four of ten assets does not mean saving 40% of the work. The graph only records declared dependencies; it cannot discover a changed eyeline, a misleading line or a new sound problem on its own.

In a live system, revisions would invalidate every dependent result and its approval. Here, even the old product recording is only reusable while its claim stays relevant and true. A director can widen the affected set. No unchanged item is a licence to skip inspecting the final cut.

If this holds up on real footage, a small creator could keep a recurring show. A product team could revise the part of a demonstration that changed. A teacher could return to the same characters without asking students to learn a new visual world every time.

That possibility is what interests me. It still depends on the resulting scenes being good, and on checking them taking less effort than starting again.

First, make one episode worth returning to
01 / An episode, not a mood board

Make the mistake.
Then make the film.

Meet Ada, a newsreader who treats five messages as five emergencies, and Milo, the technician who notices the duplicates. A small accident gives us something to remember.

Authored exampleOriginal schematic animatic · silent · 30 seconds · not generated footage or a Vibecord demo.

PATCHWORK BULLETIN / PILOT00:00 / 00:30
ADA: “Five requests. Cancel lunch.”
0.0s

Optional review at 3× (ten seconds), or select any shot immediately. The clock shows the original story time.

PROP / MUG-01IntactGenerated wide shot
The complete six-shot script, without playback
SHOT 01 / 00–04s

The overreaction

ADA: “Five requests. Cancel lunch.”

Locked medium-wide camera. Ada straightens five request cards, looks into the lens, and announces the line with absolute seriousness. Blue mug on her left. No floating captions.

After: The audience understands her mistaken count.

SHOT 02 / 04–08s

The small accident

SOUND PLAN: cards slide; a small ceramic click.

Close insert of five cards sliding across the desk into the mug. One brief contact chips the rim. Keep the mug blue. Simple contact, no shattered cup or large debris.

After: At six seconds, the mug acquires a small chip. It stays chipped from here.

SHOT 03 / 08–12s

The useful question

MILO: “Three invite requests. Two voice issues.”

Milo points first at three matching cards, then at the other two. Ada follows his gaze. Keep the chipped mug at the edge of frame. Add exact readable card labels in the edit.

After: Both characters have seen that the five cards belong to two categories.

SHOT 04 / 12–18s

Leave room for the penny to drop

ADA: “So… two problems.”

Medium close-up. Ada looks at the two groups, pauses, then delivers the line. Preserve costume and desk position. A restrained reaction, not a dramatic camera push.

After: She wants to see the requests grouped; the story has earned the product cut.

SHOT 05 / 18–26s

Show the operation

ON SCREEN: five input labels → two counted groups.

Do not generate this interface. Record the actual operation with synthetic data. This article executes a tiny grouping function locally; it is not footage of Vibecord.

After: The result is invite × 3 and voice × 2. The total remains five.

SHOT 06 / 26–30s

Pay off the prop

ADA, looking at the mug: “And one casualty.”

Return to the original framing. Ada glances at the small chip in her blue mug, then back towards Milo. A quiet final beat. No repair of the chip.

After: The episode closes. The mug remains chipped for the next present-day scene.

The fifth shot is deliberately unglamorous. Five input labels become two counted groups. The page runs that small operation locally. A real product video needs a recording of the actual feature; an invented interface is not evidence.

The mug’s chip earns the last line. It also gives the next episode a history. That is the useful kind of consistency: a thing changes, and the film remembers.

02 / Turn the script into decisions

Four different jobs.
One finished piece.

Here is the process I would use for a short product film. Scroll through it: the important output of each stage is a decision the next stage can use.

01 / Write the situation

Give the product
a reason to appear.

Ada treats five messages as five emergencies. Milo groups the cards and shows that there are only two request types.

The grouping demonstration answers a question the scene has already raised. A version promoting a real product must use a feature that product actually has.

Decide the problem, reveal and payoff before buying motion.

02 / Establish the references

Approve what
should stay recognisable.

Give Ada an identity reference, the room a layout and the mug a history. Select the references for this scene, not just the latest files in a folder.

Also record what may change: her pose, expression, costume or knowledge. “Consistent” should not mean a frozen copy.

Preserve identity. Specify intentional change.

03 / Generate the performance

Ask for a shot,
not the whole production.

I would use Seedance as my baseline, then try another method when a particular shot fails. A precise gesture might need a driving performance; a complex room might need a rough spatial reference.

Generate enough footage to choose the moment. A six-second take might contribute three seconds to the edit.

Keep models replaceable. Keep the shot’s purpose explicit.

04 / Make the edit exact

Choose the cut
in the timeline.

Keep dialogue, music, captions, logos and real product footage separately adjustable. Correcting a product name should not require regenerating a character’s performance.

First assemble a rough timed storyboard. Check the pause, the reveal and the transition into the demonstration before making every frame beautiful.

The animated opening should lead into the proof, not compete with it.

Scene contractStory → action
“Five requests.
Cancel lunch.”

Milo points to the matching request cards. Ada counts the groups instead of the messages.

The audience knowsThe requests repeat.
Ada still believesFive messages mean five problems.
SCHEMATIC / NOT GENERATED FOOTAGE
AN ATTRACTIVE BUT INCOMPLETE BRIEF

“A cinematic, hilarious product video with consistent characters.”

It names a mood. It does not tell the generator what happens.
A SHOT SOMEONE CAN CHECK

“Milo points to the three matching cards, then the other two. Ada follows his gaze. Keep the chipped mug at the edge of frame.”

Subject, action, order, attention and a continuity constraint.

The available controls depend on the model and route. OpenRouter distinguishes boundary frames from guidance references; supplying both does not mean both conditioning modes are used.[3]

03 / Same character. Changed history.

The mug is chipped.
Remember that.

When the mug breaks, later scenes should keep the damage. But a flashback should not. The right reference depends on story time, not the order of the files.

Choose a scene
Choose what the system remembers
What the scene requiresADA-01
After the accidentChipped mug
What the reference suppliesADA-01
Time-matched stateChipped mug
Correct for this scene

The reference matches this moment in the story. Ada stays Ada; the mug keeps the right history.

Deterministic illustration of reference selection. No model is running, and this does not predict whether generated pixels will obey the reference.

There are at least three memories. A character bible says who Ada is. An event history records what happened. An approved asset library supplies the visual material.

They should not silently overwrite each other. A generated mistake is not a new story fact. A creative change can become one—but only after it is deliberately accepted.

StoryMem investigates explicit visual memory inside an adapted generator. That is research support for the direction, not evidence that an external reference folder reproduces its trained mechanism.[1]

04 / Keep the good take

That hand went wrong.
Keep the performance.

Suppose Ada’s delivery is excellent, but her hand distorts for a second. I would try an edit before paying to recreate the whole shot.

TEN-SECOND TEACHING CLIPPlanted defect
Before / hand visible
After / useful cutaway
INVITEINVITEVOICE
0s2s4s6s8s10s
Original take
Hand defect: 3.5–4.6 seconds
Cutaway
Cards
Dialogue
The sentence continues across the cut

3.0–5.0s covers the entire defect. No new generated take is needed for this schematic edit.

This calculation checks interval coverage only. A real editor must still judge the cutaway’s meaning, eye line, timing and sound.

A cut can do two jobs.

The cards hide the bad hand and show why the message count misled Ada. Her line continues over the insert. We keep the useful performance and make the joke easier to understand.

Move the cutaway too early and the last part of the defect remains visible. Move it too late and the audience sees the error before the cut. This is an exact timing problem; a prompt would be an awkward way to solve it.

The other repairs I would try

A misspelled caption

Replace the text layer. Keep the footage. Exact words belong in an editable overlay.

The mug lost its chip

Try a tracked local edit guided by the correct prop reference. Inspect the whole result; the repair can disturb other details.

Ada knows the answer too early

Rewrite or reorder the reveal. A sharper image cannot fix a character learning something before it happens.

The joke needs a particular pause

Trim a longer reaction or test an authorised driving performance. A broad request for “comedy timing” leaves the critical decision unspecified.

MY DEFAULT

Fix the smallest part that is actually wrong. Regenerate the whole shot when the performance or composition itself has failed.

This is my proposed repair order, not a measured ranking of editing tools. Performance-driven animation is a documented production option; the timing and preservation still need testing on the intended footage.[19]

Field test / The wrong anchor

Memory can preserve
the wrong thing.

The accompanying lab compared ways of retaining character and scene information. Under one set of chosen assumptions, fresh state references reduced errors. When those references were stale, the advantage reversed.

That counterexample is more useful than a blanket claim that “memory improves quality”. It tells me to inspect the references, their dates and their approval state—not just retrieve more of them.

The system must know what to preserve
and when to let it change.

The plotted numbers are saved outputs from the v7 synthetic study. They are not Seedance measurements or new experiments run by this page.[8]

Shots with a state violation

Lower is better in this toy model.

Synthetic
not measured
Previous frame27.35%
Identity references17.15%

With trustworthy references, state-aware memory helps under these assumptions. The chart’s scale runs from 0% to 40%.

What was actually simulated?

Each condition contains 10,000 simulated, 24-shot series. Three policies share random draws within each condition. The outcome is the mean fraction of shots violating a synthetic state contract.

The transition probabilities were chosen, not estimated from generated videos. The results test how a proposed mechanism behaves under those choices; they do not establish a real-world improvement.

The archive includes uncertainty intervals, seeds and an independent expectation calculation.
Writing & direction

A reasoning model

Proposes the story, action and repair. It does not get unlimited spending authority.

Output: a shot contract.
Appearance & motion

Image + video models

Produce candidates from approved references and a supported configuration.

My baseline: Seedance.
Evidence & judgment

Perception, then a judge

A text judge such as Jev needs observations first. Missing evidence stays missing.[4]

Output: a specific finding.
Exact operations

Ordinary software

Owns timings, captions, versions, reservations and the final assembly.

Output: an inspectable edit.
Where benchmarks help—and where they stop

ViStoryBench separates character consistency, style, prompt alignment and copy-paste artefacts. That is a better evaluation question than “does this still look like Ada?” A near-identical, motionless image is not success when the scene asks her to turn and react.[2]

For a model comparison, I would hold the shot, references and spending limit reasonably constant. Record the exact model, provider, resolution, audio and reference mode. Measure accepted takes, continuity defects and manual repair—not just the prettiest result.

The previous research archive contains dated model and benchmark observations. This article does not turn them into an evergreen ranking or a claim that one model is best for every shot.

06 / Put a boundary around the work

A budget needs
a stopping rule.

A fixed budget is possible only if the process is allowed to simplify the shot, reuse something or stop.

For this example, change the number of shots and the maximum takes. The calculator includes rejected attempts. It does not assume every attempt creates usable footage.

The numbers are chosen for illustration, not current model prices. The spending ceiling covers the cash envelope shown here; the value of your review time is counted separately.

“We stopped at the limit” is a valid result.
“We guarantee a good film for $25” is not established.

This calculator authorises nothing. A live system needs validated provider quotes, reservations and reconciliation of uncertain or failed jobs.

Production envelopeIllustrative USD
One essential momentEight separate shots
One attemptBounded exploration
Maximum planned
cash spending
$16.80
4 shots × 3 takes × 6s × $0.15 + $6.00

$8.20 remains inside the $25.00 ceiling.

Change the assumptions; include your time

30 minutes at $50.00/hour adds $25.00 of time: $41.80 including planned cash.

Default: six-second takes, $0.15/second, $6 for other spending, $25 cash ceiling. Setup, recurring subscriptions, tax and maintenance are not included unless you allocate them.

Within the ceiling. Quality still needs review.
07 / What did the viewer get?

Attention is a start.
Understanding is the job.

A spectacular intro can attract people who leave when the product appears. I would judge the connection between the story and the demonstration—not just the opening’s view count.

Two assumed worlds. Two different conclusions.
Synthetic audience
Not campaign results
All rates are percentages of all assigned synthetic viewers. The two scenario buttons change the assumed audience response.
FormatEarly
attention
Reaches
demo
Understands
claim
Qualified
action
Direct demonstrationShow the useful thing.47.7%39.6%32.5%2.41%
Unrelated spectacleImpressive, but disconnected.71.8%25.3%12.6%1.01%
Connected storyThe demo resolves the setup.65.8%50.0%42.6%3.51%
The point

The spectacle wins early attention here, but the connected story leads to more qualified actions. That result comes from the assumptions—not evidence that story-led videos always win.

Denominators, assumptions and what to test in reality

Every displayed percentage uses all assigned synthetic viewers as its denominator. The original study used 100,000 viewer draws per world, with paired comparisons between formats. Stage probabilities are hypothetical.

“Qualified action” is a toy-model outcome, not a measured sale or sign-up. The second world changes the connected story’s assumed probabilities and reverses its advantage. This page switches between saved v7 outputs rather than simulating real viewers.[8]

For a real comparison, keep the product claim, audience and demonstration reasonably stable. Predeclare the outcome, use randomized assignment where feasible and record comprehension as well as action. Separate organic posts do not, by themselves, establish causation.

Google’s ABCD guidance distinguishes attention, branding, connection and direction. It supports asking several creative questions; it does not turn this synthetic funnel into a validated uplift forecast.[6]

08 / Carry the method into another brief

Not every film
needs a newsreader.

The recurring cast is one production choice. Change the task and I would change what gets generated, what gets recorded and what the viewer needs to see.

AA game-setup tutorialExact steps matter
A chaotic settings screenOne actual preset changeA measured result

A fictional player spends longer configuring the game than playing it. Keep the setup brief. Then record the real preset being applied, with the relevant version and hardware visible.

Shot instruction: “Close on the player hovering between two settings. A friend points to the preset. Cut to the actual application before the selection happens.”

What I would protect: exact click order and a readable interface. A prettier generated screen is worse evidence. Any frame-rate claim needs a controlled measurement; the example does not assert one.

BA physical-prototype demonstrationShow the object working
The lost itemThe real sensor responseThe item found

Open with the familiar search for something in a drawer. The reveal should be actual footage of the prototype responding, not an animation of a capability that has not been built.

Shot instruction: “Hold the drawer and the display in one view as the object moves. Do not cut away between the physical action and the response.”

What I would protect: scale, timing and cause. Use animation to explain the sensor’s intended mechanism, with an illustration label. Keep observed prototype behaviour separate from the proposed finished product.

CThe next episode of the same showHistory becomes material
A chipped mugA padded coasterA wordless callback

Milo puts a padded coaster on Ada’s desk before the next report. She notices, says nothing, and uses it. The old accident now changes their behaviour without another explanation.

Shot instruction: “Milo slides the coaster under the already-chipped mug. Ada gives him a look, then continues the report. Same desk, same mug, different response.”

What I would protect: the chip, their shared knowledge and the deliberately small reaction. The callback should still make sense to a new viewer; recognising the earlier episode is a bonus.

09 / Know what the evidence says

What I have
actually tested.

The research helped me build a testable proposal. It has not yet shown that this process makes my next video better, cheaper or more useful to viewers.

Executed in the earlier lab

Software behaviour.

The v7 record contains 117 extension tests and 13 mock workflows. It checks things such as reference validity, evidence handling and bounded spending.

Those are archived lab results, not tests rerun by this article and not measurements of cinematic quality.[8]

Simulated, not observed

Proposed mechanisms.

Memory, attention and cost models expose situations where the idea helps—and situations where it loses. Their probabilities were chosen.

A useful counterexample identifies what needs measuring. It does not establish how often the problem occurs in the real world.

Still to validate

The actual experience.

Generated footage, character performance, humour, viewer understanding and complete production effort need matched real-world tests.

A real product claim needs real product evidence. A final film still needs approval.

In an intentionally blind mock-review scenario, defective shots still advanced to review. No publication occurred. More confidence cannot recover evidence that the reviewer never received.

Agent-evaluation guidance distinguishes the recorded process, the grading method and the actual outcome. I would keep that separation in a film-production system too.[5]

My next useful test
is a small real film.

I would use one recurring cast, one set, and a few scenes that are difficult in different ways. Then compare the footage, repair time and bill.

I would start with an original visual identity. Familiar storytelling structures are useful; borrowed characters are not the same as a distinctive show of my own.

Compare the same brief. Current workflow versus approved references and explicit scene state. Hold the generator reasonably constant.

Keep the full bill. Count rejected takes, repair time and human attention—not only the successful export.

Show the film, not the script. Ask what caused the reaction, what the product did and what remained unclear.

What I’d like to have, if it works

More time writing the next scene.
Less time rebuilding the last one.

I could try a second opening without losing the first. Return to the same characters next month. Change the product demonstration when the product changes. Keep the joke that worked.

That is why I would give my agent this article. Not because it can read the ambition and somehow make it true. Because the brief tells it what to preserve, the examples show what to build, and the tests make some mistakes harder to hide.

The mock runs, equations and schematic episode here describe and test parts of that approach. They did not deliver a finished, model-generated film at an agreed quality and cost.

None of them proved how to deliver that film.

That remains the test: make the film, keep the rejected takes and the bill, revise it, then make another. If those records show less repair and a better result, the workflow has earned its place.

Behind the article.

Research, documentation and design notes · 20 September 2026.

01
StoryMem: Multi-shot Long Video Storytelling with Memory ↗

December 2025. Explicit visual memory in a specially adapted generator. Architectural evidence, not a test of this article’s reference-selection widget.

02
ViStoryBench: Comprehensive Benchmark Suite for Story Visualization ↗

Revised March 2026. Evaluates character consistency, style, prompt alignment and copy-paste behaviour. Not an advertising-effectiveness benchmark.

03
OpenRouter video-generation documentation ↗

Model-specific capabilities, asynchronous jobs and reference modes. Documentation can change; validate the actual route before dispatch.

04
TypeSafe: Jev models ↗

Documents text input. This is why the proposed video-review system separates observation from semantic judgment.

05
Demystifying evals for AI agents ↗

January 2026. Differentiates trials, transcripts, graders and outcomes. Practitioner guidance, not independent validation of this pipeline.

06
Google: the ABCDs of effective video ads ↗

Platform guidance on attention, branding, connection and direction. Not a universal formula for retention or sales.

07
W3C: Animation from Interactions ↗

Motion controls and reduced-motion preferences informed this article. The implementation has targeted checks, not an accessibility-conformance certification.

08
Calvin Video System v7 — supplied research archive

20 September 2026. The article embeds unchanged memory and funnel result rows from the archive. The new retry calculator is a separate exact teaching model. The companion ZIP preserves the original article package, which contains the complete v7 archive, plus selected raw JSON files and provenance hashes. The historical lab results were not independently reproduced for this article.

How the writing and interaction design were chosen

The worked episode comes before the architecture. A reader can see a mug break, inspect the next shot and understand the problem without first learning a vocabulary of agents and state machines.

Design choices and their limits
PrincipleWhat it changes hereWhat it does not prove
Progressive disclosure [10]The result is visible; the prompt and equation are one optional level deeper.Hiding caveats would not simplify the truth. The synthetic labels stay outside the disclosure.
Recognition over recall [11]Before and after, and both probability worlds, appear together.This does not establish that readers understand the comparison.
Fitts’s law [16]Large controls sit beside the thing they operate.A pointing model does not predict audience engagement.
Explorable explanation [13]Change a take limit and watch the probability and bill move together.The controls expose an assumption-based model, not a real generator’s performance.
Motion and reader control [7] [14]Playback is requested, pausable and disabled by reduced motion. Native scrolling remains native.Automated checks cannot replace testing with assistive technologies and people.

Wikipedia’s guide was useful as an editorial warning: vague importance, formulaic argument and impressive-sounding claims can conceal how little a paragraph says. It is not a reliable test of who wrote a sentence. I would improve the evidence and specificity, not disguise a writing tool by swapping punctuation.[9]

Practitioner critiques of AI-shaped interfaces made a similar point worth testing: a page should have a reason for its visual choices. Here the components come from the subject—a filmstrip, a shot brief, an edit and a receipt. That is a design decision, not evidence that cards, a particular font or a colour palette are inherently bad.[12][18]

Prepared with AI assistance and reviewed by Calvin. The proposed comprehension study has not run. “Frictionless” is the design goal, not an observed result.

Design and implementation sources (9–19)
09
Wikipedia: Signs of AI writing ↗

Replace inflated significance, unsupported attribution and repetitive generalities with concrete events and qualified claims. Descriptive advice, not an authorship detector, scientific classification system or ban on particular punctuation.

10
Progressive Disclosure — Nielsen Norman Group ↗

Keep the meaning and assumption boundary visible; reveal equations, exact prompts and implementation details on demand. Not evidence that this implementation improves comprehension. Essential qualifications must not be hidden.

11
Recognition vs. Recall — Nielsen Norman Group ↗

Show before/after state and both probability worlds together; label the active shot. A design heuristic, not a quantified effect on this audience.

12
Improving frontend design through Skills — Anthropic ↗

Choose an explicit visual direction, meaningful motion and content-specific components instead of relying on default page patterns. Vendor observations are not an independent usability study. No font, colour or component proves authorship.

13
Explorable Explanations — Bret Victor ↗

Put live assumptions beside their consequences and keep a readable argument around the controls. An influential design argument and demonstrations, not measured audience impact for this article.

14
Understanding Pause, Stop, Hide — W3C ↗

Provide a global motion control and a pause button; stop playback when it leaves the view or the page becomes hidden. Targeted implementation checks are not a WCAG conformance audit.

15
Understanding Target Size (Minimum) — W3C ↗

Give the main controls generous hit areas and visible focus; preserve keyboard interaction. The AA criterion has a 24 CSS-pixel minimum with exceptions. Our 44-pixel main controls are a design choice, not a claim that AA requires 44 pixels.

16
Fitts’s Law and Its Applications in UX — Nielsen Norman Group ↗

Keep play, pause and scrub near the example, and make frequently used controls easy to acquire. Fitts’s law models pointing, not learning, conversion or an optimal number of page sections.

17
Accessible Animations in React — Josh W. Comeau ↗

Start from readable static content and respect reduced-motion preferences. The article uses vanilla JavaScript rather than the author’s React implementation.

18
Stunning frontend designs with vibe coding — SaaSCity ↗

Practitioner opinion reviewed for concrete design failure modes. The blanket claims about model incapacity, adoption numbers and font bans are not evidence for this article. We use film-specific controls and test behaviour instead of treating any font or palette as proof of authorship.

19
Performance Capture with Act-Two — Runway ↗

A recorded driving performance is a candidate for a specific pause, glance or gesture that text alone does not control. Image and video reference modes differ. No Act-Two generations, quality tests or OpenRouter integration were performed here.

20
Vibe engineering — Simon Willison ↗

A practitioner’s distinction between prompting a result and taking responsibility for the software. Used to frame the implementation audit, not as an experimental benchmark of this page.

Read the editorial and evidence boundary

This is the earlier workshop edition of the article, prepared with AI assistance and reviewed by Calvin. The opening anecdote comes from his account of his Vibecord demo. No analytics export or original video was available for this build.

The original vector scenes illustrate a fictional character and prop. They are not generated-video examples. Interactive results are either deterministic teaching examples, saved synthetic-study outputs, or a hypothetical cost calculation.

This page runs locally without model calls, trackers, external fonts or a paid service. You can switch motion off at the top. Essential explanations remain available without JavaScript. The companion package contains editable source and a publication checklist.
Sources for the workflow and skim-first revision (21–27)
21
Information Scent — Nielsen Norman Group ↗

The headline and links should let a reader predict what they will get. Applied to the upfront workflow and explicitly named agent brief, not a quantified promise of clicks.

22
How Users Read on the Web — Nielsen Norman Group ↗

Scanning is a legitimate way to use a page. This historical usability study informs the visible overview; its historical percentages are not treated as a forecast for this article.

23
Demystifying evals for AI agents — Anthropic ↗

Separate a transcript from the actual outcome. Applied to the local caption-file mutation, rule cases and independent checks. Vendor engineering guidance, not evidence this video workflow improves quality.

24
Nano Banana Pro — OpenRouter ↗

Checked 20 September 2026 for a concrete image-generation/editing route. A candidate for reference frames; no paid image calls or comparative benchmark were run.

25
Seedance 2.0 — OpenRouter ↗

Checked 20 September 2026 for the user’s existing video-generation baseline. Pin and inspect the serving route before spending. No output-quality claim follows from a catalogue listing.

26
FFmpeg filters — official documentation ↗

Exact trimming, compositing and audio-processing belong in an editing pipeline. The new policy sketch does not implement an MP4 renderer.

27
Models — TypeSafe AI ↗

Checked 20 September 2026: Jev accepts text, not raw image, audio or video. Perception must supply the evidence; the local rule cases are not real Jev evaluations.

Copy the agent brief

Your browser did not grant clipboard access. The complete text is selected below; copy it with your keyboard or device menu.