← Back to writing

Three openings. One demo.

I used a storyboard, image generation and Seedance to open a Vibecord demo, and it got about 2,000 views. For the next one: record the product once, generate only the openings, keep story state and let an agent rebuild only what a revision changes.

I used a storyboard, image generation and Seedance to open a Vibecord demo. It got about 2,000 views. Now I want the next version to keep what I already liked.

Say you built a small support tool. Three players say they can’t join; two report lag. It groups the five reports into two categories. That is the useful bit the viewer needs to see. The opening can be an overdramatic newsreader, an irritated player or no fiction at all.

I would use AI to make scenes I haven’t filmed. I would record the product, not ask a model to invent its screen.

The grouping tool here is a runnable teaching example, not a Vibecord feature claim. The three finished variants are a proposed production arrangement, not delivered AI films. What follows is a worked plan, local tests and visual maths, not generated footage.

Three thirty-second edits share one product recording; only the opening changes

One product recording, three edits. A schematic, not generated footage.

Each edit is a 6s opening, the 14s product recording, a 6s reaction and a 4s end card. The outlined openings are new in each edit; the rest can stay. *Only if the reaction still fits.

  1. News report: “Five emergencies. Two actual problems.” Ada points to the reports. Milo slides the duplicates together. Cut to the real operation.
  2. Player complaint: “Why are there five reports of this?”
  3. Straight demo: “Five reports. Watch them become two groups.”

The files between the jobs are what make a revision possible

A new prompt can undo a good take. These are different jobs: a model proposes a performance, and an editor chooses its timing. I would work through them in three passes and keep what each stage produces.

Pass 1: proof, script and frames

Decide what the scene has to communicate. Record the grouping. Write the overreaction. Approve Ada and her room. A strong reasoning model is useful here because a vague idea needs choices, not just prettier adjectives.

  1. Record the proof (you and a recorder). This tiny local tool groups three “can’t join” reports and two lag reports into two categories, with no AI call. Record the real operation for a real product. The example is not a claimed Vibecord feature. Keep: demo.mp4 · brief.json
  2. Write the scene (a reasoning model). Ada mistakes five reports for five emergencies. Milo notices the duplicates. The opening gives the product recording a reason to appear. A strong reasoning model can propose alternative beats; you choose the one worth making. Keep: shots.json · script.md
  3. Approve the frames (an image model). Keep the same fox, jacket and room. Select the mug state for this story time. Reject a wrong starting image before spending on a moving version of it. Keep: references/ · state.json · approvals.json

Pass 2: sound and motion

Find the delivery before buying the take. A pause can make the joke. Time the stills and words first, then animate the approved frame. Generate enough footage to choose a moment, and keep rejected attempts in the bill.

  1. Time the sound (voice and editor). Put approved stills and dialogue on a clock. Change a pause now, while it is cheap. Use an original or authorised voice. An exact on-screen performance may need performance transfer or lip-sync work. Keep: audio/ · animatic.mp4 · timeline.json
  2. Generate takes (a video model). Use the approved starting frame and one observable action. Keep Seedance as a baseline, cap takes and record rejects. Check the actual route: unsupported controls do not become supported because the prompt mentions them. Keep: takes/ · jobs.json · cost.json

Pass 3: edit and review

Make exact changes with exact tools. Keep the product recording and captions separate from the generated scene. A text correction should change text. A new spoken line is a different request.

  1. Make the edit (timeline or code). Swap the six-second opening. Keep the fourteen-second product recording. Reuse the reaction only when its line, emotion and eyeline still fit. Write captions in a separate layer. Keep: timeline.json · captions.srt · drafts/
  2. Review the film (perception and you). Inspect the delivered motion, sound and transitions. Ask a viewer what the product did without handing them the script. Missing footage or uncertain observations remain unknown; final approval belongs to the exact export. Keep: review.json · costs.csv · approval.json

Models make candidates; the timeline keeps the decisions. This is a suggested division of labour, not a measured quality gain. For a clear screen tutorial, skip the fictional cast. Seedance supports image-led generation; the actual endpoint decides which controls are available.[1] [3]

The exact instruction for each stage
  1. Record the proof
    Inspect the product recording. State one claim the viewer can verify from it. Plan a 30-second film with a 6-second opening, 14-second demonstration, 6-second reaction and 4-second end card. Flag missing evidence. Do not invent functionality.
  2. Write the scene
    Propose three opening scenes: a newsreader overreacts to five support reports; a player mistakes duplicates for new problems; a direct explanation. Each lasts six seconds and leads into five labels grouped into two categories. Specify action, dialogue, framing and the cut. Keep the product demonstration fixed.
  3. Approve the frames
    Use the approved Ada reference and newsroom layout. Produce a medium starting frame for opening A. Preserve identity and costume; specify whether the mug is intact at this story time. No captions or product UI baked into the pixels. Return candidates; do not mark them approved.
  4. Time the sound
    Place the approved stills and voice in a 30-second rough edit. Keep speech, ambience and captions separate. Mark pauses and cut points. Report any line that does not fit the allotted time. Ask for a timing decision before generation.
  5. Generate takes
    Plan one six-second opening take from the approved first frame. Record model/provider, settings, references, estimated upper cost and job ID. Validate the quote before dispatch. Hold after an ambiguous timeout; reconcile the original job rather than silently retrying. No paid calls until authorised.
  6. Make the edit
    Assemble three drafts from the selected hooks, shared real demo, compatible reaction and end card. Keep all source takes. Captions and logos remain separate. Return an edit decision list and provenance for each interval. Regenerate no footage for a text-only correction.
  7. Review the film
    Review the actual exports against the approved brief, not the prompt alone. Identify defects with timestamps and mark unseen intervals unknown. Compare comprehension and complete production cost. Correct one caption and verify the source media hashes stayed unchanged. Do not publish automatically.
Which models, and what would make me change one?
JobStarting pointKeep it only if…
Writing and planningYour existing strong reasoning modelThe character has a reason to act, the line fits and the demo resolves the setup.
Reference imagesNano Banana Pro [4]Identity survives a new pose or costume without copying the original pose.
MotionSeedance 2.0; Fast as a matched challenger [1] [2]The take meets the same quality bar with acceptable review and repair effort.
Exact editsExisting editor / FFmpegLabels, timings and source media are correct.
Semantic checksPerception → structured observations → Jev-like questions [7]It detects labelled defects on actual footage. A text judge cannot see unobserved pixels.

Compare the exact model, provider, resolution, audio mode and references. A leaderboard preference is not your usable-take rate. A cheap draft followed by a premium rerender is two paid operations, not a guaranteed upgrade of the same performance.

OpenRouter documents separate frame and reference inputs. Supplying both does not combine their benefits: frame_images takes precedence. Validate the route or compose a suitable starting frame upstream.[3]

Keep identity, current state and story time separate

The mug chips. The next scene should remember. A flashback should not. If the agent picks reference images by “latest file” instead of story time, the chip leaks backwards.

Choosing the latest image puts a chip into scenes before the accident

Ada, the fox presenter, with her blue mug in three scenes. Story order: 07:55, 08:00, 08:10. Edit order: before, after, flashback. The mug chips between 08:00 and 08:10.

SceneMug state it needsRule: use story timeRule: use the latest image
Before, 08:00IntactIntact · correctChipped · wrong time
After, 08:10ChippedChipped · correctChipped · correct
Flashback, 07:55IntactIntact · correctChipped · wrong time

The table is filled by the page’s local state rule; these are authored states, not a test of a video model. StoryMem studies visual memory inside an adapted generator; this reference-selection rule is not that model.[5]

Save the intention separately from the output. An accidental new jacket is not automatically a character redesign.

People remember things too. Ada can believe the reports are five different problems until Milo shows her the duplicates. Don’t give her the reveal early.

The memory your agent should keep

Identity: ADA-01 · green jacket · approved silhouette
08:00 — mug intact; Ada believes five separate problems
08:10 — mug chipped; Ada has seen the grouping
Flashback at 07:55 — retrieve earlier state, not newer file

A character bible stores stable identity and motivations. An event history stores authorised changes. Approved visual references show what those facts should look like. None alone guarantees the pixels obey.

For a prop handover, record who owns it before and after. For a costume change, preserve identity while updating wardrobe. For a secret, separate world truth, character belief and what the audience has seen. A contradiction can be intentional; label it rather than forcing every character to know everything.

Pay to generate only what changes

Each new opening is generated. The product recording is not generated. The shared reaction is generated once, provided it still belongs in every version.

Sharing one reaction cuts generated seconds by a third, if it still fits

Video only, in six-second takes at US$0.1512 per generated second from a checked provider row. Three openings, up to three takes each. These are maximum dispatch costs, not a quote or a complete film cost.

Regenerate opening + reactionUS$16.3318 six-second takes · 108 generated seconds
Generate openings; share reactionUS$10.8912 six-second takes · 72 generated seconds

36 fewer generated seconds.

US$5.44 lower maximum dispatch cost under these assumptions.

Change the price; inspect the arithmetic

Separate: 2kdtc
Shared: (k + 1)dtc

k openings, d seconds per take, t maximum takes and c price per second. These are dispatch ceilings, not expected bills. Both arrangements keep the real demo. With no compatible reaction, the shared plan returns to 2kdtc.

Log the rejected takes and human repair time. Less generated footage does not prove a better film or an equal quality bar.

Images, voice, render and human time are extra.[1]

A cheaper take can be an expensive habit. Seedance 2.0 Fast’s checked row price is 60% of standard’s. It still has to produce footage you would keep.[2]

Fast costs less per accepted take only above 48% acceptance

Six-second takes. Standard is assumed to reach 80% acceptance. Change the assumed Fast acceptance; it is an assumption, not a benchmark.

Cost per accepted take against Fast acceptance. The dashed standard line is constant; the Fast line gets expensive at low acceptance.10% usable100% usable$6$0

Accent line: Fast. Dashed line: standard at an assumed 80% acceptance.

Cost per accepted takeFast: US$0.99Standard: US$1.13. Fast is cheaper under these assumptions.
Where the 48% crossover comes from
Cost / accepted take = dc / p
Fast wins when pfast > 0.6 × 0.8

Unlimited independent retries, the same quality floor and fixed price are simplifying assumptions. At 0% acceptance there is no finite cost per accepted take. Real work must stop at a budget, not retry forever.

Try the cut before another generation. Suppose her hand distorts from 3.5 to 4.6 seconds. Put the useful product cutaway over that interval and keep the voice underneath: a two-second cutaway that begins at 3.0 seconds covers the full defective interval with no new performance. That is an authored timing exercise; a real cutaway still needs to work visually and narratively.

When is an image check worth its cost?
q × avoided generation cost > check cost

At an assumed $0.10 check and a $0.9072 six-second take, the check repays itself above roughly 11% avoided waste. Include review time, false rejections and motion-only defects before using that as a production decision.

Retries are less useful when failure belongs to the shot rather than to a lucky or unlucky draw.

The same average success rate gives a different film when some shots stay hard

Four required shots, three takes each, and a 65% average first-take success. In the second case, half the shots are easier and half are persistently harder, with the same overall mean.

Every shot has the same difficulty83.9%Chance of completing all required shots.
Half easier; half persistently harder63.2%Half the shots: 90% per take. Half: 40%. Difficulty persists across retries.
Equations and what this model leaves out
Equal: [1 − (1 − p)r]n
Mixed: [1 − ½(1 − p − δ)r − ½(1 − p + δ)r]n

Here δ = min(0.25, p, 1−p), so the two shot types are equally likely and share the same overall first-attempt mean. Their difficulty persists across retries. This differs from the earlier workshop’s 60/40 mixture; its old numbers are not substituted here.

In production, diagnose the failure. Simplify the action, try performance input, cut away or stop. A model’s own confident approval is not a perfect acceptance test. Shared blind spots can affect both perception and judgment.

Toy calculations, not measured model reliability. Independent shots, a perfect acceptance check and no shared production failure are assumed.

Make the edit behave like a build

A caption change and a spoken-line change are different inputs. The dependency graph decides what is stale, so the agent does not have to guess which approved work to throw away.

A caption fix rebuilds the edit without new footage; a new line needs a new performance

References and voice feed the source take. The take, the real demo, captions and voice feed the timeline, which feeds render, review and approval. HOLD means no generation is dispatched.

RequestRebuildKeepPlanned generation jobs
Correct “too” → “two”Captions, timeline, render, review, approvalReferences, voice, source take, real demo0
Change her spoken lineVoice, source take, captions, timeline, render, review, approvalReferences, real demo2
A job times outHOLD. Reconcile the existing job. Its charge is still reserved.Everything0
The budget is exhaustedHOLD. Budget exhausted: no dispatch.Everything0
Add “saves 90%”HOLD. No evidence for the claim. Keep it out of the film.Everything0

The page’s local dependency rules fill this table, with 0 actual paid calls. It does not render a film or secure a live provider. Render-only work still costs time; a budget limit is not a quality guarantee.

What I want from this is to try a different opening without losing a good scene, return to the same characters next month, and keep the bill and repair history. The agent needs explicit dependencies, files, stop conditions and tests. Instructions ask for this behaviour; an implemented tool boundary has to enforce it.

The components exist, but the result still has to earn it

I can now test a small, explicit pipeline. That is different from finding an expiring advantage or proving that every developer needs an AI studio.

The pieces are documented. Google’s Flow announcement (20 May 2025) describes reusing ingredients by carrying subjects and scenes between clips. StoryMem (December 2025) makes historical visual memory part of a specifically adapted generator. OpenRouter’s video-generation documentation (checked 20 September 2026) describes capability checks, asynchronous jobs and usage records as real job interfaces. These are attributed source summaries, not screenshots, independent quality tests or a deadline. No social proof or limited-time offer is being invented.

  • The page demonstrates rules and arithmetic: grouping, state selection, dependency changes, timing coverage and conditional cost calculations. No model inference runs in the browser.
  • The research informs what to preserve: identity, story state, visual references, local evidence and a clear division of work. Individual studies test different tasks.
  • The next experiment needs a real film and its bill: compare matched briefs, keep failed takes, count repair minutes and ask viewers what the tool did. Views alone do not establish usefulness.
Measure the outcome without rewarding the wrong thing

For the film, record accepted output at the same quality floor, total provider spending, active repair/review minutes and viewer comprehension. Analyse matched briefs or episodes, not every frame as an independent success.

For an advertisement, first test whether the viewer can explain the claim and its evidence. Randomised exposure, when feasible, is a better test of an incremental effect than assuming people who watched were otherwise identical to people who did not. The earlier archive includes synthetic counterexamples, not real campaign lift.

Cost per accepted film =
(all spending + valued human effort + setup allocation) / accepted films

Failed and abandoned attempts stay in the numerator. A more expensive generation can be worthwhile if it saves enough repair effort; an inexpensive film that communicates nothing has not solved the product problem.

The second film is the test

If this works, I could try the joke that might fail, keep the version that didn’t, and come back to the same cast without rebuilding the room. That is the payoff I care about.

The mock runs, calculations and schematics here test parts of the workflow. They did not produce a finished, model-generated film at an agreed quality and cost. The test is to make one, keep the rejects and the bill, change something and make another. If that record shows less repair and a better result, the system has earned its place.

For coding agents: make the film, then revise it

One small assignment: three openings, one verified demonstration, editable outputs and a complete cost log. Then change a caption and prove that the source footage stayed untouched.

No credentials, paid calls or publication are authorised by the brief.

What the included code really does
python -m unittest discover -s . -p 'test_*.py' -v
python poc.py --output results

The kit edits a real subtitle file, hashes an unchanged text fixture and runs six local decision cases. It is not pretending the fixture is footage. Live generation, media evaluation, accounting persistence, authenticated approvals and rendering need further integration and testing.

Test the real boundary too: invalid parameters, reused provider IDs, ambiguous timeouts, stale references and wrong artifact approvals. The preserved research archive includes the earlier experiments and their limitations; this page does not claim to have rerun every historical test.

The complete earlier workshop is still available: a six-shot animatic, detailed shot prompts, repair examples, model comparisons, time-reuse maths, attention experiments and the earlier research. Historical claims and tests retain their original scope and date.

Sources and further discussion

Sources checked 20 September 2026. The full package preserves prior source registers, code and results. Source summaries are deliberately distinguished from executed software, imagined benefits and untested real-world outcomes.

  1. Seedance 2.0, OpenRouter. Seed provider row: US$0.1512/s. Catalogue also displays configuration-dependent from-prices. A dated example, not an authorised quote.
  2. Seedance 2.0 Fast, OpenRouter. Seed provider row: US$0.09072/s. The 0.6 ratio is this row-to-row comparison, not a measured quality or speed advantage.
  3. OpenRouter video generation. Documents asynchronous jobs, capability discovery and frame/reference modes. frame_images takes precedence over input_references when both are supplied.
  4. Nano Banana Pro, OpenRouter. Reference-image generation/editing candidate. No paid comparison of these models was run for this article.
  5. StoryMem. December 2025: historical keyframe memory in a specifically adapted video generator. A metadata ledger does not reproduce the trained mechanism.
  6. Introducing Flow. 20 May 2025: describes reusable subjects/scenes and filmmaking controls. Capability evidence, not demonstrated savings on this workflow.
  7. TypeSafe model documentation. Jev accepts text, not raw video/audio/images. A perception stage is required. No actual Jev evaluation was run here.
  8. Progressive disclosure, Nielsen Norman Group. Keep commonly needed information visible; defer secondary features. Not evidence that this page has a measured comprehension improvement.
  9. Recognition and recall, Nielsen Norman Group. Keep cues and alternatives visible. Applied here to simultaneous stages, labelled state differences and local explanations.
  10. Animated transitions in statistical data graphics. Two experiments studied graph transitions. Supports testing meaningful transitions, not a claim that more animation always helps articles.
  11. Explorable explanations, Bret Victor. A written argument with inspectable, changeable assumptions. Static reading must still work; interaction should answer a question.
  12. Animation from interactions, W3C. Allow unnecessary interaction-triggered motion to be disabled.
  13. Signs of AI writing, Wikipedia. Used to check inflated significance, vague attribution and empty claims. Stylistic clues do not establish authorship.
  14. Improving frontend design through Skills, Anthropic. Encourages deliberate design choices over generic defaults. Not a controlled usability study or an instruction to imitate an author.
  15. Target size minimum, W3C. Provides minimum target-size/spacing conditions. Automated geometry checks here are not comprehensive WCAG certification.

The accompanying design folder models skim, explore and deep-reading paths for four task-based reader profiles. Their times are hypothetical sensitivity estimates, not measured attention, comprehension or medical characteristics. The real test is time to a correct explanation of the problem, steps and proposed solution, with non-completion reported too.

No tracking. No external runtime. No paid generation.

Published 2026-09-04 · Updated 2026-09-28 · Source edition v6