← Back to writing

What will you actually do next?

A small fitted model shows how one new fact, a missing approval, moves six activity probabilities. Whether that kind of forecast helps someone finish what they meant to do, without deciding for them, is still untested.

Imagine it is 3:30 pm on the day an application closes. Your README, test results and rough budget still need turning into answers, and you are still coding. A forecast could be right that you will keep coding, and still let the important thing slip.

I do not want an AI that becomes excellent at helping me do the wrong next thing. It should know the difference between my habits and the outcome I asked it to help with.

Your plans are one input. Your behaviour is the thing to predict. A personal model learns from what happened before, reads the situation now, and assigns probabilities to what might follow. The one on this page is a fitted synthetic model: no connected accounts and no prediction about you.

One changed fact moves every probability

An agent can see your calendar without knowing whether you will work, reply, travel or rest. A useful personal forecast has to distinguish those possibilities and survive comparison with what actually happens.

Focused work: 81% → 32%.

A fictional developer has thirty minutes to finish a page. The missing approval changes the situation, not their personality. The weights were fitted to 1,353 synthetic examples; they are not personality ratings or numbers supplied by an LLM.

Time available30 minutes
Reported energy70 / 100
Last observed activityFocused work
Task blocked?yes → 1

Applying the learned weights, the blocker changes the work score by −1.81 and the communication score by +0.31; the four other scores also change. Each activity’s score is its bias plus the sum of weight × scaled input.

Focused work
32%
Communication
10%
Administration
19%
Travel
4%
Rest
29%
Leisure
6%

Primary activity over the next half-hour, not an exact task. Each dot is one rounded percentage point; the thin mark on each bar is the same situation with approval recorded. One changed input recalculates all six probabilities, and the total stays 100%.

These are predicted probabilities from a synthetic fit, not observed frequencies.

The proposed payoff is to prepare for a likely next step while it is still useful. Forecasting behaviour and deciding whether to help are different models; neither gets to turn a probability into permission.

The percentage has to show where it came from

Follow the blocked-task input through the actual fitted equation. The controls change model inputs, not a prewritten animation. The arithmetic has four steps:

  1. Encode. A verified missing approval becomes blocked = 1. A guess about your motivation does not.
  2. Scale. Subtract the mean seen during fitting and divide by its standard deviation.
  3. Weigh. The same blocker changes work, communication and the other scores differently, with one weight per activity.
  4. Normalise. Every activity gets part of the same 100%, so changing one score changes the shares.
The same evidence gives the same calculation and the same answer

The same fitted model, temperature fixed at 1. The bars show how the blocked flag changes each of the six activity scores; they are score changes, not probabilities or causal effects.

0Work−1.81Communication+0.31Administration+0.21Travel+0.25Rest+0.55Leisure+0.49

A fact becomes a feature.

The adapter reports whether the task has a known blocker. Here the fixture is explicit. Unknown cannot silently become a clear task.

A typed input
blocked = 1
“Approval missing” → blocked: true

No real language model parsed the message. A production adapter must establish the fact and its provenance.

Focused work: 31.6% at T = 1

All 20 inputs, weights and contributions for this activity

Every contribution below is wkj(xj−μj)/σj. Add the intercept to get the activity score. Most values stay fixed in this illustration; the controls in the next disclosure expose the free-time, energy, meeting and history inputs.

Focused work: intercept +0.3618 plus 20 feature contributions
InputRawStandardisedWeightContribution
hour_sin0.70711.5140-0.1620−0.2453
hour_cos-0.7071-0.3220-0.2571+0.0828
weekend0.0000-0.5998-0.1369+0.0821
meeting0.0000-0.5085-0.2169+0.1103
deadline0.70000.87560.3859+0.3379
blocked1.00001.7398-0.7839−1.3638
free_window0.50000.15910.1553+0.0247
prev_01.00001.69610.0650+0.1103
prev_10.0000-0.5884-0.0120+0.0071
prev_20.0000-0.2523-0.1221+0.0308
prev_30.0000-0.1876-0.0191+0.0036
prev_40.0000-0.5645-0.0654+0.0369
prev_50.0000-0.41650.0907−0.0378
history_age0.0800-1.6200-0.0461+0.0747
history_missing0.0000-0.03850.1311−0.0050
energy0.70000.66620.5215+0.3474
interruptions0.1500-1.2541-0.4403+0.5521
location1.00000.89930.3813+0.3429
device_active1.00000.50160.0773+0.0388
proxy0.0000-0.01490.0451−0.0007
Intercept + contributions0.891793

Weights describe relationships learned from a synthetic generator. They are not universal psychological effects. The six probabilities come from all six scores; the selected row alone is insufficient.

Change the rest of the situation, and the temperature

Focused work: 31.6%. Changing temperature alters the spread, not the evidence or top-ranked activity.

Entropy: 2.279 bits. Greater concentration is not stronger evidence.

x → standardise → score each activity → normalisezₖ = bₖ + Σⱼ wₖⱼ(xⱼ − μⱼ)/σⱼP(k | x) = exp(zₖ/T) / Σᵢ exp(zᵢ/T)
Live scores and probabilities at the selected temperature
ActivityScoreProbability
Focused work0.8920.316206
Communication-0.2720.098749
Administration0.4030.193911
Travel-1.0970.043285
Rest0.7850.284170
Leisure-0.7110.063679

zₖ is a score, not a probability. The softmax denominator makes the probabilities sum to one. Sharper probabilities are not new evidence; T = 1 is the original fitted output. Changing an input is sensitivity analysis, not evidence of a causal effect.

The full proposal estimates what, and when

There is no single score called “you”. The broader proposal estimates a changing distribution over your possible next actions and their timing:

qt(a, τ) = Pθ(Anext = a, ΔT = τ | Et)

Here a is which action, τ is how long until it starts, Et is the evidence known now and θ the learned parameters. That target is an architecture, not an integrated fitted model. It has four parts:

  1. Evidence: what can it know? Your authorised history, current situation, available tasks and missing information. Not a complete description of a person.
  2. Forecast: what might follow? Probabilities across actions and times, including the possibility that none starts within the horizon.
  3. Assistance: would help be useful? A separate estimate of benefit, checking, errors and interruption. It must include doing nothing.
  4. Feedback: what actually happened? Save the prediction first. Score it when an outcome arrives. Update future models, and leave the saved prediction alone.

The model above is a narrower implemented component: P(primary activity during the next 30 minutes | 20 synthetic inputs). Its six activities are not exact task identities, and its output is not fed into the separate timing illustration below.

The complete model contract and candidate mathematical approaches

Et contains only permitted records already available to the system. A feature function φ encodes history, context and missingness. The current activity demonstrator uses a standardised multinomial logistic model; richer sequence or hidden-state models are competing candidates, not assumed upgrades.

xt = φ(Et); pk = softmax(Wxt + b)kJoint timing candidate: qt(a,j) = S(j−1 | xt) ha(j | xt)S(j) = ∏r≤j[1 − Σa ha(r)]

ha(j) is the chance of event a in interval j conditional on no tracked event yet. The survival term prevents spending the same probability more than once. Competing-risk methods motivate this construction; the original clinical application does not validate personal-behaviour forecasts. Method reference

A known scheduled meeting may need only a rule. A learned forecast has to earn its complexity on untouched future cases. A hidden model state would be a statistical construct, not a diagnosis or a verified account of motivation.

Personal context has different lifetimes, so the input is a history rather than a profile. A recurring preference, a current blocker and an outcome reported tomorrow cannot be treated as the same kind of evidence:

  • Slower changing: your context, meaning roles, skills, preferences, goals and recurring obligations. These are candidate data, not identity scores; confirm them periodically and allow correction.
  • Changing now: your situation, meaning available time, recent activity, blockers, place and volunteered energy. Timestamp it and show missingness. Five controls above expose part of a 20-feature fit.
  • Observed later: what you actually did, when it happened, and when the system found out. An owner action is not an agent action, and an unknown outcome is not a failed task.
Tomorrow’s information cannot improve yesterday’s saved forecast

Evidence: “The approval is missing.” Forecast cutoff: 09:00. Source: a fictional owner-task record, arriving at three different times.

When the record arrivesUsed for the 09:00 forecast?Why
08:59, with permissionYesThe fictional record arrived before the prediction and passes this code-level evidence filter.
09:02, with permissionNo: too lateThe source may update a later forecast. It cannot be inserted into the evidence for a prediction already issued.
Access revokedNo: permission blocks itA model cannot outvote a revoked permission. This is an illustrative contract, not a production access-control system.

Excluding a record means unknown, not “approval received”. Optional sensitive data must earn its inclusion and remain separately authorised. This filter is a separate code demonstration; it does not retrofit a new source or an unknown-blocker value into the fitted classifier.

Save the guess, then learn from the miss

Training pairs a past situation with the activity that actually followed. The loss is larger when the model gave that observed outcome very little probability. The retained fit used 1,353 labelled records before day 60, left a gap with no fitting, and tested on 432 records from day 82: 49.8% top-choice accuracy on that future test. This revision does not rerun that training benchmark.

A lower loss on the same case shows the update works, not that future forecasts improve

Suppose the observed activity was one of the six below. The label is invented; replay one unregularised gradient step on this same example. No real observation is added, and the original model and saved probabilities stay frozen while a copy is updated.

Frozen predictionP(Communication) = 9.9%
Updated copy of the modelNot updated
Probability given to the observed outcome9.9% before learning
Loss, then after the teaching step2.315 → —

The label is hypothetical. The original probabilities stay frozen while a copy of the model is updated.

Original model and saved probabilities unchanged.

loss = −log P(observed activity)

Follow the gradient, then test generalisation separately
Training: minimise mean[−log pθ(yi | xi)] + regularisationTeaching step: w′k = wk − η(pk − 1[k=y])xb′k = bk − η(pk − 1[k=y]); η = 0.01

The original retained fit used its own training procedure. This single-step instrument teaches the gradient without regularisation and never replaces the main model. A system learning over time must predict before its label arrives and evaluate candidate updates on later cases. Delayed progressive evaluation reference

Top-choice accuracy does not establish reliable probabilities. Calibration must be checked on held-out observations. The temperature control above sharpens probabilities without new evidence; it is not a fitted calibration procedure. Calibration reference

What happens next is not the same question as when it happens. The timing model below is a separate, illustrative one: it is not fitted to you, and it does not use the activity forecast.

After an hour, some probability stays on nothing happening

Twelve five-minute intervals × three event types, plus no event: 37 outcomes. The event shares of start work (55%), reply (30%) and take a break (15%) are invented.

100%0%0 min60 min

After an hour, 6.9% of the probability remains on no event.

S(t) = (1 − h)ᵗP(first event j at t) = S(t − 1)h qⱼ
All 37 probabilities
IntervalStart workReplyTake a break
0–5 min11.0000%6.0000%3.0000%
5–10 min8.8000%4.8000%2.4000%
10–15 min7.0400%3.8400%1.9200%
15–20 min5.6320%3.0720%1.5360%
20–25 min4.5056%2.4576%1.2288%
25–30 min3.6045%1.9661%0.9830%
30–35 min2.8836%1.5729%0.7864%
35–40 min2.3069%1.2583%0.6291%
40–45 min1.8455%1.0066%0.5033%
45–50 min1.4764%0.8053%0.4027%
50–55 min1.1811%0.6442%0.3221%
55–60 min0.9449%0.5154%0.2577%

No event within 60 minutes: 6.8719%. All 37 outcomes sum to 100.0000%.

This curve is not fitted to you. Losing observation is not the same as nothing happening.

A forecast is not permission to act

The wider architecture connects these components. Authorised context feeds an evidence ledger, the ledger feeds a forecast, and a separate check decides whether help is useful and authorised. Candidate sources are broader than the 20 inputs in the running classifier: calendar, tasks, messages, Git and files, activity, goals, place and device, and optional recovery data. They are candidate sources, not connected accounts, and sensitive records are optional.

  • The ledger preserves when and why a source could be used. For example: observed 08:42, received 08:43, source decision-03, permission yes. If permission is withdrawn, preparation is blocked; permission is a gate, not an input the model can outvote.
  • A known appointment needs a rule first. “Review in 15 minutes?” An unscheduled next step needs a forecast and timing: several possibilities, not certainty.
  • Checking usefulness and authority comes after. The output is to prepare privately: the current preview, agreed changes and the unanswered approval. Sending stays locked.
  • Feedback closes the loop. Save the forecast, observe the outcome, compare, then update. A missing outcome is not a failure, and an agent’s action is not your action.

The personal-model question is narrow: does behaviour history help choose which unfinished step to prepare, and when? A scheduled deadline itself is a rule, and I have not shown that forecasting adds value to the three cases later in this article.

A lower-stakes example, kept separate, shows the mechanics. Its synthetic activity probabilities do not predict grant submission, dispute success or enrolment decisions.

  1. The plan. You have thirty minutes to finish the homepage. The last observed activity was focused work, so continuing is plausible.
  2. The missing reply. The approval is missing. Change that one input and the forecast moves, even though the free time hasn’t.
  3. The smaller job. A separate rule gathers the current preview, decisions and open question. It doesn’t need to invent a personality.
  4. The limit. Ready is not sent. The private packet is available; sending anything still requires a different permission path.
The preparation rule prepares privately only when a call is due, permitted, fresh and not yet prepared

A deterministic pre-call rule for a fictional client review, run against nine fixtures. Every case makes zero model calls and zero external actions. The same function fills this table.

CaseActionReason
Prepare a private brief when dueprepare_privateAssemble the preview and decisions. Keep the unresolved approval visible. Can we approve the logo in today’s call?
Wait when the call is an hour awaywaitToo early. This rule waits until fifteen minutes before the call.
Stop for a cancelled callskipThe call is cancelled. Do not prepare a new brief.
Stop when permission is missingblockedNo permission. Do not read the source or prepare from it.
Reject stale evidencerefresh_requiredThe source check is older than thirty minutes. Refresh before claiming the brief is ready.
Reuse an existing matching briefreuseA matching brief already exists. Do not create a duplicate.
Flag an ambiguous logo replyprepare_privateAssemble the preview and decisions. Keep the unresolved approval visible. The logo reply is unclear. Ask the owner to confirm; do not guess.
Do not send a requested messageprepare_privateAssemble the preview and decisions. Keep the unresolved approval visible. Can we approve the logo in today’s call? Sending stays blocked.
Stop once the meeting startsskipThe call has started or passed. This pre-call runner stops.

The fixtures stand in for already-verified adapter facts in a real implementation; here they are fictional, not actual permission checks. A forecast and a preparation rule are shown side by side; this is not an evaluated end-to-end assistant.

I would build the scheduled version first. Choosing a useful preparation for an unscheduled gap is a harder job for the personal model to earn.

Let code wake the model

Version checks and timers do not need a language model. Interpreting a messy reply might. The boundary should be visible.

Many checks, few reasons to call a model

A calculated scenario over an 8-hour day and 20 workdays a month. Always asking the model is a deliberately wasteful baseline; the second route lets code filter and the model interpret only relevant changes; the third uses the preparation rule alone.

Always ask the model1,920 callsMonthly cost US$6.24
Only relevant changes60 callsMonthly cost US$0.195
Rules alone0 model callsMonthly cost US$0.00*

Fewer calls under these settings. Token-only arithmetic; equivalent quality is untested.

Bring your own rates, include overhead, see when this loses
C = n × (tᵢrᵢ + tₒrₒ) / 1,000,000Monthly difference = Cpoll − Cchanges − extra overhead

After US$5.00 extra monthly overhead, the difference is US$1.04/month. US$30.00 setup takes 28.7 months to recover.

These rates are editable assumptions, not a current provider offer. Include tool charges, billable reasoning, retries, cache charges, tax, hosting and your time. No equal-quality comparison or local-model benchmark was run.

Check current provider prices. Subscription capacity is a separate constraint, not an unlimited API allowance.

Illustrative rates: US$1 / million input tokens and US$5 / million output tokens. *Rules still have engineering, infrastructure and maintenance costs.

The leverage is the job you stop asking the model to do. The cheapest acceptable baseline still has to win the comparison.

The preparation has to give back more attention than checking it costs

A prepared brief might save twelve minutes. Checking it might take eight. Both belong in the account.

Checking time decides whether the preparation saves attention

Assumptions, not results: 12 opportunities a month, 80% useful, 12 minutes saved when useful, 6 minutes of extra cleanup when wrong and 20 minutes of monthly upkeep.

Minutes per month
Useful preparation+115.2
Checking every result−24.0
Extra cleanup−14.4
Monthly upkeep−20.0
Time left for you+56.8

Under these assumptions, the prep gives back 56.8 minutes a month.

Above 6.73 minutes of review per result, this example loses time.

Change all the assumptions and follow each term
G = N[pb − r − (1 − p)h] − m

Npb is the top line. Nr is checking. N(1 − p)h is extra cleanup. m is upkeep. G is the balance. Change the inputs above: every row updates from that equation.

12 × [0.8 × 12 − 2 − 0.2 × 6] − 20 = 56.8 minutes.

This expectation omits serious harms, non-time benefits and correlated errors. A positive result makes an experiment worth considering; it does not demonstrate a saving.

Prediction accuracy is not useful-output probability. The balance needs evidence about actual benefit, checking, cleanup and maintenance.

The same preparation could follow other work: client reviews, project handovers, a study session. Each new job gets a new comparison. I would not treat success with one brief as permission to manage a life.

Feeling productive is not the outcome measure. METR’s July 2025 study assigned 246 real issues across 16 experienced open-source developers. With the early-2025 AI tools studied, work took 19% longer, although participants believed afterwards that AI had sped them up. That is a dated result in one setting, not a claim about today’s models: METR’s February 2026 follow-up identified selection and measurement problems that prevented a reliable current-effect estimate. For this project, the measure is completed, owner-intended decisions and the work around them, rather than a forecast score or a feeling of being organised.

A passing test doesn’t give you time back

The code can be right. The forecast can be accurate. The product can still waste your afternoon. Three different claims need three different kinds of evidence:

  • The code runs. Contracts and numerical checks can establish specific implementation properties. Tested in this release.
  • The forecast holds up. Synthetic worlds expose failures; your changing circumstances remain untested. Synthetic only.
  • You do less work. Compare real preparation, review and rework with the existing workflow. Not established.
Choose the world and the best model changes

Retained v0.3 results, not newly executed trials. Each bar is mean next-activity accuracy; the line across it is the 95% Monte Carlo interval across eight seeds. Within a slice, models use the same cases.

Frozen context
12.5%
Expanding history
21.9%
Recent context
39.9%
Recent trees
35.4%
Routine baseline
22.8%
Adaptive combination
37.9%

Recent context has the highest mean accuracy in this slice. Eight synthetic seeds; not a personal-performance estimate.

Adaptive-set coverage90.4%
All six activities returned18.6%

A set containing everything includes the answer, but narrows nothing. Coverage is not top-choice accuracy. The nominal set-coverage target is 90%.

Exact selected results and case counts
Exact selected result summary
ModelAccuracy95% intervalCases
Frozen context12.5%10.9% – 14.1%6,336
Expanding history21.9%21.0% – 22.8%6,336
Recent context39.9%37.9% – 41.8%6,336
Recent trees35.4%33.5% – 37.4%6,336
Routine baseline22.8%22.2% – 23.4%6,336
Adaptive combination37.9%36.3% – 39.5%6,336

Recent context performs best in the default slice; the adaptive combination is not a universal winner.

The same applies to outcomes that never reach the evaluation record.

Hiding errors from evaluation doesn’t fix them

Exact arithmetic on 100 cases: 60 correct and 40 wrong.

Reported accuracy75%60 correct / 80 reported
Actual accuracy60%60 correct / 100 actual

20 mistakes unreported. They still happened.

“It worked when we saw an outcome” is a different claim. Preserve missing results in the evaluation record.

A likely next move can still let the important one slip

Here is where an hour stops being just an hour. The chance is real; the person and afternoon are invented. Newcastle’s Circular Economy Grand Challenge advertises project grants up to A$10,000 to test a circular-economy idea, with a first-round closing date of 27 September 2026.[1] An award is not guaranteed.

The official closing times disagree, and an assistant must not smooth that over. The university announcement and the rules’ heading say 5 pm. A submission paragraph on the rules page says 11:59 pm on the same date. The scene below plans against the earlier advertised time; it does not resolve the conflict. Confirm with the organiser.[1] [2] Checked 20 September 2026: a dated example, not an evergreen offer or countdown.

A reminder is not a finished application. Today, the known deadline still leaves the work of polishing the demo, finding material, drafting answers, then reviewing and deciding. The bet is that a private draft from existing material shortens that to resolving gaps, then your review and decision. Eligibility, truthfulness and your choice still matter. Missing an application is not the same as losing A$10,000: selection and eligibility are uncertain, and the grant funds a project. The proposed benefit is keeping the application possible, not winning funding.[2]

Earlier preparation turns a 35-minute overrun into 5 minutes of spare room

Imagine it is 3:30 pm on application day, planning around 5 pm. Your README, test results and rough budget still need turning into answers. Change how long you keep coding. Illustrative arithmetic, not a forecast.

Without prep50 min assembly + 20 min review + 15 min reserve
3:305 pm7 pm

35 min over

Prepared earlier10 min assembly + 20 min review + 15 min reserve
3:305 pm7 pm

5 min spare

More codingAssemblyYour review and submit choiceReserveAccent line: the assumed 5 pm cutoff

Under these assumptions, earlier preparation turns a 35-minute overrun into 5 minutes of spare room. It does not establish a successful submission.

The arithmetic, and what it cannot establish
S = (deadline − now) − further coding − assembly − review − reserve90 − 40 − 10 − 20 − 15 = 5 minutes spare

All task durations are invented. Positive slack is not a probability of success. It assumes valid evidence, no additional delays and a completed earlier draft. Zero slack leaves no room beyond the stated reserve. A negative result means this plan does not fit.

A clock and checklist can calculate this; no behavioural model is needed. Prediction must separately demonstrate that it selects a better moment or preparation than a reminder-plus-checklist baseline. Changing this slider does not establish a causal effect.

This is not free time from nowhere. The prepared plan assumes accurate preparation was completed earlier; its build time, checking and upkeep still count in the attention balance above. Nothing is submitted, and no account is connected.

The grant is one of three documented processes with a real consequence. These are not stories about people this model has helped, and the proposed preparations are hypotheses.

The consequence is not the same as the forecast
What closesThe documented factProposed help, ending with your decisionNot demonstrated
The application window: a chance to fund a projectThe Newcastle challenge offers selected teams funding to test circular-economy ideas. Its advertised first-round cutoff is 27 September at 5 pm, with the conflict noted above.[1]You chose to apply. Map existing project material to required answers; leave gaps and eligibility questions visible. You review and decide.A submitted application or grant award.
The dispute response: customer revenueStripe says an unanswered dispute past its response deadline is lost and the funds cannot be recovered. Stripe already provides alerts and automation routes.[3]From an existing dispute alert, prepare relevant, sourced, category-relevant evidence for review: not another reminder, and never invented evidence. You accept or challenge.A timely response can still lose. No payment has been recovered here.
The census boundary: a study decisionNewcastle describes census as the financial withdrawal boundary; after it, a student generally remains responsible for the course cost. Pathways students have different financial consequences.[4]You asked to review enrolment. Assemble current enrolment facts, the applicable date and official options. You choose, not the model.Low activity is not an instruction to withdraw. No enrolment is changed.

The deadline alone does not justify personal AI. Existing alerts are the baseline. The harder claim is that personal context improves preparation or timing enough to outweigh errors and interruption.

What runs here, and what is still a research target

This article runs a 20-input, six-activity synthetic fit; an exact score decomposition; an isolated learning step; an evidence-time filter; and the retained comparison results. Predictions and maths execute locally.

Still a research target: all authorised personal sources, exact task identification, a unified fitted action-and-time model, personally calibrated predictions and demonstrated help. These are not quietly substituted for the running component. A 70% work forecast tells an agent how the model ranks a possible activity; it does not identify the exact task, estimate whether helping works or authorise execution.

A forecast may help an assistant prepare before a consequential choice closes, but a calendar rule could be enough. The learned component has to improve the comparison, not just make the story sound smarter.

The possibility is still worth testing

I do not need a machine to tell me I will probably keep coding. I need to know whether it can help me finish the other thing I meant to do, without taking the decision away.

An application still possible, a payment decision properly reviewed, a course choice made with the facts in front of me: those are reasons to test the idea. This project’s simulations did not show that a real person kept one of those options open. The unproved part is the human benefit, and the public sources establish consequences and opportunities, not that this prototype prevents a missed decision.

For coding agents: build a forecaster, then make a simpler baseline the thing to beat

Give your agent an outcome to earn. It should return a working baseline, source references, stop cases, measured workload and an honest comparison:

  1. A baseline: do the known job with a rule.
  2. A candidate: add prediction only where needed.
  3. A fair test: count misses, checking and rework.
  4. A decision: keep the rule, improve it, or stop.

Useful output is more than passed tests. Permissions and spending limits belong in trusted code outside the agent. Instructions are context; a coding agent able to rewrite its tests has not proved its own constraints.

Reusable project instructions give the experiment a place in the repository: both Codex and Claude Code document routes for loading them, with their own rules and exceptions (documentation checked 20 September 2026). Vendor capabilities do not demonstrate this workflow’s benefit. There is no artificial countdown; the useful move is to attach acceptance criteria before expanding the assistant, and rates, versions and execution conditions still need checking.

The forecasting kit

The actual fitted coefficients, a pure JavaScript forecaster, a frozen record format and executable checks: model.json, demo.cjs, tests, a forecast schema and a build brief. Synthetic only; no API key and no live collector.

Read the full model contract

The preparation kit

The earlier preparation kit remains a separate downstream example. It runs on fictional records and ordinary rules with no model calls, and it still demonstrates only low-stakes private preparation. Start in a new folder and reconcile existing repository instructions before merging anything.

  • AGENTS.md: the job.
  • policy.cjs: the same rule as the preparation table above.
  • acceptance.test.cjs: cancellation, access and boundaries.
  • EVAL_CONTRACT.md: what would count as useful.
  • STAKES_EVAL_CONTRACT.md: consequences, baselines and human choice. It specifies how a consequential version would have to be tested; it is not a live grant, payments or enrolment integration.
$ node --test acceptance.test.cjs
$ node demo.cjs ready
$ node demo.cjs revoked

No automatic installation, model call, account connection or shipping authority.

Read the agent brief
Sources and further discussion

The prior synthetic fit and saved v0.3 summaries are unchanged. The packet is fictional, the rule is local, the cost and time calculations are conditional, and no real-person benefit or live-model quality was measured. The archive contains prior releases, sources, tests and a publication plan. No human reader study has established faster understanding of this layout, and there is no claim about a universal short attention span.

  1. University of Newcastle, $10,000 to build your idea. Official university announcement. A real 2026 project-funding opportunity, not a fictional award. The announcement advertises 27 September, 5 pm. Selection is not guaranteed; this article is not an endorsement or eligibility assessment.
  2. Grand Challenge, 2026 rules and guidelines. Official competition rules. The heading and schedule say 5 pm, while the Project Submission paragraph says 11:59 pm on 27 September 2026. This conflict was unresolved on the research check. The illustrated planning cutoff does not settle the actual rule. Project grants are not personal income.
  3. Stripe, Respond to disputes. Official platform documentation. Documents a response cutoff and the consequence of not responding. Alerts and automation already exist. Owner decides whether to accept or challenge. Preparation or timely submission does not guarantee a win.
  4. AskUON, What are Census dates and when are they? Official university guidance. Defines the financial withdrawal boundary, with pathways-student exceptions. The reader must check their applicable study period and circumstances. This is not a recommendation to withdraw.
  5. METR, Early-2025 experienced-developer study. Primary randomized study report, 10 July 2025. Real task outcomes were measured, rather than relying only on perceived productivity. The result concerns a particular early-2025 setting. It is not evidence of a universal or current slowdown.
  6. METR, Changing the developer-productivity experiment. Primary research follow-up, 24 February 2026. Explains selection and time-measurement limitations in the later experiment. Read alongside the 2025 report; the latter does not establish the effect of present-day models.
  7. Communicating with Interactive Articles. Research synthesis with original interactive demonstrations. Segment the argument, keep one visual object across states, support pause/step/static inspection and details on demand. Engagement is not comprehension; this article has not been validated with readers. No universal advantage of animation is assumed.
  8. Explorable Explanations, Bret Victor. Original design essay and working demonstrations. Put manipulable assumptions beside their numerical consequences. Design precedent, not an effect-size estimate or a test of this implementation.
  9. Progressive Disclosure, Nielsen Norman Group. Usability guidance. Separate a skim route, mechanism inspection and evidence detail; keep qualifications next to figures. Hidden information can be missed. Critical target definitions and synthetic labels are never hidden solely in a disclosure.
  10. Wikipedia: Signs of AI writing. Community editorial observations. Review inflated significance, vague attribution, generic conclusions and promotional language. Descriptive observations, not an authorship detector. No attempt to disguise AI assistance.
  11. 7 Signs a UI Has Been Vibe Coded, The Fountain Institute. Practitioner critique. Challenge decoration without meaning, uniform-card layouts, weak hierarchy and default visual patterns. Aesthetic critique, not evidence that this colour palette or layout increases comprehension.
  12. WCAG 2.2: Animation from Interactions, W3C. Primary accessibility guidance. Respect system reduced motion, expose a manual switch, and allow static and step-controlled paths. Automated local checks are not full WCAG conformance or screen-reader testing.
  13. WCAG 2.2: Target Size (Minimum), W3C. Primary accessibility guidance. Keep interactive targets practical at narrow widths; test touch and keyboard routes. Target dimensions alone do not establish accessibility.
  14. Custom instructions with AGENTS.md, OpenAI. Primary product documentation. Give the agent the scoped task and acceptance requirements in a repository instruction file. Loading instructions does not enforce them; precedence and existing instructions matter.
  15. How Claude remembers your project, Anthropic. Primary product documentation. Document direct AGENTS.md loading or the included CLAUDE.md import wrapper. Version, provider, session and project settings affect loading. No real Claude session was tested.
  16. OpenAI API pricing. Primary tariff documentation. A live source for readers to check before entering their own rates. The calculator deliberately uses labelled illustrative rates, not a provider offer, quality ranking or promised saving.

Prepared with AI assistance and reviewed by Calvin. The page does not connect to accounts, call models, track behaviour or load remote fonts. External links open only when followed.

Accessibility scope: keyboard controls and a readable no-JavaScript route. These features and automated browser checks are not complete accessibility certification or a substitute for screen-reader and human testing.

Published 2026-09-06 · Updated 2026-09-28 · Source edition v0.10