← Back to writing

Build the help. Measure the burden.

A coding agent can build the interface. Give it a harder acceptance criterion: remove more work than it creates.

Start with one recurring decision, a recorded forecast and an explicit way to decline.

This is the implementation companion to Can AI give you your evening back? Its fixtures run locally; it is not a production agent. It keeps four responsibilities apart:

  1. Evidence: what was known? Source, actor, permission, event time and arrival time.
  2. Forecast: what may happen? Action probabilities or waiting time, compared with a simple rule.
  3. Policy: what may help? Benefit, burden, budget and permission for this exact action.
  4. Outcome: what actually happened? Observe it later. Missing is not failure. Prediction is not treatment effect.

Enforcing this separation does not guarantee usefulness. It makes important failure classes testable, and it prevents a confident model output from silently becoming authority.

The unit of progress is a better outcome

Do not ask the agent to “model everything about me”. Ask it to make one recurrent situation measurably easier.

The example target is to resume the chosen task in a short evening window. Draft the relevant note and next question. Keep the prototype result available but do not mistake coding habits for priorities. That means three things:

  • Observe one primary next activity within 30 minutes.
  • Store the forecast before observing the label.
  • Compare against a current-goal rule.

More tokens, more data and a nicer screenshot are not acceptance criteria. Accept only with adequate evidence of lower total burden and no unacceptable loss of quality or control:

  • Count useless offers and missed useful offers.
  • Count setup, review, errors and upkeep.
  • Keep health decisions and real execution out of this demo.
The bounded trial contract
ItemConcrete requirementReject when
TargetOne agreed action and time window.The objective quietly changes to maximising engagement.
EvidenceOnly permitted records received by issue time.Later outcomes leak into earlier features.
ForecastTimestamped, versioned, immutable record.Only the final successful forecast is saved.
DecisionA separate benefit and burden estimate.Next-action probability is treated as permission.
OutcomeObserved, unknown, censored or agent event.Missingness becomes failure or “no concern”.
AuthorityNo external action in the research build.Source text adds permissions or clinical advice.

The supplied SQLite ledger illustrates frozen forecasts and later observations. The policy code validates declared fixture fields; it does not authenticate a person, reserve real provider spend or secure a device.

Did the chosen activity actually start?

A Beta–Bernoulli toy makes the update visible. This is a binary demonstration, not a replacement for the preserved six-activity research model.

Seven confirmed starts and three non-starts give a posterior mean of 66.7%

An invented starting record of seven observed starts and three non-starts, with a Beta(1,1) prior. Independent, stable Bernoulli outcomes are a toy assumption; real personal activity is dependent and changes.

θ ~ Beta(α, β)
θ | data ~ Beta(α+s, β+f)
E[θ | data] = (α+s)/(α+β+s+f)

s: confirmed starts; f: confirmed non-starts. Unknown outcomes and agent actions do not increment either count. The posterior mean is 8/12 = 66.7%.

Toy posterior mean66.7%
Record7 starts · 3 non-starts · Beta(8,4)

Add an outcome. Unknown observations do not turn into negative labels.

When the model needs more than a count

A hierarchical model can share information across related contexts. A hidden-state or semi-Markov model can represent persistent modes and duration. A survival model can estimate time to starting or finishing. A longer sequence model may capture richer history. Those choices must be compared on future data; their complexity is not evidence of superiority.

Only some model families were implemented in the original lab: routine, observed-history, linear contextual, tree and combined variants, with separate timing examples. Do not relabel the generator’s hidden state as a learned psychological trait.

P(next action = a, waiting time = τ | evidence known at t)

This is the broader design target. A useful version may remain much narrower.

A toy posterior mean, not a personal forecast. The count display is not calibrated personal confidence.

Make the wrong request visibly fail

A forecast can say what is likely; it cannot grant permission. The gate below is a small executable check, not a claim of formal verification or production security.

Seven of eight fixtures are withheld, and none executes

A deterministic local fixture. Each request starts from a valid draft and changes one declared field. The gate runs in your browser when the page loads; the table shows what it returned.

Policy gate results
FixtureStateReason returned
Valid draftPrepare onlyPrepare a bounded draft. No execution authority.
Late evidenceWithholdRequired evidence was unavailable at the decision time.
Wrong scopeWithholdThis purpose was not authorised.
DuplicateWithholdThis event was already handled.
Clinical actionWithholdClinical action is outside this prototype.
Over budgetWithholdThe estimate exceeds the trial budget.
Agent ≠ ownerWithholdAn agent event is not evidence of owner behaviour.
Malformed permissionWithholdMalformed scope

Execution stays false in every case. The valid draft passes the declared-field checks and still returns:

{"state":"prepare","execution":false,"mode":"illustration"}

An agent commit can be legitimate project context without being a label for owner behaviour. This gate’s actor case applies to the owner-activity label path.

Source permissions and action permissions are different

A production implementation needs authenticated source identity, a purpose-scoped access policy, correction/deletion handling and a separate executor. A note that says “send this now” is still source content. It cannot enlarge authority.

For home control, allow a named, reversible action for the authorised space and period; do not extrapolate to doors, alarms or hazardous devices. Existing Home Assistant rules are an important baseline [H1].

For Apple Health, request the appropriate data-type access. No samples might reflect permissions or missing records; it is not a diagnosis. Health-related summaries in this proposal do not change care or monitor emergencies [H2].

A deterministic local fixture. It does not authenticate a person, reserve provider spend or secure a device.

The expensive model does not need to run the whole system

A large model may help build the implementation. The implementation can run mainly as ordinary code, statistics and smaller bounded calls:

  1. Rules reject waste. Deduplicate and filter before an LLM call. Measure useful cases lost.
  2. A local model estimates patterns. Run a simple forecaster without sending the entire personal history to a provider.
  3. Bounded inference interprets the hard bit. Extract or explain only the approved text needed for this target.
  4. Escalation pays for uncertainty. Escalate when the expected improvement justifies its total cost.
Filtering and routing cut the illustrative token bill from $270.00 to $6.55

USD per 30 days. Matched token counts, not matched task quality. Every rate is unbenchmarked on this proposed task [P1] [P3].

All events, higher rate$270.00
Filtered, same rate$54.00
Filter + routing$6.55

3,000 events → 600 routine calls → 60 additional escalation calls.

Change the workload and inspect the full equation

c = (tokensin·pricein + tokensout·priceout)/10⁶
C = 30E · k · (croutine + e·cescalation) + O

Escalated cases pay for both calls. Output means billed output, including any reasoning tokens. Retry and tool charges must be included in measured volume or overhead.

$6.55 tokens + $15.00 entered overhead = $21.55 per 30 days.

Official rate snapshot checked 20 September 2026 · USD per million text tokens
CandidateInputOutputEvidence
GPT-5.6 Luna$0.20$1.20[P1]
GPT-5.6 Terra$2.00$12.00[P2]
Claude Fable 5$10.00$50.00[P3]

The default model pair is a cost example, not a recommendation. API usage is not assumed covered by a chat subscription [P6].

Model review burden, break-even usefulness and setup payback
Net time per week+31.2 min
Valued time less operating cost+$20.97

Net time is the illustrative benefit less misses and upkeep. The valued figure is opportunity cost, not earned money.

u* = [r + M/N + weekly_cost·60/(vN)]/(b+r)

Break-even useful share. Undefined when N=0, v=0 or b+r=0. A value above 1 means no feasible useful share repays the entered costs. b is net of review on useful cases; r includes review on wrong cases.

Break-even genuinely useful share: 50.0%. This is not a model-accuracy target.

At 3 setup hours valued at $50.00/hour, the illustrative setup payback is 1.7 months. No cash income is predicted.

Arithmetic on entered token counts and invented workloads. No model was called or benchmarked.

Compare three workflows, not a demo against nothing

Run an authorised, predeclared comparison before enabling unattended actions. These are experimental designs, not completed personal trials.

  1. Strong baseline: rules + current goals. Use the calendar, checklist, source links and ordinary automations competently.
  2. Preparation: rules + a bounded summary. Test whether prepared context helps, without claiming to predict the person.
  3. Personal forecast: preparation + learned selection. Test whether predicting the relevant task or moment adds value over B.

Forecast value ≠ intervention effect
ΔU = E[U | do(help)] − E[U | do(no help)] − burden

The notation describes a causal target. It is not an estimator the browser has fitted. Appropriate randomisation, carryover handling and participant choice need their own design.

If B works and C does not improve it, keep B. That is progress. “More personal modelling” is not the success criterion.

Outcomes and stop conditions to agree before running
OutcomeRecordDo not substitute
Task qualityCorrectness or a goal-specific outcomeClicks, start speed or accept rate alone
Human burdenReading, corrections, recovery, upkeepToken cost alone
Missed usefulnessCases filtered out that would have matteredOnly successful retained cases
ControlUnwanted offers, overrides, reversalsMore automation as inherently good
Evidence qualityMissingness, source error, stale historyUnknown → zero or non-completion

Choose meaningful thresholds and a fixed analysis plan for the actual task. Do not keep extending the trial until it looks positive. A privacy incident, unauthorised action or unacceptable burden is a reason to stop, independent of average accuracy.

The synthetic result is a filter for bad ideas

It is not a substitute for prospective personal evidence. Preserve its counterexamples.

Relevant context helped; unrelated inputs and a changed relationship hurt

Recorded means across ten seeds; the 300 configurations are not 300 human subjects. Context, noise and shift are separate comparisons [E1].

Mean next-activity accuracy, before and after one change
ComparisonBeforeAfterWhat changed
Relevant context33.9%49.8%Same context-driven world; richer information.
Unrelated fields42.5%34.8%Short history; add 80 unrelated inputs.
Changed conditions49.8%11.7%Frozen model; the relationship changes.
Inspect the complete model-family comparison
Routine
33.9%
Task + context
46.7%
Richer context
49.8%
Richer trees
46.9%

Deadlines, blockers and recent activity influence the next move. Relevant context can help. Top-choice accuracy; higher is better. Recorded means over ten seeds, all simulator-known outcomes.

Relevant context helped in one world. A timetable won elsewhere. A tree handled nonlinear interactions better. Richer context did not invent useful signal in the independent-outcome control.

Changed circumstances can break both accuracy and uncertainty estimates. Outcome-dependent censoring can invalidate apparently careful timing estimates. Observational help outcomes can be confounded. The full original records, including those failures, are preserved in the combined archive.

A filter for bad ideas from synthetic worlds, not prospective personal evidence.

Build one thing that earns its place

More context is a candidate input. Better prediction is an intermediate result. Less total burden, at acceptable quality and risk, is the outcome.

For coding agents: start from a trace you can inspect

The attached agent kit uses standard-library Python and local SQLite. It makes no model call, sends no message and controls no device.

python agent-kit/run_demo.py --out demo-output
python -m unittest discover -s agent-kit/tests -v

The demo writes forecasts before later outcomes, rejects attempted edits, and shows why unknown or agent-authored activity is not a negative owner label. The selected task and permissions are fictional fixtures.

No live integration is connected.

What still needs production engineering?

Source authentication, a durable permission store, concurrent budget reservations, timezone changes, data correction and deletion, secret management, real adapter reliability, model evaluation, device-specific controls and independent security review. The local fixture does not solve them by naming them.

The first live version should record forecasts and drafts without acting. Enabling routine actions is a later, separately authorised task after prospective evidence and a reversible execution design.

Sources and further discussion

Each note says what the source shows and what it does not establish here.

This companion is new implementation-focused editorial work built on the existing personal-prediction research. The earlier reader-model article remains an appendix, with a new sensitivity study and an unrun human-study protocol: Reader-path methods and limitations.

Published 2026-09-07 · Updated 2026-09-28 · Source edition v7