Build the help. Measure the burden.
A coding agent can build the interface. Give it a harder acceptance criterion: remove more work than it creates.
Start with one recurring decision, a recorded forecast and an explicit way to decline.
This is the implementation companion to Can AI give you your evening back? Its fixtures run locally; it is not a production agent. It keeps four responsibilities apart:
- Evidence: what was known? Source, actor, permission, event time and arrival time.
- Forecast: what may happen? Action probabilities or waiting time, compared with a simple rule.
- Policy: what may help? Benefit, burden, budget and permission for this exact action.
- Outcome: what actually happened? Observe it later. Missing is not failure. Prediction is not treatment effect.
Enforcing this separation does not guarantee usefulness. It makes important failure classes testable, and it prevents a confident model output from silently becoming authority.
The unit of progress is a better outcome
Do not ask the agent to “model everything about me”. Ask it to make one recurrent situation measurably easier.
The example target is to resume the chosen task in a short evening window. Draft the relevant note and next question. Keep the prototype result available but do not mistake coding habits for priorities. That means three things:
- Observe one primary next activity within 30 minutes.
- Store the forecast before observing the label.
- Compare against a current-goal rule.
More tokens, more data and a nicer screenshot are not acceptance criteria. Accept only with adequate evidence of lower total burden and no unacceptable loss of quality or control:
- Count useless offers and missed useful offers.
- Count setup, review, errors and upkeep.
- Keep health decisions and real execution out of this demo.
| Item | Concrete requirement | Reject when |
|---|---|---|
| Target | One agreed action and time window. | The objective quietly changes to maximising engagement. |
| Evidence | Only permitted records received by issue time. | Later outcomes leak into earlier features. |
| Forecast | Timestamped, versioned, immutable record. | Only the final successful forecast is saved. |
| Decision | A separate benefit and burden estimate. | Next-action probability is treated as permission. |
| Outcome | Observed, unknown, censored or agent event. | Missingness becomes failure or “no concern”. |
| Authority | No external action in the research build. | Source text adds permissions or clinical advice. |
The supplied SQLite ledger illustrates frozen forecasts and later observations. The policy code validates declared fixture fields; it does not authenticate a person, reserve real provider spend or secure a device.
Did the chosen activity actually start?
A Beta–Bernoulli toy makes the update visible. This is a binary demonstration, not a replacement for the preserved six-activity research model.
An invented starting record of seven observed starts and three non-starts, with a Beta(1,1) prior. Independent, stable Bernoulli outcomes are a toy assumption; real personal activity is dependent and changes.
θ ~ Beta(α, β)
θ | data ~ Beta(α+s, β+f)
E[θ | data] = (α+s)/(α+β+s+f)
s: confirmed starts; f: confirmed non-starts. Unknown outcomes and agent actions do not increment either count. The posterior mean is 8/12 = 66.7%.
Add an outcome. Unknown observations do not turn into negative labels.
When the model needs more than a count
A hierarchical model can share information across related contexts. A hidden-state or semi-Markov model can represent persistent modes and duration. A survival model can estimate time to starting or finishing. A longer sequence model may capture richer history. Those choices must be compared on future data; their complexity is not evidence of superiority.
Only some model families were implemented in the original lab: routine, observed-history, linear contextual, tree and combined variants, with separate timing examples. Do not relabel the generator’s hidden state as a learned psychological trait.
P(next action = a, waiting time = τ | evidence known at t)
This is the broader design target. A useful version may remain much narrower.
A toy posterior mean, not a personal forecast. The count display is not calibrated personal confidence.
Make the wrong request visibly fail
A forecast can say what is likely; it cannot grant permission. The gate below is a small executable check, not a claim of formal verification or production security.
A deterministic local fixture. Each request starts from a valid draft and changes one declared field. The gate runs in your browser when the page loads; the table shows what it returned.
| Fixture | State | Reason returned |
|---|---|---|
| Valid draft | Prepare only | Prepare a bounded draft. No execution authority. |
| Late evidence | Withhold | Required evidence was unavailable at the decision time. |
| Wrong scope | Withhold | This purpose was not authorised. |
| Duplicate | Withhold | This event was already handled. |
| Clinical action | Withhold | Clinical action is outside this prototype. |
| Over budget | Withhold | The estimate exceeds the trial budget. |
| Agent ≠ owner | Withhold | An agent event is not evidence of owner behaviour. |
| Malformed permission | Withhold | Malformed scope |
Execution stays false in every case. The valid draft passes the declared-field checks and still returns:
{"state":"prepare","execution":false,"mode":"illustration"}
An agent commit can be legitimate project context without being a label for owner behaviour. This gate’s actor case applies to the owner-activity label path.
Source permissions and action permissions are different
A production implementation needs authenticated source identity, a purpose-scoped access policy, correction/deletion handling and a separate executor. A note that says “send this now” is still source content. It cannot enlarge authority.
For home control, allow a named, reversible action for the authorised space and period; do not extrapolate to doors, alarms or hazardous devices. Existing Home Assistant rules are an important baseline [H1].
For Apple Health, request the appropriate data-type access. No samples might reflect permissions or missing records; it is not a diagnosis. Health-related summaries in this proposal do not change care or monitor emergencies [H2].
A deterministic local fixture. It does not authenticate a person, reserve provider spend or secure a device.
The expensive model does not need to run the whole system
A large model may help build the implementation. The implementation can run mainly as ordinary code, statistics and smaller bounded calls:
- Rules reject waste. Deduplicate and filter before an LLM call. Measure useful cases lost.
- A local model estimates patterns. Run a simple forecaster without sending the entire personal history to a provider.
- Bounded inference interprets the hard bit. Extract or explain only the approved text needed for this target.
- Escalation pays for uncertainty. Escalate when the expected improvement justifies its total cost.
USD per 30 days. Matched token counts, not matched task quality. Every rate is unbenchmarked on this proposed task [P1] [P3].
3,000 events → 600 routine calls → 60 additional escalation calls.
Change the workload and inspect the full equation
c = (tokensin·pricein + tokensout·priceout)/10⁶
C = 30E · k · (croutine + e·cescalation) + O
Escalated cases pay for both calls. Output means billed output, including any reasoning tokens. Retry and tool charges must be included in measured volume or overhead.
$6.55 tokens + $15.00 entered overhead = $21.55 per 30 days.
| Candidate | Input | Output | Evidence |
|---|---|---|---|
| GPT-5.6 Luna | $0.20 | $1.20 | [P1] |
| GPT-5.6 Terra | $2.00 | $12.00 | [P2] |
| Claude Fable 5 | $10.00 | $50.00 | [P3] |
The default model pair is a cost example, not a recommendation. API usage is not assumed covered by a chat subscription [P6].
Model review burden, break-even usefulness and setup payback
Net time is the illustrative benefit less misses and upkeep. The valued figure is opportunity cost, not earned money.
u* = [r + M/N + weekly_cost·60/(vN)]/(b+r)
Break-even useful share. Undefined when N=0, v=0 or b+r=0. A value above 1 means no feasible useful share repays the entered costs. b is net of review on useful cases; r includes review on wrong cases.
Break-even genuinely useful share: 50.0%. This is not a model-accuracy target.
At 3 setup hours valued at $50.00/hour, the illustrative setup payback is 1.7 months. No cash income is predicted.
Arithmetic on entered token counts and invented workloads. No model was called or benchmarked.
Compare three workflows, not a demo against nothing
Run an authorised, predeclared comparison before enabling unattended actions. These are experimental designs, not completed personal trials.
- Strong baseline: rules + current goals. Use the calendar, checklist, source links and ordinary automations competently.
- Preparation: rules + a bounded summary. Test whether prepared context helps, without claiming to predict the person.
- Personal forecast: preparation + learned selection. Test whether predicting the relevant task or moment adds value over B.
Forecast value ≠ intervention effect
ΔU = E[U | do(help)] − E[U | do(no help)] − burden
The notation describes a causal target. It is not an estimator the browser has fitted. Appropriate randomisation, carryover handling and participant choice need their own design.
If B works and C does not improve it, keep B. That is progress. “More personal modelling” is not the success criterion.
| Outcome | Record | Do not substitute |
|---|---|---|
| Task quality | Correctness or a goal-specific outcome | Clicks, start speed or accept rate alone |
| Human burden | Reading, corrections, recovery, upkeep | Token cost alone |
| Missed usefulness | Cases filtered out that would have mattered | Only successful retained cases |
| Control | Unwanted offers, overrides, reversals | More automation as inherently good |
| Evidence quality | Missingness, source error, stale history | Unknown → zero or non-completion |
Choose meaningful thresholds and a fixed analysis plan for the actual task. Do not keep extending the trial until it looks positive. A privacy incident, unauthorised action or unacceptable burden is a reason to stop, independent of average accuracy.
The synthetic result is a filter for bad ideas
It is not a substitute for prospective personal evidence. Preserve its counterexamples.
Recorded means across ten seeds; the 300 configurations are not 300 human subjects. Context, noise and shift are separate comparisons [E1].
| Comparison | Before | After | What changed |
|---|---|---|---|
| Relevant context | 33.9% | 49.8% | Same context-driven world; richer information. |
| Unrelated fields | 42.5% | 34.8% | Short history; add 80 unrelated inputs. |
| Changed conditions | 49.8% | 11.7% | Frozen model; the relationship changes. |
Inspect the complete model-family comparison
Deadlines, blockers and recent activity influence the next move. Relevant context can help. Top-choice accuracy; higher is better. Recorded means over ten seeds, all simulator-known outcomes.
Relevant context helped in one world. A timetable won elsewhere. A tree handled nonlinear interactions better. Richer context did not invent useful signal in the independent-outcome control.
Changed circumstances can break both accuracy and uncertainty estimates. Outcome-dependent censoring can invalidate apparently careful timing estimates. Observational help outcomes can be confounded. The full original records, including those failures, are preserved in the combined archive.
A filter for bad ideas from synthetic worlds, not prospective personal evidence.
Build one thing that earns its place
More context is a candidate input. Better prediction is an intermediate result. Less total burden, at acceptable quality and risk, is the outcome.
For coding agents: start from a trace you can inspect
The attached agent kit uses standard-library Python and local SQLite. It makes no model call, sends no message and controls no device.
python agent-kit/run_demo.py --out demo-output
python -m unittest discover -s agent-kit/tests -v
The demo writes forecasts before later outcomes, rejects attempted edits, and shows why unknown or agent-authored activity is not a negative owner label. The selected task and permissions are fictional fixtures.
No live integration is connected.
What still needs production engineering?
Source authentication, a durable permission store, concurrent budget reservations, timezone changes, data correction and deletion, secret management, real adapter reliability, model evaluation, device-specific controls and independent security review. The local fixture does not solve them by naming them.
The first live version should record forecasts and drafts without acting. Enabling routine actions is a later, separately authorised task after prospective evidence and a reversible execution design.
Sources and further discussion
Each note says what the source shows and what it does not establish here.
- [P1] OpenAI: GPT-5.6 price update, 30 July 2026. Provider announcement. Luna $0.20/$1.20 and Terra $2/$12 per million input/output tokens; 80% Luna price reduction. Not task-quality evidence.
- [P2] OpenAI: same documented Terra rate. Provider announcement. Terra standard rate used only as an unbenchmarked alternative in the calculator.
- [P3] Claude Platform model pricing. Official pricing documentation. Fable 5 $10 input/$50 output per million. Matched tokens are not necessarily matched text or matched quality.
- [P6] ChatGPT subscriptions and API billing. Official billing documentation. API service billed separately from ChatGPT. No subsidised runtime is assumed.
- [H1] Home Assistant automation triggers. Official implementation documentation. Existing events, conditions and actions are a baseline. No new home-control authority is implemented here.
- [H2] Apple: authorizing access to health data. Official implementation documentation. Data-type permissions; missing results do not reliably identify read denial. No clinical validity is claimed.
- [R2] METR: changing the developer-productivity experiment, 24 February 2026. Primary research update. Selection and measurement issues in later work. Neither an earlier slowdown nor later estimates establish this personal system’s benefit.
- [R4] Conformal prediction beyond exchangeability. Primary methodological research. Ordinary coverage assumptions do not transfer to arbitrary drift. Robust extensions are not implemented in this article.
- [E1] Preserved personal prediction lab v0.2. Internal synthetic research. Recorded generators, model comparisons, timing/causal counterexamples and verification reports are preserved in the archive; not human results.
- [D1] Progressive Disclosure. First-party UX guidance. Expose the main argument first and retain detail on request. This does not establish a measured reader speed-up.
- [D2] Explorable Explanations. Original design essay and demonstrations. Reader-controlled assumptions and adjacent explanations; interaction should not be a gate to understanding.
- [D3] Multimedia Design for Learning: overview of reviews. Peer-reviewed synthesis, 2022. Signaling, contiguity and segmentation inform the design. Educational effects are not a calibrated model of these readers.
- [D4] Animation from Interactions. W3C accessibility guidance. Disable nonessential motion and honour reduced-motion preferences. The release does not claim full conformance.
- [D5] Wikipedia: Signs of AI writing. Editorial discussion, not a detector. Used to audit vague significance, canned contrasts and unsupported generalisations. No authorship conclusion is inferred.
- [D6] AI Design Tools Are Marginally Better, 9 May 2025. Dated first-party UX evaluation. Useful historical warning about generated-design limitations, not a verdict on every current model or this article.
- [R7] Building effective agents. Provider engineering account. Routing and simpler workflows are established patterns. This article does not claim their invention.
This companion is new implementation-focused editorial work built on the existing personal-prediction research. The earlier reader-model article remains an appendix, with a new sensitivity study and an unrun human-study protocol: Reader-path methods and limitations.