← Back to writing

Can AI give you your evening back?

The workday ends and the coordination doesn’t. You still have to remember, find, explain and check.

I want personal AI to take recurring work off my mind. It has to save more attention than it asks for.

For the person using it, the aim is fewer loose ends to carry: prepared context, a bounded choice, or a routine action you have explicitly allowed. The test is whether the whole job gets easier.

This is a proposed system. The research behind it is synthetic, and no live accounts are connected. The companion article, Build the help. Measure the burden., turns it into something a coding agent can build and test.

I’m still the coordinator

A model can produce a fine answer and leave the surrounding job untouched. I still have to remember the task, find the right material, explain what changed, review the result and carry the next dependency, and then do it all again.

That is the part I want to reduce. A faster answer matters less if I spend the evening assembling its context and supervising its output. The failure mode is paying twice: money for the calls, then attention for the supervision.

So there are two bills. The visible bill is model calls, tools and hosting, retries and maintenance. The personal bill is remembering, explaining, checking and recovering from errors. A smaller token bill does not tell us whether the personal bill improved.

If it succeeds, the advantage could be less routine coordination for the same quality of outcome. It is not proof that you will outperform someone without it. A good checklist, ordinary automation or another person may already do the job better.

Give the hour to what you chose

Take a work night. At 18:30 there is one free hour and four loose ends. A study task is due tomorrow. Your prototype tests have finished. The room is still set for work. You have one hour before you want to stop. The due task matters tonight, and your coding habit is not permission to steer you back into the prototype.

  1. 18:30, the existing problem: four tools know something, and you assemble the plan. The calendar has the hour. The task list has the deadline. The repository has a result. The home app has a setting.
  2. 18:31, use the right evidence: estimate what you may need, and check what you actually want. Match records, reject late or unapproved sources, and keep the current goal distinct from the predicted habit.
  3. 18:32, make a small part easier: a starting point is ready, and you still choose. The study note is in view. The code result is filed. A separately approved light scene could be offered. The system can also stay quiet.
  4. Afterwards, did it help? Count the work it removed and the work it created. Did you finish the chosen task? Was the suggestion wrong? How much checking did it add? Unknown outcomes stay unknown.
The same work night, with existing tools and with the proposed help

An authored scenario. At 18:30 the calendar says stop at 19:30, the prototype test run has finished, the current goal is a study task due tomorrow, and the home settings allow a desk lamp scene.

One evening, two ways
Capable baseline: existing toolsProposed personal AI, if it helps
Notice the test result.Test result filed with the last decision.
Reopen the previous decision.Study notes placed beside the due task.
Remember tomorrow’s task.A lamp scene is offered, not silently applied.
Choose what gets the hour.You choose: study, prototype or stop.

The study task is ready to resume: the last study note and tomorrow’s question, with the prototype result saved for later. You choose. No external action executes here.

A due-date rule plus a good checklist may achieve the same result. The failure to look for: it mistakes your usual coding habit for what you want tonight.

The task is familiar; rebuilding its context is the repeated work. No time-saving result is claimed by this illustration. Compare against tools you already use competently.

Between the evidence and that starting point sit three separate steps. First, code checks what is allowed and known now, and uses only current, approved evidence. Second, a model can forecast what you may resume, but a forecast is not a goal: your stated goal still decides what help is worth offering. Third, a policy weighs value and permission, then prepares a draft, offers an allowed action, or does nothing.

Challenge the proposal: what the illustrative route does
ChallengeResult
Evidence availableThe study task is ready to resume. You choose. No external action executes here.
Required source arrives lateWithheld, with no proposed action. The required source was not available at this decision time. No model call in this illustrative route.
No permission for this purposeWithheld, with no proposed action. This purpose is outside the approved scope.
Likely habit conflicts with my goalA likely habit is not your goal. Predicting the habit does not justify repeating it; the stated objective takes precedence.
Preparing it exceeds my budgetWithheld, with no proposed action. This proposal costs more than the allowed budget.

Start with one recurring moment, not a complete model of a person. The light scene is a separate, reversible permission, and no calendar or device is changed here.

Which part needs a prediction model?
  • A rule. “The deadline is tomorrow” and “the run finished” are usually straightforward checks.
  • A forecast. “Which task will I return to, and when?” requires uncertain inference. Compare it against the rule.
  • A decision. “Would preparing or interrupting help?” depends on goals, consequences and burden. Prediction alone cannot answer it.

The original lab tested a narrow next-activity forecast. It did not test this evening, a home controller, health outcomes or a learned intervention policy. Inspect the actual evidence.

Spend intelligence where it changes the decision

Use ordinary code to filter. Local statistics can estimate patterns. Pay a language model to interpret a bounded piece of text; escalate only when the task earns it.

Filtering and routing cut the illustrative token bill from $270.00 to $6.55

An illustrative token bill in USD per 30 days. Default: 100 daily events; 6,000 input + 600 output tokens per call; Luna for routine calls, then Fable for escalation. Matched token counts, not matched quality.

Every event → higher-cost model. Same assumed token counts, no filtering.$270.00
Filter → that same model. Isolates the effect of fewer calls.$54.00
Filter → cheaper model → selective escalation. Escalation pays for the cheaper call too.$6.55

3,000 events → 600 routine calls → 60 additional escalation calls.

The three bars exclude hosting, tools, retries and review. No model was benchmarked for these scenes. Filtering can miss useful events. Rates: [P1], [P3].

Inspect models, token assumptions and full monthly cost

C = 30E · k · (croutine + e · cescalation) + O

E: events per day; k: retained fraction; e: escalation fraction; c: token cost per call; O: other monthly costs. No cache discount is assumed.

$6.55 tokens + $15.00 entered overhead = $21.55 per 30 days.

Official rate snapshot checked 20 September 2026 · USD per million text tokens
CandidateInputOutputEvidence
GPT-5.6 Luna$0.20$1.20[P1]
GPT-5.6 Terra$2.00$12.00[P2]
Claude Fable 5$10.00$50.00[P3]

Rates are a dated snapshot, not a model recommendation. Real tokenisation, retries, reasoning output and tool use can alter the bill. Use authorised API access; a paid chat subscription is not assumed to cover these calls [P6].

Will the attention saved repay the attention spent?

All inputs below are invented and editable. “Useful” means extra benefit beyond the capable baseline. Predicting the next action correctly is not enough on its own.

Useful preparation returns78.0 min
Wrong offers cost16.8 min
Upkeep costs30.0 min
Net time per week+31.2 min
Valued time less operating cost+$20.97

Benefit is entered after review effort on useful offers. Miss cost includes review on wrong offers. The valued figure is an opportunity-cost calculation, not cash income or a promise of productivity.

ΔT = N[ub − (1−u)r] − M

N: offers per week; u: genuinely useful share; b: minutes returned by a useful offer; r: minutes lost on a wrong one; M: weekly upkeep.

Break-even genuinely useful share: 50.0%. This is not a model-accuracy target.

At 3 setup hours valued at $50.00/hour, the illustrative setup payback is 1.7 months. No cash income is predicted.

Setup time is priced separately. No value is assigned to unmeasured outcomes.

The cheapest architecture is allowed to lose this comparison. The right system might be a rule, or no system.

Arithmetic on entered token counts and invented workloads. No model was called or benchmarked.

The price changed; the burden of proof did not

There is a dated reason to revisit the economics. There is no honest countdown to guaranteed advantage.

On 30 July 2026, OpenAI announced that Luna’s price fell 80%. The published rate became $0.20 per million input tokens and $1.20 per million output tokens. That changes the calculation for frequent, bounded work. Provider pricing is evidence of cost, not independent proof of task quality [P1].

Measuring the benefit stayed difficult. On 24 February 2026, METR described selection and measurement problems in its follow-up developer-productivity work. Neither its earlier study nor that update establishes the value of this personal system. It is a reason to measure the complete workflow rather than borrow a productivity claim [R2].

The time-sensitive step is recording the baseline before changing the workflow. Later, a cheaper model can re-read your files. It cannot directly observe the alternative choice you would have made without its intervention.

An early advantage, if this works, is an earlier verdict: which recurring jobs are worth delegating, under which conditions. That is a hypothesis about learning. It is not a promise that everyone without the tool falls behind.

One method can cover different kinds of help

The question is always specific: which signal, which decision, which permission, and which result would make this worth keeping?

  • Home: an allowed setting at a useful time. An arrival window could help queue a room scene. Existing trigger-and-condition automation is a strong baseline, not an obsolete competitor [H1].
  • Apple Watch and Health: your records, including their gaps. A bounded summary beside your notes. No missing-data-as-zero shortcut, diagnosis or emergency inference. HealthKit needs data-type permissions [H2].
  • Git, study and admin: the unresolved decision, ready to resume. Commits, tests, notes and commitments joined by evidence, rather than by assuming an app session or agent commit means you did the work.

The ambition goes further than drafts

An assistant that asks me to approve every tiny step can leave me with the same management job. After evidence that it helps, I would want selected, reversible actions to run inside a narrow permission: file a completed test result, prepare a recurring handover, or apply one approved light scene.

  1. Research: watch quietly. Record forecasts; do not change the day.
  2. Personal test: offer a draft. Measure corrections, misses and actual usefulness.
  3. Owner decision: grant one scope. Name the action, conditions, expiry and undo.
  4. Later, if justified: finish one routine job. Within that scope; log, monitor and revoke.

That is a proposed development path, not four completed validation stages. Autonomous clinical decisions and safety-critical home actions remain outside this design.

Four more situations, worked the same way

A client call, 09:10

The approval arrived. Your call is at ten. A client has approved part of a proposal. The calendar has a call at 10:00. The latest document contains two unresolved questions.

  • Evidence: Client email: scope approved; Calendar: call at 10:00; Working document: two open questions; Current goal: leave with a clear decision.
  • Existing tools: Read the new email. Find the latest proposal. Recover the open questions. Write a call brief.
  • Proposed help: Approval linked to the correct proposal. Changed scope marked for review. Two unresolved questions kept in view. A call brief waits; nothing is sent.
  • Illustrated result: The call opens on the unresolved decision. Approved scope + exact changes + the two remaining questions.
  • Potential benefit to test: Enter the conversation ready to decide, instead of searching while someone waits.
  • Forecast role: A fixed calendar rule handles much of this. Prediction must improve which preparation is worth doing or when to offer it.
  • Baseline: Compare with a calendar reminder and a prompted summary, not an empty workflow.
  • Failure to look for: A confident summary turns a partial approval into a promise you never made.
  • Measure: Incorrect scope statements; prep minutes; clarification work after the call.
  • Scope: Client documents require the correct purpose and source permissions. No send authority.

Coming home, 17:50

The house can be ready without guessing who is home. You have explicitly shared that you are on the way. One room’s light scene is approved. Someone else may be using the rest of the house.

  • Evidence: Your shared ETA: arrival window; Room state: desk scene available; Your preference: warm desk light; Permission: only the desk scene.
  • Existing tools: Check which room you need. Open the home app. Set the approved scene. Avoid changing shared rooms.
  • Proposed help: Relevant room state checked. Only the approved scene is proposed. A stale arrival estimate cancels the plan. You can accept, undo or leave it alone.
  • Illustrated result: One allowed room setting, at a useful moment. Approved desk scene + current state + expiry time + undo plan.
  • Potential benefit to test: Remove one repeated chore without turning the household into a guessing game.
  • Forecast role: An arrival estimate might help. A deterministic presence rule is the baseline to beat.
  • Baseline: Home Assistant already supports triggers, conditions and actions [H1].
  • Failure to look for: An inferred arrival changes a shared room while someone else is using it.
  • Measure: Unwanted scene changes; manual reversals; setup and maintenance minutes.
  • Scope: Never unlock a door, disable an alarm or control a hazardous device in this prototype.

Watch / Health, 08:00

A useful summary does not need to diagnose you. You opted to review sleep and activity records beside your own notes. A planned day is visible. A missing record might be non-wear, a sync gap or limited access.

  • Evidence: Apple Health: approved record window; Apple Watch: source and sample time; Your own note: how the night felt; Calendar: today’s plan.
  • Existing tools: Check the records. Notice gaps and duplicates. Compare with your notes. Choose whether to change your plan.
  • Proposed help: Approved records summarised with gaps. Your note stays distinct from sensor data. An optional check-in is drafted. No diagnosis, treatment or emergency inference.
  • Illustrated result: The record and its gaps, in one place. Dated sleep/activity summary + missingness + your own notes.
  • Potential benefit to test: Spend less effort reconstructing what was recorded; keep any health decision with you and an appropriate professional.
  • Forecast role: A behaviour forecast is not a clinical forecast. No health benefit was tested in the lab.
  • Baseline: The Health app and a regular personal review are the comparison.
  • Failure to look for: An absent sample is misread as zero activity or evidence that you are healthy.
  • Measure: Summary fidelity; missing-record handling; unwanted pressure; review effort.
  • Scope: HealthKit permissions are per data type. Read denial cannot be reliably inferred from absent data [H2].

Back to the code, 14:30

The postcode lost its zeros. Start there. Run #184 failed while you were in a meeting. It expects "0080" and receives 80. The related commit and prior design note are available.

  • Evidence: Commit: 8f2c1d; Failing test: "0080" → 80; Prior decision: preserve leading zeros; Calendar: meeting has ended.
  • Existing tools: Read the red-build alert. Match the run to a commit. Find the previous decision. Inspect the candidate cause.
  • Proposed help: Run matched to the commit. Relevant diff kept with the failure. Previous decision included. Candidate cause offered for verification.
  • Illustrated result: The next debugging decision is ready. Failed assertion + matching diff + note. Candidate: numeric coercion.
  • Potential benefit to test: Begin with a concrete question, not a hunt across five tabs.
  • Forecast role: A webhook can assemble this without predicting a person. Test the added value of timing or selection separately.
  • Baseline: A CI notification plus run-to-commit links and a saved note.
  • Failure to look for: An agent-authored commit is mistaken for what the owner spent time doing.
  • Measure: Time to a correct diagnosis; bad causes suggested; revision and supervision effort.
  • Scope: No commit, merge, deploy or message is executed by this article.
The broader source catalogue, and what each source cannot prove

The original research catalogued 223 candidate fields across 25 domains. That was a design catalogue, not 223 validated predictors. The following 18 source summaries are preserved from the prior article; no account is connected.

Candidate sourcePossible signalNot evidence of
CalendarAvailable windows, commitments and changesA free slot is not permission to fill it.
Git & pull requestsCommit IDs, open branches, review decisionsA commit is not proof that the owner wrote or understood it.
Tests & CIRun ID, completion time, result and commit SHAGreen tests are not a release decision.
Notes & decisionsLast question, choice, reason and evidenceText in a note is data, not a new tool instruction.
MessagesExplicit requests, promised replies, dependenciesA message does not reveal someone's intentions.
Study materialsAssignment dates, last note, unresolved questionsTime on a page is not learning.
Apple WatchAvailable activity, sleep or workout samplesA sensor value is not a diagnosis or a reliable stress label.
Apple HealthAuthorised HealthKit samples and their provenanceMissing samples do not establish poor sleep or denied consent.
Your own check-inYour stated energy, intention or preferenceA momentary answer is not a permanent personality score.
Home AssistantDevice state, occupancy scope, preferred scenesPresence is not permission to change the house.
Room sensorsTemperature, humidity, light and quality flagsA sensor reading does not establish a health outcome.
Energy & devicesMeasured energy, device availability, user tariff inputA modelled saving is not a lower bill.
Household inventoryOwner-confirmed stock and shopping list changesAn image guess is not a verified missing item.
LocationCoarse place category and owner-shared journeyLocation does not establish intention.
Bills & renewalsDue dates, amounts and owner-set limitsA spending pattern is not investment advice.
Device activityForeground application, project and idle stateAn open app is not attention or productivity.
Travel & conditionsApproved trip, transit information and forecast timestampA schedule does not guarantee arrival.
What actually helpedUsed, ignored, corrected, delayed, cost and outcomeClicking an offer is not proof of benefit.

Financial commitments, browser records, device activity and household events may be useful for an authorised target. Availability is not permission. Employer, student and other people’s records do not become personal-model inputs by default.

For HealthKit, absent samples do not reliably distinguish read denial from missing records. This browser cannot directly read Apple Health. Clinical inference and high-stakes home control require a different validation effort [H2].

Useful information helped, but bad assumptions still won

These are recorded synthetic results, not measurements of my life, your life, model API quality or time saved.

Relevant context helped; unrelated inputs and a changed relationship hurt

Mean next-activity accuracy across ten seeds per cell. Six generating mechanisms × five conditions × ten seeds = 300 configurations. The seed is the replication unit; repeated forecasts are not independent people [E1].

Mean next-activity accuracy, before and after one change
ComparisonBeforeAfterWhat changed
Relevant context33.9%49.8%Same context-driven world; richer information.
Unrelated fields42.5%34.8%Short history; add 80 unrelated inputs.
Changed conditions49.8%11.7%Frozen model; the relationship changes.
Explore all model comparisons and the probability mathematics

Recorded results. Nothing is being trained in this browser.

Routine
33.9%
Task + context
46.7%
Richer context
49.8%
Richer trees
46.9%

Deadlines, blockers and recent activity influence the next move. Relevant context can help. Top-choice accuracy; higher is better. Recorded means over ten seeds, all simulator-known outcomes.

Prediction has several possible answers

The original target was one primary activity in the next half-hour, among six categories. A joint model of action and waiting time is a later design, not a result of this article.

Focused work
18.5%
Communication
18.8%
Administration
2.1%
Travel
4.6%
Rest
38.6%
Leisure
17.4%

One saved synthetic forecast. Its largest probability is only 38.6%. It is not the evening illustration’s forecast.

A confident mistake is expensive

p(y | xknown now)
L = −log p(yobserved)

Probability assigned to the action that happened; log loss is lower when that action received more probability.

01234501Probability on the observed action

p = 0.70 → log loss = 0.357

Calibration is something to test, not a synonym for confidence. Ordinary prediction-set coverage can fail under changing conditions [R4].

What passed verification, and what did not become a personal result?

The preserved research reports 188 new software tests and 75 earlier tests rerun, plus selected mutation tests and independent score replay. Those historical checks were for the synthetic lab; this article revision does not count them as newly executed tests.

The timing study exposed censoring assumptions. The intervention study showed why accurate failure-risk prediction can still target harmful interruptions. No real LLM, HealthKit collector, home controller, clinical outcome or human comprehension gain was evaluated by those experiments.

ClaimWhat the record supportsStill missing
Context can improve a forecastConditional synthetic evidencePersonal prospective accuracy
A lower-cost route saves tokensArithmetic using entered tokensMatched task quality and full burden
An intervention helpsA synthetic causal counterexample shows what can go wrongAn authorised causal evaluation
This layout is clearerImplementation checks and assumed reader-path modelsPeople correctly explaining and transferring the idea

The original lab, protocol, negative results and full earlier publishing package are in the combined ZIP. They are not rewritten as a success story [E1].

Recorded synthetic results, not measurements of people, model API quality or time saved.

I don’t need AI to predict my whole life

I need one part of it to stop costing me so much effort. The diagrams, lower prices and synthetic tests make this worth investigating. None of them proved how much better my day would be.

The next useful result is smaller and harder: one recurring job, done with less total effort, without taking choices away.

For coding agents: one recurring job and a way to say no

For a coding agent, the useful version is a job it can be held to: a data contract, a cost limit, failure cases and an outcome record, rather than another vague instruction to “be proactive”. Pick a situation, export the brief, and begin with silent forecasts and drafts, a capable baseline and a cost ceiling.

Build this first: one evidence-linked draft

A work night: Last study note + tomorrow’s question. The prototype result is saved for later.

  • Source time and arrival time are separate.
  • A likely habit cannot override a stated goal.
  • Unknown outcomes are not failures.

Make these cases fail safely: wrong scope, late data, excess cost

A forecast is not permission. Decline a draft without adequate evidence. Keep execution disabled in the trial.

  • Compare rule, rule + summary, forecast + policy.
  • Measure missed value, errors and review effort.
  • Stop when the chosen outcome does not improve.

Exports stay local. No prompt is sent to an AI service.

Sources and further discussion

Each note says what the source shows and what it does not establish here.

Original concept illustration showing personal data, modelling, forecasts and possible assistance.
Existing AI-generated concept artwork, preserved unchanged. Its broad health, autonomy and “none of them” language is an aspiration, not evidence about products or clinical effectiveness. The article’s narrower claims govern this proposal.

Prepared with AI assistance for Calvin’s editorial review. No new generated image, account access or autonomous action. This is not a live deployment or a completed human study.

Reader-path methods and limitations: a sensitivity model of planned reading paths for the v7 layout of this article, with no human participants.

Published 2026-09-08 · Updated 2026-09-28 · Source edition v7