# Accepted-outcome agent contract

Use this before an AI agent implements a domain workflow.

## Rule zero

Do not begin implementation until the workflow, authority boundary, accepted end state and evidence sources are explicit. Record missing information as **UNKNOWN** with an owner and a test. Do not silently invent business policy.

## 1. Objective

Convert a costly real-world workflow into the smallest complete outcome that an authorised operator can accept, use and recover when something fails.

## 2. Map the current work

For every step, record:

- actor and authority;
- trigger;
- input and authoritative source;
- deterministic transformation;
- human judgement;
- system or tool;
- output and handoff;
- common exception;
- current time, cost, delay and error consequence.

## 3. Define acceptance

State the observable end condition. Prefer checks against the resulting system state over claims in the agent transcript.

Acceptance must address:

- expected behaviour;
- permissions and identity;
- provenance;
- partial failure and retry behaviour;
- rollback or safe recovery;
- operator comprehension;
- expert review where domain judgement is required;
- total human review and repair burden.

## 4. Bound authority

List what the system may read, write, send, spend, deploy and approve. Keep external communication, money movement, production deployment, legal commitments, clinical decisions and equity commitments behind named human approval unless narrower authority is explicitly documented.

## 5. Compare practical alternatives

Compare:

1. current process;
2. domain expert plus frontier AI;
3. domain expert plus engineer plus frontier AI;
4. existing packaged software or service, where relevant.

Use:

`ΔV_engineer = V(SME + engineer + AI) − V(best practical alternative)`

Count discovery, expert review, rework, coordination, rollout and support.

## 6. Route models using local evidence

Public benchmarks are shortlist evidence. They are not local acceptance rates.

For each material task estimate:

`E[C_accepted] = (run cost + review wage × review hours) / local acceptance rate + latent failure probability × loss`

Record the model, harness, effort setting, tool permissions, date and representative cases.

## 7. Build the smallest complete slice

A vertical slice should traverse the real workflow and reach the accepted state. Avoid disconnected feature volume. Preserve data rights, least privilege, idempotency, observability and recovery.

## 8. Verify

Use:

- deterministic tests for code, state and permissions where possible;
- capability evals for difficult representative cases;
- regression evals for previously passing cases;
- calibrated model graders only where deterministic checks are insufficient;
- domain-expert review for consequential judgement;
- failure injection or recovery tests;
- a ledger of unresolved risks.

A passing implementation-generated test suite is not independent evidence if it encodes the same unverified assumption as the implementation.

## 9. Check economics

Separate:

- current workflow cost;
- delivery cost;
- switching and support cost;
- customer price;
- customer net value;
- provider contribution;
- the next observation that warrants another investment tranche.

## 10. Stop conditions

Recommend **STOP**, **NARROW** or **REPRICE** when:

- the expert using AI alone reaches equivalent acceptance at lower total cost;
- no authorised user or budget owner exists;
- required data or rights cannot be obtained;
- critical risk cannot be bounded appropriately;
- review and support consume the projected saving;
- the customer receives no credible surplus;
- the second build does not show measured reuse.

## Required return format

Return exactly:

1. decision summary;
2. assumptions and UNKNOWNs;
3. current workflow map;
4. accepted state and authority matrix;
5. alternatives and incremental-value estimate;
6. model-routing plan and local eval design;
7. smallest complete build plan;
8. required verification evidence;
9. customer and provider economics;
10. continue / narrow / reprice / hand over / stop recommendation.

End with: **What observation would prove this plan wrong fastest?**
