Human QA became the constraint.
OpenAI’s internal engineering account identifies QA capacity as a bottleneck and describes giving agents direct access to runnable environments and feedback.
Read the engineering account ↗The patch is ready. The decisions are still yours. I want the next run to bring back a checked result and the one thing it couldn’t settle.
Cheap, broad inspection. Code for facts. Frontier reasoning for the part that still needs it.
Six tasks. All back on your desk.
01Keep access after cancellationUnresolvedVerified patch02Stop duplicate bookingsUnresolvedVerified patch03Correct the refund totalUnresolvedVerified patch04Will buyers pay for export?UnresolvedNeeds purchase data05Can we delete old accounts?UnresolvedOwner decision06Does the signup work on iOS?UnresolvedNeeds device testSee the proposed division of labour, not a live agent run.
Scripted outcomes. “Reviewable” is not “shipped”; the two evidence gaps remain unresolved.
Get evidence and the unresolved choice together.
Build, check, gather evidence, repair or stop.
Keep the rules, failures and checks—not only the chat.
One paid month. One button. A plausible patch can still break the promise the customer bought. This is a constructed example, not a reported incident.
A customer paid for October. Turning off renewal must not remove the access already bought. That rule has to come from the product owner, not from a guess.
Date boundaries, retries, tenant isolation, interface copy, missing evidence. Code handles exact facts. A Jev-like layer flags interpretation and gaps.
Once the outcome and failure are explicit, the generator has a target. Another long opinion is unnecessary when a small executable check can answer.
A passing check belongs to a specific revision. It earns a reviewable packet, not permission to deploy. Missing rules and withdrawn authority stop the path.
| Example | Expected | Observed | Check |
|---|---|---|---|
| 6 Oct: paid account | Access | Access | Pass |
| 12 Oct: renewal off | Access | No access | Fail |
| 1 Nov: period ended | No access | No access | Pass |
The missing entitlement is observable. A confident approval does not change it.
October is an illustrative billing period. The browser executes two small candidate functions against the stated examples. There is no account, model call or deployment.
promise: cancel renewal, keep access until paid_through wrong: active = !cancelled && now < paid_through fixed: active = now < paid_through ready = tests_pass && rules_present && authority_current
The end-to-end product would also need billing events, time zones, identity, concurrency, recovery and actual UI tests. The included offline Python fixture checks a declared finite day domain. Missing product meaning is not solved by asserting the wrong rule more thoroughly.
The system is broader than regression tests: inspect the task from many useful angles, then spend effort on whichever next action could change the decision.
Proposed architecture. The complete orchestration has not been validated on live BESF work. Existing research and vendor examples already cover important parts; the combination is not claimed as a new invention.[1][6][14]
A giant tree is the wrong thing to store. Keep the distinctions that matter; explore the relevant branches.
For the cancellation change, the system should examine the paid-through date, repeat clicks, tenant scope, the interface promise and the authority to act. It should not ask a language model to subtract dates or invent a refund policy.
The source research contains 48 named questions across product, demand, marketing, software, AI, harness, economics and evidence. Their coverage and Jev accuracy remain unvalidated on real tasks.
Store task types, evidence requirements, action preconditions, measured outcomes, cost and prior counterexamples. Let a generator propose missing distinctions; test them on held-out tasks before promotion. Search can retain several candidate actions while code enforces budgets and permissions. A typed probability is an observation, not a fact or a permission. Never infer purchasing from a coding benchmark, or silently convert a simulated result into field evidence.
If the routine decisions stop returning to you, the same person might handle more useful work. If exceptions and upkeep grow instead, the system has simply moved the workload.
Hypothetical workload. Baseline: 12 owner minutes per task. Prepared result: 2 minutes to review. An exception needs the original 12 minutes plus 3 minutes of overhead. No quality uplift is assumed.
A gain only counts if accepted quality and business outcomes hold up. Deferrals, missed requirements, support and fatigue are not measured here.
N is task count; d the exception share; r the residual review; e exception overhead; u upkeep. These are editable planning inputs, not outputs of a validated capability model. Holding other defaults fixed, 80% exceptions makes the system take more owner time. A queue can also become unstable even when its average looks acceptable.
The opportunity is to turn current model access into reusable capability before the next repeated task. There is no sourced expiry date. You do not need a new platform or a larger model bill to begin.
OpenAI’s internal engineering account identifies QA capacity as a bottleneck and describes giving agents direct access to runnable environments and feedback.
Read the engineering account ↗Current Jev pricing is $0.042 per million input tokens. Shared context makes a broad question set economical; its usefulness still needs local evaluation.
Read the current model documentation ↗Anthropic’s harness work describes incremental progress and persistent artifacts between sessions. The result should leave tomorrow’s agent a better starting point.
Read the harness account ↗Attributed source notes, not social screenshots or independent replications. This is the current rationale for a bounded experiment, not evidence of a time-limited commercial advantage.
The laboratory supports some implementation claims. It does not establish autonomous delivery, safe production use or paying demand. Passing more tests does not fill those evidence gaps.
The teaching controller checks the outcome and stops on missing evidence or withdrawn permission. The runnable starter has no network or production adapter.
The complete prior laboratory is preserved in the combined ZIP. Its historical test counts are not presented as new runs of this article.
Run code-only where applicable, verified frontier, and the proposed router on the same frozen task families and permissions. Measure final state, deferrals, rework and complete cost.
Correct code does not show that anyone needs it. Simulated personas do not show that people understand this page. Real task and reader evaluation remain separate.
Use a task already authorised in your workspace. Ask for the outcome, evidence and smallest remaining choice together. The actual harness or review process must enforce the boundary.
I don’t need my agent to be more convincing. I need the next decision to cost me less attention.
Start with a known failure or repeated piece of work. Keep the simple baseline. Add broad interpretation only where it changes what happens next. This is how I would test the BESF idea before building the giant version.
No signup. No installation required to read it. The ZIP adds a Python demonstration, checks, methods and the original research.
Illustrative packet. A structured format makes missing evidence visible; it does not make the content true.
They all said it worked.
None of them proved how.
“They” are the fictional approving reviewers in this demonstration. The missing paid access is the counterexample. The check shows its rule and result; it is not a universal mathematical proof.
Web sources checked 20 September 2026. Source notes are paraphrases unless explicitly marked as a short quote. Simulations are not live-model or customer results.
Jev 1.13.0: $0.042/M input tokens, free output; text only; 64k request and 32k state-plus-longest-question limits. Provider documentation, checked 20 September 2026.
Provider cookbook reports $0.000497 for one 13-question batch versus $0.006090 for separate calls. Its speed comparison sums sequential call latencies, not concurrent wall time. The example uses Jev 1.12 and five repeats; it is not our benchmark.
11 February 2026. An internal engineering account describes human QA becoming a bottleneck and making the environment, repository knowledge and feedback legible to agents. Not a controlled BESF trial.
Standard short-context uncached input/output: $10/$50 per million tokens, checked 20 September 2026. Cache, batch, tool and long-context pricing differ. Token quantities in our examples are assumptions.
Official model ID claude-fable-5-1; standard input/output $10/$50 per million tokens, checked 20 September 2026. Equal pricing does not establish equal capability.
26 November 2025. Engineering account about incremental work, persistent progress and end-to-end feedback across sessions. Scope: their environment, not our measured outcomes.
Subscriptions and API billing are separate. No account allowance or permission was inspected. Included-call and 99%-subsidy cases are hypothetical.
Design precedent for an argument readers can manipulate and challenge. Does not prove that these articles improve comprehension.
3 December 2006. Defer secondary complexity; retain the immediate task and important limitations. Practitioner HCI guidance, not a universal reading-time law.
Design guidance for making destinations and likely value apparent. Used for specific action labels and figure captions; no conversion claim.
Nonessential interaction animation should be disableable. Implemented motion controls do not by themselves certify WCAG conformance.
WikiProject advice page, not Wikipedia policy or an authorship detector. Used to challenge inflated claims, boilerplate and unsupported significance—not to disguise AI assistance.
Context relevance, indirection and judgment limitations need domain-specific evaluation. Probabilities are neither business outcomes nor authority.
24 April 2026. Practitioner critique of decorative glow, arbitrary status, repeated cards and weak hierarchy. Useful design challenges, not an authorship detector or causal study of this article.
Prior work on model routing. Our cost equations and illustrative policy do not reproduce its experiments or establish a new routing algorithm.