The fairness test that already knew who should win
A simulation of contribution rewards appeared to answer a hard question. Its rules had access to information no real company would possess.
A shortcut crosses the boundary between a decision and its evaluation.AI-generated conceptual setting.
A company wants to reward people for the value they contribute. Hours are an imperfect measure. Visible output misses some kinds of work. A mathematical model offers the prospect of comparing different approaches before using one with real people.
Then the model supplies a piece of information the company would never have: how valuable each person’s contribution really was.
That happened in the archived simulation behind the company-design proposal. Several reward formulas could read the hidden value used to judge their performance. The comparison was supposed to test how well the rules recognised contribution. Part of the answer had already been handed to them.
Later material in the source pack withdraws the comparative rankings. The code inspection explains why.
Follow the idea
How did the fairness test get its answer?
A reconstruction of the archived simulation defect from static inspection. The old ranking was withdrawn; the private source is unavailable for public audit.
evidenceDecision
rule
outcomeEvaluation
Decision inputs and evaluation information have different jobs.
What the company could see
A reward method should work with imperfect evidence about a contribution. The evaluator holds the hidden truth back.
evidenceDecision
rule
outcomeContaminated comparison
A reliable implementation can still run an invalid comparison.
What the simulation also supplied
The archived reward formulas could read the hidden value used to score them. That gave the decision information no real company would possess.
evidenceDecision
rule
outcomeNew result unknown
Decision inputs and evaluation information have different jobs.
Take that knowledge away
Removing the shortcut reopens the comparison. The original ranking cannot say which reward method then works best.
The fictional world knew everybody’s true contribution
A simulation can create both an event and the truth about that event. That is useful: it gives the evaluator something against which to compare an imperfect decision.
But the simulated decision-maker should only receive information available to the real decision-maker. A company might see a delivered result, an investigation or evidence from colleagues. It does not receive a perfect measurement of what each contribution was worth.
The archived formulas crossed that boundary through a field called true_value. The name makes the defect easier to see, but the problem would be the same under any label.
Imagine comparing exam techniques while allowing students to consult the marking guide. Their answers might be accurate. The experiment would tell you little about which technique helps when the guide is removed.
In this case, more simulated trials would not fix the comparison. They would repeat the same information advantage. Scikit-learn’s explanation of data leakage describes the general problem: information unavailable when a real decision is made must not quietly influence that decision.
Two programs agreed on the same mistake
The archive contained a main implementation and a reference implementation. Static inspection found the hidden-value input in both.
That matters because a second implementation can look like an independent check. Agreement is reassuring when the two programs independently encode a sound premise. Here, both carried the premise that needed challenging.
They could agree about the arithmetic while failing to answer the question the arithmetic was meant to investigate.
The missing review question was ordinary enough: how would a real company know this? It did not require a more elaborate simulation. It required stepping outside the simulated world.
The uncomfortable part of withdrawing a result
A failed comparison leaves useful work behind. The software may still run reproducibly. Some checks may still establish that particular calculations behave as intended. Those facts need their own descriptions.
The ranking has to go. It cannot remain in an introduction as a promising result while a caveat several pages later explains why it is invalid.
There is also a limit to what this article can show. The source archive is private; readers cannot independently inspect it here. The account comes from static inspection and the pack’s withdrawal record, not a rerun. The public references explain the general issue rather than verify this particular archive.
A corrected comparison would have to remove the hidden information and run again. No such result has been verified for this article.
The original question remains worth asking: how could a company recognise contribution more fairly? The simulation exposed how easily a proposed answer can depend on knowing the very thing the company is struggling to discover.
