Lessons · Lesson 3 of 6
Measuring a person with a record of a supervisor
How an absence model built from payroll data learned which supervisor writes things down, why its two complaints had one cause, and what happens to an error whose entire cost lands on somebody with no route to report it.
Lesson 3 of 6 · 19 min
A real problem, solved by a competent person
Bora Sunel runs production at Erimtan. His problem was concrete and expensive. On a morning with unplanned absence, nobody spotted the gap until the line missed its first hour. A floater was then found and placed at about 07:40, and the line had already lost most of an hour. A floater is a trained operator kept spare to fill a gap.
So he built something. Three years of attendance records were sitting in the payroll export. He used them to produce, at 07:00 each day, a list of roughly 18 names likely to be absent. Supervisors used it. They liked it. Nobody asked him to build it, and nobody asked him to stop.
Two complaints, four months apart, treated as two problems
The first came from line 4, through the workers' representative. The list is unfair, the same names keep appearing, and being on it costs money. It was treated as a personnel matter.
The second came from the planning office. The list is useless on line 7, which keeps losing hours nobody predicted. It was treated as a modelling matter.
They are the same problem. The audit that found it used thirty days of gate records: the barrier log at the personnel entrance, kept for site security, with nothing to do with payroll.
| Line 4 | Line 7 | |
|---|---|---|
| Late arrivals in the gate log | 41 | 63 |
| Of those, written up by the supervisor as late | 39 | 10 |
| Share written up | 95.1% | 15.9% |
| Recorded instead as a shift swap or an agreed start | 2 | 53 |
Line 7 had 53.7% more late arrivals than line 4, and about a quarter as many written up. Nothing in that table is a falsification. A shift swap is a correct payroll entry, made by a supervisor doing exactly what a supervisor is supposed to do: keeping the line running and the pay right. Line 7's supervisor arranges cover and books it as cover. Line 4's writes down what happened.
Erimtan then measured the model against both records over sixty days. Its risk score matched how often a line's supervisor writes a late arrival up, at 0.91. It matched the gate log, meaning people actually arriving late, at only 0.14.
The model had learned the supervisors, not the operators. It was accurate about the thing it was trained on, and that thing was not absence.
The general form, which is worth more than the incident
A record is fit for the purpose it was collected for, and for nothing else. The payroll record's purpose is to pay people correctly. Paying people correctly does not require it to be comparable between lines. Two supervisors can use two different words for the same morning, and payroll is right both times.
Comparability between people is exactly what a model needs, and exactly what nobody promised. A model built from a record of who was written up learns who was written up. Every factory has this: an attendance record, a rework log, a defect-attribution sheet, a machine-downtime reason code. Each is a record of a decision by a person, not a record of an event.
Whose obligation this is, and how it reached a factory in a country that has none
Two ideas are worth naming precisely here, and both are easy to overstate.
The first is purpose limitation. Modern data-protection law says personal data collected for a stated purpose must not later be used in a way that conflicts with that purpose. Attendance collected to run payroll, then used to score a person, is the textbook shape of that question. Whether any particular law applies to any particular factory is a question for a lawyer in that country, and this page will not pretend otherwise.
The second is that consent is a weak instrument inside an employment relationship, because it is not a bargain between equals. Nobody at Erimtan asked an operator whether her attendance record could be used this way. If anybody had, the answer would not have meant much. So the honest question is whether the new use fits the purpose the record was collected for, not whether anyone objected.
Now the practical part, which most factories discover the hard way. Erimtan is not established anywhere the European regulation applies. The obligation reached it anyway, through Halbertsma's supply agreement. That agreement requires suppliers to handle workers' personal data to the buyer's own standard, and to be able to show how. The route is the contract, not the country. A factory that concludes "this does not apply to us" has read the wrong document. What applies is whatever its largest customer has already signed up to and passed down.
The imbalance that made the error invisible
Being on the list three times in a month took an operator off the Saturday overtime roster. That rule was never written anywhere. It was a reasonable supervisor's reasonable conclusion that the list meant something.
| Amount | |
|---|---|
| Saturday overtime, 8 hours at 1.5 times a base of USD 1.38 | USD 16.56 |
| Four Saturdays | USD 66.24 |
| Against a monthly wage of | USD 262.20 |
| Share of a month's pay | 25.3% |
| Cost to the factory: a floater held 0.6 hours unnecessarily at USD 1.61 | USD 0.97 |
| The operator pays this many times what the factory pays | 68.6 |
This is the governance point of the whole course, and it is arithmetic rather than sentiment. When the entire cost of a wrong answer lands on somebody with no route to report it, the error is not hidden from the people who could fix it. It is invisible to them. Nobody at Erimtan was ignoring a signal. There was no signal. The model's errors produced no cost, no complaint and no ticket anywhere a manager looks.
A system whose errors are free to the organisation running it will not be corrected by that organisation, however well-meaning it is. Something has to make the error land where it can be seen.
The route, and the number it produced
The mechanism is unglamorous. Any operator may ask her supervisor why she is on the list, and refer the answer to Tekand if it does not satisfy her. Supervisors were told to treat the question as normal and never to note who asked.
In the first quarter there were 14 challenges. 9 were upheld, which is 64.3%.
Read the number carefully. It is not evidence that the route is generous. It is the only measurement of the model's error rate that Erimtan was ever going to get, because nothing else in the building was counting. And no system survives being wrong about two named people in every three it names.
Note also what 14 measures. In the first quarter of any challenge route, the count measures how safe people feel asking, not how often the system is wrong. Erimtan's second quarter had 31 challenges and a lower upheld rate. That is the route working, not the model getting worse.
The ending nobody expected
The obvious repair is to feed it the gate log instead of the payroll record. That fails for the same reason in a different coat. The gate log is kept for site security, so using it to score a named person is the identical question wearing a different uniform.
So Erimtan stopped scoring people. The replacement forecasts, for each line, how many operators it is likely to be short tomorrow. It uses that line's own totals over time, with no name in it anywhere.
It is more accurate. Measured over sixty days, the line-level forecast was out by 1.4 operators a day, against the person-level model's 2.9. The reason is not luck. A late arrival and a shift swap leave the line short by one operator for the same period. The two words that corrupted the person-level model describe the same hole in the line, so at line level the difference simply does not arise.
The version that needed no personal data was also the better one. Sunel is clear that he would never have found it if the line-4 complaint had stayed a personnel matter and the line-7 complaint had stayed a modelling one.
Check yourselfA supervisor tells you the absence list is accurate: the people on it really do have poor attendance records. Is that a defence?Show the answer
No, because it is circular. The record is what the list was built from. The test has to come from a source collected by somebody with no stake in the outcome, and for a different purpose: a gate log, a bus manifest, a canteen count. Erimtan's model matched how often a supervisor wrote a late arrival up at 0.91, and matched people actually arriving late at only 0.14. Both of those numbers had to come from outside the system to exist at all.
Prompt · Test a list that names people before it names anybody
When anything in your factory produces a ranking, a score or a daily list of named workers or named suppliers.
Act as an auditor who has seen a workforce model turn out to be measuring its supervisors. I will describe a system that produces a list of named people. I will give you: [WHAT THE LIST IS FOR], [THE RECORD IT WAS BUILT FROM AND WHAT THAT RECORD WAS ORIGINALLY COLLECTED FOR], [WHO ENTERS THAT RECORD AND WHETHER THEY HAVE ANY CHOICE IN HOW AN EVENT IS CLASSIFIED], [WHAT HAPPENS TO A PERSON WHO APPEARS ON THE LIST, INCLUDING ANYTHING INFORMAL THAT IS NOT WRITTEN DOWN], [WHETHER THEY CAN FIND OUT WHY], [ANY INDEPENDENT COUNT OF THE SAME EVENTS THAT EXISTS ANYWHERE IN THE BUILDING]. Do the following. First, tell me whether the underlying record is a record of events or a record of somebody's decisions, and name the person whose behaviour it may actually be measuring. Second, design the cheapest possible test against the independent count I named, or tell me what to start counting if I named none, and say how many days of it would settle the question. Third, price a wrong entry twice: what it costs the person named, including the informal consequence, and what it costs the company. Give me the ratio, and say what that ratio means for whether anybody inside will ever notice an error. Fourth, draft the challenge route in three sentences: who may ask, who answers, what gets counted. Tell me the two numbers to report monthly. Fifth and last, propose a version of the same system that needs no personal data at all, and say honestly whether it would be worse.
AI can make mistakes — check anything you act on.
What to take away
- A record is fit for the purpose it was collected for, and for nothing else. Payroll never promised to be comparable between lines.
- A model built from a record of who was written up learns who was written up. Erimtan's matched the supervisor at 0.91 and the gate at 0.14.
- Two complaints in different departments can be one defect. Treating them separately is what kept it alive for four months.
- The route is the contract, not the country. The obligation reached Erimtan through its buyer's supply agreement.
- When the whole cost of an error lands on the person with no route to report it, the error is invisible rather than ignored. Erimtan's ratio was 68.6.
- Build the challenge route first and count it. A 64.3% upheld rate was the only error measurement that existed.
- The version with no personal data in it was more accurate: 1.4 operators a day against 2.9.