Lessons · Lesson 4 of 7
- 01 · Two kinds of question, and why the answer sounds the same
- 02 · The four things it did well, measured here
- 03 · The fact nobody wrote down
- 04 · When 95% right is worth less than no answer
- 05 · Ten tasks, sorted by the wrong column
- 06 · Five ways to be confidently wrong
- 07 · The four questions, and what Meskala kept
When 95% right is worth less than no answer
Work the arithmetic that decides whether an almost-reliable answer is worth having, on one factory's measured numbers, and find the two quantities that settle it — neither of which is the accuracy rate.
Lesson 4 of 7 · 20 min
One task, three ways of doing it
Meskala receives supplier order confirmations by email, around 240 a month across fabric, trim and thread. Somebody has to find the confirmed delivery date on each one and key it into the plan. It is dull, there is a lot of it, and it is exactly the shape of task an assistant is sold for.
Souad measured three ways of doing it.
- By hand. The merchandiser opens the attachment, finds the date, keys it. Timed over the month: 96 seconds a document.
- By hand, checking a proposal. The tool reads the document and proposes a date. The merchandiser confirms it or corrects it. Timed: 78 seconds a document.
- Not checking. Accept what the tool proposes. 6 seconds a document, which is the paste and the keystroke.
She measured the errors too, against the source documents, with a third person settling disputes.
- The tool got 228 of 240 right. That is the 95.00% from lesson 2.
- Re-checking 240 hand-keyed dates against the source found 3 wrong: 1.25%.
- Of the tool's 12 wrong dates, the merchandiser's check caught 7 and let 5 through.
Stop on the last of those, because everything turns on it. The check does not catch everything. It caught 58.33% of the tool's errors, and the five it missed were the plausible ones: a Tuesday in roughly the right week, in the right format, in the right field. There was nothing to catch the eye.
What a wrong date actually costs, and the number everybody uses instead
Souad traced all 12 wrong dates to what happened next.
| What happened | How many | Cost each | Cost |
|---|---|---|---|
| Caught within days by something else — a goods receipt, a supplier call | 9 | 40.00 | 360.00 |
| Caught late; a line plan had to be redone | 2 | 610.00 | 1,220.00 |
| Not caught until the line was set and standing | 1 | 8,900.00 | 8,900.00 |
| Total | 12 | 10,480.00 |
Two numbers come out of that table and only one of them is any use.
The median cost of a wrong date is 40.00. That is the number a merchandiser's instinct is built on, and it is honestly arrived at. Nine times out of twelve, that is what happened. Ask anyone on the floor what a wrong confirmed date costs and you will hear a version of "we usually catch it, it costs an hour".
The mean is 873.33, the total divided by twelve. It is four times as large as the second-worst outcome, and it is the number the arithmetic needs, because you cannot know in advance which kind of error you have just made. All the money is in the tail, and the tail is invisible from inside an ordinary week.
The three modes, costed
Labour at Meskala's loaded merchandising rate of 9.40 an hour. Error cost at the mean of 873.33. One month, 240 documents.
| Time | Labour | Errors | Cost of errors | Total | |
|---|---|---|---|---|---|
| By hand | 6.4 h | 60.16 | 3 | 2,619.99 | 2,680.15 |
| Tool, checked | 5.2 h | 48.88 | 5 | 4,366.65 | 4,415.53 |
| Tool, unchecked | 0.4 h | 3.76 | 12 | 10,479.96 | 10,483.72 |
(The error column multiplies the count by the rounded mean, so the unchecked row reads 10,479.96 where the traced total was 10,480.00. Four cents of rounding, carried through on purpose so the arithmetic can be followed.)
Read the two comparisons.
Unchecked is worse than doing it by hand by 7,803.57 a month. Nobody is surprised by that, and nobody was proposing it.
Checked is worse than doing it by hand by 1,735.38 a month. That is the result worth sitting with. The mode everybody agrees is the responsible one — a person reviews every single output, nothing is automatic — costs Meskala 1,735.38 a month more than not having the tool at all.
Take it apart. The tool saves 18 seconds a document. Over 240 documents that is 72 minutes, which is 11.28 of merchandising time. And it turns 3 errors a month into 5, which at the mean costs a further 1,746.66. The extra errors cost 154.85 times what the saved time is worth.
Why the check does not save you: reading is not agreeing
This is the mechanism, and it goes far beyond this task.
When the merchandiser works from a blank field, she is reading: find the date, decide what it says, write it down. When she works from a proposed date, she is agreeing: look at what is offered, decide whether anything is wrong with it. Those feel like the same activity. They are not, and the second one is measurably worse at catching a plausible error, because a plausible error looks fine by construction.
That is why the mode that looks most responsible is the one that fails. A person in the loop is not a control. A person doing an independent act of reading is a control. The moment you put a proposal in front of them, you have stopped buying that. And the fix is not to try harder. Souad's people were experienced and were being measured, which is close to the best case.
The general rule, which is what to take away
Strip the story out and the decision is four quantities.
a— how often the tool is right. Here 95.00%.k— the share of the tool's errors your check actually catches. Here 58.33%.e— how often the person is wrong doing it the old way. Here 1.25%.m— what an error costs on average. Here 873.33.
Plus the time saved per item, valued at the labour rate. Using the tool with a check beats the old way only when
(time saved, in money) is greater than [ (1 − a) × (1 − k) − e ] × m
At Meskala the left-hand side is 0.047 a document. The right-hand side is 7.28 a document. It loses by 7.23 a document, which over 240 documents is the 1,735.38 above. The same answer twice, by two routes, which is how you know the arithmetic is sound.
Now turn the equation around and ask what would have to change.
Hold the accuracy and solve for the catch rate. The check would have to catch 74.89% of the tool's errors for this to break even. It catches 58.33%. That gap of about sixteen points is the whole story, and it is a fact about how people check, not about the tool.
Hold the catch rate and solve for the accuracy. The tool would have to be right 96.99% of the time. It is right 95.00% of the time.
The finding: the gap that decides it is too small for any demonstration to see
1.99 percentage points. That is the distance between a tool worth having on this task and a tool that costs 1,735.38 a month.
Now ask how you would tell those two apart.
- On 20 documents, which is a generous demonstration, a tool at 95.00% produces 1.00 expected error and a tool at 96.99% produces 0.60. The difference is 0.40 of an error. You cannot see four-tenths of an error. Both demonstrations show you one mistake, or none, and both look identical.
- The batch has to reach about 51 documents before the two rates differ by one whole expected error. And one error's difference is not something to act on either.
- At 240 documents you expect 12.00 against 7.23, a gap of 4.77 errors, which is visible.
That is arithmetic on expected counts, not a sample-size calculation. Doing it properly is a statistician's job, and the same idea done properly is what an operating-characteristic curve shows in a sampling plan. Course 6.2 has that machinery for AQL — the acceptable quality limit used in inspection — and it is the same machinery here.
The conclusion is uncomfortable and it is the point of the lesson. The smallest test that can settle this question is larger than any demonstration you will be given, and it has to run on your documents. Not because vendors are dishonest, but because twenty documents cannot separate the two cases even in principle.
Check yourselfSame tool, same 95.00%, but now the task is drafting the Arabic of a supplier chase. Does the arithmetic still refuse it?Show the answer
No, and nothing about the tool changed. m did. A wrong Arabic draft is caught by the supplier's reply within hours and costs the time to rewrite it, so m is a few units of currency rather than 873.33. With m that small the right-hand side of the rule collapses and almost any time saving wins. That is lesson 5: the same accuracy is a bargain on one task and ruinous on another, and the thing that flips it is never the accuracy.