Lessons · Lesson 5 of 7
- 01 · A field rate and a document rate are different numbers
- 02 · Right about every field, wrong about which document
- 03 · A confidence score is a ranking, not a probability
- 04 · Tested in documents, failing in money
- 05 · The class it gets wrong is the class you needed
- 06 · The failure that arrives with no version number
- 07 · What to check, what to leave, and the question that decides it
The class it gets wrong is the class you needed
Read a classifier by how often it finds the rare class that costs you money, rather than by its overall accuracy, and price the fix in the currency it is actually paid in.
Lesson 5 of 7 · 17 min
The situation
Halmside's third machine job is the easiest to describe and the easiest to be misled about. Attachments arrive by email all day. Something has to decide what each one is, so it can be filed, sent to the right desk, and, for some types, extracted.
Nine types. Twelve weeks. 3,180 documents. Naji checked every one against what it actually was.
Overall accuracy: 95.2%.
That is the number on the report. There is nothing wrong with it, except that it answers a question nobody in a factory has.
The same measurement, by class
| Document type | Documents | Share | Found correctly |
|---|---|---|---|
| Supplier invoice | 968 | 30.4% | 99.0% |
| Packing list | 542 | 17.0% | 98.2% |
| Order confirmation | 431 | 13.6% | 97.4% |
| Mill test report | 388 | 12.2% | 96.1% |
| Buyer purchase order | 296 | 9.3% | 98.6% |
| Delivery docket | 244 | 7.7% | 95.5% |
| Lab-dip approval | 161 | 5.1% | 88.8% |
| Debit note | 61 | 1.9% | 63.9% |
| Purchase-order amendment | 89 | 2.8% | 41.6% |
That last column has a name: recall. It is the share of the documents of one type that the sorter actually found.
Now read the last row against lesson 2. A purchase-order amendment is found 41.6% of the time. Eighty-nine of them arrived in twelve weeks, and fifty-two were filed as something else, usually as a purchase order, because that is mostly what an amendment looks like.
So the one document type whose whole purpose is to change a value the factory plans from is the one type the sorter cannot find.
Why the rare class is always the wrong one
Two reasons, and both are arithmetic rather than bad luck.
It carries almost no weight. Getting every single amendment wrong costs 2.8 points of overall accuracy. Getting 3% of supplier invoices wrong costs 0.9 points. Anything tuned to make the overall figure large barely cares which of the two it gives up. Halmside cares a great deal: one of them is a filing annoyance, and the other is USD 4,030.06 a time.
It looks like the thing it amends. An amendment carries the same purchase-order number, the same buyer letterhead, the same style code, the same delivery terms. Every feature that identifies a purchase order is sitting on it. That is not a defect in the tool. It is a true statement about the documents, and no amount of tuning makes two nearly identical things easy to tell apart.
Moving the threshold, and what that costs
Halmside changed one rule. Anything that mentions a purchase-order number Halmside already holds, and is not laid out as a table of line items, goes to an amendment queue for a person to look at.
Twelve weeks later, measured the same way:
- Amendments found: 41.6% to 92.1%. Of 89, it found 82 instead of 37.
- Precision on the queue: 84.2% to 38.7%. Precision is the share of the queue that really is an amendment. The queue now holds 212 documents instead of 44, and 130 of them are not amendments, instead of 7.
Now price it. 123 extra documents to open and reject, at two minutes each, is 4.1 hours over twelve weeks. That is USD 46.74 at Halmside's loaded office rate of USD 11.40 an hour. Loaded means the wage plus the employer's costs on top of it.
Against that: 45 more amendments found. From lesson 2's season, 9 of 31 amendments changed a field the factory plans from, and 6 of those 9 went unnoticed until they cost money, at a mean of USD 4,030.06. Applied to 45, that is 8.71 expected costly amendments caught, worth USD 35,100.52.
What actually went wrong with it
The money was never the constraint. Twelve weeks in, Naji measured something he had not thought to measure.
The queue was worked on 47 of 60 working days. On the thirteen days it was not, four amendments sat unread for a mean of 3.1 days. A queue that is 61% noise stops being a queue and becomes a folder.
A control that is cheap in money and expensive in attention fails on the attention, and the sums above cannot see that at all, because attention does not appear in them.
The second move is what fixed it, and it took two lines. Sort the queue by whether the purchase-order number on the document belongs to an order that is currently open. Of the 82 amendments, 78 do. Of the 130 documents that are not amendments, only 31 do.
- Queue: 212 to 109.
- Precision: 38.7% to 71.6%.
- Recall: 92.1% to 87.6%.
Four amendments were given up to halve the queue, and the queue then survived a busy week. That trade is a judgement, and a defensible one. The point is that nobody could make it until somebody counted how often the queue was actually worked.
Check yourselfA classifier is offered at 97% accuracy across eleven document types. What single number would you rather have?Show the answer
The recall of whichever of the eleven types costs you most when it is missed, together with how many of them you receive. If that type is 2% of your volume, it can be found never and the headline figure still reads 95%. Ask for its precision too, because the fix is almost always to widen the net for that one class, and the price of widening it is paid entirely in wrong documents sitting in somebody's queue.
Lesson 6 asks the question that hangs over all three jobs: how do you find out that any of this has stopped working, when nothing announces it?