Lessons · Lesson 4 of 7
- 01 · A field rate and a document rate are different numbers
- 02 · Right about every field, wrong about which document
- 03 · A confidence score is a ranking, not a probability
- 04 · Tested in documents, failing in money
- 05 · The class it gets wrong is the class you needed
- 06 · The failure that arrives with no version number
- 07 · What to check, what to leave, and the question that decides it
Tested in documents, failing in money
Test a matcher in the unit its failures are paid in, split the sample by document shape rather than by supplier, and read an error rate as a spread instead of an average.
Lesson 4 of 7 · 18 min
The situation
Halmside's second machine job is matching. A supplier invoice arrives, and something has to decide two things: which goods-received docket it covers, and which purchase-order line it settles.
Three records, three different origins, and they never agree exactly. A docket says 11,240 metres where the invoice says 11,256. A lot number is written two ways. One delivery is split across two dockets and one invoice.
Naji ran an acceptance test before switching the matcher on. 200 invoices drawn at random from the year, matched by machine, every match checked by hand. Four were wrong.
98.0%. The test passed. It should not have.
The year, by supplier
| Supplier | Invoices | Share of invoices | Value USD | Share of value | Error rate |
|---|---|---|---|---|---|
| Pellworm Mills | 148 | 12.9% | 1,912,400 | 61.8% | 9.5% |
| Orsett Trims | 312 | 27.1% | 218,900 | 7.1% | 0.6% |
| Kirkgarth Thread | 196 | 17.0% | 96,400 | 3.1% | 0.5% |
| Sheldrake Labels | 224 | 19.5% | 141,700 | 4.6% | 1.3% |
| Oldknow Fasteners | 158 | 13.7% | 402,800 | 13.0% | 1.9% |
| Bramhope Packaging | 113 | 9.8% | 322,100 | 10.4% | 0.9% |
Weight those error rates by the number of invoices, and the matcher is wrong 2.07% of the time. Weight them by the money on the invoices, and it is wrong on 6.33% of the value.
The same matcher, the same year, the same errors: 2.07% or 6.33%, depending only on what you divide by. A factor of 3.06, and nothing in the test told anybody which of the two numbers they had been shown.
The sample was drawn in the wrong unit, and it was too small to argue with
Two separate faults. The second is the one that applies everywhere.
The unit. Drawing 200 invoices at random is a fair way to estimate the error rate per invoice, and 2.0% observed against 2.07% true is a good estimate. It is simply an estimate of something nobody cares about. Halmside does not lose invoices. It loses money, and the money is not spread out like the invoices are.
The size. Pellworm sends 12.9% of the invoices, so a random 200 contains about 25.7 of them. Naji's sample had 26, and 2 of those 26 were wrong.
Two errors in twenty-six. Ask what true error rate fits that observation, and the answer, at the usual confidence, runs all the way from 0.9% to 25.1%.
So the sample cannot tell a supplier that is fine from a supplier that is failing one invoice in four. And that supplier carries 61.8% of everything Halmside pays out.
Naji drew again: 60 Pellworm invoices, nothing else. Six were wrong, which is 10.0%. The range that still fits runs from 3.8% to 20.5%. That is wide, and wide is the honest answer at 60. But it no longer contains 0.9%, so the question of whether there is a problem is settled, even though its size is not.
The supplier was standing in for the real thing
Then he read the fourteen Pellworm errors of the year instead of counting them, and what he found is better than what he went looking for.
Pellworm invoices come in two shapes. 87 cover a single dye lot, which is one run of fabric through the dyeing machine. 61 cover more than one, and on those the template repeats the purchase-order number in a second column so the lot references line up underneath it.
All fourteen errors are on the multi-lot invoices. 23.0% on those 61, and nothing at all on the other 87.
So the thing that predicts an error was never the supplier. It is a document shape, and Pellworm happens to be the supplier who produces it. Two things follow at once:
- The group to sample is multi-lot invoices, wherever they come from. Any other supplier who starts issuing them inherits that error rate on their first one.
- You cannot find this by counting. Fourteen errors read one at a time took Naji under an hour, and no amount of summary reporting would have produced it.
What the fourteen actually cost
| USD | |
|---|---|
| Twelve caught by a person before anything acted on them, at a mean of USD 46.00 of handling | 552.00 |
| One closed a fabric line against goods that had not arrived; found when the cutting room opened for a lot that was not there, two line-days at USD 1,290.00 | 2,580.00 |
| One paid an invoice of USD 8,832.00 twice; recovered after 71 days, financed at 11% | 188.98 |
| Total | 3,320.98 |
Twelve of the fourteen errors cost USD 552.00 between them. Two of them cost USD 2,768.98, which is 83.4% of the whole bill.
The errors sit in one format, and inside that format the cost sits in two of them. That is why an error rate is not a measure of risk. It is an average over a spread that keeps almost all of its weight at one end, and that end is the only part that matters.
Check yourselfA matcher is offered to you with an accuracy figure from a pilot on 300 documents. What is the first thing you ask for?Show the answer
The breakdown by supplier and by document shape, with the count in each cell. A single figure over 300 mixed documents cannot tell you whether the errors are spread thinly or sitting entirely on the one supplier that carries most of your money, and those two worlds give opposite answers about whether to buy it. If the cells are too small to read, that is an answer too: the pilot has not measured the thing you need to know.
Lesson 5 turns from matching to sorting, and to the class of document that is rare, expensive, and the one the sorter gets wrong.