Lessons · Lesson 7 of 7
- 01 · A field rate and a document rate are different numbers
- 02 · Right about every field, wrong about which document
- 03 · A confidence score is a ranking, not a probability
- 04 · Tested in documents, failing in money
- 05 · The class it gets wrong is the class you needed
- 06 · The failure that arrives with no version number
- 07 · What to check, what to leave, and the question that decides it
What to check, what to leave, and the question that decides it
Set a break-even error rate for each machine job on your order, correct it for the share of errors your check can actually catch, and know which of your figures will not survive being attacked.
Lesson 7 of 7 · 16 min
The situation
Halmside now has six jobs running on its order data, and a measurement for each of them. The remaining question is the only one the office actually argues about: which of these does somebody look at, and which go through unread?
The instinct is to sort by accuracy and check the worst. That instinct is wrong, and this lesson is the arithmetic that replaces it.
The rule, and it is older than the technology
Two costs decide it. The cost of checking one item, and the cost of one wrong item going through unchecked. Set them equal, and you get a break-even error rate:
break-even error rate = cost of one check divided by cost of one error
Below that rate, checking costs more than the errors do, so you should check nothing. Above it, you should check everything.
There is no sensible middle, and that is the part that surprises people. The answer is all or none. The only question is which side of the line your measured rate sits on.
This is not a new rule invented for machine reading. It is the all-or-none rule for incoming inspection, and W. Edwards Deming set it out for materials arriving at a factory gate long before anything read a document. Every term in it means the same thing here. What is new is only that the incoming lot is a stack of PDFs.
Notice what is not in it. Volume does not appear. Neither does total spend, nor how impressive the accuracy figure sounds. A job that runs three times a year is decided by exactly the same comparison as one that runs three thousand times.
Six jobs, decided
Halmside's loaded office rate is USD 11.40 an hour, which is USD 0.19 a minute. Every error rate below is the factory's own measurement from the earlier lessons.
| Job | Measured error rate | One error costs USD | One check costs USD | Break-even error rate | Verdict |
|---|---|---|---|---|---|
| Buyer purchase-order header fields | 6.8% | 4,030.06 | 1.14 | 0.028% | Check every one |
| Supplier-invoice three-way match | 2.07% | 237.22 | 0.76 | 0.320% | Check every one |
| Packing-list quantity | 1.1% | 96.00 | 0.57 | 0.594% | Check every one |
| Mill test-report verdict | 3.9% | 1,860.00 | 0.38 | 0.020% | Check every one |
| Lab-dip routing | 11.2% | 41.00 | 0.19 | 0.463% | Check every one |
| Supplier identity on inbound mail | 0.9% | 0.60 | 0.19 | 31.7% | Leave it alone |
Read the last two rows against each other, because they are the whole lesson.
Lab-dip routing is wrong 11.2% of the time and gets checked. A lab dip is the small swatch a dyer sends for colour approval, and routing it means sending it to the right desk. Supplier identity is wrong 0.9% of the time and does not get checked.
The job with twelve times the error rate is the one you check, and the job that is almost always right is the one you leave. A misrouted lab dip costs USD 41.00 to unpick. A document filed under the wrong supplier costs USD 0.60 and thirty seconds of searching.
Why the bottom row is cheap, and it is not the error rate
Supplier identity costs USD 0.60 when it is wrong, and the reason is not that the job is easy. It is that something further down the line works the value out again anyway. The document is found by its purchase-order number the next time anyone wants it, and the misfiling corrects itself the moment it matters.
That is the third question, after how often is it wrong and what does wrong cost. It is usually the one that decides the row:
Does anything downstream produce this value independently?
Where the answer is yes, the cost of an error falls towards zero however often it happens, because the second working-out is the check, and you are already paying for it. Where the answer is no, as with a buyer's amendment, a test verdict, or a match that closes a line, the value stands alone and every consequence rests on it.
The correction that flips a row
The table above assumes a check catches the error. Lesson 1 measured that it does not. Of Halmside's 114 wrong fields, only 41, or 36.0%, are copying errors that a reviewer comparing a value against its document can see. The other 64.0% are correct readings of the wrong thing, and the reviewer confirms them.
So the honest break-even divides by the share the check actually catches:
break-even error rate = cost of one check divided by cost of one error, divided by the share of errors your check catches
At 36.0%, every break-even in the table roughly triples. Four rows are far enough clear of theirs that nothing changes. One is not.
Packing-list quantity has a plain break-even of 0.594% and a corrected one of 1.649%. Its measured error rate is 1.1%. The correction moves it from check every one to leave it alone. The check costs more than it saves, because it catches only about a third of what it is looking for.
The right response is not to stop checking. It is to change the check so that it catches more, using lesson 6's rule: compare the packing list against the goods-received docket rather than against the packing list. That is two independent origins instead of one. The catch share rises well above 36.0%, and the row comes back over its break-even.
Prompt · Price the check against the error
When you are deciding which machine-read values somebody still has to look at, and you want an answer from arithmetic rather than from preference.
Act as an operations analyst in a garment factory. I want a break-even error rate for each job a machine now does on my order data, and a verdict on each. For every job I will give you: the job, its measured error rate on my own documents, what one wrong item costs me when it is not caught, how many minutes one check takes, and whether anything downstream produces the same value independently. Here they are: [LIST ONE JOB PER LINE]. My loaded rate is [AMOUNT] an hour. From my own hand check, a reviewer comparing a value against the document it names catches about [PERCENT] of the errors that exist — if I do not know this, tell me how to measure it and use my number only once I have it. Do the following. First, for each job compute the cost of one check, then the break-even error rate as the cost of one check divided by the cost of one error, and say whether my measured rate is above or below it. Second, redo each break-even dividing by the share of errors my check actually catches, and name any job whose verdict changes between the two. Third, for any job where something downstream reproduces the value, say so and explain why that collapses the cost of an error rather than the rate of it. Fourth, rank the jobs by how much the decision would move if my cost-of-one-error figure were wrong by half, so I know which figure to go and measure properly. Fifth, tell me which of my inputs you trust least and why. Do not average anything across jobs, and do not give me a range where the arithmetic gives a number.
AI can make mistakes — check anything you act on.
What Halmside cannot measure, and neither can you
Four things, named rather than buried, because a course that ends on a tidy table has taught the wrong habit.
- Every error rate here is a ceiling. They are rates of errors that were found in the end. An error that never surfaced was counted as a correct value, in every lesson of this course and in every accuracy figure anyone will ever show you.
- USD 4,030.06 rests on six events. It is a mean drawn from a small number of amendments with one large value among them. It holds up two of the six rows, and it is the one to attack first.
- The correction assumes the catch rate stays put. The 36.0% was measured on 114 errors read deliberately, in one afternoon, by somebody who knew he was being measured. A person checking eight hundred dockets in a week catches less than that, and Halmside has no number for how much less.
- Nothing here prices the knock-on effect of not checking. A team that stops looking at a class of document stops noticing when that class changes, which is lesson 6's failure arriving through the door this lesson has just opened.
Check yourselfA job is wrong 0.4% of the time. Checking it costs one minute. Should you check it?Show the answer
Not enough information, and that is the answer. The break-even is the cost of a check over the cost of an error, so at USD 0.19 for the check you need to know what one error costs. If it costs USD 20.00, the break-even is 0.95%, and 0.4% sits under it, so leave it. If it costs USD 200.00, the break-even is 0.095%, and 0.4% is four times over, so check every one. The error rate on its own decides nothing, which is why it is the wrong thing to ask for first.
What to take back to your own desk
List every job on your order data that a machine now does. For each one, write four things: how often it is wrong, measured on your own documents; what one wrong one costs; what one check costs; and whether anything downstream produces the value again. The first is the only one that takes real work, and lesson 1 gives you the way to turn it into the unit your order is actually in.
Then check the rows above their break-even, leave the rest, and put the standing count from lesson 6 on the one supplier that carries most of your money. The whole table holds only while the documents stay as they are, and they will not.