Lessons · Lesson 1 of 7
- 01 · A field rate and a document rate are different numbers
- 02 · Right about every field, wrong about which document
- 03 · A confidence score is a ranking, not a probability
- 04 · Tested in documents, failing in money
- 05 · The class it gets wrong is the class you needed
- 06 · The failure that arrives with no version number
- 07 · What to check, what to leave, and the question that decides it
A field rate and a document rate are different numbers
Turn a per-field accuracy figure into the number that describes your whole order, and tell apart the errors a reviewer can see from the ones that look right on the page.
Lesson 1 of 7 · 17 min
The situation
Halmside Garments runs eight sewing lines. It makes woven bottoms: trousers, shorts, cargo pants.
Since March it has put all its incoming paperwork through a system that reads a document and hands back the values on it as fields. Quantities, dates, prices, codes, terms. This course calls that system the extractor.
Hadeel Zeidan runs the Ravensgate Stores account. When the extractor was installed she was told one number. She repeated it to her production manager, because it sounded like the end of an argument.
Ninety-nine point two per cent.
This lesson asks one question about that number: a rate of what? It is the plainest lesson in the course, and everything after it stands on it.
What was actually measured
Naji Kanoun runs Halmside's office systems. He did the only honest thing available.
He took six weeks of documents that staff had already typed in by hand. He ran the same documents through the extractor. Then he compared the two, field by field.
- 612 documents
- 14,208 extracted fields
- 114 of them wrong
That is 99.20% of fields correct. The figure is real. It is Halmside's own, measured on Halmside's own paperwork.
It does not transfer to you. Another factory has other suppliers, and those suppliers use other document layouts. They would give a different number. That is why this course never quotes anybody else's figure.
Now the arithmetic nobody did on the day.
A document is not a field
Ravensgate's purchase order carries 47 fields that Halmside extracts.
Say each field is right 99.20% of the time. Say one field going wrong does not make the next one more likely to go wrong. Then the chance that the whole document is clean is not 99.20%. It is 0.9920 multiplied by itself 47 times:
68.6%.
Roughly one buyer purchase order in three carries at least one wrong field somewhere on it. The rate did not change. The unit did.
Now follow it out to the whole order. Purchase order RVG-51188 is 42,000 pieces of style HM-2264, a men's woven cargo short. Nine kinds of document carry that order before the container leaves.
| Document | Fields extracted |
|---|---|
| Buyer purchase order | 47 |
| Purchase-order amendment | 11 |
| Supplier order confirmation | 26 |
| Supplier invoice | 34 |
| Packing list | 38 |
| Goods-received docket | 16 |
| Mill test report | 21 |
| Booking confirmation | 19 |
| Commercial invoice | 14 |
That is 226 extracted fields on one order. At the same per-field rate, the chance that all 226 are correct is 16.3%.
The part that makes it workable
16.3% sounds like a reason to give up. It is not, because the 226 fields are not equal.
Course 10.1 lesson 3 follows one wrong field through six decisions and prices each one. Its finding is what matters here: a field is dangerous because of how many decisions use it before anybody looks at it again, not because of what the field is.
So Halmside counted. Of the 226 fields, 61 reach a decision before any person reads them again. The other 165 are read by somebody, or worked out again further down the line, or simply never used.
At the same per-field rate, the chance that all 61 of those fields are correct is 61.3%.
The honest statement of Halmside's position is this: about two orders in five carry a wrong field that something will act on. You can work with that number. It also points your review at 61 fields instead of spreading it over 226.
Check yourselfYour supplier quotes a per-field accuracy of 98%. Your goods-received docket has 16 extracted fields. What share of dockets arrive clean?Show the answer
0.98 to the sixteenth power, which is 72.4%. So more than one docket in four carries at least one wrong field. Do this sum in front of whoever quoted you the 98%. It is the same fact in the unit you actually work in, and it usually ends the argument about whether a review step is needed.
Four ways a value is wrong, and only one of them looks wrong
Naji then did the more useful half of the work. He read all 114 errors, one at a time, and sorted them by what had gone wrong.
| What went wrong | Fields | What a reviewer sees on the screen |
|---|---|---|
| A character was read wrong | 41 | A value that does not match the document |
| The right value landed in the wrong field | 28 | A value that matches the document |
| A correct value taken from the wrong document | 12 | A value that matches a document |
| A value that was true when the document was written | 33 | A value that matches the document it came from |
41 of 114, or 36.0%, are copying errors. A digit was misread. These are the ones a review catches: put the extracted value beside the document, and the two do not agree.
The other 64.0% are correct readings. Every one of them copies something real, exactly.
Picture the reviewer. She opens the document, finds the number, sees that it matches, and ticks the row. She has confirmed the error, not caught it. She did her job properly, and the wrong value went through with her signature on it.
That finding is what this course is built on, and it changes what a review is for.
- 28 mapping errors. A correct value sitting in a field that means something else. This is course 10.1's subject arriving by machine. The repair is a clear field definition, not a second look at the page.
- 12 provenance errors. A correct value taken from the wrong document. That is lesson 2, and it is the most expensive shape in this course.
- 33 stale values. True of the document they came from, no longer true of the order. Course 10.1 lesson 5 already prices what it costs to correct a value and undo the decisions made from it. Nothing here repeats that.
The rate you were given is a ceiling
One last thing, and it belongs at the front rather than in a footnote.
Every accuracy figure in this course is a rate of errors that were found in the end. Naji's 114 are the fields where the hand-typed value and the extracted value disagreed. Where the hand-typed value was wrong in the same way, the two agreed, and both were counted as correct.
So 99.20% is a ceiling, not an estimate. The same is true of the figures in lessons 3, 4 and 5. It is true of any accuracy figure anybody ever shows you. A measured error rate can only ever count the errors your measurement was able to see.
Lesson 2 takes the smallest column in that table, 12 fields out of 14,208, and follows one of them to the money.