Lessons · Lesson 5 of 6
Buying without a benchmark
How to design an acceptance test on your own documents when every published figure was measured on somebody else's, why three true accuracy numbers from one trial ran from 99.1% to 5.0%, and why a threshold agreed after the result is not a threshold.
Lesson 5 of 6 · 19 min
Why this course quotes no published accuracy figure anywhere
Not one number in this course comes from outside Erimtan. That is a decision, not an oversight.
An accuracy figure is not a property of a system. It is the answer to three questions somebody else chose: which documents, measured against what reference, and counted at what unit of error. Change any one of the three and the figure changes, often by tens of points, with nobody misrepresenting anything. The rest of this lesson shows one trial producing three honest figures between 99.1% and 5.0%, and the whole spread comes from the third question alone.
So a published figure, however honestly obtained, is a statement about a set of documents that is not yours. The only set whose answer you can act on is the one in your own inbox. The good news is that measuring it is a week of work, and it is entirely yours to do.
Six decisions, and they are all made before anything is bought
The sample, and its strata
Erimtan's incoming documents divide sharply. 71% arrive on a buyer's own template, the same layout every time, generated by the buyer's own system. 29% are scanned, photographed, marked up by hand, or forwarded by an agent on a form of their own. A stratum is one of those groups, treated separately.
Those are two different problems, and the hard one is the small one. A random draw of 60 documents gives about 17 of the difficult kind. That is enough to produce a figure and not enough to trust it. So Erimtan drew 30 and 30, and reported per stratum, never as one number.
A blended accuracy figure is a weighted average of two different systems, and the weighting is your book's mix, which changes next season.
The ground truth, and the floor it reveals
Sixty documents at 22 fields is 1,320 values. Two clerks keyed them independently at 11.6 minutes a document: 23.2 hours, plus 1.8 hours agreeing the differences. That is 25.0 hours, or USD 615.00. Ground truth is the answer you are grading everything else against.
That is the step most factories skip, and skipping it costs more than the hours. The two clerks disagreed on 47 of the 1,320 fields, which is 3.6%. Not through carelessness. An unusual notation genuinely reads two ways, and a hand annotation genuinely is ambiguous.
So no measurement in this trial can be more precise than 96.4%. Any figure above that is inside the noise of the reference it is being measured against. A factory that has never established its own floor cannot tell a good result from an unmeasurable one.
The unit of error, which is where the whole spread lives
| Measured at | Structured | Unstructured | Both together |
|---|---|---|---|
| Fields correct | 99.1% | 95.3% | 97.2% |
| Documents with every field correct | 90.0% | 53.3% | 71.7% |
| Documents where a wrong field changed a decision | 0 | 3 | 5.0% |
Every row is true, and they are answers to different questions. There is also an arithmetic hint worth noticing. If 22 fields are each right 97.2% of the time, and the errors are independent, the document rate would be 53.5%. The observed rate is 71.7%. So the errors cluster: the same awkward document is hard in several fields at once. That is exactly what you would expect, and exactly what a field-level figure hides.
One more, because it is the number a supplier will show you. Weighted by Erimtan's actual book mix rather than by the 30-and-30 draw, the document rate is 79.4% rather than 71.7%. That is a gap of 7.7 points, from the same sixty documents, with nothing wrong with either figure. Report the unit and the weighting, or report nothing.
Escape, not accuracy
Accuracy is the wrong quantity. There is a reviewer, so what matters is how often it is wrong in a way the reviewer will not catch.
Of the 37 wrong fields, a reviewer working at the measured review time caught 31, which is 83.8%. The 6 that got through were all plausible: a date in the right format that was the wrong date, a quantity that is a real quantity for that style, a tolerance that is a normal tolerance.
A reviewer catches the implausible and misses the plausible, and the plausible errors are the ones that change a decision. Three of the six changed one, and all three were in the unstructured stratum.
The threshold, worked out and written down before the documents were drawn
| Line | Amount |
|---|---|
| Documents a year | 3,180 |
| Clerk keying time today | 13.8 minutes |
| Review time with the tool, structured / unstructured | 2.1 / 5.6 minutes |
| Time saved a year | USD 13,931.39 |
| Less the annual fee | USD 8,940.00 |
| Net before any error | USD 4,991.39 |
| Escaped error reaching a real consequence, 1 in 8 at | USD 780.00 |
| The other 7, caught later at 0.9 hours | USD 22.14 |
| Expected cost of one escaped decision-changing error | USD 116.87 |
| Break-even escape rate | 1 in 74.5 documents |
| Threshold written down and signed | 1 in 60, and review time under 4.0 minutes |
The threshold is deliberately stricter than break-even, and the reason is worth saying out loud. A threshold set at exactly break-even buys a year of work for nothing.
The comparison, which is against what you do now and not against perfection
The same sixty documents were keyed by Erimtan's own clerks under normal conditions and scored the same way: 98.9% of fields and 88.3% of documents.
Two things came out of that half of the exercise, and Tekand says it is the half she would keep if she could only run one.
The first is that nobody had ever measured the existing process either. Every argument in the building had been comparing a measured candidate against an unmeasured incumbent. That is how a system worse than what you have can win a meeting.
The second is that the errors sit in different places. The clerks mis-key quantities and dates under time pressure. The tool misreads unusual notation and handwriting. Of the 37 fields the tool got wrong, the clerks got 34 right. Of the 14 the clerks got wrong, the tool got 11 right. The overlap is 3. A process where each checks the other is better than either alone. That is a different purchase from the one that was proposed, and it was only visible because the incumbent was measured.
What the numbers decided
On the unstructured stratum: 3 escaped decision-changing errors in 30 documents, or 1 in 10. That is six times the threshold. Review time was 5.6 minutes against a limit of 4.0. It fails both tests, and it fails them before anybody argues.
On the structured stratum: 0 escapes in 30, and 2.1 minutes. It passes the time test. Here is the part that is almost always missed: it cannot be shown to pass the error test. With no failures in 30 documents the true rate could still be about 3 in 30, one in ten, which is the very rate being ruled out. A clean result on a small sample is not evidence of a small rate.
The sample size you need is set by the threshold you chose, and almost every pilot is too small to demonstrate the thing it was run to demonstrate. Erimtan extended the structured trial to 180 documents and got 1 escape, which is inside 1 in 60 with something to spare.
The decision: buy it for the structured stratum, keep the clerks on the unstructured one, and re-test at renewal under the clause of lesson 4. Same technology, same supplier, narrower scope, different answer. That is the shape course 24.1 sets out for a capture purchase, and this course does not repeat it.
Check yourselfA supplier's trial reports 97% accuracy on your documents. What are the three questions before you react to it?Show the answer
Which documents: the whole mix, or the easy stratum? Measured against what reference, and how far apart were two of your own people on the same fields? And counted at what unit: a field, a whole document, or a decision that changed? Erimtan's own trial gave 97.2% of fields, 71.7% of documents and 5.0% of documents where a decision moved, all from the same sixty documents. Until those three answers are on the page, 97% is not a number you can act on in either direction.
Prompt · Design my acceptance test and fix the threshold first
Before a trial, a pilot or a demonstration, and specifically before anybody sees a result.
Act as a sceptical evaluator who has watched a threshold be agreed after the result. Design an acceptance test on my own documents. I will give you: [WHAT THE SYSTEM IS SUPPOSED TO DO], [MY DOCUMENT TYPES AND THE SHARE OF MY VOLUME EACH ONE IS], [WHICH OF THEM ARRIVE ON A FIXED TEMPLATE AND WHICH ARE SCANNED, MARKED UP OR RETYPED BY SOMEBODY ELSE], [DOCUMENTS A YEAR], [HOW LONG MY OWN PEOPLE TAKE PER DOCUMENT TODAY], [MY LOADED HOURLY COST], [THE ANNUAL FEE], [WHAT A WRONG VALUE HAS ACTUALLY COST ME WHEN ONE GOT THROUGH, WITH HOW OFTEN THAT HAPPENED]. Do the following. First, tell me the strata, how many documents to draw from each, and where to draw them from, and insist on reporting per stratum rather than blended. Second, tell me how to establish ground truth, how many hours that will cost, and why the disagreement rate between two of my own people is the ceiling on everything else I will measure. Third, define three units of error — field, document, and decision changed — and tell me which one the threshold will be written in. Fourth, work the threshold out from my figures: the annual saving, the annual fee, the expected cost of one escaped error that changes a decision, and the break-even escape rate. Then propose a threshold slightly stricter than break-even, and say why. Fifth, tell me the smallest number of documents needed to demonstrate that rate, and warn me if my planned sample is too small to show it even with a clean result. Sixth, insist that my current process is measured on the same documents in the same way, and tell me what to do if the errors turn out to be in different fields.
AI can make mistakes — check anything you act on.
What to take away
- An accuracy figure is a statement about a set of documents, a reference and a unit of error. None of the three is yours until you run the test.
- Split the sample into strata. A blended figure is a weighted average of two different systems, weighted by a mix that changes.
- Establish your ground truth and read its disagreement rate. Erimtan's two clerks differed on 3.6% of fields, so nothing above 96.4% was measurable.
- Measure escape, not accuracy. A reviewer catches the implausible and misses the plausible, and the plausible ones move decisions.
- Work out the threshold before the trial and sign it. Erimtan's break-even was 1 in 74.5 and it wrote down 1 in 60.
- Measure what you do now, on the same documents. Nobody ever has, and their errors may not even be in the same fields as yours.
- A clean result on 30 documents cannot demonstrate a rate of 1 in 60. The threshold sets the sample size, not the calendar.