Lessons · Lesson 6 of 6
Running it: the refusals nobody counted
Price the approval that remains, size the audit that can actually detect a change, and set the triggers that redraw the line.
Lesson 6 of 6 · 19 min
An approval is worth its refusals
The boundary is drawn. Five rows still carry an approval, and an approval that never refuses anything is a form.
Serdana counted. In the first two weeks the reviewers refused 13 of 89 proposals, or 14.6%. In weeks nine to twelve they refused 4 of 214, or 1.9%. Course 10.3 gives a rule of thumb from a completely different kind of gate, and it lands in the same place: a gate that is approved more than about nine times in ten has stopped being a gate. Serdana's approval queue crossed that line in week six.
The refusal rate is only a proxy, though, and the desk went one better. What an approval is actually worth is the share of the wrong actions it catches. That can be measured directly, by comparing what the audit found later against what the reviewers had refused at the time.
| First quarter | Third quarter | |
|---|---|---|
| Wrong actions the audit found | 9 | 11 |
| Of those, refused at the approval | 4 | 1 |
| Catch rate | 44.4% | 9.1% |
| Expected loss a month on that row | 1,323.92 | 1,323.92 |
| Removed by the approval | 588.41 | 120.36 |
| Reviewers' hours a month | 210.60 | 210.60 |
| Order clock the wait costs | 2,294.08 | 2,294.08 |
At its best, the approval on that row removed USD 588.41 a month and cost USD 2,504.68. By the third quarter it removed USD 120.36. The reviewers' hours, which is the number every meeting is about, are 8.4% of what the control costs.
An approval is worth the catch rate times the loss it can prevent. An approval that refuses nothing removes nothing, however conscientiously it is clicked.
Count the refusals. If they are falling, either the system has got better, in which case the row should move, or the reviewers have stopped reading, in which case the row should move for the opposite reason. Both answers lead to the same action, and neither of them is to leave it alone.
The audit that could never have worked
Every autonomous row is sampled: one action in eight, re-checked against the evidence. It is a sensible-looking control, and it does not do what most people assume.
allocations a month 268
sampled at one in eight 33.5
error rate 1.9%
wrong actions the sample expects to find a month 0.64Two-thirds of one finding a month. A doubling of the error rate would take many months to show through that much noise. To expect five findings a month, which is the fewest that would make a change visible quickly, the sample would have to be 263 of the 268. That is all of them, which is the thing the automation was for.
So say plainly what the random sample is for. It keeps the recorded rate honest over a long window, and it can turn up a class of error nobody would have thought to look for. It is not a detector, and a desk that believes it is one has a monitoring plan with a hole in the middle.
The two things that do detect are both cheap, and neither is a sample.
The cap, which cost nothing and once saved twelve thousand
On 3 November, Barsana Mills changed the layout of its packing list. The agent began reading a column that was no longer the one it needed, and every allocation built from that supplier's documents was wrong from the first one.
Serdana's caps are unglamorous: no more than 12 actions an hour across the whole desk, no more than 2 on any one order in a day, no single trim order over USD 900.00, and no more than USD 3,600.00 of them in a week. The hourly cap fired 41 minutes in, and the alert is what told anybody that something had changed.
uncapped, replayed on the same inputs 61 wrong allocations USD 15,860.00
capped 12 wrong allocations USD 3,120.00
prevented USD 12,740.00What did the cap cost? Over nine months it delayed 34 correct actions by an average of 27 minutes, which cost 0.00 days of order clock. That is not luck. Correct work is not bursty and failures are. A cap set from your own normal hourly volume almost never binds on real work, and it binds immediately on a new failure mode.
The second detector is targeted rather than random. Re-check every action whose input document changed shape, and every action on an order that later slipped. It is the sample aimed at where the drift actually enters.
When the line has to be redrawn
The boundary is a function of terms that all move. Serdana's four triggers:
- The audited error rate on any row moves by more than half its recorded value.
- Anything new subscribes to a field the agent may write.
- An action type's monthly count doubles.
- Six months pass without any of the above.
The first redraw came from none of them. In September the fabric-variance work described in lesson 2 finished, and the honest cost of one wrong allocation went from USD 260.00 to USD 456.20, because five wrong allocations that nobody had ever complained about had been paid for as cutting waste.
value of removing the wait 268 x 0.04 x 214.00 USD 2,294.08
expected loss at 456.20 268 x 0.019 x 456.20 USD 2,322.97
net USD -28.89The fabric-allocation row came off. It was the highest-volume row on the desk, the one with the lowest cost per error, the one everybody had been most comfortable with. And it was removed by a stock-variance investigation that nobody had connected to the agent at all. Nothing about the system had changed, nothing about its accuracy had changed, and the answer was different.
Where nine actions ended up
autonomous 3 rows net after ownership USD 3,827.42
split, preparation only 1 row the milestone tick USD 1,787.33
daily digest of what was written USD -48.30
with a person, permanently 5 rows
total USD 5,566.45Two things about that scoreboard are worth more than the total.
The first is that 32.1% of the money comes from a row that was never automated: the milestone tick, where a person still makes every decision and a machine does the fetching. The second is that of the nine actions the desk originally handed over, six ended up with a human. And the desk is better off than it was when all nine were approved, and better off than it would have been with all nine autonomous.
Check yourselfYour approval queue has a 2% refusal rate and everybody says the system is working well. What are the two possible readings, and how do you tell them apart?Show the answer
Either the proposals really are almost always right, or the reviewers have stopped reading them. The refusal rate alone cannot distinguish those, and they call for opposite explanations of the same number. Measure the catch rate instead. Audit a closed period, count the wrong actions, and check how many of them the reviewers had refused at the time. If the catch rate is high, the approval is doing real work and the low refusal rate means the system is good, so move the row and keep sampling. If the catch rate is low, the approval is a form, and it is costing you clock to remove nothing.
Check yourselfYou are asked to write the one-page rule that governs your agent. What goes on it?Show the answer
The action list, with a side for each row and the four numbers behind it. The three classes that stay human whatever the numbers say. The caps, in your own units, taken from your normal hourly volume. What the stop does to the queue when it is released. Who owns each field the agent may write, and who gets the daily digest of what it wrote. And the four triggers that force a redraw, with a date next to the last one. It fits on a page because everything on it is a number or a name. If your version needs a paragraph explaining why the system is trustworthy, that paragraph is doing the work a measurement should be doing.
What you own at the end of this course
An action inventory rather than a feature list. A cost for one wrong action that reaches past the undo button. Two clocks that decide reversibility, and a census of who reads what you write. A four-term inequality with an ownership cost to clear. Three classes that no arithmetic moves, and the habit of automating the preparation instead of the decision. Caps taken from your own volume, a stop whose behaviour is decided in advance, an audit that admits what it cannot detect, and four triggers that redraw the line.
And one sentence to take into the meeting where somebody asks whether the system is accurate enough: it is the wrong first question, and the right one is which actions.