Lessons · Lesson 4 of 6
The inequality, and the wait that bought nothing
Draw the boundary as a calculation on four measured terms, and find the row that passes on its error rate and fails on the honest version of it.
Lesson 4 of 6 · 21 min
The four terms, and where each one comes from
value of removing the approval = n x D x V
expected cost of removing it = n x p x Cn comes out of the system's own log, and it is the only term nobody argues about. V is the value of a day of order clock, and Serdana uses USD 214.00 from its own late-delivery history, as lesson 1 said. That leaves the three terms that carry the argument:
- `D` is the days of order clock the approval wait actually costs. Not the wait. The part of the wait that reaches the ship date.
- `p` is how often the action is wrong, measured by replaying a closed period.
- `C` is what one wrong action costs all the way to the end, which lesson 2 spent its whole length on.
Three of the nine rows have a D of nearly nothing. Understanding why is the difference between a boundary and a guess.
The wait that bought nothing
Rescheduling tomorrow's loading sequence has a mean approval wait of 6.9 hours. Removing it recovers 0.03 days of order clock. That is not a small fraction of the wait. It is almost none of it.
The reason is not in this course. Serdana's loading sequence is fixed at the 07:15 meeting and is not touched afterwards. So an action taken at 22:00 and an action taken at 04:40 are consumed at exactly the same moment. Course 24.1 derives this properly: when the interval between the moments a decision can actually be taken is longer than the latency you removed, the latency does not appear in the answer at all. Removing an approval from an action whose consumer meets once a day buys you the hours between the two meetings, which is to say nothing.
`D` is not the wait you removed. It is the part of the wait that ends before the next thing that could have used it.
Two more rows are in the same position, for different reasons. Answering a buyer's status query has a D of 0.00, because the benefit goes to the buyer's clock and not to Serdana's. That is a decent reason to answer quickly. It is not a reason to remove the approval. Allocating a fabric receipt line has a D of 0.04, because the cutting room draws in the morning.
The row with no errors at all
Releasing an order from credit hold came back from the replay with a perfect record: 43 cases, 0 wrong. On the face of it that is the safest row on the desk, and its D of 1.40 days is the largest.
Zero is where the arithmetic goes wrong, and it goes wrong in the reassuring direction.
An observed rate of zero is not an error rate of zero. It is an error rate you have not yet measured. The standard treatment of a zero numerator gives an upper bound of roughly three divided by the number of cases:
cases with no failure 43
upper bound on the rate 3 / 43 6.98%
expected loss a month 7 x 0.0698 x 6,300 USD 3,076.74
value of removing the wait 7 x 1.40 x 214 USD 2,097.20
net at the point estimate USD 2,097.20
net at the bound USD -979.54The same row is worth two thousand dollars a month and minus a thousand dollars a month. Which one you believe is a decision about statistics, not about garments. Serdana draws its boundary against the bound, on the reasoning that a term you have failed to measure should not be allowed to buy you anything.
That gives a genuinely useful answer instead of an argument. Solve for the number of consecutive clean cases at which the bound falls far enough to clear:
need 7 x (3 / N) x 6,300 < 2,097.20 so N > 63.1Sixty-four clean cases. Serdana has forty-three and books seven a month, so three more months of clean history settles it. The row goes on the human side until then, with a date attached rather than a shrug. This is the honest form of "not yet": a number, a rate and a date.
The nine rows
| Action | Wrong | Cost of one wrong | Value of removing the wait | Expected loss | Net |
|---|---|---|---|---|---|
| Chase an unanswered approval by mail | 3.1% | 40.00 | 273.92 | 79.36 | 194.56 |
| Move a plan date on the critical path | 7.4% | 1,880.00 | 2,719.94 | 5,703.92 | -2,983.98 |
| Allocate a fabric receipt line | 1.9% | 260.00 | 2,294.08 | 1,323.92 | 970.16 |
| Book container space | 2.8% | 310.00 | 3,466.80 | 156.24 | 3,310.56 |
| Answer a buyer status query | 1.9% | 2,400.00 | 0.00 | 2,416.80 | -2,416.80 |
| Mark a milestone complete | 4.3% | 1,150.00 | 2,465.28 | 4,747.20 | -2,281.92 |
| Reschedule tomorrow's loading sequence | 1.2% | 540.00 | 141.24 | 142.56 | -1.32 |
| Issue a top-up trim order | 3.6% | 1,420.00 | 1,412.40 | 562.32 | 850.08 |
| Release an order from credit hold | 6.98% | 6,300.00 | 2,097.20 | 3,076.74 | -979.54 |
Read the two rows that share an error rate, which lesson 1 promised. Allocating a receipt line and answering a buyer query are both wrong 1.9% of the time. One is worth USD 970.16 a month and the other minus USD 2,416.80. Nothing about the technology distinguishes them. C differs by a factor of 9.23, and D differs by everything.
The third test: what a row costs to own
A positive net is not enough, because an autonomous action is not free to run. It has to be watched, sampled, reviewed and occasionally cleaned up after. Serdana measured that as a fixed cost per action type plus a variable one:
share of the monthly boundary review 1.5 h x 34.00 51.00
share of the incident allowance 24.5 h a year x 34.00 69.42
subscriber re-check when anything changes 0.8 h x 34.00 27.20
daily look at the action queue 5 min x 21 days x 11.50 20.13
fixed, per action type, per month 167.75
audit of one action in eight 11 min x 11.50 an hour 2.11 eachApplied to the rows with a positive net:
chase an unanswered approval net 194.56 to own 184.63 clears by 9.93
allocate a fabric receipt line net 970.16 to own 238.44 clears by 731.72
book container space net 3,310.56 to own 172.50 clears by 3,138.06
issue a top-up trim order net 850.08 to own 170.65 clears by 679.43Four of nine. The chasing mail is the action every desk automates first, and the one that feels least consequential. It clears by USD 9.93 a month, which is not a margin. It is a rounding error. Serdana kept it, on the grounds that it rides on work already being done for the other three rows, and wrote down that it would be the first row dropped if anything changed. That is a defensible answer, and it is not the same answer as "it is obviously fine".
The three settings, priced
approve everything no gain, no loss costs 14,870.86 of clock and 455.75 of hours
approve nothing gain 14,870.86 expected loss 18,209.06 net -3,338.20
the line, four of nine gain 5,325.36 cost of owning them 766.22 net 4,559.14Approving nothing is worse than approving everything. Both are worse than the line by thousands a month. And the argument most desks have, should we trust it or not, is an argument about which of the two losing settings to adopt.
The one-off work behind the line was a subscriber census of 14.5 hours and a replay harness of 31.0 hours, at the manager rate: USD 1,547.00. Against USD 4,559.14 a month, it pays back in 10.32 days.
Prompt · Price one approval
When a team is arguing about whether to trust a system, to replace the argument with four numbers.
Help me price the approval on ONE action, using my numbers only. Ask me for these five, one at a time, and do not go on until I have given each: how many times a month the action happens; the mean hours a proposal waits for approval; how many days of order clock that wait actually costs, which is not the same as the wait; how often the action would be wrong, measured by replaying a closed period rather than from complaints; and what one wrong action costs all the way to the end, including the work other people did on the wrong number before it was corrected. Then compute the value of removing the approval as count times clock-days times my value of a day, and the expected cost as count times error rate times cost of one wrong action, and show the net. Three rules. If I report zero errors, tell me the number of clean cases I would need before that row could be trusted, using three divided by the number of cases as the upper bound on the rate. If the clock-days figure is near zero, say that removing the wait buys nothing, and ask what downstream step consumes the result and how often it runs. And if my answer to the error rate came from an incident log rather than a replay, say that the figure is too low and explain why.
AI can make mistakes — check anything you act on.
Check yourselfAn action comes back from the replay with a positive net of USD 60 a month. What do you do with it?Show the answer
Nothing yet. Sixty dollars will not cover the cost of owning it, and an autonomous row that nobody can afford to watch is worse than an approved one. Either leave it approved, or fold it into a row already being watched, so it carries only the marginal audit cost rather than a share of the fixed one. Then look at which term is small. If it is D, ask whether anything downstream could consume the result sooner, because that is usually free to change. If it is n, leave it alone permanently, since a rare action is where a human's judgement is cheapest.
What you own at the end of this lesson
A boundary you can compute rather than defend: four measured terms, an ownership cost that a positive net has to clear, an honest treatment of a row with no observed errors, and a number of clean cases at which that row can be looked at again.
Next: the three classes of action that no arithmetic moves, and why one of them looks completely harmless.