Lessons · Lesson 3 of 6
The threshold nobody derived
Set an alert threshold from the process's own spread instead of from the target, and work out the message rate before the rule goes live rather than after.
Lesson 3 of 6 · 18 min
The rule that seemed obvious
When the feed went live, Chandana asked for an alert. The specification was one line, and every factory writes this line: send a message when a line's hourly efficiency falls below the target.
The target is 82.0%. The rule went into a group that all nine supervisors, Chandana and Iresha belong to, because that is how factories do it: one group, everybody sees everything, nothing gets missed.
By the third week the group was muted on every phone in it.
The arithmetic nobody did first
Iresha did the calculation afterwards, and it takes ten minutes.
Halewatte runs nine lines, eight hours, twenty-six days: 1,872 line-hours a month. Over March, the hourly efficiency of a Halewatte line had an average of 79.6% and a standard deviation of 6.8 points. The standard deviation is a measure of spread: it says how far a typical hour sits from the average. Those two numbers are the whole answer, and both were available before the rule was written.
A threshold at 82.0% sits 0.35 standard deviations above the average. So the rule does not fire on exceptions. It fires on 63.8% of all line-hours — every hour that is anything less than clearly good.
| Threshold | Alerts a month | Alerts a day | One every | |
|---|---|---|---|---|
| As specified | 82.0% | 1,194 | 45.9 | 10.5 minutes |
| Derived from the process | 66.0% | 43 | 1.6 | roughly a shift and a half |
The derived threshold is the average less two standard deviations. It is not a softer target and it concedes nothing: the target is still 82.0%, the line is still short of it, and the Monday sheet still says so. The threshold on an alert answers a different question from the target. Not "is this good enough" but "is this unusual enough to interrupt somebody".
What the tight rule bought and what it cost
It is not enough to say there were too many messages. The tight rule genuinely catches more real events, and an honest comparison has to say how many.
Iresha took the maintenance log — an independent record, written by mechanics who never see the efficiency feed — and counted every line-hour in March that contained a stop of fifteen minutes or more. There were 44.
| Real events caught | Messages sent | Share of messages that were a real event | |
|---|---|---|---|
| Below 82.0% | 43 | 1,194 | 3.6% |
| Below 66.0% | 38 | 43 | 88.4% |
So the tight rule caught five more real events and cost 1,151 more messages to do it. Those five stops averaged twenty-four minutes. At Line 6's USD 1.3567 an attended minute they are worth USD 195.36 of standing time between them.
That is the entire case for the tight rule, and on paper it wins: USD 195.36 is worth having.
Except that the tight rule did not catch forty-three events. It sent 1,194 messages into a muted group and caught none of them, because a rule that is not read is a rule that does not exist. On 14 April the bartack on Line 6 seized at 10:40. The alert went out at 10:47, into the same stream as the other forty-four that day. Kumudu saw it when she opened the group at 13:37 to send a photograph of a fabric fault.
The budget, in minutes
There is a way to know in advance whether a rule is affordable, and it is arithmetic rather than judgement.
Iresha timed what an alert actually costs Kumudu when she reads one and does something about it: 6.4 minutes, from reading it to being back at her station. She also measured how much of Kumudu's shift is already committed to things that cannot be interrupted — the hourly count, the changeover, the two quality walks, chasing bundles of cut work down the line: 38% of a 480-minute shift, leaving 297.6 free minutes.
Divide. At 6.4 minutes an alert, the absolute ceiling is 46.5 alerts a shift, and that is a shift in which she does nothing else whatsoever.
The rule as configured delivered 45.9 a day. Reading and acting on all of them would have taken 294.0 minutes — 98.8% of every free minute she had.
Nobody decided that. It was the arithmetic consequence of a one-line specification that felt cautious when it was written.
A working budget is a tenth of her free attention, which is 4.7 alerts a shift. The tight rule is 9.9 times over it. The derived rule uses 35% of it, and leaves room for the alerts that are not about efficiency at all.
Cutting the alert list down
The discipline for doing this properly is not new, and it does not come from manufacturing analytics. Oil refineries and chemical plants have had it for decades, and its central rule is the one that matters here: an alert exists to prompt a defined response, from a defined person, within a defined time. An alert with no response is not an alert. It is a notification, and it belongs somewhere quieter.
That is lesson 1's sentence again, pointed at the alarm list instead of the screen. Applied at Halewatte it produced four rules, not one.
| Alert | Threshold | Who | Does | Within |
|---|---|---|---|---|
| Line stopped | no count for 12 minutes | Kumudu | walk to the line, find the cause, call the mechanic | 5 minutes |
| Hour badly short | hourly efficiency below 66.0% | Kumudu | name the cause on the shift log | end of the next hour |
| Line down over an hour | any single stop past 60 minutes | Chandana | decide whether to move the work | 15 minutes |
| Absence uncovered | roster short at 08:00 | Kumudu | pull from the spare pool | 30 minutes |
Two of those four are not about efficiency at all. The stopped-line rule fires on the absence of a count, which is faster and leaves no room for argument. A line that has produced nothing for twelve minutes is stopped, and no ratio has to be worked out or trusted to know it. It reaches Kumudu about thirty minutes before an hourly efficiency figure could.
Check yourselfYour alert threshold is the target and you are told loosening it is 'accepting poor performance'. What is the reply?Show the answer
That the target has not moved and the report still measures against it. An alert threshold is a decision about when to interrupt a person. Its correct value comes from the spread of the process and the capacity of that person, not from the commercial target. Setting the two equal guarantees the alert fires on ordinary variation, which is how a channel gets muted, which is how the genuinely bad hour arrives unread.
Prompt · Set an alert threshold from my own data
Before you switch on any alert, and the week after somebody mutes the group it sends to.
Act as a control engineer who has rationalised alarm lists and has no interest in making my system look busy. I want an alert threshold derived from my own process rather than from my target. Data: [PASTE AT LEAST FOUR WEEKS OF THE MEASURE, ONE OBSERVATION PER LINE, WITH ITS TIMESTAMP AND WHICH LINE OR MACHINE IT CAME FROM]. My target for this measure is [VALUE] and my current alert threshold is [VALUE]. The alerts go to [WHO], and they all go to [ONE GROUP OR SEPARATE PEOPLE]. That person's shift is [MINUTES] long and roughly [PERCENT] of it is already committed to things that cannot be interrupted. When they act on an alert it takes them about [MINUTES] end to end. I also have an independent log of real events - [MAINTENANCE, QUALITY, SOMETHING NOT DERIVED FROM THE SAME FEED] - for the same period: [PASTE IT]. Do the following. First, give me the average and standard deviation of the measure, and say how many readings you had. Second, calculate the message rate my CURRENT threshold produces: a month, a day, and one every how many minutes of a shift. Third, do the same for thresholds at one, two and three standard deviations below the mean, in a table. Fourth, score each candidate threshold against my independent log: how many real events it catches, how many messages it sends, and what share of its messages were a real event. Fifth, multiply the message rate by the minutes each alert costs and compare that against the receiver's uncommitted minutes, and say plainly whether the rule is affordable. Sixth, recommend one threshold, and write the alert as a row with four columns: threshold, who, does what, within how long. Seventh, tell me which of my alerts should not be efficiency-based at all - in particular whether an absence-of-count rule would reach the person sooner. Do not recommend a threshold equal to my target, and do not tell me to send alerts to more people.
AI can make mistakes — check anything you act on.
What to do on Monday
- Take four weeks of the measure your alert watches. Work out its average and its standard deviation. Two cells in a spreadsheet.
- Count how many readings fall below your current threshold, and divide by the number of days. That is your message rate, and you now know it before anybody is annoyed by it.
- Time one alert from start to finish for the person who receives them. Multiply by the rate. Compare against their free minutes.
- Set the threshold from the spread, then check it still catches the events you care about by testing it backwards against an independent log — maintenance, quality, anything not derived from the same feed.