Lessons · Lesson 6 of 6
Did anything actually change?
Tell a real improvement from an ordinary good week, using the process's own history, and see what reacting to a single point costs beyond the reaction.
Lesson 6 of 6 · 18 min
The good week
On 12 May, Iresha finished rebalancing Line 3: three operations put in a different order, one folder changed, a helper moved. A folder is the metal guide on a machine that turns a raw edge before the needle reaches it. The following week the line reported 83.1% against 79.8% the week before.
Chandana asked her to do the same on Line 5.
The problem with that entirely reasonable request is that a single week is not evidence, and the reason is in the fourteen weeks that came before the change.
| Week | Week | Week | |||
|---|---|---|---|---|---|
| 1 | 78.2% | 6 | 80.1% | 11 | 79.3% |
| 2 | 81.4% | 7 | 78.8% | 12 | 83.4% |
| 3 | 79.0% | 8 | 82.6% | 13 | 77.8% |
| 4 | 83.2% | 9 | 76.9% | 14 | 80.2% |
| 5 | 77.5% | 10 | 80.5% |
The average of those fourteen weeks is 79.9% and their standard deviation is 2.1 points. Nothing was done to Line 3 in any of them: the same style, the same operators, the same shifts.
Three of those fourteen weeks came in at or above 82.6%, and one of them at 83.4% — higher than the week that followed the rebalance. The good week after the change is 1.5 standard deviations above the average, and a line that varies this much naturally produces a week like that about once a quarter for no reason at all.
83.1% is not evidence that anything changed. It is evidence that Line 3 is a line that does 83.1% sometimes.
What would count as evidence
This is not a counsel of despair, and the answer is not "wait a year". There are two things a factory can read off a weekly series, and both are available with a spreadsheet.
The first is the limits. Take the average and the standard deviation of a settled period, and draw the average plus and minus three standard deviations. For Line 3 that is 73.6% to 86.2%. Any single week inside that range is the line being itself. A single week outside it is a real event and worth a question on its own.
The second, and much the more useful in a garment plant, is the run. A single point inside the limits says nothing. A sequence of points all on the same side of the centre line says a great deal, because the chance of a long one-sided run under ordinary variation gets small quickly. This is the older discipline of reading a control chart, and its rules are written down: one point beyond three standard deviations; two points out of three beyond two standard deviations; and a long run of points in a row on one side of the centre. Each of those is a signal that the process has moved rather than merely wobbled.
| Week after | Efficiency | Above the old centre |
|---|---|---|
| 1 | 83.1% | yes |
| 2 | 81.0% | yes |
| 3 | 82.4% | yes |
| 4 | 84.0% | yes |
| 5 | 82.2% | yes |
| 6 | 83.5% | yes |
| 7 | 81.9% | yes |
| 8 | 82.8% | yes |
| 9 | 83.3% | yes |
Not one of those nine weeks is outside the limits. Week four, the best of them, is 1.9 standard deviations above the average, which on its own proves nothing.
But nine weeks in a row above the centre line is a signal, and it fired in the week of 14 July. The improvement it confirms is the average of those nine against the old average: 2.8 points. That is real, and it is a good deal smaller than the 3.2 points the first week appeared to promise.
The cost of the reaction, which is not the cost of the mistake
Halewatte had already learned this once and had not noticed it had.
In January, Line 7 reported 83.6% in one week against an average of 79.4%. Chandana concluded the line had spare capacity and moved two operators to Line 9, which was behind on a delivery. It was a defensible decision made on the best number available. Line 7's own standard deviation over the preceding quarter was 2.4 points, so 83.6% sits 1.8 standard deviations above its average — the same size of event as Line 3's good week in May, and just as meaningless on its own.
| Week | Efficiency | Against the average |
|---|---|---|
| 1 | 75.8% | 3.6 below |
| 2 | 77.1% | 2.3 below |
| 3 | 78.9% | 0.5 below |
| 4 | 79.6% | 0.2 above |
Net over the four weeks, the shortfall is 6.2 points-weeks. Line 7 attends 100,800 operator-minutes a week, so that is 6,249.6 earned minutes, 284 shirts, USD 610.76 at Halewatte's cut-and-make rate of USD 2.15 a shirt.
USD 610.76 is not a large number and it is not the finding.
The finding is that Line 7's January good week was ordinary variation, and moving two operators on the strength of it did something worse than cost USD 610.76: it destroyed the experiment. After week one, the line no longer has the staffing it had, so the four weeks that follow cannot tell anybody whether the original good week meant anything. The factory spent four weeks and USD 610.76 buying the permanent inability to answer its own question. That cost happens every time, it is invisible, and it appears on no report.
The failure that runs the other way
The same rule catches the thing nobody looks for.
Line 8 spent seven weeks in a row below its centre line of 80.6%, averaging 77.8%. Every single week was inside the limits, and every single week was explained away on its own — a slow week, an absence, a new operator, a Monday. Seven separate reasonable explanations, each of which was probably even true.
Seven weeks in a row on one side is a signal. It fired, and the cause was found in an afternoon: a fabric lot with a heavier finish that added roughly nine tenths of a minute to the collar-run operation, arriving three days before the first of the seven weeks. Nobody had connected the two because nobody was reading the sequence.
That decline cost 19,756.8 earned minutes, 898 shirts, USD 1,930.78 — three times what the January over-reaction cost, and it was invisible for exactly as long as each week was read on its own.
And where the chart belongs
The control chart is a descriptive instrument. It compares a period against a history. It cannot say what to do in the next hour, and the earliest it can speak is at the end of a week. By the test in lesson 1 it fails all five blanks for a supervisor and passes all five for a manager: when a line runs nine weeks on one side of its centre, Iresha looks into it, within a week.
So it does not go on the shop-floor screen. It goes on the Monday sheet, one small chart per line, thirteen weeks wide.
That is the course closing on itself. A number on a wall is a request for a decision. The chart is a request for a decision too — just not a decision anybody standing on Line 3 at half past ten is in a position to make.
Check yourselfYour best line just had its best month. What do you do?Show the answer
Write down the average and standard deviation of the twelve months before it, and see how many of those twelve were within a point of this one. If two or three were, nothing has happened yet and you should change nothing, because changing something now removes your ability to find out. If none were close, you have a real event, and the question worth asking is what was different, before it stops.
Prompt · Did anything actually change?
The morning after a good week, and before anybody moves an operator on the strength of it.
Act as a quality engineer who reads control charts for a living and is comfortable telling somebody their improvement is not yet visible. I need to know whether a change I made has actually done anything. The change: [WHAT WAS DONE, WHICH LINE, WHICH DATE]. The measure: [WHAT IT IS AND HOW IT IS CALCULATED, INCLUDING WHAT IS IN THE BOTTOM OF IT]. History BEFORE the change, from a period when nothing was deliberately altered: [PASTE AT LEAST THIRTEEN WEEKLY OR DAILY FIGURES]. Everything since the change: [PASTE IT]. Do the following. First, work out the average and standard deviation of the BEFORE period only, and state the limits at three standard deviations either side. Do not work them out again using anything after the change, and say why not. Second, tell me how many of the before-period figures were as good as or better than my best figure since the change. Third, apply the published tests for a special cause and say which, if any, have fired: a point outside the limits, a run of consecutive points on one side of the centre line, a sustained trend. Name the test and the date it fired. Fourth, if nothing has fired, tell me how many more periods at the current level would be needed before one would, and what I should NOT do in the meantime. Fifth, if something has fired, give me the size of the improvement as the mean since the change against the old mean, and compare that with what the first good period appeared to promise. Sixth, check the OTHER direction: is there any run of consecutive periods below the centre line anywhere in my data, on any line, that was dismissed one period at a time? Seventh, warn me about anything in the measure that could have moved for recording reasons rather than real ones, and name the independent count that would settle it. Do not congratulate me on a single good period.
AI can make mistakes — check anything you act on.
What to do on Monday
- Take thirteen weeks of a weekly figure for one line, from a period when nothing was deliberately changed. Work out the average and the standard deviation.
- Draw the average, and the average plus and minus three standard deviations. Fix them; do not work them out again.
- Plot every week from now on against those fixed lines, and mark a signal on a point outside the limits, or on a long run of points on one side.
- Write the rule down before the next good week arrives, because the moment to decide what counts as proof is while nobody has an answer they are hoping for.