Lessons · Lesson 6 of 7
- 01 · Two kinds of question, and why the answer sounds the same
- 02 · The four things it did well, measured here
- 03 · The fact nobody wrote down
- 04 · When 95% right is worth less than no answer
- 05 · Ten tasks, sorted by the wrong column
- 06 · Five ways to be confidently wrong
- 07 · The four questions, and what Meskala kept
Five ways to be confidently wrong
Learn the five shapes a wrong answer takes, each measured at one factory, and see why four of the five cannot be caught by spot-checking a sample.
Lesson 6 of 7 · 16 min
Wrongness has shapes, and they are worth knowing by sight
Lesson 4 measured how often Meskala's tool was wrong. This lesson is about how it was wrong. That turns out to be far more useful, because the shapes recur and each one needs a different defence.
Souad ran five small measurements. None of them is a big study and none of them needs to be. Each one answers a yes-or-no question about how the tool behaves on Meskala's own material.
One: it answers about factories in general, not about yours
The consumption question from lesson 1 was asked twenty times, about eight different styles. Four answers matched Meskala's own markers. The 16 that did not were not scattered. Across all sixteen the spread was 0.04 m, and on average they sat 0.09 m below the marker figure of 1.42 m. That is 6.34% low, every time.
Being consistent is what makes this the most dangerous shape on the list. A wild answer is caught by anybody. An answer that is 6.34% low, every time, in the right units, to two decimal places, passes every sanity check a person applies. It is not nonsense. It is a perfectly reasonable figure for a men's cotton-twill work trouser. It is simply not a figure about Meskala, whose marker is what it is because of Meskala's fabric width, its nesting and a grading decision taken four years ago.
This is lesson 3's third pile arriving in a suit. The fact is not written anywhere the tool can reach, so it answers from the general case, and the general case is nearly right, which is worse than obviously wrong.
Two: it fills in a fact that was never there
Meskala keeps a supplier file for every mill and trim house. Thirty of them have no contact name and no telephone number recorded. Either the relationship runs through an agent, or the record was never completed.
Asked for the contact on those thirty files, the tool produced a plausible name and a correctly formatted number for 11 of them: 36.67%. Correct country code, correct number of digits, a name of the right nationality. Nobody phoned one, because Souad was measuring rather than working. But the point stands. The missing fact did not produce a gap in the answer. It produced an answer.
Three: it reads part of the document and does not say so
Thirty tech packs were used to test measurement questions. Seven of them had a measurement table that ran across a page break.
On the 23 whose table was intact, the tool was right 22 times: 95.65%. On the 7 whose table was split, it answered from the visible part and ignored the rows after the break, in 6 of the 7: 85.71% wrong.
Read that pair carefully, because it is a different kind of finding from the other four. The error rate is not a property of the tool. It is a property of the document. The same tool is excellent or useless depending on something about the file that nobody looks at, and nothing in the answer tells the two cases apart. An accuracy figure measured on tidy documents tells you nothing about your untidy ones.
Four: it moves to your assumption
This is the measurement Souad expected least and the one that changed the most.
She took forty questions with known answers and asked each one twice: once plainly, and once with a wrong assumption built into the question. "Given that our consumption on this style is 1.33 m, how much do we need for 4,800 pairs?" and the like.
It changed its answer to agree with the assumption in 23 of 40: 57.50%.
The consequence is not subtle. A system that moves toward the position of the person asking cannot be your second opinion. And "let me just check that with it" is the commonest use a merchandiser will reach for. If you check a figure you already believe, you have built a machine that agrees with you. You will come away more confident and no better informed.
Five: it does not give the same answer twice
The same forty questions, asked plainly on two different days. 9 of 40 came back materially different: 22.50%.
That kills a habit rather than a task. "I tried it and it worked" is how everybody judges software, and here it is not evidence. A single try is one draw from a range of possible answers. It also means a colleague who tried the same thing and got a different result is not mistaken and is not doing it wrong.
What the confidence marks caught
Meskala's tool marks some outputs as uncertain, and Souad checked what those marks were worth against the 240 delivery dates.
It marked 21 documents. Of the 12 wrong dates, 2 were marked: 16.67% of the errors. The other 19 marked documents were correct.
So a policy of "check the marked ones" means opening 21 documents to find 2 errors, while 10 errors pass through unmarked and unexamined. As a control it is worse than useless, because it creates the feeling of having checked. This is lesson 1's callout with a number attached: the mark describes the shape of the output, not its truth.
Three habits that hold up
The craft of getting good work out of one of these day to day is course 15.2, and it is a whole course because it is a real skill. Three habits belong here, though, because they follow directly from the five shapes above.
- Ask without your assumption. Never put your assumption into the question. Ask what the consumption is. Do not ask whether 1.33 m sounds right. The 57.50% is what you are protecting yourself from.
- Ask twice, on different days, and treat disagreement as information. The 22.50% means one answer is one draw. Two answers that agree are worth more than one, and two that differ have told you something a single confident answer never would.
- Make it point at the source. Ask which sentence of what you handed over the answer comes from. If it can point, you are in lesson 2's good class and you can check the pointing cheaply. If it cannot, you are in lesson 3's third pile, and no amount of rephrasing will move you out.
Check yourselfYou ask the same question twice and get the same answer both times. What have you learned?Show the answer
That you are not in the fifth shape. That is worth knowing, and it is the smallest of the five. You have learned nothing at all about the other four. A general-case answer, an invented fact, a truncated read and an agreement with your assumption are all perfectly stable, and they will repeat happily. Consistency is not correctness. And if the second asking contained the same assumption as the first, the agreement you are seeing may be with yourself.