Lessons · Lesson 7 of 7
- 01 · Two kinds of question, and why the answer sounds the same
- 02 · The four things it did well, measured here
- 03 · The fact nobody wrote down
- 04 · When 95% right is worth less than no answer
- 05 · Ten tasks, sorted by the wrong column
- 06 · Five ways to be confidently wrong
- 07 · The four questions, and what Meskala kept
The four questions, and what Meskala kept
Turn the whole course into a test you can run in a meeting, then read one factory's honest verdict — including what its own measuring cost and what it returned.
Lesson 7 of 7 · 12 min
The test
Everything in this course comes down to four questions. Ask them in this order, about any proposed use, before anybody demonstrates anything.
1. Where does the answer live? Inside what you would hand over; or in a record you own but have not connected; or written down nowhere. The first is free today. The second is a connection project with a price. The third is not a purchase at all. It is a decision to record something, and lesson 3 showed that at Meskala it was also where 11 of the 14 decision-changing answers were.
2. What detects a wrong answer, how long does it take, and what has already happened by then? Answer this in one sentence, out loud. If the sentence is "nothing", stop here. You have no way of knowing whether it is working, and no amount of saving is worth an error you cannot see.
3. Is checking much cheaper than doing? Not slightly. If checking means re-doing the work, the time saving is small no matter what the list says, and the arithmetic will not clear the cost of the errors that get through.
4. What share of its errors would your check actually catch — measured, on your documents, on enough of them? This is the quantity nobody quotes and the one that moved Meskala's decision. Lesson 4 showed the break-even sitting at 74.89% against a measured 58.33%, and the whole difference between a good tool and an expensive one hiding inside 1.99 percentage points that no demonstration can resolve.
Prompt · Run the four questions on a proposal
Before a demonstration, on any written proposal, pitch or internal business case. It reads the document and reports which of the four questions the document answers and which it says nothing about. That is a restating task: every fact it needs is inside what you hand over. The silences are the finding, not its opinion of the tool.
Read the proposal below. It suggests using an AI system for a task at a garment factory. Do not tell me whether the tool is good, and do not judge the technology. You cannot, and neither can the document. Report only what the document does and does not say, against these four questions: 1. WHERE THE ANSWER LIVES. Does the document say which files, records or documents the system reads to produce its answer? Quote the sentence, or write NOT STATED. 2. THE DETECTOR. Does it say what would detect a wrong output, how long that takes, and what would have happened by then? Quote, or NOT STATED. 3. THE COST OF CHECKING. Does it say how long checking one output takes, compared with how long producing it takes today? Quote, or NOT STATED. 4. THE CATCH RATE. Does it say what share of the system's errors a human reviewer actually caught in a measured trial, and on how many items, on whose documents? Quote, or NOT STATED. Then list every accuracy or performance figure in the document, and for each one say whether the document states the documents it was measured on and how many. End with the questions I should ask, phrased so they cannot be answered with a demonstration. Here is the proposal:
AI can make mistakes — check anything you act on.
What Meskala kept
Of the ten tasks on the list:
- Four are running. Summarising a buyer thread, drafting a supplier chase in Arabic, translating a tech-pack comment for the floor, routing incoming mail. Together they are worth 60.00 a month. All four are in the class where the answer is inside the input, and something cheap and fast detects a wrong one.
- Five were dropped, not because anyone disliked them but because the arithmetic refused them. Two of the five are worth revisiting if the underlying data changes. The supplier-history question becomes a completely different proposition once the goods-receipt records are connected, because the fact moves out of lesson 3's third pile and into the first.
- One was declined as unmeasurable. Choosing which exception to work first has no detector, so it has no error rate, so there is no arithmetic to do. Meskala did not run it, and did not pretend the empty column was a zero.
60.00 a month is not a stirring number, and it is the honest one. It is what these systems were worth to one factory's merchandising desk in the month it was measured, on the tasks where they belonged.
What the measuring cost, and what it returned
Souad's measurement was not free. Across the month she spent 41 hours on it — designing the comparisons, running them, settling disagreements — at a loaded 14.20 an hour, which is 582.20. Imane spent 18 hours doing work twice so the comparisons had something to compare against, at 9.40, which is 169.20. Total: 751.40.
Set that against the 720.00 a year the four surviving tasks are worth. The measurement cost 104.36% of the first year's saving. On the face of it, it did not pay for itself.
Now set it against the choice that was actually on the table, which was adopting the list as written. Running the nine costable tasks loses 23,798.11 a month, or 285,577.32 a year. Running the four earns 720.00. The gap between the two courses of action is 286,297.32 a year, and the thing that stood between them was 751.40 of measuring.
The measurement returned 381.02 times its cost, and not one unit of that return is a saving. All of it is a loss that did not happen. That is why it will never appear in any account, no report will show it, and nobody will thank Souad for it. That is the shape of most good work in risk, and it is worth saying plainly so you recognise it when you are doing it.
What this course deliberately did not do
Four things follow from here, and each is a course rather than a paragraph.
- Getting good work out of one of these day to day — what to hand over, how to frame an ask, how to check what comes back without leading yourself — is course 15.2.
- Applying it to your own order data at volume — extraction, matching, and what to do when an extracted number and a counted number disagree — is course 15.3.
- Systems that take actions rather than answer questions — where the approval boundary sits, and what a wrong automatic action costs against the wait it removes — is course 15.4.
- What you owe your buyers and your workers, and how to judge a purchase without a benchmark — is course 15.5.
None of them changes the four questions above. They are what you do after the four questions have said yes.
Check yourselfSomeone shows you a demonstration on twelve of your own documents and it gets all twelve right. What have you learned, and what is your next sentence?Show the answer
You have learned that it is not catastrophically bad. That is worth something, and it is not much: at any rate above roughly 95.00% a twelve-document run is likely to come back clean, so the demonstration cannot separate a tool that pays from one that costs 1,735.38 a month. The next sentence is a request, not an objection. Run it on a batch big enough to see the difference, on our messiest documents rather than our tidiest, and let our own people key the same batch on their own, so we can measure what our check would catch.