Lessons · Lesson 5 of 7
- 01 · Two kinds of question, and why the answer sounds the same
- 02 · The four things it did well, measured here
- 03 · The fact nobody wrote down
- 04 · When 95% right is worth less than no answer
- 05 · Ten tasks, sorted by the wrong column
- 06 · Five ways to be confidently wrong
- 07 · The four questions, and what Meskala kept
Ten tasks, sorted by the wrong column
Apply lesson 4's rule to a whole week's work at once, and watch the ranking by time saved come apart from the ranking by what the work is actually worth.
Lesson 5 of 7 · 16 min
The list everybody makes
Every factory that looks at this makes the same list: the things a merchandiser does that an assistant could help with, ranked by how much time each would save. Meskala's had ten items on it, and it was a good list. Every item is a real task somebody really does.
Souad then did the thing nobody does. She added two columns: how many extra wrong answers would get through in a month, and what one of them costs.
Be precise about "extra". The comparison is not against perfection. The merchandiser makes mistakes too, and those mistakes are already in the cost of running the factory. So the number in the table is the errors the tool-plus-check leaves behind minus the errors the person makes doing the same work by hand. On the delivery-date task from lesson 4 that is 5 minus 3, which is 2.
| Task | A month | Minutes saved each | Value of the time | Extra wrong a month | Cost of one | Cost of the extras | Net a month |
|---|---|---|---|---|---|---|---|
| Summarise a buyer email thread | 60 | 4.0 | 37.60 | 2 | 6.00 | 12.00 | 25.60 |
| Draft a supplier chase in Arabic | 45 | 6.0 | 42.30 | 3 | 9.00 | 27.00 | 15.30 |
| Translate a tech-pack comment for the floor | 80 | 3.0 | 37.60 | 2 | 12.00 | 24.00 | 13.60 |
| Route incoming mail to an owner | 300 | 0.5 | 23.50 | 12 | 1.50 | 18.00 | 5.50 |
| Draft a quality-alert note to the buyer | 12 | 15.0 | 28.20 | 1 | 400.00 | 400.00 | −371.80 |
| Say whether a supplier has missed a date before | 25 | 10.0 | 39.17 | 4 | 260.00 | 1,040.00 | −1,000.83 |
| Build a size grid from a PO | 30 | 9.0 | 42.30 | 1 | 1,480.00 | 1,480.00 | −1,437.70 |
| Key a confirmed delivery date | 240 | 0.3 | 11.28 | 2 | 873.33 | 1,746.66 | −1,735.38 |
| Answer a consumption question | 20 | 12.0 | 37.60 | 9 | 2,150.00 | 19,350.00 | −19,312.40 |
| Choose which exception to work first | 20 | 8.0 | 25.07 | — | — | — | unknown |
Four of the ten pay. Five lose. One cannot be judged at all, and that is the most interesting row on the table. We come back to it.
Running the four that pay is worth 60.00 a month. Running all nine that can be costed loses 23,798.11 a month.
The reversal
Now sort the same ten tasks the way the original list was sorted, by minutes saved, and watch what happens.
The task at the top of the time ranking is the Arabic chase draft, at 270 minutes a month. It is genuinely one of the good ones: second best by net. So far so encouraging.
The second item down that list is building a size grid from a PO, also 270 minutes. It loses 1,437.70 a month. The third is the supplier-history question, which loses 1,000.83. The fifth is the consumption question, at 240 minutes a month of apparent saving, which loses 19,312.40 a month. That is more than three hundred times the entire programme's honest value.
And the task ranked ninth of ten by time saved — routing mail, at half a minute an item — is fourth by net, and one of the four Meskala kept.
The two orderings share almost nothing. Working down a list sorted by time saved, you would adopt one good task, then two losers, then another loser worth nearly twenty thousand a month, before you reached your second winner. The column the list was sorted on is not related to the answer.
There is a reason for that, and it is not coincidence. Minutes saved is roughly volume multiplied by how fiddly the task is. The cost of a wrong answer is set by what depends on the answer further downstream. Those two things have nothing to do with each other, so a ranking built on the first tells you nothing about the second. And the tasks that feel most worth automating are often the ones where a person is doing something slowly because it matters.
What the four that pay have in common
It is not that they are low-stakes. Look at the third row. A tech-pack comment mistranslated for the floor can put a wrong stitch into a whole bundle. That is not trivial. It pays anyway.
The four winners share two properties, and the second is the one missing from every list like this.
One: the check is genuinely much cheaper than the work. Not slightly. A lot. Reading a five-line summary against a thread you were going to skim anyway is nothing. Re-finding a delivery date in a PDF is nearly the whole job, which is why that row saves 18 seconds an item and could never pay.
Two: something else finds a wrong one quickly and cheaply. The supplier's reply finds a bad Arabic draft the same afternoon. The wrong owner bounces a misrouted email within the hour. The line supervisor queries a nonsense instruction at the first bundle. In each case a detector already exists, it costs nothing, and it is fast.
So add a column to your own list, before any of the others:
What detects a wrong answer, how long does it take, and what has already happened by then?
Answer that in a sentence for each task. Half your list will fall over as you write it.
The row that cannot be judged, and why it is the dangerous one
The last row is "choose which exception to work first". It saves 160 minutes a month of deciding, it is the kind of thing an assistant is very willing to do, and Souad could not put a number in any of the error columns.
Not because the task is safe. Because nothing detects a wrong answer. If the queue is ordered badly, no event fires, no supervisor queries it, no report turns red. The order Meskala worked is never compared against the order it should have worked, by anybody. There is nothing to compare against, so there is no measurement.
It is tempting to read an empty error column as a low one. It is the opposite. A task whose wrongness has no detector is a task you cannot manage, cannot measure and cannot improve. A systematic bias could run for a year in it without producing a single visible symptom. Meskala did not run it. Not because it decided the task was harmful, but because it could not tell, and "we cannot tell" is a legitimate answer that is not used enough.
Note also what this row is not. It is not an action. The assistant would be suggesting an order, and a person would still be working the queue. A system that took the action itself — reordering the queue, sending the chase, releasing the cut — is a different subject with a different boundary, and course 15.4 is where that is worked out.
Prompt · Put the missing column on your task list
When somebody hands you a list of tasks an assistant could help with, ranked by time saved. It writes the questions you have to answer yourself. It cannot know your error costs or your detectors, and any figure it offers for those is invented. Use it to build the empty table, then fill the last three columns from your own records.
I am going to give you a list of tasks at a garment factory that somebody has proposed handing to an AI assistant, ranked by the time each would save. Do not rank them, do not score them, and do not estimate any cost — you have no way to know my error costs and any number you produce for them would be invented. Instead, build me a table with one row per task and these columns, filling in only the first two from what I give you and leaving the rest blank: - Task - Times a month, and minutes saved each - WHAT DETECTS A WRONG ANSWER (blank for me to fill) - HOW LONG THE DETECTOR TAKES (blank) - WHAT HAS ALREADY HAPPENED BY THEN (blank) - COST OF ONE WRONG ANSWER THAT GETS THROUGH (blank) Then, for each task, ask me one specific question that would help me fill in the detector column — naming the person, system or event you think most likely to notice, phrased as a question rather than an assertion. Flag any task where you cannot think of a plausible detector at all, because that is the one I most need to look at. Here is the list:
AI can make mistakes — check anything you act on.
Check yourselfYour own list has a task at the top saving 300 minutes a month. What are the two questions that decide it, in order?Show the answer
First: what detects a wrong answer, how fast, and what has happened by then. If the answer is "nothing", stop. You have no way of knowing whether it is working, and 300 minutes of saving is not worth a risk you cannot measure. Second: is the check much cheaper than the work? If checking means re-doing, the time saving is small whatever the list says, and lesson 4's arithmetic will not clear the cost of the errors that get through.