Lessons · Lesson 2 of 7
- 01 · Two kinds of question, and why the answer sounds the same
- 02 · The four things it did well, measured here
- 03 · The fact nobody wrote down
- 04 · When 95% right is worth less than no answer
- 05 · Ten tasks, sorted by the wrong column
- 06 · Five ways to be confidently wrong
- 07 · The four questions, and what Meskala kept
The four things it did well, measured here
Sort the tasks that worked from the one that did not, and find the single property that separates them — one you can test on any new proposal in about a minute, without running anything.
Lesson 2 of 7 · 16 min
How Meskala measured, and why the method matters more than the result
Souad Amrani runs systems and quality at Meskala. When the factory decided to look seriously at an assistant, she refused to judge it by using it. Using something tells you whether you like it. It does not tell you whether it works.
So she measured, and the method was the same every time. Take real Meskala material. Run the tool on it. Have the person who normally does that work do it on their own, without seeing the tool's output. Then compare the two. Where they disagreed, a third person decided which was right by going back to the source document.
That last step is the one people skip, and skipping it turns a measurement into a demonstration. If the person checks the tool's answer instead of producing their own, you have not measured the tool. You have measured how convincing it is. That is a different thing, and as lesson 4 shows, a much bigger one.
| What was measured | On | Agreed with the person |
|---|---|---|
| Cut a long buyer email thread to five lines | 60 threads | 58 |
| Turn an English chase into an Arabic draft | 45 drafts | 41 |
| Pull the confirmed delivery date out of a supplier confirmation | 240 confirmations | 228 |
| Route an incoming email to one of seven owners | 300 emails | 271 |
| Answer a question about how Meskala itself works | 20 questions | 4 |
As rates: 96.67%, 91.11%, 95.00%, 90.33% — and then 20.00%.
Say again what these are. They are one factory's results, on that factory's documents, in one month, with one tool. They are not a claim about any system's general ability. They are on the page for the pattern, which holds even though the numbers do not.
The property that separates the top four from the bottom one
Look at what each task actually needs.
To cut a thread to five lines, every fact you need is inside the thread. To draft the Arabic, every fact is inside the English. To pull a delivery date, the date is printed on the document. To route an email to one of seven owners, the seven owners were given in the instruction and the subject matter is in the email.
To say what Meskala's consumption is for MT-2280, the fact is in a marker file in the cutting room. Nothing in the question, and nothing the system was ever fitted to, contains it.
That is the whole distinction. You can test it in a minute, with no software:
Can you point at the place, inside what you handed over, that the answer has to come from?
If you can, you are in the class that works. If you cannot, you are in the class that produces 1.33 m and means it.
Notice what this rule is not. It is not "easy tasks work and hard ones do not". Summarising a forty-message thread with three changes of mind in it is not easy. It is not about stakes, or creativity, or how clever the question is. It is about where the raw material of the answer is sitting when you press send.
Where "did well" is still not "always"
96.67% is not 100%, and the two summaries that failed teach you more than the fifty-eight that did not.
Both dropped the same kind of fact: a condition buried in the middle of a long thread. In one, Kessendal's merchandiser had written that she could live with the revised ship date if the pre-pack ratio changed at the same time. The summary carried the ship date and dropped the condition. In the other, a mill had offered a price subject to taking the full dye lot.
That is not a random error. It is a systematic one, and the two behave very differently under checking. A random error is caught by spot-checking a sample, because it is as likely to be in your sample as anywhere else. A systematic error is not, because it always happens in the same place — and that place is the middle of a long document, which is exactly the part a person checking a summary skims.
So the honest reading of that row is not "it summarises at 96.67%". It is this: it summarises reliably, and it loses conditions, so somebody has to read the middle of anything where a condition would change the decision. That is an instruction you can use. A percentage is not.
Two boundaries, so you know what this course is not doing
Extraction is a whole course of its own. Pulling fields out of your own documents at volume — what a field is worth, how a mismatch is settled, what happens when the extracted number and the counted number disagree — belongs to course 15.3, which builds it properly on order data. This lesson only shows that extraction sits in the class where the answer is present in the input.
Getting good work out of it day to day — what to hand over, how to frame the ask, how to check what comes back — is course 15.2. Here we are drawing the map, not walking it.
Check yourselfA vendor offers to have the system 'learn your factory' so it can answer the consumption question. What is the first thing to ask?Show the answer
Which file it reads. The consumption fact lives in Meskala's marker files. A system can only answer from them if somebody connects those files and keeps them connected as markers are revised. That is a data-plumbing project with a cost and an owner. It is a perfectly reasonable thing to buy, but it is a different purchase from the assistant. Lesson 3 shows that at Meskala it would have moved 21 of 90 questions, not the 35 that mattered most.