Lessons · Lesson 2 of 3
Coverage is not assurance
Why a programme that hits every target can be getting worse, and how to measure one without corrupting it.
Lesson 2 of 3 · 36 min
The year that hit every target
A programme can hit every target it was set and still be getting worse. This lesson shows how. Everything the team was measured on went up. The number nobody was working out was how often the same problem came back at the same place. The cause was not dishonesty. The cause was a target.
Ourique's FY25 compliance report ran to two pages, and both of them were green.
| Measure | Target | Achieved |
|---|---|---|
| Audits completed against plan | 100% | 100% |
| Findings closed within sixty days | 90% | 94% |
Forty-six audits planned, forty-six done. Twenty-one verification visits planned, twenty-one done. Two hundred and fourteen findings raised, two hundred and one closed inside the window. The manager took it to the board and the board was pleased.
There was a third number. Nobody was measured on it because nobody was working it out: the repeat-finding rate. That is the share of findings raised at an audit that had also been raised at that site's previous audit.
In FY24 it was 18%. In FY25 it was 41%.
The programme was closing findings faster than ever, and the same findings were coming back more than twice as often. Those are not two problems. They are one problem, and it starts with somebody doing exactly what the procedure told them to do.
A correct action with a bad outcome
Ourique's closure procedure said a finding may be closed on documentary or photographic evidence accepted by the coordinator. That sentence was written to save money, and it does save money. A verification visit costs USD 780. Verifying all 214 findings would have cost 214 × 780 = USD 166,920. That is more than the whole audit line and nearly the whole compliance budget. Nobody can visit every closure. The procedure was not stupid.
The coordinators were measured on closing findings inside sixty days. A photograph arrives by email in two days. A verification visit takes three weeks to book. So the coordinators closed what they could close quickly. They were working correctly, under a correct procedure, towards a target their manager had set.
| Closed on | Findings | Share | Recurred at the next audit | Recurrence |
|---|---|---|---|---|
| A photograph or a document | 179 | 83.6% | 61 | 34.1% |
| A verification visit | 35 | 16.4% | 3 | 8.6% |
A finding closed on a photograph came back four times as often as one closed on a visit. That is the whole of the rise in the repeat rate, and a metric bought it.
Notice what the metric did not do. It did not make anybody dishonest. Every one of those 179 photographs was a real photograph of a real thing at a real site. The blocked aisle was cleared, photographed, and blocked again the following Tuesday, because nothing about the reason it was blocked had changed.
Rewriting the rule, and paying for it
Here is Ourique's new closure rule, written in October and in force for FY26.
- A zero-tolerance or high-class finding closes on a verification visit. Nothing else.
- A repeat finding closes on a verification visit, paid for by the supplier. This is in the vendor agreement, so it is a contract term rather than an argument.
- A records finding may close once on a remote review. That means a full payroll export for three consecutive months, checked against the site's own production output records, reviewed by an Ourique coordinator at USD 90. It may close this way once. A repeat records finding at the same site needs a visit. The rule states its own limit out loud: a remote review cannot detect a second set of books, and pretending otherwise is how a programme fools itself.
- A physically visible low-class finding may close on a photograph. It must be dated, with the site's own reference visible in the frame.
Now cost it, because a rule nobody has costed is a wish. Apply the new rule to FY25's actual findings and 68 of the 214 would have needed more than a photograph. Of those, 22 are repeats or high-class and need a visit, and 24 are records findings that qualify for a remote review. The rest are already covered.
| Item | Count | Unit | Cost |
|---|---|---|---|
| Verification visits | 22 | USD 780 | USD 17,160 |
| Remote payroll reviews | 24 | USD 90 | USD 2,160 |
| Required | USD 19,320 | ||
| Budgeted in the FY26 band treatments | USD 11,076 | ||
| Drawn from the reserve | USD 8,244 |
The reserve was USD 15,480. After this it is USD 7,236, and that is the honest position to take into the year. The closure rule is affordable, but it eats more than half the reserve. If two escalations arrive before June, the programme has to choose between them.
You cannot bolt a verification rule onto a plan without pricing the plan again. Most programmes find this out in month seven.
The audit that found nothing
Fahmy Knitwear, 10th of Ramadan, Egypt. Band B, score 52. Three announced audits in three years, and each one came back with zero findings.
For three years that was reported as the base's best result. It is not a result. It is a measurement, and the thing it measures is the audit.
Here is what the audit was. One auditor, one day, at a site with 340 workers in two buildings. Eleven days' notice. Eight worker interviews, held in a meeting room with a glass wall onto the corridor. A payroll sample of twelve records, picked by the site from a list the site prepared. A site tour following a route the site prepared.
That design will find a blocked fire door and a missing chemical data sheet. It cannot find much else. Three zero-finding reports in a row are the correct output of it.
Ourique's own FY25 numbers make the same point without any story. Across 46 announced audits, four came back with zero findings — 8.7%. Across nine unannounced visits, none did.
Note what this does to the ranking in lesson 1. Fahmy scored 8 out of 20 on audit history. That is a low, comfortable number, produced by three audits that could not see. Input four inherits the blindness of the audits that feed it. After the re-visit, history was rescored to 15 and Fahmy moved from band B to band A. Nothing about the factory changed that quarter. What changed was the evidence.
What a visit is worth, in numbers
Ourique tracks one figure that most buyers never work out. When a later unannounced visit confirms that a non-compliance existed at the time of an earlier audit, did the earlier audit record it?
Across FY25, an announced audit had recorded it 45% of the time. An unannounced visit had recorded it 72% of the time. Those are Ourique's own hit-rates, measured on its own base. They are not a published figure and they do not transfer to yours. Work out your own; the method is what matters.
Now the arithmetic that decides a plan.
| A year of | Chance the issue is found | Cost |
|---|---|---|
| One announced audit | 45% | USD 1,450 |
| Two announced audits | 69.8% | USD 2,900 |
| One announced and one unannounced | 84.6% | USD 3,350 |
Two announced audits: 1 − 0.55 × 0.55 = 69.8%. One of each: 1 − 0.55 × 0.28 = 84.6%. So for USD 450 more than a second announced audit, you buy 14.8 percentage points of detection.
What to measure instead
If a metric drives behaviour, the fix is not to stop measuring. The fix is to measure things whose cheap version is also the real version.
Ourique replaced its two numbers with four.
- Repeat-finding rate, by site and by finding class. The cheap way to improve it is to fix the cause, because there is no cheap way to make a finding not come back.
- Closure evidence mix. The share of closures resting on a visit or a remote review rather than a photograph, reported by class. This is the number that would have caught FY25 in its second quarter instead of at the year end.
- Detection rate. The share of non-compliances confirmed later that the earlier audit had recorded. It is the only measure here that is about the programme's own instrument rather than about suppliers.
- Time from finding to verified closure, not to closure. The sixty-day target survives with the word verified added, and the target stretches to ninety days, because you cannot book a visit in two.
Notice what has been given up. Stretching the target makes the headline number worse in year one, and a manager who has to explain that to a board will want the old measure back by March. That is the price of a metric you cannot satisfy cheaply, and you pay it in the meeting where the target is set, not afterwards.
Notice too what is missing. Audits completed against plan is still tracked as a budget control, because you need to know whether the plan you costed was actually carried out. But it is no longer a performance measure. It is a measure of activity, and you can satisfy any activity measure by doing the activity badly.
The same money, twice
Here is the comparison the argument always comes down to. Take USD 49,300, which is what the FY25 tier-one sweep cost, and spend it two ways.
| Programme X | Programme Y | |
|---|---|---|
| Design | Every tier-one factory, one announced audit | The seven band-A sites, four visits each; everyone else desk-reviewed |
| Band-A visits | 7 announced audits | 14 unannounced audits and 14 spot visits |
| Other sites | 27 announced audits | 96 desk reviews |
| Verifications | Inside the plan | 10 |
| Cost | USD 49,300 | USD 48,920 |
| Sites touched | 34 of 103 | 103 of 103 |
| Detection at a band-A site | 45% | 92.2% |
Programme Y costs USD 380 less. It touches every site in scope, though 96 of them only on paper. And on the seven sites that carry the exposure it detects at 92.2% — 1 − 0.28 × 0.28 — against 45%.
Programme X has better coverage in the sense a board understands: thirty-four real audits, thirty-four reports, thirty-four files. Programme Y has better assurance, and assurance is the thing the programme exists to produce.
The sentence to carry out of this lesson: auditing every supplier once a year and auditing the risky third four times a year cost the same and are not the same programme. One of them buys documents. The other buys the truth about seven factories.
Check yourselfYour closure rate rose from 79% to 94% in a year and your repeat-finding rate rose with it. What is the first thing you look at?Show the answer
The evidence type behind the closures, split by finding class. A rising closure rate and a rising repeat rate together are the signature of closures accepted on evidence that cannot support them. Usually that means photographs against records findings. Split the closed findings into photograph-closed and visit-closed, work out the recurrence of each, and you will usually find the whole rise sitting in one column.
Prompt · Test whether your closures are real
When the closure rate is rising and you want to know whether the programme is improving or only reporting that it is.
Act as an auditor of my own compliance programme, not of my suppliers. I want to know whether my closed findings are actually closed. Here is my finding register for the last two years, one row per finding, with the site, the audit date, the finding text, its class, the closure date, and the evidence the closure was accepted on — photograph, document, remote review, or verification visit: [PASTE THE REGISTER]. Do the following. First, calculate my repeat-finding rate for each year: the share of findings raised at an audit that had also been raised at that site's previous audit. Second, split the closed findings by evidence type and give me the recurrence rate of each type, so I can see whether one column is carrying the whole problem. Third, list every finding that was closed on a photograph but is a RECORDS finding — wages, hours, contracts, age documentation, deductions — because a photograph cannot evidence a record, and tell me which of those sites needs a visit now. Fourth, name any site with two or more consecutive audits returning zero findings and more than two hundred workers, and tell me what a zero-finding result at that site actually measures. Fifth, propose a closure rule that states which finding classes may close on which evidence, and include the limit of each evidence type in the rule itself. Sixth, price the rule: how many extra verification visits and remote reviews it implies against my register, at [VISIT RATE] and [REVIEW RATE], and what that total is. Do not tell me the programme is improving because the closure rate rose. Tell me what the closures were made of.
AI can make mistakes — check anything you act on.