Lessons · Lesson 3 of 3
What being measured does to a supplier
Price three behaviours a badly built scorecard buys, and decide what a supplier scorecard should never be allowed to decide on its own.
Lesson 3 of 3 · 40 min
The part of this that is not yours to manage
A scoring system looks like an instrument for watching, and the natural worry is whether it is accurate. The more useful worry is what it does to the people at the other end. A published rule is also an instruction. Firms judged by it will read it closely and act on whatever the wording rewards. This lesson draws a line around what a number of this kind is actually good for.
Two courses in this academy have already shown a measure changing the thing it measures.
Course 5.6 follows a line supervisor who stopped an improvement that was working. His own scorecard made him do it: the cost of the improvement landed on a line he was measured on, and the benefit landed somewhere he could not see. Course 6.6 follows a compliance team whose coordinators closed findings on photographs. They did it correctly, under a correct procedure, toward a target their manager had set — and drove the repeat-finding rate from 18% to 41% while every headline number improved.
In both cases the fix is available to you. The supervisor works in your factory. The coordinator works in your team. You can change the procedure, move the cost, re-base the target, explain.
A supplier is not that. It is a separate company with its own margin, its own ability to sell to somebody else, and a merchandiser whose bonus depends on your score. It can read your formula. It will optimise against your formula, honestly and rationally, in ways you did not intend and cannot forbid. You have exactly one lever, and it is the formula.
This lesson prices three of those behaviours on Undercliffe's own season.
One: a hard on-time metric buys you short shipments
Undercliffe fixed its definitions for SS27, along the lines lesson 1 sets out, and published the new delivery weighting. Of forty delivery points, 25 are for on time and 15 for in full, with in full measured at order level against a 5% tolerance.
Read that as a supplier would. On time is worth two thirds more than in full, and the first 5% of a shortfall is free.
PO UC-7442. Nedra Apparels, 9,000 cotton chinos, ex-factory 14 September. On 11 September the nominated waistband label runs short by enough to stop 430 pieces. Nedra ships 8,570 on the fourteenth, which is 95.2% and inside the tolerance. The balance of 430 pieces follows on 6 October.
On the scorecard this is one purchase order, delivered on time and in full. In the warehouse it is two deliveries.
| Line | Working | Cost |
|---|---|---|
| Sea freight, less than container load | 1.85 cbm at USD 92, plus documentation USD 185 and haulage USD 240 | USD 595.20 |
| Customs entry | second entry | USD 65.00 |
| Inbound booking at the distribution centre | second slot | USD 118.00 |
| Goods-in and putaway | 12 cartons at USD 1.40 | USD 16.80 |
| Merchandising and finance administration | 2.5 hours at USD 31 | USD 77.50 |
| Invoice and credit reconciliation | one shipment split across two invoices | USD 137.00 |
| Per split delivery | USD 1,009.50 |
Less than container load means your goods share a container with other shippers' goods, so you pay by the cubic metre (cbm) rather than for the whole box.
Now the season. In AW26, before the weighting was published, 3 of 61 POs arrived as a split delivery, which is 4.9%. In SS27 it was 14 of 58, which is 24.1%. At AW26's rate, SS27 would have produced three. Eleven of them are new, and eleven splits cost USD 11,104.50.
And the score went up. Undercliffe's SS27 delivery score was 49 of 58, or 84.5%, against AW26's 67.2%. Strip out the splits — count a PO as a failure if the whole quantity did not arrive on the date — and SS27 is 38 of 58, or 65.5%.
The correction is small, and it is not a lecture to suppliers. Count a split delivery as a failure against the original PO, whenever the balance lands. Measure fill at colourway level with no tolerance, and leave the tolerance in the contract where it decides payment. Then add one plain number that is not a score at all: deliveries per purchase order. It costs USD 1,009.50 every time it exceeds one, so it deserves a line of its own.
Two: a gate that decides next season buys you late news
Undercliffe's SS27 vendor manual carried a sentence its suppliers read very carefully: a supplier below 70% receives no new development for a season.
That sentence is not unreasonable. It is also an instruction to every supplier merchandiser about when to put a risk in writing.
Undercliffe tracks one number that most buyers never calculate: warning lead time, the days between a supplier's first written notice that a date is at risk and the confirmed ex-factory date. In AW26 the mean was 34 days. In SS27, after the gate was published, it was 17.
Price the difference on one order.
PO UC-7469. Selmani, 5,200 pieces of a taped-seam technical shell, FOB USD 41.20, order value USD 214,240. Ex-factory 12 August, and the range lands in store on 26 September.
| Flagged | What is still available | Cost |
|---|---|---|
| 34 days out | Re-place the fabric with a second mill at USD 0.74 a piece more | USD 3,848 |
| 34 days out | Or move the store date by ten days, which this range can absorb | nil |
| 17 days out | Air the whole order to hold the store date | USD 14,321.92 |
Now the air arithmetic, because that is where the money is. Actual weight is 5,200 pieces at 0.62 kg, so 3,224 kg. Volumetric weight is 260 cartons at 0.072 cbm, which is 18.72 cbm, and at the usual air convention that is 3,120 kg. An airline charges on whichever is greater, so this shipment is charged on actual weight. At USD 4.35 a kilo plus USD 0.48 fuel and security, that is USD 15,571.92, plus origin handling USD 420 and customs USD 190. Against the sea leg it replaces at USD 1,860, the net cost of the seventeen days is USD 14,321.92.
Nothing dishonest happened. The supplier raised the flag the day its mill confirmed the delay in writing, which is precisely what a supplier does once a flag is a scored event. Before the gate, the merchandiser telephoned at the first sign. The gate did not make anybody lie. It changed the moment at which a suspicion becomes a notice, and that moment is worth USD 14,321.92 on one order out of fifty-eight.
Course 6.6 makes the same observation about a buyer's own coordinators, and reaches a fix you can implement yourself. Here you cannot, because the person whose behaviour changed does not work for you. So the fix has to be in the formula.
- You cannot gate on the score and expect early warning unless early warning is worth something on the card. Undercliffe now runs two cards. The first scores outcomes: delivery, quality, cost. The second, worth 8 points of the hundred, scores conduct when things go wrong. It looks at warning lead time, at whether the supplier proposed at least one priced recovery with the notice, and at whether its own weekly reported dates matched what actually happened.
- Eight points is deliberately large. Undercliffe's bands are ten and twelve points wide, so eight points moves a supplier most of a band. A warning credit worth two points is a gesture, and a supplier will not trade a season of placement for it.
- What you then do about a supplier that is genuinely failing — escalation, expediting, and when a claim is worth less than the supplier — is course 27.6. This lesson stops at the measurement.
Three: a target nobody can reach is not a target
| Band | Delivery score | Consequence |
|---|---|---|
| Grow | 92% and above | first refusal on new development |
| Hold | 80% to 91.9% | current volume maintained |
| Watch | 70% to 79.9% | no new development, quarterly review |
| Exit | below 70% | placement wound down |
Selmani's raw AW26 score was 18.2%. The distance to the bottom of the next band up is 51.8 points. To cross it, Selmani would have to convert six of its nine failures — and eight of those nine were Undercliffe's approvals, amendments and nominations.
There is no action available to Selmani that reaches the next band. So Selmani does not attempt one, and it is right not to. A second cutting shift, a fabric buffer, an extra planner: all of it lands it at 45% and in exactly the same band, having spent the money. A target with no reachable next step is not a target. It is a label, and a labelled supplier manages nothing.
Now the adjusted score, 90.9%. On eleven POs, one failure is 9.1 points, so a single order is the whole distance between Hold and Grow. That is a gradient somebody can act on, so price it from Selmani's side. Holding a 900-metre greige buffer at its own cost at USD 4.64 a metre ties up USD 4,176 for six weeks. Greige is fabric straight off the machine, before dyeing and finishing. Three days of a second cutting shift costs USD 1,340. Set that against first refusal on new development, which in AW26 meant nine new styles and USD 1.94m of placement.
That is a supplier making an investment decision, which is the only thing a scorecard has ever usefully caused. It became possible the moment the score stopped being contaminated, which is why lesson 2 comes before this one.
The movement that was not a factory
One more, briefly, because it will happen to you and it looks exactly like performance.
Undercliffe's ex-factory date came from the supplier's own weekly report until the second quarter, when it moved to the forwarder's cargo-receipt date at the container freight station. In the quarter the source changed, base on-time fell 6.1 points. No factory did anything differently. The second source is the better one, because it is a third party's record rather than the measured party's own, and the fall is the size of the gap between what suppliers reported and what the forwarder saw.
The rule is the one lesson 1 ends on, and this is the reason that rule exists: a definition or a source changes only between periods, and the previous period is restated on the new basis and published alongside. Otherwise half of every scorecard movement you will ever explain to a board is a data-source movement wearing a supplier's name.
What a scorecard is actually for
Four things, and they are worth having.
- Allocating attention. You have fourteen suppliers and time to visit three. The card tells you which three, and that is a good enough reason to build one.
- Making a conversation evidential. "Your last four purchase orders, here are the dates" is a different meeting from "you are always late", and it is a much better meeting for a supplier who is doing well.
- Trend at base level. Undercliffe's base going from 67.2% to 90.2% under attribution was the most important thing the exercise produced, and it was a finding about Undercliffe.
- Triggering a look. A score is a reason to go and find out. It is never itself the finding.
And here is what it must never decide on its own.
- Dropping a supplier. A number cannot see capability you never tested, difficulty you imposed, or what a replacement costs. That decision is course 27.6's subject, and it needs more than a card.
- The size of a claim. A score is an average across a season. A claim is about one shipment, one cause and one contract. Course 7.5 is the same conversation from the factory's chair.
- Whose fault one specific slip was. The card aggregates. Fault is particular, and it is established from dated documents, not from a percentage.
- Price against performance. A supplier at 96% that is 9% expensive is not thereby the right supplier. Let the weights inform the trade. Do not let a weighted total make it silently.
- Anything a gate should decide. Compliance, safety, a restricted substance, an unauthorised subcontractor: these stop an order. They are not points.
The last sentence of the course is the one to argue about with your own team. A scorecard is a hypothesis about a supplier, expressed as a number, and every quarter it invites you to go and test it. The moment it stops being a hypothesis and becomes a verdict, it stops improving — and your suppliers, who are cleverer than your formula, start managing it instead of the orders.
Check yourselfYour on-time percentage improved 19 points this season and your suppliers say nothing has changed. Name three explanations before you accept the improvement.Show the answer
One: the definition or the data source moved, and nobody restated the prior period. Two: the metric bought a behaviour — split deliveries counted as complete, a tolerance absorbing shortfalls, a date quietly revised — so the same shipments now score differently. Three: the mix of orders got easier, because a season with fewer new styles and longer lead times raises everyone's number. Only after all three have been eliminated is it a real improvement. The cheapest way to eliminate the first is to have published the restatement before anybody asked.
Prompt · Read your scorecard for the behaviour it is buying
Before you publish a scorecard or a placement rule, and once a year afterwards, to find out what a rational supplier will do about it.
Act as the merchandising director of a garment factory that is about to be measured by this scorecard, and whose bonus depends on the score. Your job is to tell me, without embarrassment, exactly how you would raise the score at the lowest cost to your factory. Include the ways that cost the buyer money. Here is the card: dimensions and weights [PASTE], the exact formula for each dimension including what counts as on time, what counts as in full, tolerances, measurement point, and the data source for each date [PASTE], the placement bands and what each one triggers [PASTE], and any gate that is expressed as points rather than as a stop [PASTE]. My current scores by supplier are [PASTE]. Do the following. First, for each dimension, name the cheapest action that improves the score without improving the underlying service, and say who pays for it. Second, price at least two of those actions from my side — what the behaviour costs the buyer per occurrence and per season — using the freight, warehouse and administration figures I give you: [PASTE]. Third, tell me what the card makes it rational for a supplier to STOP doing, especially around telling me bad news early, and estimate what that costs me on one order that goes wrong late. Fourth, find every target on this card that a supplier at the bottom cannot reach by any action available to it, and say what such a supplier will do instead. Fifth, name any gate I have written as points, and say what score a supplier could still achieve while breaching it. Sixth, propose the smallest set of changes that removes the behaviours you named, and tell me what each change costs me to administer and what it makes worse. Be concrete and be uncharitable; a polite answer here is a useless one.
AI can make mistakes — check anything you act on.