The Supplier Scorecard: What to Measure Besides Price
Price is the only supplier attribute that arrives as a number, so it's the only one most stores manage. Everything else — whether the order ships complete, whether the file is readable, whether the cost you agreed in March is still the cost in September — shows up as time, apologies and stock-outs, none of which land in a spreadsheet. This is a way to put those on the same page as the price, in dollars, so the comparison stops being a matter of opinion.
Six things worth scoring
A scorecard that tracks fifteen metrics gets filled in once and abandoned. Six is about the limit of what a small buying operation will actually maintain, and these six cover the ways a supplier relationship costs money.
- Price level. Delivered cost on the lines you actually buy, not list, not the headline discount.
- Cost stability. How much your costs moved over twelve months, how many separate increases it took, and how much notice you got.
- Fill rate. The share of ordered lines that ship complete, first time.
- Lead-time reliability. Average days, and — more important — the spread around that average.
- Data quality. Whether the price file is machine-readable, whether part numbers stay stable, whether changes are flagged.
- Terms and support. Payment terms and cash discount, freight-free threshold, return and stock-rotation allowance, whether a human answers.
Measuring cost stability properly
"Their prices went up about 7%" is almost always the average of the percentage changes in the file, which is the wrong average. A file with 4,000 lines contains thousands you never buy, and one line you buy heavily counts the same as one you've never ordered. Weight the change by what you actually spend:
cost drift = Σ(annual qty × cost change) ÷ Σ(annual qty × old cost)
Three lines from a real basket, after one price file:
| Line | Annual qty | Old cost | New cost | Change |
|---|---|---|---|---|
| A | 240 | $8.50 | $9.35 | +10.0% |
| B | 600 | $42.00 | $42.84 | +2.0% |
| C | 900 | $2.10 | $2.31 | +10.0% |
The plain average of the three percentages is (10 + 2 + 10) ÷ 3 = 7.33%. Weighted by spend, the numerator is (240 × $0.85) + (600 × $0.84) + (900 × $0.21) = $204 + $504 + $189 = $897, and the denominator is $2,040 + $25,200 + $1,890 = $29,130. Drift is 897 ÷ 29,130 = 3.08%.
Less than half the headline figure, because the line carrying 86% of the spend barely moved. Run it the other way — line B up 10% and A and C flat — and the plain average is a soothing 3.33% while the weighted drift is $2,520 ÷ $29,130 = 8.65%. The unweighted number isn't conservative or aggressive; it's just unrelated to your bank account.
Fill rate and lead time convert to money
Fill rate is measured on lines, not orders and not dollars. Of 200 lines ordered in a month, if 24 arrive short, cut or substituted, the line fill rate is 88%. Order fill rate — the share of orders that arrive completely intact — looks far worse and is harder to act on; dollar fill rate flatters suppliers who ship the expensive things and skip the cheap ones.
Lead time needs two numbers. A supplier averaging six days with a spread of ±4 forces you to plan for ten; one averaging six days with a spread of ±1 lets you plan for seven. The average is identical and the working capital is not, because safety stock is bought against the spread, not the mean.
Two suppliers, same 300 lines
Supplier A quotes about 2% under Supplier B across a basket worth $180,000 a year — 200 order lines a month, average gross profit of $14 on the items involved. A's fill rate is 88%, B's is 97%. A's lead time is 6 days ±4, B's is 3 days ±1, so A needs roughly 8 extra days of cover. A sends a PDF that has to be retyped or converted; B sends a CSV that imports.
| Supplier A vs B, per year | Amount |
|---|---|
| Price advantage — 2% of $180,000 | +$3,600 |
| Lost sales — 216 extra short lines, half never recovered, × $14 | −$1,512 |
| Backorder admin — 216 lines × 12 min at $25/hour | −$1,080 |
| Extra safety stock — 8 days × $493/day × 20% carrying | −$789 |
| Price-file handling — 16 extra hours a year at $25 | −$400 |
| Net | −$181 |
The workings: 200 lines a month at 88% leaves 24 short, at 97% leaves 6, so A produces 18 extra short lines a month — 216 a year. Daily cost of goods from this supplier is $180,000 ÷ 365 = $493, so eight extra days of cover is $3,945 of additional inventory, and at a 20% annual carrying rate that's $789. Sixteen hours is roughly eighty extra minutes a month spent turning a PDF into rows.
A is cheaper by 2% and more expensive by 2.1%. That's the whole argument, and it's the argument you can't have without measuring. Note also how fragile the conclusion is: at a 92% fill rate instead of 88%, A wins comfortably. The scorecard isn't there to convict anyone — it's there to make the sensitivity visible.
Terms are worth more than they look
A 2/10 net 30 cash discount means 2% off for paying twenty days early. Annualised, that's:
(2 ÷ 98) × (365 ÷ 20) = 37.2% per year
On $180,000 of annual purchases, taking that discount is worth $3,600 — the same as the entire price advantage Supplier A was offering, available from a supplier who never lowered a price. If cash allows it, terms are usually the cheapest concession to ask for and the easiest for a supplier to grant, because it costs them financing rather than margin.
The other terms worth recording: the freight-free order threshold (it sets your minimum sensible order and therefore your average stock level), the return-to-vendor and stock-rotation allowance, and whether a price increase requires written notice. That last one is the difference between reacting to an announcement and discovering it in a file — the sequence to run when it happens is in the five-step response to a price increase.
Putting it together
Score each dimension 1–5, weight the dimensions, and keep the weights fixed once you've set them. Changing weights after seeing the scores is how a scorecard becomes decoration.
| Dimension | Weight | A | B | A wtd | B wtd |
|---|---|---|---|---|---|
| Delivered price level | 30 | 5 | 3 | 150 | 90 |
| Cost stability | 20 | 2 | 4 | 40 | 80 |
| Fill rate | 20 | 2 | 5 | 40 | 100 |
| Lead-time reliability | 15 | 2 | 5 | 30 | 75 |
| Data quality | 10 | 1 | 5 | 10 | 50 |
| Terms and support | 5 | 3 | 4 | 15 | 20 |
| Total (max 500) | 100 | — | — | 285 | 415 |
A scores 57%, B scores 83% — on a basket where A quotes lower prices. The weights are a starting point; a store selling fast-moving consumables would push fill rate above price level, and one buying long-lead specials would push lead-time reliability up.
Two rules keep this honest. First, a dimension only enters the scorecard if you can point at where the number came from — an order log, a delivery date, a price file, a stopwatch. Second, a score of 1 on any dimension is a conversation regardless of the total, because a supplier who is excellent at everything except shipping complete orders is still a supplier who doesn't ship complete orders.
Where the data actually comes from
None of this requires software you don't have, but it does require someone writing things down as they happen — reconstructing a year of fill rates in December is not possible.
- Fill rate: when a delivery is checked in, record lines ordered and lines received complete. Two numbers per delivery.
- Lead time: order date and receipt date. Both already exist; they just aren't usually in the same place.
- Cost drift: every price file, kept and dated. This is the one most stores can't produce, because Shopify holds a single cost per variant with no history and each import silently overwrites the last. A dated folder of the original files is the cheapest possible fix.
- Data quality: log the minutes spent making each file usable. The habit of timing it is what turns "their file is a nightmare" into $400 a year. What to look for in the files themselves is in reading a supplier price list.
- Terms: read the agreement once a year and write the numbers on the scorecard.
What to do with a bad score
A low total is not automatically a reason to leave. Three responses in rough order of cost:
Show them the scorecard. Fill-rate and lead-time data is unusual enough from a small account that it changes the conversation, and it's the rare complaint a rep can escalate internally with evidence attached. Suppliers generally do not know what their service costs you, because nobody has ever told them in dollars.
Price the gap instead of closing it. If a supplier is structurally slow but genuinely cheap, the honest answer may be to keep buying and hold more stock — as long as the carrying cost is charged to that supplier's score and the affected SKUs are priced with it in mind.
Add an alternative for the lines that matter. Not a wholesale switch, which is expensive and slow, but a second source on the subset that drives the damage. The mechanics — finding a common key, comparing offers with different pack sizes, and what splitting volume costs you in lost tier discounts — are in second-sourcing a SKU. And if a line performs badly on every dimension including sales, the exit arithmetic is in when to drop a product line.
Reviewed once a quarter, a scorecard takes about an hour and mostly confirms what you suspected. Its real value is the quarter it doesn't: when the supplier everyone likes turns out to be the one whose costs drifted 9% while the pleasant one drifted 2%, and the reason nobody noticed is that nobody had written the number down.
Cost drift is the row you can measure today
Load your latest supplier price list and your Shopify product export into the free Supplier Price Margin Checker to see which lines moved, by how much, and what it does to margin — matched by SKU, entirely in your browser.
Open the free checker ↗