Most shippers have a carrier scorecard. Far fewer have one that does anything — the document gets built, distributed once or twice a quarter, and then quietly ignored by the carriers it was supposed to influence. There are two reasons, and neither is the choice of KPI. The first is that the numbers came from the carriers themselves, so by the third quarter nobody trusts the file: two carriers are reporting 97% and the line is still starving. The second is that scores produce no consequence, and carriers work that out faster than anyone admits. Fix those two and almost any sensible metric set will work; leave them and no metric set will. The prize is not a tidier report — it is that a rejected truckload raises the price of that move by nearly 15% on average, which makes tender acceptance a cost line rather than a courtesy. Book a 30-minute session and bring your current scorecard — we'll run it against the trust and consequence tests in Fleet Rabbit and show you which of your carriers would dispute their own numbers.
2026 GUIDE · CARRIER SCORECARDS
Carrier Scorecards for Manufacturing Inbound
Four to six metrics that matter, data both sides accept without argument, a cadence carriers can anticipate, and consequences real enough that the score changes behaviour.
75%
Carrier A
Same volume next quarter
95%
Carrier B
Same volume next quarter
If this is your programme, the scorecard is theatre. Behaviour only changes when scores produce outcomes — and carriers establish which kind of programme they are in within about two quarters.
Why Scorecards Die
They almost always die the same way, and the failure is visible in the first two quarters if you know what to look for.
The data came from the carrier
Someone builds a spreadsheet, asks each carrier to submit their own on-time numbers, and by the third quarter nobody trusts the tab — because two carriers are reporting 97% while the plant keeps chasing late loads.
The ruleOnly worth building if it runs on data you control, not numbers handed over in a PDF once a quarter
Nothing happened as a result
Scores were published and volume allocation did not move. Carriers are commercially rational — once they establish that the score has no consequence, they optimise for the things that do.
The ruleCarriers need to believe the scores translate into real decisions, and belief comes from repetition
Everyone defined the metrics differently
One of the biggest challenges in scorecarding is that each shipper, carrier and receiver defines their KPIs differently — so the same shipment is on time in one system and late in another, and the review becomes a definitional argument.
The ruleDefine on-time precisely, in writing, before the first score is issued
It was assembled by hand
Pulled together from carrier portals and spreadsheets manually. Done once, then abandoned — because the effort per cycle exceeds the attention anyone has for it in month three.
The ruleThe data has to land in one place automatically or the habit does not survive
The signal that it is working
Two consecutive review cycles producing numbers the carrier does not dispute. That is the first real indication the programme has a foundation — before that, everything else you build on top is standing on a definitional argument waiting to happen.
Choose Four to Six Metrics
Resist the temptation to track fifteen. Pick the ones that genuinely affect your operation and your line, and accept that everything else is context rather than score.
← Swipe to see all columns →
Weight them to your operation rather than equally. A plant running tight sequenced windows weights on-time delivery far above billing accuracy; a site with thin margins and high invoice volume may weight the reverse. What matters is that the weighting reflects your priorities and is published, so a carrier improving the wrong thing is a failure of communication rather than of effort.
Bring one carrier's last quarter to a 30-minute call
We'll build their scorecard live in Fleet Rabbit from your gate events, appointment records and invoice data — no carrier-supplied figures — and you will see immediately whether your current numbers and ours agree. Where they diverge is usually where the last review went sideways.
Data Both Parties Trust
The credibility of the whole programme rests here. Five sources, ranked by how hard they are to argue with.
StrongestYour own gate and dock eventsBarrier timestamps, door assignment, unload start and finish. Generated by you, on your site, and effectively unarguable.
StrongTransport and EDI feedsTender, acceptance, status and despatch messages pulled structurally rather than reported. Consistent by construction.
GoodInvoice and claims recordsObjective, dated, and already reconciled by finance. The easiest metrics to defend in a review.
UsefulPublic safety and compliance dataSafety ratings and inspection records, where available. Neutral third-party input that neither side authored.
WeakestCarrier-supplied performance figuresThe route by which most scorecards lose credibility. Use them for context if you like, but never as the score.
Governance, briefly
Assign ownership for data validation, set the refresh cadence, and publish a dispute-resolution pathway before the first scorecard goes out. A carrier who believes a number is wrong needs a route to say so that is not the quarterly review — otherwise the review becomes the route, and the review has other work to do.
Cadence Carriers Can Anticipate
Three rhythms doing three different jobs. Most operators settle on monthly scoring with a quarterly trend view, and the weekly layer is what makes the monthly one credible.
Weekly
Exception checksCatches SLA breaches while there is still time to act on a specific shipment. Not a score — a working queue.
Monthly
The score itselfTracks trends and carrier mix shifts, and catches problems fast enough that a bad month can still be discussed as a bad month.
Quarterly
Formal reviewAligned with business reviews and routing guide audits. Tells you whether a bad month was a blip or a decline, and it is where consequences get applied.
Annually
Audit the scorecard itselfMetrics drift. Ask whether it still measures what matters, whether the weights still reflect priorities, and whether consequences are still being applied.
On low-volume lanes, guard against noise. Use rolling ninety-day windows where volume is volatile and discrete monthly windows for steady-state operations, and make sure the sample is large enough to mean something. Annotate catastrophic single events rather than letting them define a quarter — where appropriate, exclude them from trend calculations after a root cause analysis, and say so on the scorecard.
Benchmark Against Peers, Not Absolutes
A number in isolation is close to meaningless, and this is where a lot of otherwise sound programmes tolerate underperformance for years.
The isolation trap
A carrier scoring 92% on-time delivery looks fine on its own. If the rest of your carrier base averages 96%, that 92% is a problem — and it will not read as one on a report that shows only their own figure.
The aggregation trap
A national 96% on-time rate can hide a 78% rate on a specific lane in a specific quarter. Granularity by lane, by facility and by period is what makes the metric actionable rather than reassuring.
The stale-benchmark trap
A peer benchmark calibrated against a different carrier mix can quietly mislead for years. Recalibrate when the mix changes, not only when someone questions it.
What good looks like
Each carrier's score shown against the peer average for the same lane type and period. Same effort to produce, considerably harder to argue with, and it makes the conversation about the gap rather than the number.
The Consequence Ladder
This is where most programmes fall apart. Four mechanisms, and the mechanism matters far less than applying it consistently.
RewardTier-based volume allocationHigher-performing carriers get first-tender priority on premium lanes. The cleanest positive consequence, and the one carriers respond to fastest.
CommercialPerformance-based rate adjustmentApplied at contract renewal, using the score as the evidence base rather than as an opinion.
WarningProbationary statusBelow a stated threshold, with a defined improvement window and stated exit criteria. Probation without an exit is just a label.
TerminalRemoval from the routing guideFor sustained underperformance — and the scorecard is what provides the documentation to support that decision when it is questioned later.
Where the score has to appear
A scorecard nobody opens changes nothing. Surface it at booking time so routing decisions see it, review it in carrier business reviews, and bring the lane-level data to every rate negotiation. Once carriers begin to anticipate the cadence they adjust to it — and those who do not adjust become candidates for replacement, with the documentation already in place.
Deployment brief
Twenty minutes on your carrier mix and your last two quarters
On the call we'll map which of the six metrics you can already produce from data you control, show what the peer-benchmarked view of your base looks like in Fleet Rabbit, and identify the carriers where a consequence has been overdue for more than two cycles. You keep the metric definitions and the weighting model either way — most teams find at least one carrier scoring well on a metric that has quietly stopped mattering.
Frequently Asked Questions
How many metrics should a carrier scorecard have?
Four to six, chosen because they genuinely affect your operation rather than because they are available. For most plants the shortlist is on-time pickup, on-time delivery, tender acceptance, claims ratio, billing accuracy and dwell. Resist tracking fifteen — a scorecard with too many metrics dilutes the ones that matter and takes longer to produce, which is how the habit dies in month three. Weight them to your own priorities and publish the weighting so carriers know what to improve.
Why do carriers dispute the numbers?
Usually because the definitions differ. One of the biggest challenges in scorecarding is that each shipper, carrier and receiver defines KPIs differently — a load can be on time to the appointment and late to the required date, and both parties are technically correct. Define on-time precisely in writing before the first score is issued, and use data you control rather than figures the carrier supplies. Two consecutive cycles producing numbers the carrier does not dispute is the first real signal the foundation is sound.
How often should we review?
Weekly for exceptions, monthly for the score, quarterly for the formal review. Weekly checks catch SLA breaches while there is still time to act on a specific shipment; monthly reviews track trends and carrier mix shifts; formal quarterly reviews align with business reviews and routing guide audits. Monthly-with-quarterly-trend is the cadence most operators settle on, because a monthly view catches problems fast while the quarterly trend tells you whether a bad month was a blip or a decline.
Is 95% on-time a good score?
Impossible to say without the peer figure. A carrier scoring 92% looks fine in isolation, but if the rest of your base averages 96% then 92% is a problem — and a report showing only their own number will never reveal it. Aggregation hides things too: a national 96% rate can conceal a 78% rate on one lane in one quarter. Benchmark against peers on comparable lane types, and break the number down by lane, facility and period.
What makes carriers actually change behaviour?
Consequences applied consistently. If a carrier scoring 75% receives the same volume next quarter as one scoring 95%, the scorecard is theatre and carriers work that out quickly. The mechanisms available are tier-based volume allocation with first-tender priority on premium lanes, performance-based rate adjustments at renewal, probationary status with a defined improvement window, and removal from the routing guide for sustained underperformance. Which one you choose matters far less than applying it every cycle.
How do we handle low-volume lanes?
Carefully, because small samples produce noisy signals that damage credibility. Use rolling ninety-day windows where volume is volatile and discrete monthly windows for steady-state operations, and set a minimum shipment count below which you report but do not score. Annotate single catastrophic events and, where appropriate, exclude them from trend calculations after a root cause analysis — with the exclusion visible on the scorecard rather than applied silently.
Where should we start?
With the three or four metrics you can already produce from data you control — gate events, tender records, invoices — and nothing that requires a carrier to send you a figure. Get two clean cycles out, then add metrics rather than starting broad and narrowing under pressure. Surface the score at booking time so it enters routing decisions rather than living in a quarterly deck.
Book a working session with your carrier list and we'll show which metrics are already available in your own data inside Fleet Rabbit before you build anything.
Trusted Data, Real Consequences
Four to six metrics defined precisely, scored from data you control rather than figures you were sent, benchmarked against peers on comparable lanes, reviewed on a rhythm carriers can anticipate, and tied to outcomes applied every single cycle.
Bring your current scorecard to the call · Works alongside existing TMS and EDI feeds · Free tier available
August 21, 2026By Emily Parker
All Blogs