In short: a rating is a prediction, and without checking it against what actually happened it's an opinion with decimals. The truth is the sales team's checkboxes (Qualified / Sale / Bad) or the CRM statuses. The tool is the calibration curve: if the sale rate doesn't rise as the rating rises, the rating is wrong. Recalibration is monthly, not a project.
A rating is a prediction. Without checking it against what actually happened, it is an opinion with decimal places. This lesson is the quality control of lead scoring: how you measure whether it is right, how you correct it and who supplies the truth.
What is ground truth in lead scoring?
The truth is what the business knows about every lead weeks later: whether they qualified, whether they bought, whether they were rubbish. Three sources:
- Checkboxes in the spreadsheet (✔️ Qualified · 💰 Sale · ✖ Bad): sales tick each lead. It is the main source and the cheapest to maintain — if the team actually does it. The agreement is compulsory.
- Synchronised CRM statuses (webhook or read): the same thing, with no double work when the CRM already captures it.
- The sale amount: the column that turns "sale" into euros.
With no truth there is no recalibration; the rating stays on its initial version for ever.
What is the calibration curve and how do you read it?
Group the leads with an outcome by rating band (0-20, 21-40, 41-60, 61-80, 81-100) and work out the qualification rate (and sale rate) of each band:
| Band | Leads | Qualified | % |
|---|---|---|---|
| 81-100 | 120 | 78 | 65% |
| 61-80 | 210 | 92 | 44% |
| 41-60 | 260 | 70 | 27% |
| 21-40 | 180 | 22 | 12% |
| 0-20 | 230 | 6 | 3% |
If the curve rises with the band, the rating predicts. If it is flat (every band qualifies the same) or it inverts, the rating is wrong and you need to review the weights or the signals. The report makes this visible every month; it is the first table to look at.
How do you adjust the rating by network and campaign?
Qualified leads and sales by network and by campaign (by ID): does Demand Gen qualify better than Display? Does generic campaign X close like brand does? With that data, layer C's source ranking is reordered: in the first version the report suggests the new order and the client applies it in the configuration; with more data, block D picks it up on its own and, in future versions, adjusting the weights can be automatic with guard rails.
What is lift per signal?
For each layer A/B signal (viewed pricing, confirmed their email, returning visitor, mobile number…): the qualification rate with the signal against without it. A lift of 2× ("the ones who confirmed their email qualify twice as often") justifies its weight; a lift of 1× ("downloading the PDF predicts nothing") reduces it. This is supervised lead scoring: the weights come from the business, not from an opinion. It needs a sample (hundreds of leads with an outcome); before that, the default weights.
Why is the sales team's manual scoring useful?
The checkboxes are not only there to upload conversions: they are the truth that feeds everything above. Which is why the design insists they be easy (centred checkboxes, next to the lead, in the same sheet where sales already work) and that ticking "Bad" be as natural as ticking "Sale": the negatives teach as much as the positives. A natural extension: a Bad/Good/Sale button in the scoring panel itself, stored by lead ID, for teams that do not live in the spreadsheet.
What goes in the monthly recalibration report?
A 📊 tab with: the calibration curve, qualification and sales by network and campaign (with a suggested reordering), lift per signal, duplicate leads and leads with no GCLID (how much traffic cannot be attributed), cost per qualified lead and per sale by campaign (the real cost per customer from the start of the module), and the rating distribution of the month's leads (is the average quality coming in improving?). It is the document that proves whether the system works — and the one that reveals when sales stopped marking.
Which mistakes are made when recalibrating?
- Recalibrating on too little data (dozens of leads): noise.
- Changing weights by hand on a hunch instead of on lift.
- Ignoring the "Bad" ones: without negatives, the curve cannot be drawn.
- Not reviewing the source ranking for a year.
- Treating calibration as a project rather than as a monthly report.
💡 Ninja trick: in Lead Rating the calibration curve, the performance by network/campaign and the lift per signal all live in the report tab, along with the suggested reordering of the ranking; the scoring checkboxes are the truth that feeds it. And there is one extra circuit that closes the fraud module: qualified leads from a Display placement add positive points to its Site Score (the "behavioural phase" of the trust signals). The Shield learns from real customers.
What you should remember
- The truth is the checkboxes (Qualified / Sale / Bad) or the CRM statuses; without it there is no recalibration.
- A calibration curve by band: it must rise; if it is flat, the rating is wrong.
- Performance by network/campaign reorders the source prior; lift per signal adjusts the weights; both need a sample.
- A monthly report, not a project; the negatives teach.