Lead Scoring That Sales Actually Trusts
A lead scoring model that works uses two independent axes — fit and intent — rather than one blended number, fits its weights to closed-won data instead of intuition, and applies decay to time-sensitive signals. A single score cannot tell you whether to nurture or to call, which is why blended models get ignored.
- Two axes, always. A single blended score cannot distinguish right-company-wrong-time from wrong-company-right-now.
- Fit weights come from closed-won analysis. Intent weights come from observed conversion. Neither from a workshop.
- Apply decay to behavioural signals or a 60-day-old page view scores the same as yesterday's.
- Cap the model at 8–10 inputs. A score nobody can explain is a score sales ignores.
- Run it in parallel for a quarter before it routes anything. Trust is earned with evidence.
Why one score fails
The standard model produces a single number, often out of 100. A lead scores 72. What should happen?
You cannot say, because 72 could be a perfect-fit enterprise account that has done nothing, or a student who read nine blog posts. Those require opposite actions — patience and nurture in one case, disqualification in the other — and a blended score has averaged away the only distinction that matters.
| Low intent | High intent | |
|---|---|---|
| High fit | Nurture. Right company, wrong time. Do not route to sales. | Route now. This is the only true MQL. |
| Low fit | Suppress. Stop spending here. | Investigate — usually support, a student, or a competitor. |
Two axes, four quadrants, four different actions. This single change resolves most lead quality disputes on the spot, because the argument about whether a lead was good becomes a question about which quadrant it was in.
Building the fit score
Fit is objective, slow-moving, and derivable from enrichment. It should be computed from attributes, never from behaviour.
- 01Start from accounts that renewed and expanded
Not all closed-won — closed-won includes customers who churned, and those should score lower, not higher. Renewal is the only honest evidence of fit.
- 02Compare attribute distributions against churned accounts
You are looking for attributes where retained and churned accounts separate clearly. Attributes common to both are not predictive, however intuitive they feel.
- 03Weight by observed separation, not by intuition
If company size separates strongly and industry does not, weight size heavily and industry lightly — even if everyone in the room believes industry matters. The data is describing your business, not the market.
- 04Include negative weights
Attributes that predict churn should subtract. Most models only add, which means a bad-fit account with lots of activity can outscore a good-fit one.
Building the intent score
Intent is behavioural, fast-moving, and must decay. The most common failure is treating a 60-day-old pricing page visit as equivalent to yesterday's.
| Signal | Relative weight | Decay window |
|---|---|---|
| Pricing page visit | Very high | 7 days |
| Demo or contact request | Very high | 14 days |
| Comparison page view | High | 7 days |
| Multiple people from one domain | High | 14 days |
| Documentation read | High | 14 days |
| Repeat sessions in a week | Medium | 10 days |
| Content download | Low | 30 days |
| Email open | Very low | Do not score |
The last row matters. Email opens have been unreliable since privacy-protection features began pre-fetching images, and scoring them adds noise that looks like signal. Score clicks if you must score email at all.
Apply linear decay to the end of each window rather than a binary flag. A signal at day 6 of a 7-day window should carry almost no weight — the buying signals taxonomy covers the decay behaviour of each family in more depth.
Keeping it explainable
Cap the model at eight to ten inputs per axis. Beyond that nobody can explain why an account scored what it did, and a score sales cannot interrogate is a score sales will route around.
Practically, this means every routed lead should arrive with its top three contributing factors visible. Not the number — the reasons. A rep who can see enterprise size, hit pricing twice this week, three people from the domain will act on it. A rep who sees 84 will not.
Proving it before it routes anything
- 01Score in the background for a full quarter
The model computes and records but changes no routing. Nobody is affected, and you accumulate evidence.
- 02Compare scores against actual outcomes
Of the leads that became opportunities, what did the model score them? Of those it scored highly, how many converted? This is the only honest test.
- 03Show sales the disagreements
The leads the model scored high that sales rejected, and vice versa. Discussing specific records is what converts scepticism into input — and their objections usually improve the model.
- 04Roll out on one segment first
Route by score for one team or region. Compare against the rest for a cycle before extending.
Skipping this is why so many scoring models are ignored. A model imposed without evidence is an opinion with arithmetic attached, and sales teams correctly treat it as one.
Maintaining it
- Refit fit weights every two quarters against recent renewals and churn. Your ICP drifts, usually upward, and the model should follow.
- Review intent weights quarterly against observed conversion. Signal value changes as your site and market change.
- Track score-to-opportunity conversion by decile. If the top decile is not converting materially better than the fifth, the model is not working.
- Watch for score inflation. Adding signals without removing any pushes everything upward until the threshold is meaningless.
The third item is the health check worth running monthly. It takes minutes, it is unambiguous, and it is the number to show anyone who asks whether the scoring model earns its complexity. What happens after scoring is covered in lead routing.
Want this diagnosed on your own numbers?
The RADAR™ Scan scores your revenue engine in 2 minutes — 12 questions, a 0–100 score, and your gate verdict. No email required.
Run your RADAR™ Scan→Questions this raises.
How do you build a lead scoring model?
Why does single-score lead scoring fail?
How do you set lead scoring weights?
How many inputs should a lead scoring model have?
How do you get sales to trust a lead scoring model?
Related guides.
Assignment models compared, the fallback rule nobody builds, and why alerting the manager rather than the rep is what makes an SLA real.
GTM EngineeringFour signal families ranked by predictive strength, each with a decay window — plus how to weight them without building a scoring model nobody trusts.
GTM EngineeringEvery lead type in common use, what each actually means, and the two-axis model that replaces the whole confusing taxonomy.
Lead GenerationFirst we build your pipeline. Then we build the machine that scales it.
Every engagement starts with the RADAR™ Reveal — a 2-week audit with a scored report, gate verdict, and roadmap. Yours to keep, whatever you do next.