Every lead scoring guide ends with the same validation step: run the model over historical leads and check that the ones who converted scored higher. That test cannot fail, because the old score decided who a rep called, and a lead nobody called cannot convert. You are measuring your own routing. The only honest check is a holdback: route a random slice of leads against what the score recommends and compare. Everything else in a scoring build is calibration; this is the part that tells you whether the model knows anything.

Lead scoring is one of the few growth systems that gets built, shipped, and then validated with a test that cannot fail. The scoring mechanics are well covered everywhere; what follows is about the validation step, why the standard one is circular, and how to route leads once you know the score is real, whether you run it through automated journeys or by hand.
What lead scoring is meant to do
It converts a queue of inbound leads into handling decisions: who gets called today, who gets nurtured, who gets pointed at self-serve, who gets nothing.
The mechanics are not in dispute. Combine fit attributes with behavioural signals, weight them, decay the behavioural half over time, the way any buying signal loses value with age so a whitepaper download from last year stops counting, and set thresholds. Most guides cover this well and the better ones separate fit from intent rather than summing them.

The problem is what happens after the model ships.
The validation step everyone recommends is circular
The standard advice is to run the new model over historical leads and confirm that leads who converted score higher than leads who did not. That comparison is contaminated before you start.
Your previous scoring, formal or informal, decided which leads a rep called, called first, or never called. A lead that nobody contacted had a structurally lower chance of converting no matter how good a prospect it was. So the conversions in your history are partly a record of your own routing.
- The outcome is downstream of the treatment. Score drove routing, routing drove contact, contact drove conversion.
- Bad models pass anyway. Any model correlated with the old one inherits its apparent accuracy, including the old one's mistakes.
- The errors are invisible. A high-quality lead the old model buried never converted, so it looks like a correct low score rather than a miss.
This is ordinary selection bias, and it is worth being precise about it: the test is not weak, it is uninformative. It cannot distinguish a good model from a model that agrees with your past behaviour.
What an honest check looks like
Hold back a random slice of inbound and route it against what the score recommends, then compare outcomes, run as a proper experiment between that slice and the scored population.
A lead the model scored low still gets worked in the holdback. A lead it scored high sometimes waits. That feels wasteful and it is the only way to learn whether the low scores were right, because it is the only condition where a low-scored lead gets a fair chance to convert.

Read two things. Whether high-scored leads outperform the holdback's high-scored leads, which tells you the routing helps. And whether low-scored leads in the holdback convert at a rate you can live with, which tells you what the model is throwing away.
Keep the slice small. Five to ten percent of inbound is usually enough at normal volumes, and it shrinks once the model has been through a few cycles.
Score to a destination, not a rank
A number between one and 100 is not a decision. Somebody still has to choose where to cut it, and that choice carries more weight than any of the weights inside the model.
Name the routes first, then set thresholds that fill them. Four or five is usually enough, and each needs an owner and a response time, which is the same routing discipline behind signal-based selling or it is not a route.
| Fit | Intent | Route | Why not just a score |
|---|---|---|---|
| High | High | Call today, named owner | A rank puts these behind a bigger number from a worse lead |
| High | Low | Nurture on the fit case | Averaging drops these into a middle band and they get ignored |
| Low | High | Self-serve, or disqualify fast | Summing hides that intent is real and fit is not |
| Low | Low | No treatment, revisit on new signal | A rank still leaves them in a queue somebody feels obliged to work |
The two middle rows are why fit and intent should stay separate. Collapsed into one number they produce the same score and need opposite handling.

Why the standard thresholds are borrowed
Fifty points for a marketing qualified lead and 75 to 100 for a sales qualified one appear in guide after guide. They are conventions passed between articles, not values derived from anyone's data.
A threshold is only meaningful when it maps to a route your team actually staffs and a holdback has shown that leads above it behave differently from leads below it. Until then it is a number that makes a dashboard look calibrated.
Keeping the score from going stale
A score computed at form fill describes the lead at the moment they were least informed about you, and stops being true almost immediately.
Intent moves. Fit rarely does. So the behavioural half has to recompute against live activity while the fit half can sit still, and a lead that crosses a route boundary needs to actually move route rather than wait for the next list export. The docs cover behavioral segmentation for the recompute.

A holdback only works if routing and outcome are recorded against one profile, which is the reason the agentic GTM platform keeps scoring and routing in the same place.
The lead qualification recipe stands up fit and intent scoring with the routes attached, so the holdback has something to be compared against on day one.
Where to start
Name your routes before you weight a single attribute. Then stand up a holdback on day one rather than adding it later, because a model that has run unchecked for two quarters has already contaminated the data you would use to check it.

Keep the historical comparison if you like; it is a reasonable sanity check that nothing is inverted. Just stop treating it as evidence. Good lead scoring models are the ones whose owners can say what the model gets wrong, and that answer only ever comes from leads the score told you to ignore. Start with Intempt if you want the routes and the holdback to run off the same live behaviour.
Frequently asked questions. Answered.
Lead scoring ranks inbound leads by how likely they are to become customers, usually by combining fit attributes such as company size and role with behavioural signals such as pricing page visits. The score then decides how the lead gets handled.






