Skip to main content
Intempt

Lead scoring models cannot be validated the way everyone validates them

Sid Chaudhary
Sid Chaudhary·5 min read

Published: March 9, 2026

TL;DR

Every lead scoring guide ends with the same validation step: run the model over historical leads and check that the ones who converted scored higher. That test cannot fail, because the old score decided who a rep called, and a lead nobody called cannot convert. You are measuring your own routing. The only honest check is a holdback: route a random slice of leads against what the score recommends and compare. Everything else in a scoring build is calibration; this is the part that tells you whether the model knows anything.

YouTube video player

Lead scoring is one of the few growth systems that gets built, shipped, and then validated with a test that cannot fail. The scoring mechanics are well covered everywhere; what follows is about the validation step, why the standard one is circular, and how to route leads once you know the score is real, whether you run it through automated journeys or by hand.

What lead scoring is meant to do

It converts a queue of inbound leads into handling decisions: who gets called today, who gets nurtured, who gets pointed at self-serve, who gets nothing.

The mechanics are not in dispute. Combine fit attributes with behavioural signals, weight them, decay the behavioural half over time, the way any buying signal loses value with age so a whitepaper download from last year stops counting, and set thresholds. Most guides cover this well and the better ones separate fit from intent rather than summing them.

Fit attributes assembled from firmographic data for a lead score

The problem is what happens after the model ships.

The validation step everyone recommends is circular

The standard advice is to run the new model over historical leads and confirm that leads who converted score higher than leads who did not. That comparison is contaminated before you start.

Your previous scoring, formal or informal, decided which leads a rep called, called first, or never called. A lead that nobody contacted had a structurally lower chance of converting no matter how good a prospect it was. So the conversions in your history are partly a record of your own routing.

  • The outcome is downstream of the treatment. Score drove routing, routing drove contact, contact drove conversion.
  • Bad models pass anyway. Any model correlated with the old one inherits its apparent accuracy, including the old one's mistakes.
  • The errors are invisible. A high-quality lead the old model buried never converted, so it looks like a correct low score rather than a miss.

This is ordinary selection bias, and it is worth being precise about it: the test is not weak, it is uninformative. It cannot distinguish a good model from a model that agrees with your past behaviour.

What an honest check looks like

Hold back a random slice of inbound and route it against what the score recommends, then compare outcomes, run as a proper experiment between that slice and the scored population.

A lead the model scored low still gets worked in the holdback. A lead it scored high sometimes waits. That feels wasteful and it is the only way to learn whether the low scores were right, because it is the only condition where a low-scored lead gets a fair chance to convert.

Behavioural activity signals tracked alongside fit for scoring

Read two things. Whether high-scored leads outperform the holdback's high-scored leads, which tells you the routing helps. And whether low-scored leads in the holdback convert at a rate you can live with, which tells you what the model is throwing away.

Keep the slice small. Five to ten percent of inbound is usually enough at normal volumes, and it shrinks once the model has been through a few cycles.

Score to a destination, not a rank

A number between one and 100 is not a decision. Somebody still has to choose where to cut it, and that choice carries more weight than any of the weights inside the model.

Name the routes first, then set thresholds that fill them. Four or five is usually enough, and each needs an owner and a response time, which is the same routing discipline behind signal-based selling or it is not a route.

FitIntentRouteWhy not just a score
HighHighCall today, named ownerA rank puts these behind a bigger number from a worse lead
HighLowNurture on the fit caseAveraging drops these into a middle band and they get ignored
LowHighSelf-serve, or disqualify fastSumming hides that intent is real and fit is not
LowLowNo treatment, revisit on new signalA rank still leaves them in a queue somebody feels obliged to work

The two middle rows are why fit and intent should stay separate. Collapsed into one number they produce the same score and need opposite handling.

Fit and intent normalised as two separate scores rather than summed

Why the standard thresholds are borrowed

Fifty points for a marketing qualified lead and 75 to 100 for a sales qualified one appear in guide after guide. They are conventions passed between articles, not values derived from anyone's data.

A threshold is only meaningful when it maps to a route your team actually staffs and a holdback has shown that leads above it behave differently from leads below it. Until then it is a number that makes a dashboard look calibrated.

Keeping the score from going stale

A score computed at form fill describes the lead at the moment they were least informed about you, and stops being true almost immediately.

Intent moves. Fit rarely does. So the behavioural half has to recompute against live activity while the fit half can sit still, and a lead that crosses a route boundary needs to actually move route rather than wait for the next list export. The docs cover behavioral segmentation for the recompute.

Leads segmented by route as their scores recompute

A holdback only works if routing and outcome are recorded against one profile, which is the reason the agentic GTM platform keeps scoring and routing in the same place.

The lead qualification recipe stands up fit and intent scoring with the routes attached, so the holdback has something to be compared against on day one.

Where to start

Name your routes before you weight a single attribute. Then stand up a holdback on day one rather than adding it later, because a model that has run unchecked for two quarters has already contaminated the data you would use to check it.

Route-level reporting comparing scored leads against a holdback

Keep the historical comparison if you like; it is a reasonable sanity check that nothing is inverted. Just stop treating it as evidence. Good lead scoring models are the ones whose owners can say what the model gets wrong, and that answer only ever comes from leads the score told you to ignore. Start with Intempt if you want the routes and the holdback to run off the same live behaviour.

Frequently asked questions. Answered.

Lead scoring ranks inbound leads by how likely they are to become customers, usually by combining fit attributes such as company size and role with behavioural signals such as pricing page visits. The score then decides how the lead gets handled.

Your GTM. Hired.

You set the strategy. Agents run the plays. Seven AI agents across design, marketing, sales, and analytics. One customer context, tracked from first pixel to final dollar.

Start for free

More to read

How to Use AI for Sales Prospecting Without Tool Sprawl (2026)

How to Use AI for Sales Prospecting Without Tool Sprawl (2026)

Find, qualify, and reach prospects without switching between six tools. Lead scoring, personalized email, and follow-up in one flow.

12 best Claude Code skills for SDRs (list building, cold email, reply handling)

12 best Claude Code skills for SDRs (list building, cold email, reply handling)

12 free Claude Code skills for the SDR desk: list building, list cleaning, cold openers, sequence repair, reply classification, and the follow-up nobody does.

9 Best Sales Engagement Tools in 2026 (and the One Company That Owns Two of Them)

9 Best Sales Engagement Tools in 2026 (and the One Company That Owns Two of Them)

Nine tools, eight companies: Clari owns both Groove and Salesloft, and Salesloft's own site never says so. Only three of the nine publish a rate card. Outreach moved to outreach.ai and now prices tiers in AI credits. Verified against every vendor's own pages in July 2026.

Will AI Replace Sales Jobs? What's Actually Changing for SDRs in 2026

Will AI Replace Sales Jobs? What's Actually Changing for SDRs in 2026

The gtm-skills SDR pack just grew from 7 to 12 skills, closing the one gap left in the whole catalog: reply handling. Here's the honest answer on whether AI replaces the SDR job, and which 12 skills actually run the desk.

AI Cold Outreach: Website Data vs. Scraped Intent (What Actually Works)

AI Cold Outreach: Website Data vs. Scraped Intent (What Actually Works)

98% of marketers call intent data fundamental to demand gen, but only 24% report exceptional ROI. Here's the honest gap and what actually moves the reply rate.

Will AI Replace Sales Jobs? The Honest Answer

Will AI Replace Sales Jobs? The Honest Answer

22% of teams fully replaced SDRs with AI, but most headcount change comes from attrition and capacity-scaling, not layoffs. Here's the real data, which roles are actually exposed, and how to read a vendor's replacement claims.

Lead scoring validation is circular, and a fix | Intempt