Every scored feed has a number it does not show you: the cutoff below which a candidate stops being a signal and starts being the morning you did not have to spend. Bombora draws its surge line at 60. TechTarget's classic playbook starts its MQL line at 100, with a demo request worth the full hundred on its own. Open Pulse draws its line at 55.
The number is the least interesting part. What matters is what sits underneath it: what the points measure, what happens to the candidates that miss, and how a team moves the line when their pipeline disagrees with it.
The seven parts of the score
Every signal Open Pulse keeps carries a score out of 100, built from seven components. The weights are fixed:
Purchase intent, 30. Evidence the poster is ready to buy, not vocabulary. A "looking for recommendations" thread that mentions a budget outranks a post that merely uses buying words.
Fit, 20. How well they match who you sell to. This is where the commercial profile you supply does its work: your ideal customer, the geographies you serve, the floor below which a deal is not worth pursuing, and the things you never want surfaced. Fit is 20 of the 100 points, which means a perfect-intent post from the wrong company can score at most 80.
Worth saying plainly, because it is the opposite of what most tools do: when you have not supplied a commercial profile, the classifier is not told that you have not. The block is omitted rather than filled with "not supplied" lines. A prompt full of blanks teaches a model that the whole section is noise, and it then discounts the lines that are filled — so the cost of an empty profile is paid quietly, in a fit score guessed from the product description alone.
Urgency, 15. How soon this matters: an outage being worked around beats a vague future project.
Recency, 10. How fresh the post is. The standard listening window is 30 days (90 for Google Reviews), so recency scores how close to today the post sits inside that window.
Commercial value, 10. What the opportunity is plausibly worth.
Evidence quality, 10. How much the source actually gave the classifier to read. This is the quiet one, and the one most scoring models skip. Reddit and Hacker News hand over the full post, its author and its engagement counts. LinkedIn, X and the job boards hand over a search-engine snippet: enough to spot a topic, not always enough to judge intent. A source that cannot hand over the post body is the single biggest cause of a bad label, so the score discounts what it could not properly read rather than pretending a snippet was a conversation. The discount is a hard ceiling rather than a nudge: a snippet caps evidence quality at 6 out of 10, and the cap lifts only when the page behind the snippet has actually been fetched.
Multi-signal, 5. Whether this company has come up before. The roll-up view already aggregates repeats into one row per rival, and this component nudges companies with a pattern.
One detail that matters for a post about the bar specifically: the bar is applied to six of the seven. Multi-signal is added in a second pass, once the batch is stored and repeats can be counted, and it can only ever raise a score. Anything failing the bar on six components would have failed on seven too, so the alternative — classify everything, store everything, score, then delete what the bar rejects — would mean writing rows in order to remove them.
The ordering is a stance: intent leads at 30, fit is the second voice at 20. Behaviour predicts buying better than demographics do, but a hot signal from the wrong company is not a hot lead. The category's practitioners agree on the shape, if not the exact numbers. House of Martech's 2026 framework runs a 100-point composite of behavioural signals (50), firmographic fit (30) and third-party intent (20). Scalarly's model allocates 25 to demographic fit, 25 to firmographic fit, 40 to behavioural engagement and 10 to negative-signal deductions, with an MQL line at 65. Different allocations, same argument: combine what they are doing with whether they are yours.
Why 55 and not 50, 60 or 70
An honest answer first: 55 is a starting point, not a law of nature. It is the line we ship because it separates the signals a team can act on this week from the mentions that are merely interesting. But every scoring threshold is a bet about where your morning goes, and the bet should be settled against your own outcomes, not ours.
The category is explicit about this calibration. House of Martech (published 2 May 2026) proposes tiers of 70 to 100 for immediate sales routing, 40 to 69 for active nurture, and below 40 for long-cycle nurture or suppression, then gives the test: a well-calibrated model should convert 25 to 45 percent of MQLs into SQLs. Under 25 percent, your line is too loose; over 45 percent, it is too tight and you are leaving pipeline on the table. Scalarly (February 2026, updated May 2026) sets the MQL line at 65, roughly the top 15 to 20 percent of the database, and recommends validating against your last 100 closed-won and 100 closed-lost deals before going live: won deals should score higher than lost ones, or your criteria are wrong before you start.
Two caveats, because this industry manufactures certainty. The 25-to-45 percent benchmark is a consultancy's rule of thumb, not a law: a starting instrument, not a verdict. And Bombora's widely quoted figure that only 15 percent of buyers are in market at any given time is the vendor's own measurement of its own data. It may be right, but it is vendor-authored, and vendor-authored numbers deserve the label.
Open Pulse's version of the same discipline is the quality report: per-query and per-source yield, visible in the interface, with kept, high-intent and rejected counts against each search you are running. If your bar feels wrong, the first question the report answers is whether the searches returned nothing or the filters threw everything away. Those are different problems with different fixes, and a score with no diagnostics cannot tell them apart.
What happens below the bar
Here is the part most scoring products never show you: the reject pile. In Open Pulse, a candidate that clears every filter and still scores under 55 is not silently dropped. It is recorded as a rejection, with its reason.
The cheap gates run first, in this order: blocked domain, a review above the star ceiling you set, wrong page shape (a profile or a docs page rather than a post), too old for your window, excluded by one of your own keywords, off-topic on lexical overlap, then duplicate. What survives that goes to the reranker, and what the reranker puts out of contention is recorded as ranked out. Only then does a model get called, and the two rejections it can produce are noise first and below the score bar second. Everything cheap happens before any model runs, so filler costs nothing, and the recorded reason says exactly which gate fired.
The two kinds of rejection are kept on different clocks, and the difference is deliberate. A cheap-gate rejection is kept for 14 days: there are thousands per run and it is a debugging artefact. A candidate that actually reached the classifier and was scored is kept for a year — there are roughly twenty-five of those per run, and they are the negative class, the set the model decided against rather than the set a regex did. A below-the-bar rejection is always in the second group, which is to say the ones most worth arguing with are the ones kept longest.
This is why a quiet listener is diagnosable rather than mysterious. If the rejections are all "too old" and "duplicate", your searches found nothing new this week. If they are all "below the score bar", the searches are finding candidates and the classifier does not rate them: time to look at the bar, the exclusions, or the commercial profile. A score nobody can audit is a superstition. A rejection with a reason is an instrument reading.
Three deliberate exceptions. Competitor rows and content rows skip the bar, because low-scoring intel is still intel — a rival's launch or a publishable article is worth filing at a low score behind its own tab. Urgency alerts, which interrupt a person rather than waiting for the run summary, sit far above the bar at a default of 85 with immediate or high urgency, capped at five a day: a threshold that fires forty times in a morning gets muted by lunch, and a muted channel is worse than no channel.
And the third is the one that sounds like a bug and is not: a small random share of below-the-bar candidates is kept anyway. Without it, every candidate is shown with probability zero or one, and a probability of zero means no amount of later analysis can say what a different bar would have produced. The rated examples fed back into the prompt are drawn from rows that cleared the bar, so a system with no holdout only ever confirms the bar it already has. Those rows are not marked in the interface, deliberately — a signal visibly flagged as a sample would be read differently, and the labels it exists to produce would be biased again.
Three ways to move the bar
The bar moves, because it is yours to tune:
One: change it directly. Every listener carries a sensitivity setting, on every plan rather than as an upsell: Strict keeps only what it can evidence and puts the bar at 70, Balanced is the default at 55, Broad names weak opportunities rather than discarding them and drops to 40. An exact number can be set through the API and wins over the preset, so a workspace that tuned its bar by hand does not have it overwritten by a preset it never chose. The seven weights themselves are the same for everyone today. Alternative weightings — fit-led for a narrow market, urgency-led for work won by responding first — are designed and sitting in the code, but nothing reads them yet, and a post about honest scoring is the wrong place to describe a setting you cannot currently change.
Two: rate signals. Every signal carries a thumbs up or down, and this is not cosmetic. Rated signals are fed back into that listener's next classification as worked examples, so the definition of relevant becomes yours rather than one a model guessed at on day one. This is the feedback loop the category's best practitioners insist on: Scalarly calls the accept-or-reject feedback from sales "the single most important mechanism for improving your scoring model over time". Here it is built into the product instead of left as process advice.
Three: watch the yield, then prune. A query that produces candidates but no kept signals for three runs in a row is disabled automatically, reported in the run log and named in a notification, so it can be edited rather than quietly abandoned. Scores should decay with disuse: TechTarget's playbook suggests subtracting 10 points after 30 days of no action, 25 after 60, 50 after 90. The principle travels even where the exact numbers do not — stale engagement is not engagement.
Subtract points, or record rejections
The category has two philosophies for candidates that should never reach a rep. Negative scoring docks leads for disqualifiers: competitor domains, geographies you do not serve, personal email addresses, job titles with no purchase influence. House of Martech calls it "the part most teams skip" and notes the gap costs SDRs time they do not have.
Open Pulse takes the second philosophy: reject and record. Candidates land in their own tabs — competitor, content, noise — rather than being penalised numerically, and every rejection keeps the name of the gate that stopped it. A negative score tells you a lead is bad; a rejection reason tells you which gate thought so, which is the information you need to fix the gate. The one place the first philosophy survives is inside fit: a geography you do not serve, or something on your never-surface list, is scored down by the classifier rather than gated out, because those are judgements about a post rather than facts about a URL.
A score that routes nothing is a report
The final test of any scoring model is whether it triggers action. A score that ends in a dashboard is a report with extra steps. Open Pulse's score routes in three places: the pipeline rule, which puts strong signals onto the deal board by score and class with a daily cap; the urgency alerts at 85; and the run summary email, which carries the highest-scoring signals inline rather than a link to go and look.
The design in one sentence: seven measured components, a visible bar at 55, a recorded reason for everything below it, and three ways to move the line when your pipeline disagrees.
If you want to see what your own 55 looks like before committing to anything, the free report runs a capped, real scan against your site and shows you the scores it produces. One per email address, no account needed. For what an individual score is made of once it clears the bar, the anatomy of a buying signal takes one apart; for the reasons candidates get thrown out before they are ever scored, there is a taxonomy of the false positives.
References
- House of Martech, "Lead Qualification Framework 2026: Behavioral + AI Scoring", published 2 May 2026.
- Scalarly, "B2B Lead Scoring Model 2026: Template + MQL/SQL Thresholds", 1 February 2026, updated May 2026.
- TechTarget, "10 lead scoring best practices to improve sales efficiency", date not shown on page.
- Bombora, "Company Surge Intent: how intent scores are calculated", vendor documentation, date not shown on page.
- RollWorks help centre, "Bombora Company Surge Intent", summarising the 0 to 100 surge score and the 60+ surging threshold, date not shown on page.
- lead-spot.net, "B2B Intent Data: What It Is and When It Works", September 2026.
See it on your own market
Paste your website, review the plan it proposes, and read what comes back tomorrow morning.
Questions about anything here? Email support@openpulse.cloud.