Open the local prospects screen and you see rows of Google reviews, each with its stars. The instinct is to sort by the stars: one-star businesses first, because one star means unhappy, and unhappy means they need what you sell.
That instinct is worth questioning, because the star is not a measure of need. It is a measure of a visit, compressed into a single number, by a person who was thinking about something specific when they tapped it.
Consider the one-star review that is entirely about parking. The lot was full, the customer circled twice, they walked in annoyed and tapped one star. The business is a dental clinic. The parking has nothing to do with whether that clinic is a good prospect for the software you sell. If your pipeline treats one star as a hot lead, your morning starts with a car park.
Stars summarise an experience. A score judges the text
A star rating is the customer's own summary of their visit. It is one number for everything: the product, the service, the parking, the weather, the mood they arrived in. It is fast to give and fast to read, which is why it works for choosing a restaurant and why it misleads as a buying-intent reading.
The score is something else. Every signal carries a score out of 100 built from seven components:
| Component | Points |
|---|---|
| Purchase intent | 30 |
| Fit | 20 |
| Urgency | 15 |
| Recency | 10 |
| Commercial value | 10 |
| Evidence quality | 10 |
| Multi-signal | 5 |
The question it answers is not "how was the visit" but "does this text show evidence that someone has a problem I can solve". Vocabulary counts for nothing. A review full of angry words about the wrong thing scores low; a calm, specific description of the exact problem you solve scores high.
Each listener also has a score bar, 55 by default, with Strict at 70 and Broad at 40 as the other presets. Below the bar a candidate is recorded as rejected with its reason rather than stored as a signal. So the one-star parking review is not only judged differently. It is judged, filed under its reason, and kept out of your way. The score bar post covers what happens underneath it.
The research: ratings and words disagree more than people think
This is not a suspicion. Researchers have measured how often the number and the text part ways.
A University of Moratuwa team analysed 16,156 tourism attraction reviews from 2010 to 2023, running a transformer-based sentiment pipeline on the text independently of the assigned stars. Incongruence occurred in 18.6% of reviews: nearly one in five, the text said something the stars did not. The mismatches fell into directional patterns, with what the authors call "Conservative Rater" and "Obligatory 5-Star" behaviours accounting for most of them. Their conclusion is the one that matters for anyone building on reviews: star ratings are not interchangeable with textual sentiment and should be validated before being treated as ground truth.
An independent analysis of more than 212,000 Google Maps restaurant reviews found the same shape a different way. Ratings and textual sentiment were related but not interchangeable, and the biggest divergence sat at three stars, which split nearly evenly between positive and negative text. A three-star review is a coin toss the number cannot resolve. Only the words tell you which half you are looking at. That analysis is an independent repository rather than peer-reviewed work, so treat it as corroborating shape rather than as a measured constant.
Where the two do agree, the agreement is loose. Work published in the Journal of Marketing Analytics found a Spearman correlation of 0.671 between sentiment-derived star categories and actual ratings: strong, significant, and a long way from identity. Roughly, the stars point in the right direction. A score that has to decide whether to interrupt your Monday has to do better than roughly.
The parking problem: one number, many aspects
The parking review is not an edge case. It is the normal shape of the problem: a review has aspects, and the star has no room for them.
That same large-scale restaurant analysis noticed what every support team already knows. Unhappy customers explain exactly what went wrong, naming the failure in detail. Satisfied customers rely on short affective language, "awesome", "loved it", without saying much more. So the text of a bad review carries the aspect information you need, and the text of a good review often carries almost none. The star flattens both into a number. Reading the words recovers what the number threw away.
There is a second reason to distrust the number on its own: some of it is not real. Estimates circulating in 2025 put roughly 30% of online reviews as fake or manipulated, and one vendor reports that 46% of the fake reviews it identified carried five stars. Every one of those figures comes from companies that sell review management, which is to say interested parties measuring a problem they sell the solution to. Treat them as directional, not precise. The underlying point survives the caution: the star distribution on any review platform is shaped by incentives, and text written by a real person about a real problem is harder to fake at scale.
Two numbers, drawn differently
In Open Pulse the local prospects objective reads Google reviews for the businesses you sell to over a 90-day window, longer than the 30 days used for conversations because reviews arrive slowly and stay true for longer. Every signal carries the review's star rating and its themes, plus the 100-point score. The star is the customer's number. The score is the classifier's judgement of the text against your business. Neither is allowed to swallow the other.
Signals are classified by the text, not the stars. The two classes shown by default are pain confirmed and at risk; healthy and noise are stored behind their own tabs rather than thrown away. A one-star review about parking lands where it belongs: judged on its words, themed around what it actually complained about, and filed accordingly. It is not deleted. It is not hidden. It is simply not presented as a lead.
The roll-up applies the same discipline to counting: one row per place rather than one per review. Eleven one-star reviews at the same dental clinic is one row with eleven signals underneath, not eleven rows and a misleading sense of scale. The raw reviews stay one click away, because the day the grouping is wrong is the day you need the originals.
How to read a local-prospect signal
- Read the themes and the text, not the stars. The themes tell you which aspect the review is actually about. If the aspect is parking the score will already be low, but knowing why it is low is what keeps you from second-guessing the filter.
- Check the class. Pain confirmed means the text showed evidence of a problem shaped like the ones your business solves. Healthy and noise are filed, not lost, and the excluded list sits next to what was kept so you can audit the filter's work.
- Look at the score components when a signal surprises you. A signal that reads hot but scores low will usually show you why: weak fit, stale date, thin evidence. The components are the filter showing its working.
- Rate the signals. Thumbs up and down are not cosmetic. Rated signals become worked examples for that listener's next classification, so the definition of relevant becomes yours rather than one the model guessed at on day one.
What the score does not know
Honesty requires the other side of the ledger. The star rating is at least the customer's own number. The score is a model's judgement, built from a text that may be sarcastic, mixed, or partly fictional. Sentiment analysis still fails on sarcasm, on understatement, and on reviews that praise one aspect and bury the complaint in the middle.
The evidence-quality component is an admission of exactly this. It measures how much the source actually gave us to read, because a thin snippet deserves less confidence than a full review, and the classifier is capped on that component when all it received was a fragment.
So the honest claim is not that the score is right where the star is wrong. It is that the two are measuring different things, and collapsing them into one number loses the only one of the two that was ever about your business.
References
- Abaiyan, de Silva et al., "Fault of Our Stars: Behavioral Drivers of Rating-Sentiment Incongruence", arXiv, June 2026. Source for 16,156 reviews and the 18.6% incongruence rate.
- "How written product reviews influence consumer impressions of star ratings", Phys.org, March 2023. Stars as fast gut reaction, text as the slower conclusion.
- "Social value in online reviews: the role of sentiment and social language in consumer engagement", Journal of Marketing Analytics. Source for the 0.671 Spearman correlation.
- restaurant-review-nlp, an independent analysis of 212,000+ Google Maps restaurant reviews. Not peer-reviewed; cited for the shape of the three-star split and the negative-reviews-are-more-specific pattern.
- Fake-review prevalence figures (roughly 30% of reviews, 46% of fakes carrying five stars) come from review-management vendors including Shapo and WiserReview. Vendor-authored and not independently verifiable; cited as directional only.
- Open Pulse, "The score bar: why 55, and what happens below it", on what the bar does to everything underneath it.
See it on your own market
Paste your website, review the plan it proposes, and read what comes back tomorrow morning.
Questions about anything here? Email support@openpulse.cloud.