Skip to content
Open Pulse

Blog · Explainer

Your thumbs move the bar

The thumbs under every signal are not a like button. Each rating is inlined into that listener's next classification as a worked example, so the definition of relevant becomes yours rather than one a model guessed at on day one. What relevance feedback has done for search since 1971, what it does for spam today, and what it can and cannot fix here.

By Viraj Bandara

· 10 min read · For anyone tuning a listener through its first weeks

Under every signal in Open Pulse there is a pair of buttons that most people read as sentiment. Thumbs up, thumbs down. It looks like a like button. It is not. Every rating is fed back into that listener's next classification as a worked example, so the definition of "relevant" quietly becomes yours rather than one the model guessed at on day one. The bar does not move because you asked it to. It moves because you taught it.

This is not a new idea. It is one of the oldest ideas in search.

The oldest trick in information retrieval

In 1971, Rocchio described a method called relevance feedback, associated with the SMART information retrieval system built between 1960 and 1964. The setup is simple: a user runs a query, marks the returned documents relevant or not relevant, and the system revises the query to move closer to the relevant documents and farther from the irrelevant ones. In the textbook formulation, the revised query is the old query plus the average of the relevant documents minus the average of the non-relevant ones.

The textbook result is equally simple: relevance feedback improves recall, and often precision too. The underlying assumption, as the standard summary of the algorithm puts it, is that most users have a general conception of which documents should be denoted relevant or irrelevant. That is the whole trick. The machine cannot guess what matters to you, but it can do excellent arithmetic on what you tell it.

There is a footnote worth carrying forward. In lecture notes on relevance feedback, Cornell's Lillian Lee points out that the negative weight in the formula is sometimes set to zero, because feedback about irrelevant documents may not always be reliable. Positive feedback, in practice, is usually more valuable than negative feedback. Keep that in mind. It comes back.

Gmail runs this at the largest scale in the world

You have probably used relevance feedback this week, maybe this morning. Every time you report an email as spam in Gmail, Google's own help page says plainly: "As you report more spam, Gmail identifies similar emails as spam more efficiently." Reporting a message also sends a copy to Google, which "may analyze it to help protect users from spam and abuse."

This is not a side channel. It is the training data. In February 2019 Google announced that TensorFlow-based protections were blocking an additional 100 million spam messages a day in Gmail, on top of what it said it already blocked: 99.9 percent of spam, phishing and malware. Both numbers are Google's own, reported via ZDNet, so read them the way a vendor would like them read. The direction of travel is the point, not the decimal.

The spam case also shows why per-user calibration matters more than one global answer. One person's newsletter is another person's junk, and no vendor can ship a single definition of spam that survives contact with a billion inboxes. The filter has to learn your version. Your reports are how it does that.

What a thumbs up becomes inside Open Pulse

Now the product. When you rate a signal in Open Pulse, up or down, that rating is not stored as a mood. It is loaded at the start of the listener's next run and inlined into the classification prompt under a line that tells the model this is the user's own judgement on earlier signals and to calibrate to it. Each example carries the verdict, the title and the first 160 characters of the excerpt. The twelve most recent ratings ride along, newest first, so a listener you rated six months ago is not still being taught by a judgment you have since changed your mind about.

It costs tokens on each call and no extra calls, which is the whole reason the mechanism can be this direct. The next run does not only consult the search plan, the commercial profile and the score weights; it also consults your record of what was actually worth your time.

This is the part worth saying carefully, because it is where the honest version differs from the marketing version. Your ratings do not retrain a global model. They calibrate one listener: the standing question you wrote, with the keywords, exclusions and queries you edited. The definition of "relevant" becomes yours, for that question, in that workspace. The mechanism is narrow by design. A rating on a competitor-dissatisfaction listener should not move the bar on your find-customers listener, because those two listeners disagree about what deserves your attention.

What this solves is the cold start. On day one, the classifier is working from your website crawl, your commercial profile and the search plan it proposed. That is a good starting position: fit is 20 of the signal's 100 points, and without a profile that fit score is guessed from your product description alone. But it is still a guess made by a machine that has never seen you reject anything. The first ten ratings are the cheapest improvement available to you, because they are the only information the classifier cannot derive from your website.

Thumbs down does the harder work

Remember the footnote about negative feedback being unreliable? In the classical formulation, a "not relevant" judgment is ambiguous: the document might be wrong, or the user might just not have needed it that day. Later versions of the method often down-weight or ignore it entirely.

We kept it, because the ambiguous half of your judgment is where the bar actually lives. A thumbs up says "more like this", which mostly confirms the classifier's existing direction. A thumbs down draws a line: this looked like a buying signal and it was not, and here is where the near miss was.

The rejected candidates the product keeps, with their reasons, are the raw material of that line. Every dropped candidate records why, from a fixed list of eleven: wrong page shape, excluded by one of your own keywords, off topic, a blocked domain, too old, undated, a duplicate, ranked out by the reranker, a review above your star ceiling, below the score bar, or noise. The cheap gate drops thousands a run and those are kept for 14 days; the couple of dozen a run that actually reached the classifier are kept for a year, because those are the set a model made a real decision about. The excluded list sits in the interface next to what was kept, so the boundary is something you can inspect rather than something you are asked to take on trust.

That said, the classical caution applies to you too. A thumbs down that means "right company, wrong week" is a different statement from one that means "never this topic". The ratings work best when they are decisive: up for what you would act on, down for what you would not, and no rating at all for the shrug. An ambiguous rating is an ambiguous worked example, and the next classification inherits the ambiguity.

What ratings cannot fix

Honesty requires the limits next to the capabilities.

Ratings calibrate; they do not rewrite. They work inside the listener's intent: its objective, its search plan, its score bar. If the search plan is returning nothing but noise, ten thumbs down will not conjure a better query. That is what the quality report is for. It shows per-query and per-source yield: which searches earn their cost and which have never produced a kept signal. A query that produces candidates but no signals for three runs in a row is disabled automatically, and the disablement is reported in the run log. If a listener feels quiet, the rejections tell you whether the searches returned nothing or the filters threw everything away. Ratings are the second lever, not the first.

Ratings are also local, and they are not the score. They move one listener's worked examples, not your workspace's, and they do not touch the seven weights or the bar. Those are separate controls for a reason: the bar is a policy decision about how selective a listener should be, and it is set per listener as a sensitivity, Strict at 70, Balanced at 55 or Broad at 40, with an explicit number still settable through the API and winning over the preset so a hand-tuned bar is never overwritten by a preset nobody chose. The seven weights themselves are the same for everyone today. Alternative weightings, fit-led for a narrow market and urgency-led for work won by responding first, are designed and sitting in the code, but nothing reads them yet. Your thumbs say what you liked; the bar says how selective the listener is. They answer different questions, and the score bar post is the one about the second.

And the cold start has one shortcut no rating can replace: the commercial profile. Fit is a fifth of every signal's score. If you skipped onboarding, fill it in before you rate anything, because a classifier guessing at fit will spend your first ten ratings confirming what a profile would have told it.

Making the first week count

If you want the ratings to do their work quickly, the routine is short:

  • Rate early. The first runs are when the classifier knows least about you, so each rating carries the most information. Rate before you start editing the search plan, so you can see what the plan was doing unprompted.
  • Rate both sides. The ups are confirmations; the downs draw the boundary. A rating history of nothing but thumbs up teaches the classifier that everything is relevant, which is the same as teaching it nothing.
  • Fix the profile first. Fit is a fifth of the score. Rate after, not before.
  • Watch the quality report. If your ratings and the report disagree, if you keep thumbing down signals that cleared the bar with room to spare, that is information about the bar, not the signals. Raise it, and let the ratings work within the new shape.
  • Remember it is per listener. Five listeners means five definitions of relevant. That is the point of the design, and it is why rating on one listener and expecting another to improve is expecting the wrong thing.

The buttons are small. The arithmetic behind them is fifty years old, battle tested on spam inboxes and search engines, and it works because it asks you for the one judgment a machine cannot make for you: this, not that. Press them like they matter. They are the training set.

If you want worked examples to start rating today, run a free report: a capped, real scan of the public conversations about the problem you solve, on a live page that updates as it runs. Take the first signal that is worth your time and the first that is not. That is your training set's first page.

References

  • "Rocchio algorithm", Wikipedia. Relevance feedback as described by Rocchio (1971), associated with the SMART information retrieval system (1960-1964). Source for the query revision formula and the recall result.
  • Lillian Lee, "CS630 lecture notes", Cornell University, 28 February 2006. Relevance feedback in the vector space model. Source for the note that the negative weight is sometimes set to zero because feedback about irrelevant documents may not always be reliable.
  • ZDNet, "Google eliminates more spam from Gmail with TensorFlow", February 2019. Reporting Google's own announcement: an additional 100 million spam messages blocked a day, against a claimed 99.9% already blocked. Vendor-authored numbers; treat as directional.
  • Google, "Report spam in Gmail", Gmail Help. Source for both quotations, including that reported messages may be analysed to protect users from spam and abuse.

See it on your own market

Paste your website, review the plan it proposes, and read what comes back tomorrow morning.

Questions about anything here? Email support@openpulse.cloud.

Questions

Frequently asked questions

What does relevance feedback actually mean?

It is a technique from information retrieval, first described by Rocchio in 1971: a user marks results relevant or not relevant, and the system revises its notion of the query to sit closer to the relevant results and farther from the irrelevant ones. The textbook result is better recall and often better precision. The modern consumer example is a spam button.

Does rating a signal retrain the model for everyone?

No. A rating calibrates one listener in one workspace. It is stored against the signal and read back at the start of that listener's next run, then inlined into the classification prompt as a worked example. Nothing crosses into another listener, another workspace or a shared model.

How many of my ratings does the classifier see?

The twelve most recent, newest first, each carrying the verdict, the title and the first 160 characters of the excerpt. The cap is there because the examples ride inside every classification call for that run: they cost tokens on each call rather than an extra call, and an unbounded history would be paid for on every candidate forever.

Is a thumbs down worth as much as a thumbs up?

In the classical literature, no. Negative judgments are ambiguous, so the negative term is often down-weighted or dropped entirely. We keep it, because a rejection is the only thing that tells the classifier where the near miss was. The caveat travels with it: rate down what you would never want, not what merely arrived in a bad week, and leave the shrug unrated.

My listener is too noisy. Should I rate more, or change the bar?

Check the quality report first. If most searches are producing candidates that never become signals, the search plan is the problem and ratings will not fix it. If the signals are real but too marginal to act on, that is the bar, which is a per-listener setting, Strict at 70, Balanced at 55 or Broad at 40, not something ratings adjust.