Skip to content
Open Pulse

Blog · Explainer

Retrieved, not recalled: why your competitor list should come from sources

Ask an assistant who your competitors are and it will answer with total confidence, because a fabricated company name reads exactly like a real one. What the hallucination research actually measures, why a phantom rival is the most expensive kind of wrong, and how Open Pulse records where every name on a watchlist came from instead of pretending the problem does not exist.

By Viraj Bandara

· 9 min read · For founders and product marketers who were about to paste an ai-generated competitor list into a strategy doc

Ask any AI assistant "who are my competitors?" and it will answer with total confidence. Clean formatting, plausible names, maybe a little table with columns. It looks like research.

The question nobody asks is the one that matters: did this answer come from the world, or from the shape of words?

There are two ways to produce a list of companies. One is to go and look. The other is to generate a sequence of tokens that resembles a list of companies. These are not two versions of the same thing. One of them can include a company that does not exist, and you will never catch it by reading the list, because a fabricated company name reads exactly like a real one.

Recall is a guess wearing a list format

A language model does not store facts the way a database does. It learns patterns from text and predicts the next likely token. When you ask for competitors, it produces the names that are most plausible given the shape of your question. Plausibility is not truth.

This has been measured carefully in the adjacent case of software packages. In "We Have a Package for You!", presented at USENIX Security 2025, researchers generated 576,000 code samples across 16 large language models and checked every package the models recommended. They found 205,474 unique non-existent package names.

The headline figure usually quoted is 19.7%, nearly one in five. That number is worth unpacking, because the paper splits it sharply: commercial models hallucinated at least 5.2% of packages on average, while open-source models averaged 21.7%. If you are asking a frontier assistant, the rate is a great deal better than one in five. It is not zero, and that is the part that matters, because the errors are not random noise. When the same prompt was repeated ten times, the paper reports that a large share of invented names came back on every single run. Repeatable errors look deliberate, and deliberate looks trustworthy.

The same convergence shows up with brand names. Palo Alto Networks' Unit 42 team probed 913 global brands across 685,339 adversarial prompts and found models consistently inventing the same domains for real companies. They identified 13,229 already-registered malicious domains that had started life as model hallucinations, plus roughly 250,000 hallucinated domains still unregistered and waiting to be claimed.

The mechanism is identical. The model fills a gap with the most plausible thing it can generate, and if the gap has the shape of a company, it generates a company. A competitor list has exactly that shape: a row, a name, a one-line description. Nothing in that shape resists fabrication.

The courtroom receipts

If you think professionals would catch this, the courts have news.

In Mata v. Avianca (2023), attorneys filed a federal brief citing six judicial decisions that did not exist. ChatGPT had invented them, complete with realistic citations, and when the lawyers asked the model to confirm the cases were real, it confirmed they were. Judge Kevin Castel sanctioned the lawyers $5,000 each.

That was the first case, not the last. In February 2025, lawyers from Morgan & Morgan were sanctioned in Wyoming federal court after a filing cited eight nonexistent cases. The firm's own response called it "a cautionary tale for our firm and all firms, as we enter this new age of artificial intelligence".

The underlying rate is not anecdotal. Stanford researchers Dahl, Magesh, Suzgun and Ho found that when major models were asked specific, verifiable questions about federal court cases, legal hallucinations appeared between 58% and 88% of the time.

The pattern in every incident is the same: fluent, specific, and wrong. And the person who pasted it was a professional whose job is verification. If a fluent fake can survive a law firm, it will survive a founder pasting a competitor list into a strategy doc on a Friday afternoon.

A wrong name on a watchlist costs more than a wrong fact

A wrong fact costs you a correction. A wrong competitor costs you a quarter.

Track a rival that does not exist and you spend research hours on a phantom. Track one that was acquired two years ago and your pricing intel is a fossil. Track one the model merged from two real companies and your positioning fights a hybrid nobody competes with. Each corrupts every decision the list feeds: who you benchmark against, whose launches you watch, whose customers you try to win.

Here is the specific nastiness of a phantom rival, and it is why this problem needs a design rather than a disclaimer. A company that does not exist returns zero complaints for ever, and in a roll-up that looks exactly like a real rival nobody happens to be complaining about. The failure is not loud. It is a quiet empty row that you read as good news.

Stale is the same failure, slower. A model's memory is frozen at its training cutoff; your market is not. A recalled list is a snapshot of a market as it was when the weights were set, served as if it were today. You cannot interrogate a snapshot, and you cannot ask it when it was taken.

What Open Pulse actually does, which is not "we solved it"

The honest version of our design is not that hallucinated names never reach your watchlist. It is that every name carries where it came from, and the ones that came from a model are marked as such.

Each competitor on the watchlist has a provenance, and there are exactly three values:

ProvenanceWhat it means
userSomebody in your workspace typed it. Authoritative.
retrievedFound in published comparisons of your market. Corroborated by a source.
modelProposed from the model's own knowledge and found nowhere else.

A model competitor is not silently dropped and it is not silently kept. It is added with a dashed warning border and it says, in as many words, "Proposed by the model and not found in any published comparison: treat with care." It is a visible badge rather than only a tooltip, because a tooltip says nothing at all on a touch screen.

That is the whole design, and the reasoning is the same as the Unknown row in Territory: a system that aggregates uncertain labels into a decision surface can either propagate the uncertainty or launder it into certainty. Dropping the model's suggestions would throw away real leads. Keeping them unmarked would hand you a phantom that reads as a quiet rival. Marking them is the only option that leaves you able to tell the difference.

The two halves of a watchlist

Beyond provenance, the competitor tab joins two things that answer different questions about the same companies, on one key.

Intel covers what the rival did, swept and filed under six kinds: news, funding, product, people, pricing, and entrant, that last one being a company we were not watching that turned up in your space.

Chatter covers what their customers say, fed by competitor-dissatisfaction listeners reading Reddit, Hacker News, YouTube, G2 and Trustpilot. The chatter is rolled up one row per competitor, with the previous window and a per-day series beside it, because a rival with 40 complaints is a fact and a rival whose complaint rate has doubled since last month is the one to aim at this quarter.

Watchlist size is a plan limit: none on Go, five on Growth, twenty-five on Pro, unlimited on Enterprise.

How to audit any AI-generated market map

You do not have to take anyone's design claims on trust, ours included. If a tool hands you a competitor list, run these four checks before you act on it.

  1. Ask for a source per name. Not a footnote at the bottom, a source next to each row. A retrieved list can name the article, the review site or the dataset that surfaced each company. A recalled list cannot, because the source is the model's weights, and "the model remembers" is not evidence.
  2. Spot-check one verifiable fact per row. Founding year, headquarters, headcount band. These exist on the public web. If three rows are wrong on checkable facts, the uncheckable ones are suspect too.
  3. Look for the honest gaps. A good list can say "not corroborated" and still show you the row. That gap is information. A list that never admits one is not a list without gaps; it is a list that cannot distinguish a gap from a guess.
  4. Ask when each entry was last verified. A list without dates has no memory of its own freshness. The answer you want is a schedule, not a promise.

A tool that passes all four is a research tool. A tool that fails the first is a story generator. Both have uses. Only one should feed your strategy.

References

See it on your own market

Paste your website, review the plan it proposes, and read what comes back tomorrow morning.

Questions about anything here? Email support@openpulse.cloud.

Questions

Frequently asked questions

Can AI really invent a company that does not exist?

Yes, and it does so in a way that is hard to catch by reading. The measured case is software packages: one study generated 576,000 code samples and found 205,474 unique package names that do not exist. A company name is an easier thing to fabricate than a package name, because nothing about it can be checked by trying to install it.

Why not just drop any competitor the model suggested?

Because some of them are real and useful, and dropping them silently is its own kind of lie. A model suggestion is a lead worth checking, not a fact worth trusting. The useful middle is to keep the row and mark its origin, so you can see at a glance which names are corroborated and which are the model's idea.

How do I tell a hallucinated rival from a quiet one?

That is exactly the problem provenance solves, and without it you largely cannot. Both produce an empty row. A competitor that no source corroborates and that returns nothing week after week is far more likely to be a phantom than a real company having a quiet quarter, but only if you can see which of the two you are looking at.

Does a competitor list need re-checking, or is once enough?

Re-checking, on a schedule. Companies pivot, rebrand, shut down and get acquired, so a list verified once is accurate exactly once. A list that cannot tell you when each entry was last confirmed is giving you a snapshot without a date on it.

What should I ask a vendor about their competitor data?

Four things: where each name came from, whether you can see that per row, what happens to a name nothing corroborates, and how often the list is re-verified. A vendor that can answer all four is doing retrieval. A vendor that treats the question as a technicality is probably doing recall and calling it research.