Competitors Appear in AI Answers and You Don’t. Most Gap Audits Will Not Tell You Why
A list is not a diagnosis
Every AI visibility tool now ships a citation gap report. Semrush, Peec, Profound, Scrunch and others will hand you a ranked list of domains that cite your competitor and not you, usually with an impact score attached.
The problem is that two companies with identical gap lists can need completely different work, and the list cannot tell them apart. One is absent because a crawler directive blocks retrieval. One is absent because the source is a competitor’s own customer case study, which no amount of content will ever change. The report looks identical in both cases.
A useful audit does three things a tool report does not. It disambiguates what kind of absence you have. It separates gaps you can close from gaps you cannot. And it treats the result as a hypothesis rather than a cause.
First, disambiguate the absence
“They appear and we don’t” describes at least five different failures, each with a different fix.
Not mentioned at all. Your brand does not appear in the answer text. This is a visibility failure and it is the only one most teams measure.
Mentioned but not cited. The model names you and links to someone else’s page about you. Your competitor’s comparison post is doing your positioning. This is an attribution failure, and the fix is a page worth citing, not more mentions.
Cited on one engine, absent on another. Only about 11% of cited domains appear across both ChatGPT and Perplexity for the same query, and roughly 71% of cited sources show up on a single platform. Engine-specific absence is normal, not a crisis, and diagnosing it requires per-engine data.
Present but positioned badly. You appear as the cheap option, the legacy option, or the one with the caveat. Share of voice looks fine. The answer is still costing you the deal.
Present in branded prompts, absent in category prompts. The model knows who you are and does not consider you a member of the category. This is the most expensive failure and the slowest to fix.
Report these separately or the audit averages them into something meaningless. Note too that share of voice, mention rate and citation rate use different denominators, so a brand can move on one and not the others.
Then classify by eligibility, not by impact score
This is the step that saves a quarter of wasted work.

Figure 1. Two stage classification. The quadrant split is the easy part; the eligibility split on the gap quadrant is what stops a team spending a quarter on sources that were never open to them.
Source: Storylake framework, created internally
Sort every source into four quadrants: cites both of you, cites them only, cites you only, cites neither. Then take the “cites them only” quadrant, which is your actual gap, and split it again on a question no tool asks.
Could this source cite you, in principle, today?
Winnable gaps are sources where you are eligible and simply absent. A category review site you have no profile on. A comparison listicle that omits you. A forum thread where nobody mentioned you. This is your work queue.
Structural gaps are sources where you are not eligible and cannot become eligible through content. The competitor’s own domain. Their customer’s case study. A partner directory that requires a certification you do not hold. A publisher operating under a licensing deal. These will sit at the top of an impact-scored gap list forever, because they are high value and permanently closed.
The eligibility test takes about thirty seconds per source and a human can do it. It is the cheapest filter in the whole method and almost nobody applies it.
The “cites you only” quadrant is worth as much as the gap. It tells you what already works, and it is the thing to defend when a competitor runs this audit on you.
Rigour the tool reports skip
Never aggregate engines. Given how little overlap exists between platforms, a combined gap list averages incompatible retrieval systems.
Work at URL level. On ChatGPT, 99% of Reddit citations point to individual threads rather than subreddit pages or brand profiles. “Get on Reddit” is not an action. “This thread” is.
Run each prompt multiple times. Outputs are probabilistic, so a single run tells you what happened once. Five runs per prompt per engine is a reasonable floor.
Expect brutal concentration, and take it as good news. One agency reported that for a B2B client tracking more than 300 prompts across thousands of responses, two Reddit threads carried the vast majority of citations. If your gap list has four hundred rows, most of them do not matter.
Diagnosing cause, and the denominator trap
Once you have a winnable gap, work through causes cheapest first.
Crawler access. The cheapest check and the most commonly missed. Confirm retrieval agents can actually fetch your pages, and distinguish training crawlers such as GPTBot from live retrieval agents such as OAI-SearchBot and ChatGPT-User, because blanket rules catch both. Test from outside using the agent’s user agent string and confirm you get a full page back.
But here is where the audit will lie to you if you are careless.

Figure 2. The same question, two competent studies, opposite answers. The difference is the denominator: raw counts reward size, normalised rates reveal the effect.
Source: Left: BuzzStream, 4 million citations across 3,600 prompts, March 2026, via PPC Land. Right: Cloro, 1,058 domains, July 2026.
In March 2026, BuzzStream analysed 4 million citations across 3,600 prompts and concluded that blocking crawlers rarely stops AI systems citing you. Their evidence was concrete: cnbc.com blocks three OpenAI agents simultaneously and still appeared 1,298 times in the dataset. Yahoo blocks Google-Extended and appeared close to 30,000 times.
In July 2026, Cloro cross-referenced crawler rules against citations for 1,058 prominent domains and reached the opposite conclusion. The median GPTBot-blocking domain earned 0.003 ChatGPT citations per Google organic appearance, against 0.417 for domains that allow it. Sites blocking PerplexityBot showed a median Perplexity propensity of zero, against 1.167 for the rest. Blocking a provider’s own crawler depressed citations in that provider’s engine specifically.
Both studies are competently executed. They disagree because BuzzStream counted raw citations and Cloro normalised by how visible each domain already was. Enormous domains get cited despite blocking, because they are enormous. Normalise for prominence and the effect reappears.
Your gap audit has exactly the same failure mode. If you count raw citations, you will conclude that big competitors win because they are big, which is true and useless. Normalise by something (prompts entered, pages eligible, existing organic presence) or you are measuring size.
Lexical match. Cornell researchers found that retrieval systems often use lexical similarity to the query as a proxy for accuracy, so text that reads like the question gets treated as though it answers it. Compare the cited competitor page against your equivalent. Does theirs mirror the phrasing buyers actually use, while yours uses internal product language?
Verification signals. Research on AI-cited health sources found that publishers without inherent institutional authority compensate through visible review statements, structured data, content depth and recency. Check whether the cited page carries signals yours lacks.
Commercial arrangement. Some sources appear because of a licensing agreement rather than merit. That is a structural gap wearing a winnable disguise.
What the audit cannot tell you
A gap is a correlation. Observing that a competitor is cited by a review site you are absent from does not establish that getting listed would produce citations. Retrieval timing varies, and competitors ship many things at once.
So treat every diagnosed cause as a hypothesis with a test attached: change one thing, hold the prompt set constant, re-baseline after a fixed interval, and accept that some tests come back null. Re-run the baseline rather than comparing against an old one, because retrieval shifts underneath you. Semrush tracking recorded ChatGPT’s Reddit citations falling from roughly 60% of prompt responses to around 10% in mid September 2025 before recovering. A delta measured across that window would have been pure noise.
The audit spec
Reproducible version, one category, one competitor set.
- Prompt set. 40 to 60 prompts across three tiers: category (no brands named), comparison (“X versus Y”, “alternatives to X”), and branded. Tier them explicitly, because absence in category prompts and absence in branded prompts are different diseases.
- Engines. Three minimum, run and reported separately. Never merged.
- Runs. 5 per prompt per engine. Fresh sessions, memory and personalisation off, locale fixed and logged, timestamped to the hour, model version recorded where exposed.
- Capture. Full response text verbatim, every cited URL, and brand mentions with position in the answer. URL level, not domain level.
- Classify. Every source into one of four quadrants. Split the gap quadrant into winnable and structural using the eligibility test.
- Label the absence type. One of the five failure modes per prompt, not a single visibility score.
- Diagnose. For the top winnable gaps only, work the cause ladder: crawler access, lexical match, verification signals, commercial arrangement.
- Normalise before comparing. Choose a denominator and state it on every chart.
- Test. One change, one hypothesis, fixed interval, re-baselined.
The output is not a visibility score. It is a short list of specific URLs, each with a named absence type, an eligibility verdict, and a cause hypothesis you can falsify. That is considerably less impressive on a slide and considerably more useful on a Monday.











