Why ChatGPT Recommends Your Competitor: A Provenance Audit
I traced 14 circulated stats about why AI recommends your competitor back to source. Seven held up. One had no primary at all. The full provenance audit.
I took the pages that rank for this question, pulled fourteen of the statistics they circulate, and tried to walk each one back to a primary source with a disclosed methodology. This is what survived.
If you have typed some version of "why does ChatGPT recommend my competitor instead of me" into a search box, you have already met the answer. It arrives fast, it arrives confident, and it arrives with numbers attached. Authoritative list mentions carry 41% of the weight. Awards carry 18%. Reviews carry 16%. Structured data lifts your visibility by up to 40%. Perplexity has 13 citation slots. Google's AI Overviews have about 8.
I wanted to know where those numbers came from. So I took the pages that rank for this question, pulled out fourteen of the statistics they circulate, and tried to walk each one back to a primary source with a disclosed methodology.
Seven of the fourteen survived the walk. Six partially survived, which usually meant a real primary existed but the methodology was hidden, or the number had been mis-attributed or mis-dated somewhere on its way into circulation. And one did not survive at all.
The most-repeated claim in the entire set is not in that last group, and working out where it actually came from took the longest.
An LLM names a competitor because the retrievable, third-party evidence connecting that competitor to the category is stronger and better positioned. Not because there is a knowable brand-ranking formula that someone has cracked.
The signal hierarchy (authoritative list mentions 41%, awards 18%, reviews 16%) comes from Onely's 1 December 2025 article, "How ChatGPT Decides Which Brands to Recommend," written by Bartosz Góralewicz. Onely states it without a citation. The figures are not Onely's own. The identical five-factor breakdown, 41% authoritative list mentions, 18% awards and accreditations, 16% online reviews, 14% customer examples and usage data, 11% social sentiment, sits in First Page Sage's "Generative Engine Optimization (GEO): Explanation & Algorithm Breakdown." That page has been carrying those figures since at least January 2025, which I can date from an Internet Archive capture of 20 January 2025 rather than from the page's own timestamp. That is ten months before Onely published. First Page Sage does disclose a sample, 11,128 commercial queries across four chatbots, but never explains how raw queries became percentage weights. Onely reproduces three of the five factors and names nobody, while citing First Page Sage elsewhere in the same article for its cost benchmarks.
That matters more than it sounds, because those weights are doing enormous work. They tell you where to spend. They imply a formula exists and that someone has measured it. And when you check what the platforms themselves publish, the implication falls apart.
OpenAI's ChatGPT Search help documentation states that ChatGPT may use an approximate location based on your IP address to return relevant local results, with device location off by default, and that it may rewrite prompts using that location. The Memory documentation states that memory personalizes your experience from chats and files, and that a given response can be shaped by custom instructions, past chats, files, and memories.
That is nearly the extent of it. OpenAI also documents, for shopping results only, how product and merchant results are selected: availability, price, quality, and whether the seller is the maker or primary seller. The other platforms publish more than nothing too, but never a formula. Perplexity says it ranks sources by authority and relevance. Microsoft lists Bing's ranking factors in general order of importance. Google publishes an eligibility rule rather than a ranking one. Anthropic publishes nothing on the subject at all. Not one of them publishes a weight, and not one publishes a citation-slot count.
This adjudicates the loudest contradiction inside the corpus. Onely publishes precise weights without naming where they came from. A competing page, seo-agency.com.sg, states flatly that OpenAI has published no brand-ranking formula. Both can be true, and both are. The weights are a vendor's estimate, not a platform disclosure, and the vendor never shows how the estimate was made. What happened to them in transit was not invention. It was the quiet removal of the attribution, until a vendor's own estimate reads like a measured fact about ChatGPT. I had this wrong in an earlier draft of this audit, and the correction is the more interesting finding.
Here is the whole audit. Read it as a source-checking exercise, not a stat dump: the interesting column is the last one.
| # | Claim as circulated | Original publisher | Date | Primary link | Methodology disclosed? | Match? | Verdict |
|---|---|---|---|---|---|---|---|
| 1 | "80% of LLM citations don't rank in Google's top 100" (via launchcodex) | Ahrefs (Louise Linehan, Xibeijia Guan) | 11 Aug 2025 | ahrefs.com/blog/ai-search-overlap/ | Yes: 15,000 long-tail prompts; ChatGPT/Gemini/Copilot/Perplexity | Yes. Ahrefs bucketed cited URLs as top 10, top 100, or neither; the 80% is the "neither" bucket, so "don't rank in the top 100" restates it accurately | TRACEABLE |
| 2 | Semrush: Google top-10 domains are cited across AI, so SEO improves AI visibility | Semrush AI Visibility Index / "Most-Cited Domains in AI" | Sept 2025 (2,500 prompts), then 2026 index (126M prompts); 3-month study 230,000 prompts / 13 weeks | semrush.com/blog/ai-mode-comparison-study/ | Yes | Overstated by the circulators, not contradicted. Semrush's AI Mode study (5,000 keywords, 150K citations, Jul 2025) finds substantial overlap with Google's top 10: Perplexity 91% domain, AI Overviews 86%, AI Mode 54%. Overlap is real but partial and platform-dependent | TRACEABLE |
| 3 | "BrightEdge found 60% overlap between Perplexity citations and Google top-10" (via mersel.ai) | BrightEdge (Jim Yu) | 3 Apr 2024 | globenewswire.com BrightEdge release; Search Engine Land art. 439029 | Partial: "Generative Parser," 9 industries; no sample size or prompt count | Yes (60%; healthcare 82%, restaurants 27%) | TRACEABLE (primary exists; methodology thin; data is early-2024) |
| 4 | "28% of ChatGPT's most-cited pages have zero organic visibility" (via onely.com) | Onely attributes to Seer Interactive; figure actually traces to Ahrefs (28.3%) | 2025 | ahrefs.com/blog/ai-seo-statistics/ | Partial | Mis-attributed. Seer's cited study is about recency, not this figure | PARTIALLY TRACEABLE |
| 5 | Onely signal hierarchy: list mentions 41%, awards 18%, reviews 16% | First Page Sage (GEO Algorithm Breakdown); relayed uncited by Onely (Bartosz Góralewicz) | First Page Sage carrying these weights by Jan 2025 (Wayback capture 20 Jan 2025); Onely 1 Dec 2025 | firstpagesage.com/seo-blog/generative-engine-optimization-geo-explanation/ | Partial: 11,128 commercial queries across 4 chatbots disclosed; the weighting method is not | Origin exists and verifiably predates Onely by 10 months. Onely reproduces three of five factors and names no one | PARTIALLY TRACEABLE |
| 6 | SparkToro: under a 1-in-100 chance of the same recommendation list twice | SparkToro (Rand Fishkin) + Gumshoe.ai (Patrick O'Donnell) | 27 Jan 2026 | sparktoro.com/blog/new-research-ais-are-highly-inconsistent... | Yes: 600 volunteers, 12 prompts, ChatGPT/Claude/Google AI, 2,961 runs, November to December 2025 | Yes (same list under 1/100; same order roughly 1/1,000) | TRACEABLE |
| 7 | Profound: brand visibility fell 31%, over 85% of brands down, 6 or 7 mentions down to 3 or 4 | Profound | approx. Oct 2025 | tryprofound.com/blog/chatgpt-entity-update | Partial: "millions of prompts," brands scored 5% to 95%, before and after Oct 18 | Yes | TRACEABLE |
| 8 | "44.2% of LLM citations from first 30% of text" (attributed to Position Digital AND BrightEdge) | Kevin Indig (Gauge dataset) | Feb 2026 | searchengineland.com/chatgpt-citations-content-study-469483 | Yes: 1.2M ChatGPT responses, 30M citations, 18,012 verified | Number right; both circulated attributions FALSE | TRACEABLE (to Indig) |
| 9 | "Structured data gives up to 40% higher visibility" (onely, then gofishdigital) | No primary; "up to 40%" traces to Princeton GEO paper (different interventions) | 2024 | arxiv.org/abs/2311.09735 | GEO paper: yes. Schema link: none | No. The GEO paper's 40% is about statistics, quotes, and citations, not schema | UNTRACEABLE (as stated) |
| 10 | mersel.ai: Bain 85%, BrightEdge 4.4x, 61% CTR drop, 73%/34% | Bain & Co (Day One list); Semrush; Seer Interactive | 2024 to 2026 | forrester/bain/seer primaries | Mixed | Name-dropped, unlinked, and two of the four mis-attributed. The 61% CTR drop is Seer Interactive (Nov 2025); the 4.4x conversion figure is Semrush (July 2025), not BrightEdge | PARTIALLY TRACEABLE |
| 11 | FirstPageSage: GEO CAC $559 vs SEO $612; 89 vs 127 days | FirstPageSage | 2026 | firstpagesage.com/seo-blog/generative-engine-optimization-customer-acquisition-cost-cac-benchmarks/ | Partial: n=127 B2B companies, Jan 2024 to Jul 2025, minimum 8 per sector, and a CAC attribution definition; but unaudited proprietary client data, with no per-industry n and no dispersion | $559 and 89/127 days confirmed; the "14.4% premium over SEO" claim cannot be reconciled with the same page's $559 against $612 table, and the page does not say which is authoritative | PARTIALLY TRACEABLE |
| 12 | Idea Grove client 47.5% to 52.9% overall, 44.6% to 57.9% on ChatGPT in 2 months; 1,000-person survey | Idea Grove / TrustSignals (Scott Baradell) | 16 Aug 2026 | trustsignals.com/blog/why-does-chatgpt-recommend-your-competitors-instead-of-you | Client program: no. Survey: partial (n=1,000 US consumers, 98% verify, 78% reviews) | Published by primary; author explicitly cautions against generalizing | PARTIALLY TRACEABLE |
| 13 | SE Ranking (Nov 2025) plus Forrester (2025) | SE Ranking; Forrester Buyers' Journey Survey | SE Ranking's AI Mode study, 15 Dec 2025 (2,328,533 pages, 295,485 domains); Forrester 2024 (89%) and 2025 (94%) | seranking.com/blog/how-to-optimize-for-ai-mode/; forrester.com | Yes | Forrester "89% in 2025" is mis-dated. 89% is the 2024 wave; 2025 is 94%. And the 2.3M-page methodology belongs to SE Ranking's AI Mode study, not to its AI-statistics listicle, which is a secondary aggregation | PARTIALLY TRACEABLE |
| 14 | Fixed citation slots: ChatGPT 3 to 4, Perplexity 13, Google AIO about 8 | Search Engine Land (James Allen, powered by Rankscale.ai), via uof.digital, relayed by Onely | 12 May 2025 | searchengineland.com/how-to-get-cited-by-ai-seo-insights-from-8000-ai-citations-455284 | Yes: ~8,000 citations, 57 queries, four engines, data spreadsheet published | Traceable but distorted twice. The primary reports averages, not slots. And the "about 8" is Gemini's average, which Onely moved onto Google AI Overviews; the primary puts Google AI Overviews at 3 to 4 | TRACEABLE (distorted in transit) |
One entry deserves calling out on its own. The claim that 44.2% of citations come from the first 30% of a page is attributed, inside a single launchcodex article, to both Position Digital and BrightEdge. Both attributions are wrong. The number originates with Kevin Indig's analysis of 1.2M ChatGPT responses and 18,012 verified citations, published via Search Engine Land in February 2026, which described the "ski ramp" pattern as statistically indisputable with a P-value of 0.0.
The number is real. The sourcing around it is not. That gap is the whole story of this category.
It helps to sort every claim in this space into three tiers, because the tier determines how much weight it can carry.
Documented by the platforms themselves. The personalization inputs described above: IP-based approximate location, prompt rewriting using that location, and memory shaped by custom instructions, past chats, files, and memories. Plus one narrow disclosure, for shopping results only, where OpenAI names the factors behind product and merchant selection: availability, price, quality, and whether the seller is the maker or primary seller. Qualitative, unweighted, and scoped to commerce. Nothing resembling a general brand-ranking formula for conversational answers.
Inferred from credible third-party measurement. Ahrefs on citation and ranking overlap (15,000 prompts) and its 75,000-brand correlation study. SparkToro on recommendation instability (2,961 runs). Profound on entity-update effects (millions of prompts). Kevin Indig on positional bias and domain concentration (1.2M responses). BrightEdge on Perplexity overlap. Semrush on most-cited domains (230,000 prompts). These are measured and disclosed, though most of the publishers sell a related product.
Real numbers, laundered in transit. Onely's 41/18/16 hierarchy, which is First Page Sage's, with the attribution removed and the weighting method never shown. The fixed-citation-slots framing, which hardens Search Engine Land's measured averages into slots and moves Gemini's number onto Google AI Overviews. And launchcodex's mis-attribution of the 44.2% figure to two publishers who did not produce it. Only one claim in the whole set has no primary at all: the structured-data-gives-40% schema claim.
Worth naming the pattern underneath all of it: publishers overwhelmingly sell the product their finding supports. Onely sells GEO. First Page Sage sells GEO, and the weights everyone quotes come from the company selling the service those weights tell you to buy. Profound sells AI-visibility tracking. Idea Grove and TrustSignals sell digital PR. The two most neutral sources in the set, Ahrefs' controlled schema test and the academic benchmarks, are the two that most undercut the prescriptions everyone is selling.
SparkToro, working with Gumshoe.ai, ran 600 volunteers through 12 prompts across ChatGPT, Claude, and Google AI, for 2,961 total runs between November and December 2025. Their finding, in their words, is that there is under a 1 in 100 chance that ChatGPT or Google's AI, asked the same question 100 times, returns the same list of brands in any two responses. Getting the same list in the same order is closer to 1 in 1,000.
Sit with that before you buy anything. If the output is that unstable, then any vendor promising you a fixed position is selling against the primary evidence. There is no position to hold. There is only a probability of appearing, measured across many runs.
Before you can fix your absence, you have to know which machine produced it. There are three.
Parametric recall from training data. The brand-to-category association was learned during pre-training. It is static until the next model update, and it favours brands with long-standing, widely repeated third-party presence.
Live web retrieval at answer time. ChatGPT triggers a search when prompt probability is low (Ahrefs describes this as a classifier threshold), then uses query fan-out and reciprocal rank fusion. That is why cited pages frequently do not rank for the literal prompt, and why Ahrefs finds only about 12% of AI citations sitting in Google's top 10.
Per-user memory, personalization, and location. Documented by OpenAI. IP-based location and memory or custom instructions change results per user. This is the obvious explanation for "my competitor showed up for her but not for me," and not one of the ten competing pages I read raises it.
You can separate them from the outside without any tooling. Run the prompt logged out in a fresh or incognito session, with web search off, then on. Run it across different IP addresses and locations. Run it in an account with memory disabled. If results change with web search on and off, it is retrieval. If they change by location or account, it is personalization. If they stay stable across all three, it is parametric recall.
Absence is not one problem. It is four, and they take different fixes.
Mode 1: the engine does not know the brand exists. Ask "What is [brand]?" with web search off. If it hallucinates or denies knowledge, you have a training-data and entity gap.
Mode 2: it knows the brand but not the category association. "What does [brand] do?" comes back correct, but "best [category] tools" omits you. The entity exists. The category edge is missing.
Mode 3: associated, but below the mention cut-off. The brand appears occasionally across repeated runs, not consistently. Given SparkToro's instability finding, run each prompt 60 to 100 times, the range SparkToro used and recommends, and compute a visibility percentage. SparkToro is explicit that the number of runs needed for statistical soundness is still an open question. A low but non-zero rate is a cut-off problem, not an absence problem.
Mode 4: named, but described generically. You appear with vague descriptors while competitors get specific ones. That is a positioning and evidence-consistency problem, or as Idea Grove frames it, a question of whether the internet clearly understands what you do.
Each finding below carries an evidence strength, because the difference between strong and weak evidence is the difference between a budget decision and a guess.
Reviews. Weak to moderate. SE Ranking (November 2025) found about 34.5% of AI Overviews cite at least one review platform, with G2 prominent. verified-reviews.com claims products with 500 or more recent reviews outrank stale ones. But no study cleanly separates review-as-training-data from review-retrieved-live, or count from recency from wording. Idea Grove's survey of 1,000 people shows 78% of consumers say reviews increase trust in an AI-recommended brand, which is a demand-side finding, not a mechanism finding.
Memory and personalization. Strong for existence, weak for magnitude. OpenAI documents location and memory personalization. No public controlled test quantifies how much brand lists actually shift per user. That is a genuine gap in the literature and a strong candidate for original research.
Engine disagreement. Strong. Ahrefs shows Perplexity aligning closest to Google at 28.6% top-10 overlap, while ChatGPT and Gemini hover around 8%. Ahrefs puts that down to Perplexity being built to cite, and treats the alignment as surprising rather than explained, since Perplexity runs its own index via perplexitybot instead of drawing on Google or Bing. Profound, running 100,000 prompts across both engines, found only 11.0% of cited domains appear in both ChatGPT and Perplexity, with 37.4% ChatGPT-only and 51.6% Perplexity-only. Architecturally distinct retrieval produces distinct winners, which is exactly why a brand can win on Perplexity and lose on ChatGPT.
Can a smaller player win? Moderate. Yes, on narrow prompts. Kevin Indig found top-cited pages often carry fewer backlinks than the pages they beat, and the GEO paper shows domain-specific wins. But Indig also found the countervailing concentration: roughly 30 domains capture 67% of citations within a topic, and the top 10 capture 46% in product-comparison topics. Niche specificity can win a narrow prompt. Broad category prompts belong to a small domain oligopoly.
Timelines. Weak. The corpus spans "2 weeks" to "12 months" with no shared definition of what counts as a result. FirstPageSage claims 89 days from its own unaudited client database, without defining what counts as a result, and Onely repeats it. No source cleanly separates time-to-first-citation from time-to-stable-visibility from time-to-pipeline. Treat every timeline promise you are given as unfalsifiable until someone defines the finish line.
Cost. Moderate. Profound Lite is $499 per month, per Profound's own funding announcement. FirstPageSage's GEO CAC of $559 comes from its own unaudited client database, 127 B2B companies between January 2024 and July 2025, and it contradicts the 14.4% premium claimed on the same page. Otterly, Profound, and SE Ranking's SE Visible are the named monitoring tools. Digital PR and review-generation pricing are largely undisclosed across the corpus, which is its own signal.
What does not work. Strong. C-SEO Bench (Puerto et al., NeurIPS 2025, arXiv:2506.11097) tested conversational-SEO tactics and found most of them "largely ineffective," and frequently negative in their effect on document ranking, the opposite of what practitioners expect, while traditional retrieval-ranking strategies proved significantly more effective. The retrieval-ranking baseline measured roughly 7.6 times more effective than the best C-SEO method in the retail domain. The "Statistics" tactic decreased rankings in 19 of 24 settings. Only 3 of 54 method-by-domain cases showed significant positive effects. And on publishing volume, the practitioner quoted inside Onely's own article reports that daily publishing lifted AI Overview and Copilot visibility for two to three weeks before dropping sharply, which is consistent with high-cadence content decaying and directly against the high-cadence prescription.
Schema. Strong against. One controlled test and one observational study found no independent lift. Ahrefs tracked 1,885 pages that added JSON-LD, using a matched difference-in-differences design against 4,000 controls: AI Overviews down 4.6%, AI Mode up 2.4%, ChatGPT up 2.2%, with the latter two statistically indistinguishable from zero, concluding that schema had no clear positive or negative effect. Fischman's SSRN study (1,006 pages, 730 citations) found schema presence null (OR 0.678, p = .296) with rank position dominant (OR 0.762 per position, p < .001). The one exception was attribute-rich Product and Review schema, cited at 61.7% versus 41.6% for generic schema types (p = .012), an advantage most pronounced on low-authority domains (DR 60 or below), where the gap was 54.2% versus 31.8%. Schema is the most over-prescribed and least-evidenced tactic in the set.
Backlinks and domain authority. Moderate. Onely claims near-zero influence. Semrush prescribes backlink-gap analysis. Ahrefs' 75,000-brand study is the arbiter: web mentions correlate at 0.664, far above backlinks at 0.218, and the top three correlates are all off-site (web mentions 0.664, brand anchors 0.527, brand search volume 0.392). A December 2025 follow-up put YouTube mentions strongest at roughly 0.737. So earned brand mentions matter considerably more than backlinks for AI recommendation. Onely is closer to right than Semrush here, but "near-zero" overstates it.
Publishing what failed is the point of an audit, and so is publishing where I was wrong. Two claims have no primary at all — one inside the fourteen (the schema-gives-40% claim), and one I met along the way but did not number, the secondhand Search Atlas December 2024 schema study. Four more that I had filed as unsourced turned out to have one I had simply not found, and those corrections are below too.
- Onely's claim that Wikipedia accounts for 47.9% of ChatGPT's top-10 sources. A primary does exist: Profound's "AI Platform Citation Patterns," 5 June 2025, built on 680 million citations. The defect is scope, not absence. It is Wikipedia's share within ChatGPT's ten most-cited sources, which is not overall citation volume, and Onely passes it along through blissdrive without naming Profound.
- "Structured data gives up to 40% higher visibility" as a schema claim. No primary. The 40% belongs to the Princeton GEO paper's non-schema interventions.
- "Perplexity returns 13 citation slots" and "Google AIO about 8" as fixed numbers. A measured primary does exist one hop past uof.digital: Search Engine Land with Rankscale.ai, 12 May 2025, roughly 8,000 citations across 57 queries and four engines. It reports averages, not slots. And the "about 8" is Gemini's average, which Onely moved onto Google AI Overviews; the primary puts Google AI Overviews at 3 to 4.
- FirstPageSage's 14.4% GEO-cost-premium claim, which cannot be reconciled with the $559 against $612 figures in its own table on the same page, and the page does not say which is authoritative.
- The Search/Atlas December 2024 schema study referenced secondhand. No primary located.
Diagnose before you fix. Run the four-mode diagnostic: 60 to 100 runs per prompt, which is SparkToro's own range, and establish a visibility-percentage baseline. Add the logged out versus logged in and across-locations controls too, which are my recommendation and not SparkToro's, since that study deliberately imposed no such controls. The thresholds that change the plan are simple. Zero percent (0%) visibility points to Mode 1 or Mode 2, an entity or category problem that needs third-party evidence. Intermittent visibility points to Mode 3, a consistency problem. Present but vague points to Mode 4, a positioning problem.
Spend on earned third-party evidence, not schema or content tricks. The strongest evidence (the Ahrefs schema difference-in-differences test, C-SEO Bench, and Fischman's observational study) says schema produces no independent AI-citation lift and that conversational-SEO tactics run from ineffective to harmful. Redirect the budget toward being clearly and consistently described in the sources engines actually retrieve: review platforms, industry best-of lists, credible press, and the top Ahrefs correlate, web and brand mentions. The one narrow exception is attribute-rich Product and Review schema (pricing, ratings, specs) if you are a lower-authority domain, where it showed a measured edge.
Treat visibility as probabilistic, not positional. Given the SparkToro instability finding, never promise or track a fixed rank. Track share of voice across many runs, and react to multi-week trends rather than single results.
Publish provenance-first. Lead with the audit, not the conclusion. When a finding happens to support the seller who published it, say so in line: Onely to GEO, FirstPageSage to GEO, Profound to tracking, Idea Grove to PR. That disclosure costs you nothing and it is the only thing that separates a claim from a sales asset.
Frequently asked questions
- Why does ChatGPT recommend my competitor instead of me?
- Because the retrievable third-party evidence linking that competitor to your category is stronger and better positioned than yours. Not because of a published ranking formula. Platforms publish some criteria, Perplexity on authority and relevance, Bing on its ranking factors, OpenAI on shopping results, but not one of them publishes a weight or a citation-slot count.
- Is there a known weighting for lists, awards, and reviews?
- Not a published one, and not a platform's. The 41% / 18% / 16% split is First Page Sage's own estimate, drawn from 11,128 commercial queries across four chatbots, and it never explains how those queries became percentages. It has been on that page since at least January 2025 and reaches most readers through an uncited December 2025 article that does not name First Page Sage. No platform has confirmed any of it.
- Does schema markup improve AI citations?
- A controlled test and an observational study both say no independent lift. The only measured exception was attribute-rich Product and Review schema on low-authority domains.
- How many times should I test a prompt before believing the result?
- 60 to 100 runs. That is the range SparkToro ran and the minimum it recommends, precisely because the same question asked 100 times has under a 1 in 100 chance of returning the same brand list twice. SparkToro adds that the run count needed for statistical soundness is still an open question.
- Can a small brand outrank a large one in AI answers?
- On narrow prompts, yes. On broad category prompts, roughly 30 domains capture 67% of citations within a topic, so the odds tighten considerably.
- How do I tell whether it is retrieval, training data, or personalization?
- Run the prompt logged out with web search off, then on, then from different locations and an account with memory disabled. If results move when web search toggles, it is retrieval. If they move by location or account, it is personalization. If they hold steady across all three, it is parametric recall from training data.
- If not schema, where should the budget go?
- Toward being clearly and consistently described in the sources engines actually retrieve: review platforms, industry best-of lists, credible press, and web and brand mentions, which is the top correlate in Ahrefs' 75,000-brand study at 0.664 against 0.218 for backlinks.
Four honest limits, because an audit that hides its own weak points is just a longer advertisement.
Many of the primaries above are vendors publishing findings that support the product they sell. I have flagged that conflict wherever it appears. Where you can, prefer the two neutral controlled tests: the Ahrefs schema difference-in-differences study, and the academic benchmarks C-SEO Bench and GEO.
Dates matter more than usual here. BrightEdge's 60% Perplexity-overlap figure is from early 2024, and the platform has changed materially since.
Figures in this field decay quickly. Ahrefs' own AI Overview and top-10 overlap fell from 76% in mid-2025 to 38% in 2026, inside a single year. Treat any single number as a snapshot, not a constant, including the ones in this post.
And the Fischman schema study is a non-peer-reviewed preprint whose author discloses AI assistance in drafting. I still rate the "schema does not independently help" conclusion as strong, but only because Ahrefs' independent controlled test corroborates it. One preprint alone would not carry that weight.
Primary sources are marked (P), secondary (S).
- Ahrefs, "Only 12% of AI Cited URLs Rank in Google's Top 10," 11 Aug 2025, ahrefs.com/blog/ai-search-overlap/ (P)
- Ahrefs, "90+ AI SEO Statistics," 2025, ahrefs.com/blog/ai-seo-statistics/ (P)
- Ahrefs, "An Analysis of AI Overview Brand Visibility Factors (75K Brands)," Aug 2025, ahrefs.com/blog/ai-overview-brand-correlation/ (P)
- Ahrefs, "We Tracked 1,885 Pages Adding Schema. AI Citations Barely Moved," 11 May 2026, ahrefs.com/blog/schema-ai-citations/ (P)
- SparkToro, "AIs are highly inconsistent when recommending brands," 27 Jan 2026, sparktoro.com (P)
- Profound, "ChatGPT's entity update," Oct 2025, tryprofound.com/blog/chatgpt-entity-update (P)
- Kevin Indig, via Search Engine Land, Feb 2026, searchengineland.com/chatgpt-citations-content-study-469483; domain-concentration follow-up art. 472349 (P/S)
- Princeton GEO paper, Aggarwal et al., KDD 2024, arxiv.org/abs/2311.09735 (P)
- C-SEO Bench, Puerto et al., NeurIPS 2025, arxiv.org/abs/2506.11097 (P)
- Fischman, "Does Schema Markup Predict AI Citation?," SSRN, Apr 2026, papers.ssrn.com/sol3/papers.cfm?abstract_id=6284518 (P, non-peer-reviewed preprint)
- BrightEdge Perplexity research, 3 Apr 2024, globenewswire.com; Search Engine Land art. 439029 (P/S)
- Semrush, "Most-Cited Domains in AI: A 3-Month Study," 2025, semrush.com/blog/most-cited-domains-ai/ (P)
- Semrush, "How Google's AI Mode Compares to Traditional Search and Other LLMs," 21 Jul 2025, 5,000 keywords / 150,000+ citations, semrush.com/blog/ai-mode-comparison-study/ (P)
- Semrush, "We Studied the Impact of AI Search on SEO Traffic," 21 Jul 2025, semrush.com/blog/ai-search-seo-traffic-study/ (P)
- Profound, "Answer Engine Citation Overlap Strategy," 100,000 prompts across ChatGPT and Perplexity, tryprofound.com (P)
- Profound, "AI Platform Citation Patterns," 5 Jun 2025, 680 million citations, tryprofound.com (P)
- Search Engine Land, James Allen, "How to get cited by AI: SEO insights from 8,000 AI citations," 12 May 2025, powered by Rankscale.ai, searchengineland.com/how-to-get-cited-by-ai-seo-insights-from-8000-ai-citations-455284 (P)
- Seer Interactive, "AIO Impact on Google CTR," Nov 2025, seerinteractive.com (P)
- Onely, "How ChatGPT Decides Which Brands to Recommend," 1 Dec 2025, onely.com (S; the weights it publishes are First Page Sage's, uncredited)
- First Page Sage, "Generative Engine Optimization (GEO): Explanation & Algorithm Breakdown," continuously updated; carrying the 41/18/16 weights by at least Jan 2025 per Internet Archive capture 20 Jan 2025, firstpagesage.com/seo-blog/generative-engine-optimization-geo-explanation/ (P, vendor; the origin of the 41/18/16 weights)
- FirstPageSage GEO CAC benchmarks, 2026, firstpagesage.com/seo-blog/generative-engine-optimization-customer-acquisition-cost-cac-benchmarks/ (P, vendor)
- Idea Grove / TrustSignals (Scott Baradell), 16 Aug 2026, trustsignals.com/blog/why-does-chatgpt-recommend-your-competitors-instead-of-you (P, vendor)
- OpenAI Help Center, "Searching the web with ChatGPT," "Memory FAQ," and the shopping-results documentation, help.openai.com (P)
- Perplexity Help Center, on how sources and product listings are ranked, perplexity.ai (P)
- Microsoft, "How Bing delivers search results," support.microsoft.com (P)
- Google Search Central, "AI features and your website," developers.google.com/search (P)
- launchcodex, "Why your competitors are already showing up in ChatGPT," 2026, launchcodex.com (S)
- SE Ranking, "How to Optimize for AI Mode," 15 Dec 2025, 2,328,533 pages / 295,485 domains, seranking.com/blog/how-to-optimize-for-ai-mode/ (P)
- SE Ranking, "70+ AI Search Stats," Dec 2025, seranking.com/blog/ai-statistics/ (S, aggregation)
- Forrester, "B2B Buyer Adoption of Generative AI" (2024) and "B2B Buyers Make Zero-Click Buying Number One" (2025), forrester.com (P)
- Academic corroboration on bias and instability: Hutter et al., "Lost but not only in the middle: Positional bias in RAG," ECIR 2025; "Source Coverage and Citation Bias in LLM-based vs. Traditional Search Engines," arXiv:2512.09483 (P)
Related reading
- Cited but Not Chosen: How AI Visibility Wins No ClientsI earned ~60 AI citations and 22% share of authority on a new site — and zero clients. Why being cited isn't being chosen, and what a services firm should do.
- GEO Monitoring Tools Compared (2026): Profound, Peec & MoreA vendor-neutral 2026 comparison of GEO monitoring tools — Profound, Otterly, Peec and more — with pricing, a scorecard, and a build-vs-buy verdict.
- AEO/GEO Playbook 2026: Your Next B2B Buyer Is a MachineHalf of B2B buyers now start in an AI chatbot, but AI sends only 1% of your traffic. The evidence-led AEO/GEO playbook for getting cited, not clicked.