How Do I Get My Business Cited in ChatGPT Answers?
ChatGPT read 548,534 pages to answer 15,000 prompts and cited 15%. A labelled evidence audit of what drives AI citation, what is unsourced, and what it costs.
ChatGPT read 548,534 pages while answering 15,000 prompts. It cited 15% of them.
[PRIMARY-VERIFIED] That figure comes from AirOps' study of 12 March 2026, "The Influence of Retrieval, Fan-out, and Google SERPs on ChatGPT Citations," reported by Search Engine Land on 13 March 2026 under the headline "Only 15% of pages retrieved by ChatGPT appear in final answers." In AirOps' own phrasing, ChatGPT left 85% of retrieved pages uncited.
That 85% is repeated constantly, and almost never with a link back to the study that produced it. Following it to source changed how I understand the entire question. Almost every article about getting cited in ChatGPT is written as though there is one gate. There are two, and they are governed by different things.
Getting read is gate one. Getting quoted is gate two. Five out of six pages that clear gate one die at gate two.
This is a reference document, not a tactics listicle. Compiled 30 August 2026. Every claim carries a label, because in this field the label matters more than the number.
Five labels appear throughout, and they mean exactly what they say.
[PRIMARY-VERIFIED] means I read the original source: a peer-reviewed paper, a platform's own documentation, a live pricing page. [SEMI-PRIMARY] means the publisher generated the data from their own systems and published the analysis themselves, so it is originating but not independent. [SECONDHAND] means a real source exists but I reached it through a report of it, or the methodology is undisclosed. [UNSOURCED-IN-CIRCULATION] means the claim is repeated widely and has no locatable origin at all. [CONTRADICTED] means the claim conflicts with its own claimed source. [NO EVIDENCE FOUND] means I looked and found nothing either way.
Anything published before January 2026 also carries [PRE-2026, potentially stale], because this field moves fast enough that a two-year-old measurement is a historical artefact rather than a fact. Commercial interest is flagged inline, because most of the research here was published by companies selling the solution their research recommends.
[PRIMARY-VERIFIED] The GEO paper (Aggarwal et al., KDD '24) models a generative engine as a pipeline. A query-reformulation model splits the query, then a search engine retrieves a ranked source set S made up of s1 through sm, then a summariser condenses each source, then a response generator produces text whose sentences carry inline citations Ci, where Ci is a subset of S.
Read that notation carefully, because the separation is formal, not rhetorical. A source can belong to the retrieved set S and never appear in any sentence's citation set Ci. The paper measures visibility with Position-Adjusted Word Count, which is words attributed to a source weighted by an exponentially decaying citation-position factor and reported alongside its Word Count and Position sub-metrics, plus Subjective Impression, seven aspects scored by a model in a method akin to G-Eval and normalised to the same mean and variance. Its test engine was two-step: top-5 Google sources fed into GPT-3.5-turbo.
[PRIMARY-VERIFIED] The AirOps study fills in what happens between the two gates in production. 89.6% of prompts generated two or more fan-out sub-queries, expanding 15,000 prompts into 43,233 queries. And of the cited pages that ranked in any top-20 SERP, 32.9% appeared only in the SERP for a fan-out query, not for the starting prompt.
That single statistic dissolves the most common assumption in this category. Ranking for the question the user typed is not the only pathway. Among cited pages that ranked in any top-20 SERP, a third were reached only through a query the user never typed, generated by the model on the way to an answer. Across all citations that works out to about 18%, just under one in five, and it is a floor rather than a ceiling, because 44% of cited pages appeared in no top-20 SERP at all and the study cannot assign those a pathway either way.
[SECONDHAND] What gets quoted is a passage, not a page. Vendor consensus (Averi, Onely, AirOps, 2026) is that engines lift self-contained "answer capsules" of roughly 40 to 60 words. Onely, in "LLM-Friendly Content: 12 Tips to Get Cited," states that listicles account for 50% of top AI citations and that tables increase citation rates 2.5x. Neither figure is Onely's own research. Onely attributes the 50% listicle figure to a third party whose current page no longer contains it, and the 2.5x tables figure carries no source at all. Neither survives one hop back, and Onely sells SEO services.
Here is the honest line between what is documented and what is inferred. The GEO paper documents that adding quotations, statistics and source citations raises the position-adjusted word count of a retrieved source, which means properties of the passage genuinely move the citation outcome. Everything else about "extractability," meaning tables, capsules, and server-rendered HTML, is third-party inference. No platform documents it.
[SECONDHAND] Vendors measure these as technically distinct, and the distinction is the reason two credible studies can report wildly different numbers for the same engine.
A citation is a linked source. A mention is the brand named without a link, and 5W PR reports that ChatGPT mentions brands roughly 3.2x more often than it cites them with links. A recommendation is the model actively endorsing a brand in prose.
No platform documents these definitions. They are industry measurement conventions. When someone tells you ChatGPT cites brands 0.59% of the time and someone else says 68.8%, check which of the three they are counting before you conclude either is wrong.
[PRIMARY-VERIFIED] Every agent below is sourced to the platform's own documentation. The Allow column is my recommendation rather than a documented fact, and two rows carry inline exceptions.
| Agent | Owner | Function | Governs citation? | Allow? |
|---|---|---|---|---|
| OAI-SearchBot | OpenAI | Surfaces sites in ChatGPT search; opt-out means the site "will not be shown in ChatGPT search answers" | YES. This is the citation-governing agent | Allow |
| GPTBot | OpenAI | Training crawl | No (training only) | Optional |
| ChatGPT-User | OpenAI | User-triggered fetches and GPT Actions; "not used to determine whether content may appear in Search" | Indirect | Allow |
| OAI-AdsBot | OpenAI | Validates ad landing pages | No | If advertising |
| PerplexityBot | Perplexity | Search index (citation) | Yes | Allow |
| Perplexity-User | Perplexity | User-triggered fetches; not used for crawling or training; Perplexity's docs say it generally ignores robots.txt rules | Indirect | Allow |
| ClaudeBot | Anthropic | Training | No | Optional |
| Claude-SearchBot | Anthropic | Search indexing; blocking "may reduce your site's visibility" | Yes | Allow |
| Claude-User | Anthropic | User fetches | Indirect | Allow |
| Googlebot | Search index (feeds AI Overviews and Gemini grounding) | Yes (Google surfaces) | Allow | |
| Google-Extended | A robots.txt token, not a fetching bot; governs Gemini and Vertex training plus grounding; no effect on Search ranking | Governs Gemini training and grounding, not Search citation | Optional | |
| Bingbot | Microsoft | Bing index (historically feeds ChatGPT Search) | Indirect (via Bing) | Allow |
The strategic distinction that most robots.txt files get wrong: allowing GPTBot is not the same as appearing in ChatGPT search. GPTBot is the training crawler. OAI-SearchBot is the one that decides whether you can be shown in an answer.
This is not a theoretical risk, though the documented version is narrower than the one in circulation. [PRIMARY-VERIFIED] One major vendor really did flip the default. On 1 July 2025 Cloudflare launched a one-click block for AI bots and began blocking by default for new domains. Cloudflare's own account scopes that preset to single-purpose bots crawling for model training, and its documentation says the setting excludes mixed-purpose bots used both for training and for search. So the preset did not, on Cloudflare's own account, block every OpenAI agent. [UNSOURCED-IN-CIRCULATION] The wider claim that several vendors blocked all OpenAI agents by default has no primary behind it. If you did nothing wrong and disappeared anyway, check your CDN regardless. The load-bearing allow is OAI-SearchBot.
[SEMI-PRIMARY, COMMERCIAL] Vercel and MERJ analysed more than 500M GPTBot fetches (December 2024, [PRE-2026]) and found that no major AI crawler executes JavaScript. GPTBot fetched JS files in roughly 11.5% of requests and ClaudeBot in roughly 23.84%, but neither executes them. Only Googlebot with Gemini, and AppleBot, render.
The consequence is blunt. Client-side-rendered content, JS-injected FAQ and accordion blocks, tab-hidden content, and JS-loaded prices are invisible to ChatGPT, Claude and Perplexity. Server-side rendering is a prerequisite, not an optimisation.
You can test this in one line. Run curl -A "GPTBot" <url> and confirm your body text is present in the initial HTML. If the text you want quoted is not in that response, no amount of content strategy will help you, because gate one never opens.
[SECONDHAND] CDN and WAF blocking is the second silent failure. Cloudflare bot-fight rules and similar WAF configurations can block AI crawlers even when robots.txt permits them. [PRIMARY-VERIFIED, vendor doc] Perplexity's own crawler documentation carries explicit WAF allow-listing instructions and publishes live IP endpoints. Detect it in server logs by looking for the agent's user-agent string alongside verified IP ranges from the platform's published JSON files.
One case I want to label honestly. The underlying brief reports that three of twelve competing pages on this topic had FAQ content that failed to render to a crawler. I could not independently fetch-and-diff those specific pages within budget. The mechanism itself, JS-injected FAQPage schema disappearing from the server response, is independently documented (jangwook.net, 2026). So the mechanism is real, and the specific count is [SECONDHAND].
No, with one documented exception. [PRIMARY-VERIFIED] OpenAI documents no registration, listing, directory or submission mechanism for general web and editorial citations in ChatGPT answers. Those are earned through crawlable content plus OAI-SearchBot access. The exception is commerce: merchants can apply to submit a product feed, which governs product-discovery results and nothing else. It does not affect editorial citation.
The People Also Ask box surfaces "How do you register your business on ChatGPT?" and there is no product behind the question. Pages selling "ChatGPT business registration" concede the point themselves. One of them (sociolabs.in) admits there is not a direct sign-up button yet.
[PRIMARY-VERIFIED] "ChatGPT Business" is an enterprise seat product with Standard and Premium seats, Premium priced at $100 per user per month billed annually, or $125 monthly, offering centrally managed workspaces. It has nothing to do with visibility or citation. Buying it will not get you cited.
[PRIMARY-VERIFIED, vendor assertion] There is no paid-placement mechanism for organic citations. OpenAI introduced separate, labelled ChatGPT ads in 2026 that appear below answers and do not alter the organic citation.
[CONTRADICTED] The circulating claim that Gemini, Perplexity and Copilot offer no registration or listing mechanisms at all is the thing contradicted here: all three do operate them, and my earlier reading of this was wrong. Perplexity runs a free merchant programme taking a product feed, Google runs Merchant Center and Business Profile, and Microsoft runs Merchant Center listings plus Bing Places. What none of them offers is a way to register for general editorial citation. Those programmes govern shopping and local surfaces only, and for editorial citation the path is identical across all of them: crawler access, third-party authority, and structured, extractable content. Custom GPTs exist, but they do not affect organic citations.
This one costs people money, because they buy the wrong product.
[PRIMARY-VERIFIED] A local-SEO citation, a term popularised by David Mihm around 2008, though Mihm himself credits the concept to others before him, and defined by BrightLocal and Whitespark, is any online mention of a business's NAP (Name, Address, Phone number) across directories, review sites, and the Google Business Profile. It is a local-ranking relevance signal, and Whitespark and BrightLocal rank it among the top signals for local pack results.
[PRIMARY-VERIFIED] An AI-answer citation, as defined in the GEO paper, is a source quoted or linked inside a generated answer.
The two share the word "citation" and nothing else.
[SECONDHAND] Does local-citation work bear on AI citation at all? Indirectly and weakly. Consistent NAP and entity data across directories, Google Business Profile, and Bing Places improves the accuracy of what AI says about a local business and adds crawlable corroboration. But local-citation volume is not a documented AI-citation driver. If a vendor sells you a directory-listing package as an AI visibility service, they are selling a 2008 product with a 2026 label.
This is the only peer-reviewed primary in the field with per-tactic numbers, which is why it gets quoted constantly and why the quoting is so often wrong.
[PRIMARY-VERIFIED] The real scope. GEO-bench consists of 10,000 queries drawn from nine sources and spanning 25 domains (MS MARCO, ORCAS-1, Natural Questions, plus others and synthetic queries). The evaluation engine fetched the top 5 Google sources per query and generated answers with GPT-3.5-turbo, sampling 5 responses at temperature 0.7. The methods were then validated on Perplexity.ai, where the paper reports gains of up to 9% on position-adjusted word count and up to 37% on subjective impression, both of them Statistics Addition cells. Nine content methods were tested: Authoritative, Statistics Addition, Keyword Stuffing, Cite Sources, Quotation Addition, Easy-to-Understand, Fluency Optimization, Unique Words, and Technical Terms.
[PRIMARY-VERIFIED] The real per-tactic figures, on the position-adjusted word count metric, against a no-optimisation baseline:
- Quotation Addition, roughly +41%. The single best method on the word-count metric.
- Statistics Addition, roughly +31% (roughly +23% on the subjective metric). Never the top method.
- Cite Sources, roughly +28% on position-adjusted word count, 24.6 against a 19.3 baseline, measured with one randomly chosen source per query optimised. In the paper's separate all-sources-optimised scenario the same method moves a rank-5 page +115.1% while the rank-1 page drops 30.3%. These are two different experiments, not one curve.
- Keyword Stuffing: about 8% below baseline. The word "obsolete" is not the paper's, and I should not have attributed it. The paper files keyword stuffing under non-performing methods and says such methods offer little to no improvement. It measures 17.7 against a 19.3 baseline on position-adjusted word count, and 10% worse than baseline on Perplexity.
- Effects vary by domain and by starting position. The safe headline is "up to 40%."
One reading corroborates this and one does not, and the difference is instructive. wpconsults.com (14 June 2026) reproduces the table with Quotation around +41%, Statistics around +31%, and Cite Sources around +28% rising to +115% at rank 5. Secondary write-ups mostly read the table one column too far left. Position-Adjusted Word Count is reported as three sub-columns, Word, Position and Overall, and the paper's headline is computed on Overall: Quotation Addition 27.2 against a 19.3 baseline, about +41%. The plain Word sub-column gives 27.8 against 19.5, about +43%, which is what aisearchglossary.com quotes as the position-adjusted figure. Keyword Stuffing sits below baseline either way.
Notice the Cite Sources result, because it is the most useful and least repeated finding in the paper. The same tactic that lifts a rank-5 page by up to 115% costs a rank-1 page roughly 30%. Optimisation advice that ignores your starting position can actively harm you.
BuildRocketLabs (quotations +41%, statistics +33%, cite sources +28%, keyword stuffing at a 9% decline, scope stated as "10,000 queries / 25 domains") is the closest to correct, and the only rendering that reports the negative keyword-stuffing result. Its one deviation is small, and it is a metric-selection difference rather than an error: Statistics Addition is +31% on the headline position-adjusted word count column, 25.2 against 19.3, while the +33% figure matches the adjacent word count column, 25.9 against 19.5. The scope is accurate. GEO-bench is 10,000 queries spanning 25 domains and drawn from nine sources, and the paper states both.
Adriel (citing sources +40%, statistics +40%, quotations +28%, "up to 115% for lower-ranked sites") is [CONTRADICTED]. It swaps the ranking of the tactics, inflates statistics to +40%, and mis-assigns the +115%, which belongs specifically to Cite Sources at rank 5.
CompleteSEO ("statistics +37%, similar gains for quotes and citations") is [SECONDHAND], mis-scoped rather than wrong. The +37% is a real cell: it is Statistics Addition's Subjective Impression gain on Perplexity.ai, 33.9 against a 24.7 baseline, measured on the paper's validation subset rather than on the main benchmark, where statistics is roughly +31% on position-adjusted word count. The "similar gains for quotes and citations" half is contradicted: on Perplexity, Cite Sources scored below baseline on that same metric. I had this adjudication backwards in an earlier draft.
Apollo Digital ("up to 40% overall, no per-tactic figures") is incomplete but not wrong.
[CONTRADICTED] Where the "+37% statistics" chain forked. A secondary paper written in peer-reviewed style (Oruesagasti, arXiv:2603.12282, 5 March 2026) re-tabulated GEO's results as Cite Sources +40.0%, Statistics +37.0%, and Quotation +22.0%, complete with significance stars. Those are that author's reconstruction. They do not match the original position-adjusted figures, and they have propagated as though the original paper said them. The other branch of the chain, CompleteSEO's, is legitimate but mis-scoped, and is dealt with above.
| # | Circulating claim | Verdict | Tag |
|---|---|---|---|
| a | Wikipedia 7.8% of ChatGPT citations | Correct as share of total citations (Profound 680M dataset) | [PRIMARY-VERIFIED, COMMERCIAL] |
| b | Wikipedia 47.9% | Correct as Wikipedia's share within ChatGPT's top-10 most-cited sources (same Profound dataset). MarketerHire misframes this as "share of responses." | [PRIMARY-VERIFIED, COMMERCIAL] / MarketerHire framing [CONTRADICTED] |
| c | ChatGPT cites brands 0.59%; Perplexity 13.05%; Grok 27% | Superlines 34,234-response study. Metric: share of responses containing a link to the tracked brand. | [PRIMARY-VERIFIED, COMMERCIAL] |
| d | ChatGPT cites sources 87%, names brands 20.7% | Semrush ghost-citations study, June 2026. Note the metric: mention rate is the brand name appearing in the answer text, not a citation | [PRIMARY-VERIFIED, COMMERCIAL] |
| e | 60/40 parametric-recall vs live-retrieval | Contently states it with no source | [UNSOURCED-IN-CIRCULATION] |
| f | 76.4% of ChatGPT citations updated within 30 days | Originates with Ahrefs, October 2025. The "ConvertMate" and "Digitaloft Research" attributions are both false, passed along by secondary write-ups. Caveat: Wikipedia is over half the analysed set | [PRIMARY-VERIFIED, COMMERCIAL] |
| g | +28% at 2 months; 2x at 3 months (SE Ranking) | Does not match SE Ranking's own 129,000-domain study (which found freshness matters but FAQ schema underperforms) | [SECONDHAND], likely misattributed |
| h | Tables 81% vs prose 23%; FAQ 2.6x | Contently states it with no source (Onely separately claims tables around 2.5x) | [UNSOURCED-IN-CIRCULATION] |
| i | Only 12% of AI-cited sources in Google top 10; 88% elsewhere | Ahrefs (12% overlap) is [PRIMARY-VERIFIED]. Moz (88% of AI Mode citations outside top 10, roughly 40k queries, Feb 2026) I reached only through a summary and could not retrieve a stable canonical URL for | [PRIMARY-VERIFIED, COMMERCIAL] (Ahrefs) / [SECONDHAND] (Moz) |
| j | Google rank is the strongest predictor | Fischman / Growth Marshal, 730 citations, 75 queries: rank OR 0.762 per position | [PRIMARY-VERIFIED, COMMERCIAL author] |
| k | Comparative 2.4x; how-to 42.8% | Semrush ghost-citations study, June 2026: comparative queries 43.3% mention rate, 2.4x informational's 18%; how-to 42.8%. Adriel restates these without attribution | [PRIMARY-VERIFIED, COMMERCIAL] |
| l | Generic schema null; only Product/Review schema wins | Fischman: generic schema OR 0.678, p=.296 (null); Product/Review with concrete fields 61.7% vs 41.6%, p=.012 | [PRIMARY-VERIFIED, COMMERCIAL author] |
1. Does schema help? Resolved, moderate-high confidence. Generic schema does not independently drive citation. Fischman's corrected model returns a null result (OR 0.678, p=.296). SE Ranking found pages without FAQ schema averaged more citations than pages with it (4.2 versus 3.6). Ahrefs' 1,885-page test found AI Mode up 2.4% and ChatGPT up 2.2%, both statistically indistinguishable from zero. What does help is Product and Review schema with populated concrete fields (price, aggregateRating, specs), at 61.7% versus 41.6%, especially for lower-authority domains. "Zero schema is a citation killer" is [CONTRADICTED]. What would settle it further is a controlled experiment adding attribute-rich schema to unranked pages, which nobody has run.
2. The GEO numbers. Resolved, high confidence. See the section above.
3. How much of ChatGPT is Wikipedia? Resolved, high confidence. The apparent 6x gap between 47.9% and 7.8% is a denominator artefact, not a contradiction. 47.9% is Wikipedia's share within the top-10 cited sources. 7.8% is its share of all citations. Both come from Profound's 680M dataset. The reconciliation is confirmed.
4. Bing or OpenAI's own index? Unresolved and evolving, moderate confidence. OpenAI's crawler documentation never names Bing. Multiple 2024 sources (Yoast, Medium/thekeyword) state that ChatGPT Search relies on Bing's index, with OpenAI's VP of Engineering reportedly confirming it at an AMA [PRE-2026]. But Search Engine Land reported in 2026 that OpenAI now operates its own web index, codenamed "labrador," for instant mode, with only around 1.5% of its URLs appearing in Bing's top 20 for the same fan-out queries. Practical verdict: Bing indexation remains valuable but is neither sufficient nor exclusive, and the documented governing lever is allowing OAI-SearchBot. Bing Webmaster Tools setup is still worth doing, because it is cheap and still feeds part of the pipeline, but it is not the whole game.
5. Does Google rank predict AI citation? Partially resolved, moderate-high confidence. Both camps are right about different things. Fischman shows rank strongly predicts per-page citation odds: position-1 pages were cited 43% of the time, position-7 pages 5%. Moz and Ahrefs show that most citation volume lands on pages outside the original query's top 10, because query fan-out pulls from many sub-query SERPs. Rank raises your odds on any given sub-query. The aggregate still spreads wide. Those two facts are compatible, and treating them as a contradiction is how most articles on this topic go wrong.
6. How strong is freshness? Likely overstated, moderate confidence. The 76.4% figure has a broken provenance chain, attributed to ConvertMate in one place and Digitaloft in another, and Contently's own article is internally inconsistent, calling 76.4% dominant while its ConvertMate factor breakdown assigns freshness a 10% weighting. [PRIMARY-VERIFIED, COMMERCIAL] The direction is well supported, with Ahrefs finding AI-cited content is 25.7% fresher than organic results, measured on time since publication rather than time since last update. The specific number should be treated as unreliable.
7. How often does ChatGPT cite brands at all? Unresolved, low-to-moderate confidence, because of the metric confusion described earlier: linked citation versus mention versus source-family share. If the low figure of roughly 0.59% linked brand citation (Superlines) is correct, then the strategically correct move is to target Perplexity and Gemini first and ChatGPT last for link citations. That conclusion is drawn below, and notably the original page reporting 0.59% did not draw it.
8. Is llms.txt worth doing? Kill it, high confidence. [PRIMARY-VERIFIED] Google's John Mueller stated that no AI system currently uses llms.txt (17 June 2025, [PRE-2026, but reaffirmed]), and Gary Illyes confirmed in July 2025 that Google does not support it. Google has since put it in writing: its generative-AI optimization guide (May 2026, llms.txt note added 15 June 2026) states Google Search doesn't use llms.txt or other AI text files, which "will neither harm nor help" visibility. Ahrefs studied 137,210 domains in 2026 and found 97% of llms.txt files received zero requests during the observation month. Of the 3% that were read, 96% of requests came from bots, and Slackbot fetched them more often than PerplexityBot did. No major engine documents reading the file. This is the clearest "stop doing this" finding in the entire evidence base.
| Format / page type | Evidence | Tag |
|---|---|---|
| Content with quotations, statistics, cited sources | GEO paper: quotations +41%, statistics +31%, cite sources +28% (position-adjusted word count) | [PRIMARY-VERIFIED] |
| Comparison pages and tables | Onely "tables 2.5x"; widely reported as the most extractable format | [SECONDHAND] |
| Listicles | Onely: "50% of top AI citations" | [SECONDHAND, COMMERCIAL] |
| FAQ question-and-answer blocks | Reported as ready-made retrieval units, BUT SE Ranking found FAQ-schema pages underperformed (4.2 without vs 3.6 with) | [SECONDHAND / CONTRADICTED on schema] |
| Documentation and canonical pages | ChatGPT skews to first-party docs on technical queries (see the engine-priority section) | [SECONDHAND] |
| Original research and proprietary data | Multiple vendors; consistent with the GEO statistics finding | [SECONDHAND] |
| Brand websites and product pages | tryanalyze.ai: 68.8% of ChatGPT citations | [SECONDHAND, COMMERCIAL] |
| Pricing pages | Valuable but frequently fail, because JS-rendered prices are invisible to crawlers | [SECONDHAND] |
| Glossary pages | Asserted, no controlled evidence | [UNSOURCED-IN-CIRCULATION] |
[PRIMARY-VERIFIED, COMMERCIAL] The engine-specific split matters as much as the format ranking. ChatGPT concentrates in brand websites and product pages at 68.8%. Perplexity and Google AI Mode spread roughly evenly across websites (35% to 38%), lists (33% to 34%) and editorial (24% to 25%). That comes from tryanalyze.ai's analysis of 22,295 answers.
Read those two facts together and the implication is concrete. If you are targeting ChatGPT, your own site is the asset. If you are targeting Perplexity or Google AI Mode, editorial placement and list inclusion carry a third of the weight each, and your site alone will not get you there.
Assume no press, no reviews, no rankings, no Wikipedia entity. Here is what is genuinely available.
[PRIMARY-VERIFIED] Allow OAI-SearchBot, PerplexityBot, and Claude-SearchBot, then publish server-rendered, extractable pages that directly answer buyer questions. This is the floor, and it is free.
[SECONDHAND] Get indexed in Bing via Bing Webmaster Tools and IndexNow. It is still part of ChatGPT's pipeline.
[PRIMARY-VERIFIED, COMMERCIAL] Seed third-party mentions on the community platforms engines cite disproportionately. Profound found Reddit accounts for 46.7% of Perplexity's top-10 source share. SE Ranking and Semrush both place Quora and Reddit high in Google's AI surfaces. One scope limit: treat seeding as a lever for engines like Perplexity that lean on community sources. For Google's AI surfaces, Google's May 2026 guide warns that seeking inauthentic mentions across the web "isn't as helpful as it might seem" — genuine participation and earned coverage are the only durable version of this play.
[PRIMARY-VERIFIED] What is genuinely gated: Wikipedia, because of the notability bar. You cannot create your own page from scratch. High-DR editorial coverage is reachable but takes time.
[SECONDHAND] The shortest documented path to a first citation is weeks, not days. Perplexity crawls established domains 1 to 3 times a week, and a crawled page is "likely being considered for citation within 24 to 48 hours," which is a vendor claim from Conbersa. Perplexity is the most reachable engine from zero, because its crawl is query-driven and it cites broadly.
The honest headline, and I want to be plain about it: anyone promising 30-day ChatGPT results is overpromising. Vendor consensus (gugubrand, 2026) is first signals in 2 to 3 months.
[SECONDHAND, evidence-weighted] For a small business seeking link citations, the order is not the order of audience size. It is the order of citation probability.
Perplexity first. Among the engines with meaningful buyer audiences it has the highest linked brand-citation rate, roughly 13% against ChatGPT's roughly 0.59% (Superlines). Grok cites at roughly twice Perplexity's rate, 27.01%, and is set aside here on audience-size grounds rather than on the evidence. Superlines' larger all-time dataset uses a broader definition, any response containing at least one citation, and puts Perplexity at 15.43% against ChatGPT's 2.78%. The two ChatGPT figures are not two readings of one measurement, so do not treat them as mutually reinforcing. It crawls broadly, it is the most reachable engine from zero, and it cites sources transparently, which means you can trace exactly which page won and fix yours.
Google AI (AI Overviews and AI Mode) and Gemini second. Both are grounded in Google's index, so existing SEO partially transfers, and Gemini has no separate crawler at all. It uses Googlebot.
ChatGPT last for link citations, because it has the lowest linked brand-citation rate. But it has the largest audience, so brand mentions without links and accurate entity presence still matter enormously. Nobody in the source set drew this staged conclusion. It follows directly from the citation-rate evidence.
One pattern worth investigating further. On this exact query class, ChatGPT reportedly cited 30 URLs, all of them first-party vendor documentation (OpenAI Help Center, Google Developers, including non-English locale variants), and zero third-party articles. That suggests that for how-to and technical queries, ChatGPT retrieves a canonical documentation entity rather than adjudicating a field of competing articles.
If that generalises, the strategic consequence is significant: for technical topics, being the canonical documentation, or being cited by it, beats out-writing the listicles. [NO EVIDENCE FOUND] of a formal study confirming that it generalises. I am flagging it as a genuine, actionable gap that nobody has closed.
[SECONDHAND, corroborated across multiple vendors and observed platform behaviour] No platform offers an editable business listing the way Google Business Profile does. What exists is feedback, and feedback is not editing.
ChatGPT: thumbs-down, then report, then select the harmful or incorrect option and state the correct fact. OpenAI reviews flagged output but does not hand-edit facts about individual businesses.
Perplexity: the feedback icon. Because it shows sources, you can identify and fix the specific cited page, which is the more useful path.
Google AI Overviews: the feedback icon under the answer, then mark it inaccurate and supply the correct fact and a source URL.
Copilot: the Bing feedback tool.
What actually works, as opposed to what feels like action: fix the sources the engine reads. Your own site first, stating facts explicitly and literally ("Founded in 2021," not "a few years ago"), then Google Business Profile, Bing Places, Wikipedia and Wikidata, Crunchbase, LinkedIn, and the directories. Then publish one plain, crawlable page stating your facts, get it recrawled, and re-test until the answer flips. Timeline: months, not days. The feedback buttons take seconds and have no downside, but they do not rewrite answers by themselves.
[PRIMARY-VERIFIED] Legal context. In Walters v. OpenAI, the first US defamation suit over a ChatGPT hallucination, a Georgia court granted OpenAI summary judgment in May 2025 [PRE-2026]. EU GDPR accuracy rights, the basis of noyb's complaint, cover personal data rather than company facts. Practically, the legal route is not a route yet.
[PRIMARY-VERIFIED], from vendor pricing pages retrieved 30 August 2026.
Monitoring tools. Otterly.ai: Lite $29/mo, Standard $189/mo, Premium $489/mo, Enterprise from $1,000/mo. Scrunch AI: Starter $250/mo on annual billing, Growth $417/mo. Nightwatch: Starter €79/mo, with AI tracking included on all plans.
[PRIMARY-VERIFIED via Profound's public pricing page, verified 1 August 2026 by Trakkr] Profound: Starter $99/mo (ChatGPT only, 50 prompts) and Growth $399/mo (adds Perplexity and Google AI Overviews), both shown as billed yearly. Profound does not publish an annual total, so any yearly figure is derived rather than quoted; Enterprise custom. [SECONDHAND] The $2,000 to $5,000+ per month band appears on no Profound page; it comes from review aggregators, several of them competitors.
[SECONDHAND, some vendor self-interest] Peec AI runs roughly $80 to $420/mo. Semrush's AI Visibility Toolkit is $99/mo per domain billed annually, sold standalone rather than as an add-on, with extra domains at $99 each. Semrush's core plans run $139 to $549/mo. Ahrefs Brand Radar splits in two: Custom Prompts from $50/mo, free on Lite and above, up to $699/mo for all models, and the AI Visibility Index at a flat $199/mo covering all platforms. Ahrefs base plans start at $29 for Starter and $129 for Lite.
[SECONDHAND, agency self-interest, flagged] GEO and AEO agency retainers cluster around $2,000 to $10,000/mo for mid-market, with a $1,000 to $2,500 entry floor, $15,000+ at enterprise, and one-time audits at $1,500 to $5,000. For context, Ahrefs' 2024 survey of 439 agencies found the most common classic SEO retainer is $500 to $1,000/mo. The premium is real and it is being charged on thin evidence.
[SECONDHAND, agency self-interest, flagged] Digital PR, which the evidence suggests is the highest-impact lever, runs $3,000 to $15,000/mo in retainers, with a market-average monthly contract of approximately $5,458 (BuzzStream survey, via Reporter Outreach). Cost per earned link runs $150 to $1,000+ depending on publication authority. Newswire distribution runs $600 to $3,000 per release. Senior PR billing sits around $278/hr (PR Council 2025 Labor Rate Report).
[SECONDHAND] Review generation: entry tools $15 to $89/mo (RepliFast from $15, ReviewTrackers around $89 per location); full reputation platforms such as Birdeye and Podium around $299 to $599/mo per location.
[SECONDHAND] Internal hours: roughly 4 to 6 hours per month for a minimal small-business DIY effort, and roughly 20 to 40 hours per month to run an active programme that actually acts on the monitoring data. The second number is the one people underestimate. Monitoring you never act on is a subscription, not a strategy.
[PRIMARY-VERIFIED, COMMERCIAL] This is the part of the evidence base with the least ambiguity and the most consequence.
Muck Rack's "What Is AI Reading?" study (May 2026, 25M+ links across ChatGPT, Claude and Gemini, 17 industries) found that earned media accounts for 84% of all AI citations, that paid and advertorial content accounts for just 0.3%, and that journalism alone represents 27% of cited sources.
Stacker and Scrunch ran a citation-decomposition study (December 2025, 8 articles across 944 prompt-platform combinations on 5 LLMs). There was no un-distributed control arm, so read it as a measurement study rather than a controlled test. Across those 944 runs the brand's own URL was the sole citation 7.6% of the time. Third-party republished versions were cited in a further 27.5%, made up of 19.2% where only the syndicated version was cited and 8.3% where both were, taking total citation coverage to about 34%. Stacker's own headline puts that as a 325% lift, 8% to 34%. Their expanded March 2026 study (87 stories, 30 distinct brands, 2,600+ prompts, 8 platforms) revised the lift down to a median 239%, which matters, because the effect shrank as the sample grew tenfold.
Same content, more hosts, roughly four times the citation coverage the brand's own domain produced alone. If you take one budget decision from this document, take that one: distribution beats production.
Most people "test" by asking ChatGPT about their own domain once and reacting to whatever comes back. That produces a feeling, not a number. This protocol produces a number you can compare month over month.
Prompt set: 20 prompts. 10 category or buyer prompts ("best [category] for [use case]") and 10 branded prompts ("what does [business] do / cost / where is it").
Runs: 5 runs per prompt per engine. Non-determinism is real (wpconsults, 2026), so a single run tells you almost nothing.
Engines: ChatGPT with web search ON, Perplexity, Gemini, Google AI Mode.
Controls: logged out, fresh session, memory cleared, web search ON. Record the model name, version, and date.
Record per run: whether the brand was cited with a link (Y/N); whether the brand was mentioned without a link (Y/N); the position of the citation; which competitor sources appeared; and whether the cited URL was yours.
Compute: citation rate equals cited runs divided by total runs, per engine. Then compare against the published benchmarks: Superlines' ChatGPT 0.59% and Perplexity 13.05% linked brand citation, and tryanalyze.ai's ChatGPT 68.8% brand-site source share.
Repeat monthly, so you can separate the effect of your content changes from ordinary model drift. Without a monthly baseline you will credit yourself for drift and blame yourself for it in equal measure.
[SECONDHAND] ChatGPT's citations on this query class reportedly included multiple non-English locale variants of the same vendor documentation (OpenAI Help Center, Google Developers), which suggests that locale-duplicated canonical docs are retrieved as a single entity across languages.
[NO EVIDENCE FOUND] of any rigorous study measuring locale, language, or regional variation in AI citation behaviour. This is a real and unclosed gap. Academic citation-language studies, such as Google Scholar language distribution work, are unrelated to generative-engine citation and should not be borrowed as evidence for it.
| Claim | Repeated by | Where the chain breaks | Primary source? |
|---|---|---|---|
| "85% of retrieved pages never cited" | Many | Traceable to AirOps, Mar 2026; downstream repetitions rarely link it | YES. AirOps, Mar 2026 (15% cited / 85% not, 548,534 pages) |
| 60/40 parametric vs retrieval split | Contently | Stated with no source | None found |
| 76.4% freshness within 30 days | Contently, Averi, ZipTie, needle.sh | ConvertMate and Digitaloft attributions are both false | Found: Ahrefs, October 2025 |
| Tables 81% / prose 23%; FAQ 2.6x | Contently | No source given | None found (Onely's 2.5x is separate and carries no source of its own) |
| Brand named in 20.7% of responses | Adriel | Circulated without attribution | Found: Semrush ghost-citations study, June 2026 |
| Comparative 2.4x / how-to 42.8% | Adriel | Circulated without attribution | Found: Semrush ghost-citations study, June 2026 |
| "+28% at 2 months / 2x at 3 months" (SE Ranking) | Adriel | Does not match SE Ranking's own study | Contradicted by claimed source |
| "+37% statistics" (GEO) | CompleteSEO, many | Is the Perplexity overall figure, or the Oruesagasti re-tabulation, not GEO's statistics tactic | Original paper says roughly +31% |
Publishing the failures is the point, so here they are.
The exact subjective-impression per-tactic figures from GEO Table 1, beyond the position-adjusted word count metric. The position-adjusted figures are verified; the full subjective table was not fully extracted within budget.
Whether ChatGPT's canonical-documentation retrieval pattern on technical queries generalises. No formal study exists.
A rigorous locale and language citation-variation study. None found.
The 76.4% freshness figure originates with Ahrefs, October 2025, not with ConvertMate or Digitaloft, both of which are false attributions passed along by secondary write-ups. Treat the number cautiously anyway, because Wikipedia is a little over half the analysed set.
Independent fetch-and-diff confirmation that three specific competing pages' FAQ blocks fail to render. The mechanism is documented; those exact pages I did not verify.
| Organisation | Title / item | Date | URL | Primary/secondary | Commercial interest |
|---|---|---|---|---|---|
| Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan, Deshpande (IIT Delhi / Princeton / independent) | "GEO: Generative Engine Optimization," KDD '24, arXiv:2311.09735v3 | 28 Jun 2024 | https://arxiv.org/abs/2311.09735 ; https://arxiv.org/html/2311.09735v3 ; https://arxiv.org/pdf/2311.09735 | PRIMARY (peer-reviewed) | None |
| OpenAI | "Overview of OpenAI Crawlers" | live (retrieved Aug 2026) | https://developers.openai.com/api/docs/bots | PRIMARY (platform doc) | Vendor |
| OpenAI Help Center | "Publishers and Developers FAQ" | live | https://help.openai.com/en/articles/12627856-publishers-and-developers-faq | PRIMARY | Vendor |
| OpenAI Help Center | "Sign up for ChatGPT Business"; "ChatGPT Business: General FAQ" | live (2026) | https://help.openai.com/en/articles/8980713-how-do-i-sign-up-for-chatgpt-business | PRIMARY | Vendor |
| Perplexity | "How does Perplexity follow robots.txt?"; "Perplexity Crawlers" | live | https://www.perplexity.ai/help-center/en/articles/10354969 ; https://docs.perplexity.ai/docs/resources/perplexity-crawlers | PRIMARY | Vendor |
| Anthropic | Crawler opt-out doc (art. 8896518); crawler docs update | 2024 / Feb 2026 | https://support.anthropic.com/en/articles/8896518 | PRIMARY | Vendor |
| Search Engine Journal | "Anthropic's Claude Bots Make Robots.txt More Granular" | Feb 2026 | https://www.searchenginejournal.com/anthropics-claude-bots-make-robots-txt-decisions-more-granular/568253/ | Secondary | Trade press |
| Google / Search Engine Land | "Google introduces Google-Extended" | 2023 [PRE-2026] | https://searchengineland.com/google-extended-crawler-432636 | Primary concept / secondary reporting | Trade press |
| John Mueller (Google) | "No AI system currently uses llms.txt" | 17 Jun 2025 [PRE-2026] | https://bsky.app/profile/johnmu.com/post/3lrshm4gggs2v (his own post); write-up: https://www.seroundtable.com/google-ai-llms-txt-39607.html | PRIMARY (author's own post) | None (write-up: trade press) |
| Ahrefs | "97% of llms.txt Files Never Get Read" (137k sites) | 2026 | https://ahrefs.com/blog/llmstxt-study/ | Semi-primary (network data) | SEO SaaS |
| Ahrefs | "Only 12% of AI Cited URLs Rank in Google's Top 10" | 2026 | https://ahrefs.com/blog/ai-search-overlap/ | Semi-primary | SEO SaaS |
| Moz | ~40,000-query AI Mode study (88% of citations outside the top 10) | Feb 2026 | reached via a LinkedIn summary; I could not retrieve a stable canonical URL, so treat the figure as [SECONDHAND] | Secondary as reached | SEO SaaS |
| Profound | "AI Platform Citation Patterns" (680M citations) | Aug 2025 update [PRE-2026 data window Aug 2024 to Jun 2025] | https://www.tryprofound.com/blog/ai-platform-citation-patterns | Semi-primary | AI-visibility SaaS |
| 5W PR | "State of AI Citations 2026" | 2 Jun 2026 | https://www.5wpr.com/research/state-of-ai-citations-2026/ | Secondary synthesis | PR agency |
| Kurt Fischman / Growth Marshal | "Does Schema Markup Predict AI Citation?" (SSRN; Zenodo DOI 10.5281/zenodo.18728697) | 22 Feb 2026 | https://papers.ssrn.com/sol3/papers.cfm?abstract_id=6284518 ; https://www.runmarshal.com/research/does-schema-markup-predict-ai-citation | Semi-primary study | Author is a GEO agency |
| Vercel / MERJ | "The Rise of the AI Crawler" | Dec 2024 [PRE-2026] | https://vercel.com/blog/the-rise-of-the-ai-crawler | Semi-primary (network data) | Hosting vendor |
| AirOps | "The Influence of Retrieval, Fan-out, and Google SERPs on ChatGPT Citations" (548,534 pages / 15,000 prompts / 43,233 queries) | 12 Mar 2026; Search Engine Land write-up 13 Mar 2026 | https://www.airops.com/report/influence-of-retrieval-fanout-and-google-serps-in-chatgpt ; https://searchengineland.com/chatgpt-retrieved-vs-citations-study-471606 | PRIMARY (report read directly) | AI content SaaS |
| Superlines | 34,234-response brand-citation study | 2026 | https://www.superlines.io/articles/ai-search-statistics/ | PRIMARY (vendor's own write-up) | AI-visibility SaaS |
| tryanalyze.ai | Source-family mix (22,295 answers) | 2026 | https://www.tryanalyze.ai/blog/state-of-ai-search-source-family-mix | Semi-primary | AI-visibility SaaS |
| SE Ranking | "How to optimize for ChatGPT" (129,000 domains) | 24 Nov 2025 | https://seranking.com/blog/how-to-optimize-for-chatgpt/ | Semi-primary | SEO SaaS |
| Novastacks | "What Drives AI Citations?" | 26 Jan 2026 | https://www.novastacks-ai.com/blog/what-drives-ai-citations | Secondary | SEO SaaS |
| Muck Rack | "What Is AI Reading?" (25M+ links, 17 industries) | May 2026 | https://media.muckrack.com/documents/What_Is_AI_Reading__May_2026.pdf | PRIMARY (report PDF) | PR SaaS |
| Stacker + Scrunch | Earned-distribution citation study | Dec 2025 [PRE-2026]; expanded Mar 2026 | https://machinerelations.ai/research/earned-vs-owned-ai-citation-rates-2026 | Semi-primary | Content/AI SaaS |
| Oruesagasti (Interamplify) | "Algorithmic Trust, UK iGaming," arXiv:2603.12282 | 5 Mar 2026 | https://arxiv.org/pdf/2603.12282 | Secondary (re-tabulates GEO) | Agency author |
| Contently | "How to Get Your Brand Cited in ChatGPT" | 29 Apr 2026 | https://contently.com/2026/04/29/how-to-get-cited-in-chatgpt/ | OBJECT OF STUDY | Builds Radarly (disclosed) |
| MarketerHire | "How to Get Cited by ChatGPT: 9 Tactics" | 2026 | https://marketerhire.com/blog/how-to-get-cited-by-chatgpt | OBJECT OF STUDY | Marketplace |
| Onely | "LLM-Friendly Content: 12 Tips" | 2026 | https://www.onely.com/blog/llm-friendly-content/ | OBJECT OF STUDY | SEO agency |
| BrightLocal / Whitespark | "What is NAP / a Local Citation" | 2023 to 2024 [PRE-2026] | https://www.brightlocal.com/learn/what-is-nap/ ; https://whitespark.ca/blog/what-is-a-local-citation-for-local-seo/ | PRIMARY (definitional) | Local-SEO SaaS |
| Search Engine Land | "Inside ChatGPT's retrieval stack" (own index / "labrador") | 2026 | https://searchengineland.com/chatgpt-retrieval-stack-index-cache-pages-485036 | Secondary | Trade press |
| Otterly.ai / Scrunch / Nightwatch | Vendor pricing pages | retrieved 30 Aug 2026 | https://otterly.ai/pricing ; https://scrunch.com/pricing/ ; https://nightwatch.io/pricing/ | PRIMARY (pricing) | Vendors |
- Can I register my business with ChatGPT?
- Not for editorial citation. OpenAI documents no registration or submission mechanism for general web citations in ChatGPT answers, and there is no paid-placement route to an organic citation. Commerce is the exception: OpenAI, Perplexity, Google and Microsoft all run merchant or product-feed programmes, but those govern shopping and local surfaces only. Editorial visibility is earned through crawlable content plus OAI-SearchBot access.
- Which crawler actually decides whether I appear in ChatGPT?
- OAI-SearchBot. OpenAI's own documentation states that opting out means the site will not be shown in ChatGPT search answers. GPTBot is the training crawler and does not govern citation, so allowing GPTBot while blocking OAI-SearchBot leaves you invisible in answers.
- Does schema markup get me cited?
- Generic schema does not independently drive citation. Fischman's model returns a null result (OR 0.678, p = .296), and Ahrefs' 1,885-page test found effects statistically indistinguishable from zero. What does help is Product and Review schema with populated concrete fields such as price and aggregateRating, cited 61.7% against 41.6%.
- Is llms.txt worth adding?
- No. Google's John Mueller stated no AI system currently uses it, and Ahrefs studied 137,000 sites and found 97% of llms.txt files received zero requests in the observation month. Of the 3% that were read, 96% of requests came from bots, and Slackbot fetched them more often than PerplexityBot. This is the clearest stop-doing-this finding in the evidence base.
- Why is my page read but never quoted?
- Because retrieval and citation are two different gates. AirOps found ChatGPT read 548,534 pages while answering 15,000 prompts and cited 15% of them, so five of every six pages that get retrieved are never quoted. Clearing gate one only makes you eligible for gate two.
- Which AI engine should a small business target first?
- Perplexity, because among engines with meaningful buyer audiences it has the highest linked brand-citation rate, roughly 13% against ChatGPT's roughly 0.59%, it crawls broadly, and it shows its sources so you can see which page won. Google AI second, since existing SEO partially transfers. ChatGPT last for link citations, though its mentions still matter most by audience size.
- Is my own site or third-party coverage the better investment?
- Third-party coverage, by a wide margin. Muck Rack found earned media accounts for 84% of all AI citations and paid content for 0.3%. Stacker and Scrunch found the brand's own URL was cited alone 7.6% of the time, with syndicated third-party versions accounting for a further 27.5%, taking total citation coverage to about 34%.
Most of the research above was published by companies selling the remedy their research recommends. Onely sells SEO. Profound, Scrunch, Otterly and tryanalyze sell AI-visibility tracking. Muck Rack sells PR software. Fischman's schema study is authored by a GEO agency. Stacker and Scrunch sell distribution, and their study found distribution works.
I have flagged every one of those inline, and I am not dismissing the work. Vendors have the data because they built the pipes. But the two findings I would bet money on are the two with the least commercial convenience attached: the peer-reviewed GEO paper, which has no commercial interest at all, and the null schema results, which nobody selling GEO wanted to publish.
The rest is a snapshot dated 30 August 2026. Re-run the test protocol in ninety days and half of it will have moved.
Related reading
- Why ChatGPT Recommends Your Competitor: A Provenance AuditI traced 14 circulated stats about why AI recommends your competitor back to source. Seven held up. One had no primary at all. The full provenance audit.
- GEO Monitoring Tools Compared (2026): Profound, Peec & MoreA vendor-neutral 2026 comparison of GEO monitoring tools — Profound, Otterly, Peec and more — with pricing, a scorecard, and a build-vs-buy verdict.
- AEO/GEO Playbook 2026: Your Next B2B Buyer Is a MachineHalf of B2B buyers now start in an AI chatbot, but AI sends only 1% of your traffic. The evidence-led AEO/GEO playbook for getting cited, not clicked.