← Blog
GEO·19 min read

How to Track Whether AI Mentions Your Brand or Competitors

Two tools scored the same brand 45 and 22. I ran two instruments over 40 buyer questions in one day; they agreed on 18% of sources. The method that survives.

TathagataFounder, ParaphrasePublished August 30, 2026· Updated August 31, 2026
Two tools.One brand.Two answers.SCORED 45 AND 2245TOOL A22TOOL B=one brandsame 40 questions, same day

Two tools can score the same brand 45 and 22. I ran two instruments over the same 40 buyer questions on one day; they agreed on 18% of cited sources. Here is the method that survives that, and what it costs.

Last updated: 30 August 2026

The question, verbatim from Quora: "How do you track whether your brand is actually being recommended by AI when traditional SEO tools cannot see inside a private chat session?"

The premise is correct. The shared public results page is gone, and what replaces it is not one hidden number you are failing to find. It is a sampling problem over a surface that is non-deterministic, personalised and different on every engine.

Which has an uncomfortable consequence. Every vendor sells a visibility score and none publish a method you could reproduce. The score is real in that it was computed, and unauditable in that you cannot check it.

What you can build instead is a directional, per-engine mention rate, relative to your own prompt set, read as a trend. Less satisfying than a score. Actually true.

That last sentence is not a figure of speech. A July 2026 preprint (arXiv 2607.13304) decomposed where the variation comes from, across 12,933 responses about 20 Central and Eastern European brands in 8 languages and 3 models. Its measured outcome is sentiment polarity rather than mention presence, which the author names as future work, so read the shares as the shape of the noise rather than as mention-rate numbers. Query language accounted for 26.5% of the variance in a single response. Brand identity accounted for 1.5% (ICC 0.0146). The paper puts it plainly: "a single AI answer carries almost no brand-discriminating signal." On the subset where the same prompt was actually re-run, a second fit isolated pure resampling at 34.8% of the variance, a separate decomposition rather than a third slice of the same pie.

WHY ONE ANSWER TELLS YOU NOTHINGYour brand explains 1.5% of the answerTwo variance decompositions, run on two populations. They are not slices of one pie.FULL CORPUS, 12,933 RESPONSESWhich language the query is in26.5%Which brand is being asked about1.5%The remaining 69.3% is one undivided residual: interactions and resampling together.A SEPARATE FIT, REPLICATED SUBSET, 7,173 RESPONSESRe-running the identical prompt34.8%arXiv 2607.13304, 2026. Full corpus n=12,933; replicated subset n=7,173. Brand identity ICC 0.0146. Outcome measured is sentiment polarity.
The thing you are trying to measure explains 1.5% of what you see. Everything else is the instrument and the weather.

The same study tests the field's five-run convention and finds five is a ceiling, not a target. Five is where the paper starts, not what it derives, and "a repeat past the fifth reduces it by only 0.0003." HALC (arXiv 2507.21831) landed there independently: "5 coding repetitions already led to robust evaluation metrics." So: five runs per prompt, then spend the next unit of budget the way the paper's own rule spends it, on another language first, then another engine, with a sixth repeat dead last. If you only sell in English, that collapses to five runs, then another engine.

WHERE TO STOP SPENDINGFive runs, then switch enginesHow much each additional repeat of the same prompt reduces error.5 runspast here, only 0.00031 run10 runsAdding a languageor an engine cutserror far more thanadding a repeat.arXiv 2607.13304; HALC (arXiv 2507.21831) reached five repeats independently. Curve is illustrative of the decay shape.
Five runs is a ceiling, not a target. The sixth repeat buys almost nothing; another language or engine buys a lot.

For scale: small models show 50% to 80% answer consistency at low temperature, and Atil et al. found accuracy swings of up to 10% across five identical runs of six models configured for determinism at temperature zero (arXiv 2408.04667v2, September 2024). A later revision, running 10 repeats of five models, pushes that to 15%. The authors do not isolate a cause experimentally. AirOps reports only 30% of brands stay visible from one answer to the next and 20% survive five runs, a vendor benchmark with undisclosed method.

Three words that are not synonyms. A mention is your brand named in the answer text. A citation is your URL listed as a source, which requires the engine to return source URLs at all. A recommendation is your brand named as a positive answer, usually scored 0 to 5 by prominence. Log all three separately. Collapsing them is how two honest people get different numbers for the same brand.

Hold these constant and record them every run: logged out, memory cleared, web search on or off, fixed location, model version, time of day. One row per run, carrying the answer text, brands mentioned, their positions, and citations returned.

ToolEnginesPromptsPrice, per third-party trackers, Jun-Jul 2026
Profound10, but tier-capped: Starter is ChatGPT only, Growth 3, Enterprise up to 950 to 100+$99/mo ChatGPT only; $399/mo for 3 engines
Peec AI3 of 6 included; a 4th model is a paid add-on50 / 150 / 350 by tier$95, $245, $495/mo; extra models $30 to $140
Otterly.AI4; Gemini, Claude, AI Mode are paid add-onsUser-definedLite $29, Standard $189; add-ons push it past $300
Semrush AI Visibility625 base, +50 for ~$60/mo$99/mo standalone, one domain
Ahrefs Brand RadarSix prompt-based indexes; Claude is custom-prompts only and Grok collection is paused; Reddit and YouTube free in beta; shows source URL per mentionVendor-stated corpus$199/mo per index or $699/mo for all, sold standalone; add an Ahrefs base plan from $129 and it is $328 to $1,148 all-in
WHAT THE SCORE COSTS$29 to $699 a monthEntry tier to working tier, per month.Ahrefs Brand Radar$199-$699Otterly.AI, with add-ons$29-$300Peec AI$95-$495Profound$99-$399Semrush AI Visibility$99Vendor pricing and third-party trackers, June to July 2026. Brand Radar is sold standalone; add an Ahrefs base plan and it runs $328 to $1,148.
Each bar is the span from entry tier to realistic working tier. None of these vendors publishes a reproducible method behind the score you are paying for.

They disagree, and they concede it. Pixis gives the example: one tracker reports a brand at 45, another at 22, and attributes the gap to "the measurement surface, not a change in the brand." No independent head-to-head that names the tools has been published. The closest is a June 2026 test by GTECH, a Dubai agency that sells no tracker of its own: it graded 12 platforms against 600 manually verified prompt-answer pairs across ChatGPT, AI Overviews, Gemini, Claude and Perplexity, and found mention-detection accuracy averaging 81% and ranging from 67% to 94%, but it reports the tools by tier, not by name. So here is a named one from my own data.

On 24 August 2026 I ran the same 40 B2B buyer questions through two instruments in one window: a six-engine citation log and a five-engine rank tracker. On Perplexity, the only engine where both returned sources on all 40, instrument A found 531 domain-query pairs and instrument B found 675. They agreed on 184. That is 18% agreement on cited sources, same brand, same questions, same day. On two questions they agreed on nothing.

It was worse elsewhere. Instrument A returned sources for ChatGPT on 35 of 40 questions, AI Overviews on 24, Gemini on 22. Instrument B returned zero for all three on all 40, logging fallback_responses for every ChatGPT row and no_index_match for every AI Overview row. One of the two would have told you the engines cite nobody at all.

MY OWN HEAD-TO-HEAD, SINCE NOBODY NAMES THE TOOLS18% agreement, same brand, same dayCited domain-query pairs found on Perplexity, 40 questions, 24 August 2026.Instrument A531Instrument B675184 agreedOn two of the forty questions, the two instruments agreed on nothing at all.Paraphrase Labs, 24 Aug 2026. Perplexity was the only engine where both instruments returned sources on all 40 questions.
Same brand, same questions, same day, two instruments. They agreed on 18% of cited sources. The instrument is part of the number.

The instrument is not a window onto the truth. It is part of the number.

Design the prompts around real purchase moments, not your own name: "best X tools for enterprise," "X alternatives," "how do I choose an X." Never "what is [brand]".

If you already use Apify, the parts exist: apify/chatgpt-search-scraper returns ChatGPT Search answers with cited sources at $0.005 per search on the $29 Starter plan, apify/perplexity-search-scraper at $0.013 on Starter and less on higher tiers, and community actors such as santhej/ai-rank-tracker-pro cover five platforms at $0.05 per check. First-party actors also exist for Gemini and Microsoft Copilot at $0.005, plus Google AI Overviews at $0.002 and AI Mode at $0.005. No first-party actor exists for Claude or Grok. All those are Starter-tier rates; the free plan pays $0.20 a search, so its $5 credit buys about 25 searches, not a thousand.

The architecture is boring on purpose. A schedule fires the actors against your stored prompt set, answers land in a dataset, a parsing step does brand matching and citation-domain extraction, and a scoring layer computes mention rate and share of voice. Apify plans run $0 (a $5 credit that does not roll over), $29, $199 and $999, and most pay-per-event actors, including all three above, bundle platform usage into the per-result price rather than billing it on top. Some switch usage on separately, so check each actor's pricing section before you budget. The cost that does stack regardless is post-run storage: reading and writing your dataset after a run incurs standard platform usage.

One thing to be blunt about. An actor existing in a store does not mean the platform permits it. OpenAI's terms bar users from "Automatically or programmatically extract[ing] data or Output" from the Services, with the API governed by separate Business Terms, so it remains the sanctioned route. Perplexity's bar any "robot, spider, crawlers, scraper, or other automatic device." Scraping either UI breaches their terms regardless of tooling. The compliant routes are the OpenAI API's web_search tool, which does browse and returns url_citation annotations, and Perplexity's Sonar API, which returns citations natively. Neither is a mirror of the consumer UI, so treat them as a related answer surface, not a substitute for the one your buyer sees.

Anthropic's position, from its May 2025 web search API announcement, is that "every web-sourced response includes citations to source materials." The condition is the search. My own run corroborates it from the other side: across all 40 questions, Claude returned zero sources. Not a low number. Zero, forty times out of forty. You cannot earn a Claude citation directly, only a brand-mention effect over time.

Grok gates citations the same way: the citations list is "always returned by default" once search tools fire, though xAI notes separately that enabling inline citations "does not guarantee that the model will cite sources on every answer." The full source list always comes back; the in-text markers do not.

EngineCan you get citations?How
ChatGPTYesAPI web_search returns url_citation; UI scraping breaches terms
ClaudeOnly with web search invokedNo source URLs in a default chat turn
GeminiConditionallyGrounding returns groundingChunks (source uri and title) and groundingSupports (text segment to source) whenever the model decides to search
AI Overviews / AI ModeNo API at allSearch Console report, impressions only
PerplexityYesSonar API returns citations natively
Copilot / BingYesBing Webmaster AI Performance report, incl. Citation Share
GrokOnly when search tools fireCitations come from the API's Live Search; grok.com sign-in is X, Google, Apple or email
A MODEL, NOT A MEASUREMENTWho will tell you they cited youWhether a citation is retrievable, and by what route.PerplexitySonar API returns them nativelyYESChatGPTAPI web search tool returns url_citationYESCopilot / BingBing Webmaster AI Performance reportYESGeminiGrounding returns source chunks when it searchesCONDITIONALGrokOnly when its search tools actually fireCONDITIONALClaudeNothing unless web search is invokedCONDITIONALAI Overviews / AI ModeNo API. Search Console impressions onlyNO API
Three engines you can measure directly, three only when they choose to search, and one you cannot measure at all.

Google's Generative AI performance report launched 3 June 2026. Read its limits: impressions only, no clicks, CTR or queries; it blends AI Overviews and AI Mode into one impressions figure you cannot split, with Discover reported separately in its own generative AI report; it reached a subset of sites, backfilled to about 18 May 2026; and AI Mode follow-ups count as new queries. Bing's AI Performance report hit public preview in February 2026 and expanded on 16 June. Beyond Google and Microsoft, no first-party citation dashboard is open to an ordinary site owner. The nearest exception is gated: Perplexity's Publishers' Program partners "receive data analytics to help track trends and content performance", though Perplexity has never specified what that contains.

Share of voice survives the variance problem because you and your competitor are measured by the same flawed instrument on the same day. The absolute number stays wrong. The ratio stays useful.

Microsoft's Citation Share is the only first-party competitive metric that exists: "the percentage of citations attributed to your site out of all citations shown across all sites for that same grounding query." Everything else is your own sampling or a vendor's undisclosed composite.

Two rules for reading it. A single-run change is almost always noise, since on the resampled subset the repeat itself accounts for roughly 35% of the variance, so report only what persists across five runs and several weeks. As Pixis puts it, a single snapshot is noise and the slope is the signal. An absence needs the same discipline: with 50% to 80% answer consistency, one "not mentioned" means nothing.

There is no universal benchmark. Read it against your own baseline and category.

Ahrefs Brand Radar has the largest real-prompt corpus, though the vendor's own figure for it is inconsistent across its own pages, marketed variously as 320 million, 300 million and 150 million prompts, with third-party reviews citing 199 million to 400 million. Treat it as a moving vendor-stated figure, not a datapoint. Independent testing has flagged gaps between what it reports and what manual inspection finds.

While we are on circulating numbers: Profound's $1 billion valuation is company-stated, and I had this the wrong way round. It announced a $96M Series C at a $1B valuation on 24 February 2026, led by Lightspeed Venture Partners with Sequoia, Kleiner Perkins and others participating, taking total funding past $155M. Older write-ups still quote the August 2025 Series B, $35M led by Sequoia at a $58.5M total, when the company declined to disclose a valuation and had previously said it was valued above $100 million. Category "average price" claims are equally soft: one tracker says $337 a month, another an $89 median.

The test that would settle this has been half-published, twice. GTECH graded 12 platforms against a manual baseline and published the divergence, 67% to 94%, but would not say which tool was which. Visiblie ran eight named tools on the same 50 prompts across four models for 30 days and put every one between 94.3% and 97.3%, a suspiciously narrow spread from a vendor with a tool in the race. Nobody has published the version that is both named and disinterested: fix one prompt set, one brand and one window, push it through Tool A, Tool B and a manual logged-out baseline, and publish the divergence. Until someone does, treat any two vendors' scores as non-comparable.

The fees that matter are per search, not per token. OpenAI's web search tool costs $10 per 1,000 calls plus retrieved content as input tokens, on GPT-5.6 Terra's standard rates of $2 and $12 per million, noting that the bare gpt-5.6 alias routes to a costlier tier. Perplexity Sonar runs $1 per million tokens with a $5 to $12 request fee per 1,000; Sonar Pro is $3 and $15 with a $6 to $14 fee. Grok is $2 and $6 per million on grok-4.6, with Web and X Search each $5 per 1,000 calls. Anthropic's web search tool is $10 per 1,000 searches on the Claude API plus token costs for search-generated content. Google's Grounding with Google Search is $14 per 1,000 requests on Gemini 3 after a free monthly allowance.

The arithmetic on 240 calls, assuming a short answer and default search settings per call: forty ChatGPT calls run about $0.80, of which $0.40 is search fees alone; forty Perplexity Sonar Pro calls about $1.30, mostly output tokens, because Perplexity bills retrieved content through the request fee rather than as input; forty Grok calls about $0.60. That is $2.70 for the three engines with a priced search API. On API fees alone the full pass stays in single digits.

For a real figure with the method attached, here is what my own runs cost. The 40-query baseline across six engines, 240 measurement points, cost $15.90 on 24 August 2026. The teardown of 36 competitor pages cost $0.02. The question-mining pass cost $0.65. The gap between that and the modelled few dollars is Apify compute and actor per-result fees on the engines I could not reach through a priced search API. That is $16.57 for the entire dataset this article is written from. I have not found another page in this category that publishes a single cost figure, which is why I am publishing mine.

PUBLISHING MY OWN NUMBER, SINCE NOBODY ELSE DOESThe whole dataset cost $16.57Actual spend, 24 August 2026.40-query baseline, six engines$15.90Question-mining pass$0.65Teardown of 36 competitor pages$0.02Total$16.57Paraphrase Labs, 24 Aug 2026. Against $95/mo for the cheapest three-engine tracker and $399/mo for Profound's three-engine tier.
The entire dataset this article is written from cost $16.57. The cheapest three-engine tracker is $95 a month, and Profound's three-engine tier is $399.

Break-even is not close on cash. It tips the other way the moment you price your engineer's maintenance time, or the moment you need real-prompt demand data and multi-year baselines you cannot generate yourself.

What tools track brand vs competitor mentions in AI answers?
Dedicated tools (Profound, Peec AI, Otterly.AI) and SEO-suite modules (Semrush at $99/mo standalone, Ahrefs Brand Radar at roughly $828 to $1,148/mo all-in for all models, or from about $328/mo for a single-model index). All track competitor share of voice. None publish a full reproducible methodology, so read their numbers as internal trends, not cross-comparable truth.
How do I compare my AI visibility against competitors?
Use one shared unbranded prompt set and compute share of voice, your mentions divided by all brand mentions, per engine, over at least eight weeks. Microsoft's Bing Webmaster Citation Share is the only first-party competitive metric. Everything else is your own sampling.
Can I monitor brand mentions in ChatGPT and Claude?
Yes for ChatGPT, through the API's web_search tool, which returns url_citation annotations; UI scraping breaches OpenAI's terms. Claude is limited: its default chat answers from training data with no source URLs unless web search is invoked, so tracking only works on browsing-enabled answers.
How do I track mentions in AI Overviews and LLMs?
AI Overviews and AI Mode have no citation API. Search Console's generative AI report shows impressions only and blends AI Overviews and AI Mode into a single figure, with Discover in its own separate report. For the assistants use each engine's API: Perplexity, Gemini and Grok return citations, ChatGPT through its web search tool, Claude only when browsing.
Which tool is best for competitor benchmarking in AI search?
No defensible best exists. No vendor discloses methodology and the tools disagree, with one tracker scoring a brand 45 and another 22. Compare on capability and price, then benchmark everyone on one shared prompt set you control.
How many times should I run each prompt?
Five per prompt per engine. A 2026 variance study found that on the subset where prompts were re-run, the repeat itself accounts for roughly 35% of the variation, and that a repeat past the fifth reduces error by only 0.0003. Spend the next unit of budget on another language first, then another engine, rather than a sixth repeat.
Is a single 'we are not mentioned' result meaningful?
No. Brand identity accounted for just 1.5% of the variance in a single response, and answer consistency runs between 50% and 80%, so one absent reading is close to noise. Read the trend across five runs and several weeks instead.

  1. arXiv 2607.13304, variance decomposition of non-determinism in LLM brand answers.
  2. arXiv 2507.21831 (HALC); arXiv 2509.09705; Atil et al., arXiv 2408.04667v2, Sept 2024, on non-determinism.
  3. OpenAI Terms of Use (effective 1 Jan 2026); Perplexity Terms of Service; docs.x.ai pricing and citations, 21 Aug 2026.
  4. Anthropic, Introducing web search on the Anthropic API, 7 May 2025; Citations API, Jan 2025.
  5. Google Search Central, AI features (10 Dec 2025); Generative AI report (3 June 2026).
  6. blogs.bing.com, AI Performance in Bing Webmaster Tools, 16 June 2026.
  7. Apify Store actor listings and pricing, retrieved 30 Aug 2026.
  8. Tool pricing via That Marketing Buddy (5 June 2026), get-ryze, ayzeo, menra, tryanalyze (July 2026).
  9. Profound Series C, $96M at $1B, 24 Feb 2026; Series B via PR Newswire, Fortune, Adweek, 12 Aug 2025.
  10. AirOps State of AI Search 2026; Pixis, seoClarity, Otterly.AI on metric definitions.
  11. Paraphrase Labs AI Visibility Baseline, 24 Aug 2026: 40 B2B buyer questions, six engines, 240 measurement points, 1,515 citations across 1,028 domains. Costs $15.90, $0.02, $0.65.
#geo#aeo#ai-search#measurement#tools