How to test whether ChatGPT, Claude and Perplexity recommend you
Run a fixed panel of buyer prompts several times a week, track how often you're mentioned and cited, and test page changes against similar unchanged pages.
Updated 28 September 2026 · 7 min read
Write a fixed panel of 30 to 60 prompts your buyers would ask, run each one several times in each assistant every week, and record whether your brand is mentioned and which of your pages are cited. Track the mention rate over time rather than any single answer, because AI recommendation lists rarely repeat. To test a change, edit one group of pages and compare its citation rate with a matched group you left alone.
To test whether ChatGPT, Claude and Perplexity recommend you, build a fixed prompt panel: 30 to 60 questions your buyers would really ask, such as "best invoicing tool for freelancers in the UK". Run each prompt several times in each assistant every week, in fresh sessions, and log two things: whether your brand is mentioned, and whether one of your pages is cited as a source. Your mention rate across hundreds of answers is the number to track. Any single answer means very little, because the same prompt returns a different list almost every time.
To find out whether a change to your site made a difference, don't compare before and after on the whole site. Change a group of pages, leave a similar group untouched, and compare how often each group gets cited over the next few weeks. That's the same page-group method used for Google SEO split tests, applied to AI answers.
Why one-off checks mislead
SparkToro and Gumshoe had 600 volunteers run 12 prompts through ChatGPT, Claude and Google's AI tools 2,961 times. They found there's less than a 1 in 100 chance that ChatGPT or Google's AI returns the same list of brands twice across 100 runs, and roughly 1 in 1,000 that two lists come back in the same order. Their conclusion was that ranking positions in AI answers are close to meaningless, but "visibility % across dozens to hundreds of prompts run multiple times is a reasonable metric" (SparkToro, January 2026).
So a screenshot of ChatGPT recommending you, or not, proves nothing. A mention rate that moves from 18% to 30% over six weeks across 360 answers a week does.
Build a prompt panel
Write prompts the way buyers type them, not the way you'd describe yourself. Cover the stages of a purchase:
| Prompt type | Example (for an invoicing tool) |
|---|---|
| Category | "What's the best invoicing software for freelancers in the UK?" |
| Use case | "How do I chase late invoices without annoying clients?" |
| Comparison | "[Competitor] vs [your brand] for a two-person agency" |
| Alternatives | "Cheaper alternatives to [competitor]" |
| Brand check | "Is [your brand] any good? What are the downsides?" |
Freeze the wording once the panel starts. Changing "best" to "top" halfway through breaks your trend line. Add new prompts as a separate set if you need them.
Run every prompt the same way each time: a fresh chat, logged out or with memory and personalization off, the same language and country, and web search on if the assistant offers it. Record the model name the assistant shows, because a model change can shift results overnight.
What to record for each answer
| Field | Why it matters |
|---|---|
| Mentioned (yes or no) | The main visibility metric |
| Recommended or just listed | Being named as the pick counts for more than being one of eight |
| Your pages cited (URLs) | Tells you which pages the assistant trusts, and feeds page-group tests |
| Competitors mentioned | Shows who you're losing answers to |
| Wrong facts about you | Old pricing or missing features are fixable problems |
Skip position within the list. SparkToro's data says order is close to random.
How many runs you need
A panel of 40 prompts, run 3 times in each of 3 assistants, gives 40 × 3 × 3 = 360 answers a week. If you're mentioned in 18% of them, the standard error is √(0.18 × 0.82 ÷ 360) = 0.020, so the 95% margin of error is about ±4 points. A week-on-week move from 18% to 21% is noise. A move from 18% to 26% that holds for three weeks is probably real.
If you can afford more runs, spend them on repeats of the prompts that matter most to revenue, such as category and alternatives prompts, rather than on dozens of loosely related questions.
How to test a change with page groups
AI assistants read one version of each page, just like Googlebot, so you can't split visitors. You split pages. Outtest runs AI SEO tests this way, as changed pages against similar unchanged pages, and you can run the same design by hand.
- Pick a set of similar pages, for example 30 comparison or use-case pages built on one template.
- Record four weeks of citations to each page from your prompt panel.
- Split the pages into two matched groups of 15, balancing pages that are already cited often.
- Change one group only, for example adding a two-sentence direct answer at the top, a table of sourced numbers, and clear product facts such as pricing and who it's for.
- Keep running the panel for four to six weeks and compare citation counts between groups.
Worked example
A B2B software company runs the steps above on 30 use-case pages, with 360 panel answers a week.
| Citations to the 30 pages | Changed group (15 pages) | Control group (15 pages) | Changed group's share |
|---|---|---|---|
| 4 weeks before (1,440 answers) | 22 | 24 | 48% (22 of 46) |
| 4 weeks after (1,440 answers) | 41 | 27 | 60% (41 of 68) |
A before-and-after view says the changed pages nearly doubled their citations (22 to 41). But the control group rose too, from 24 to 27, perhaps from a model update. If the change had no effect, the changed group would have grown in line with control, to about 22 × (27 ÷ 24) = 25 citations. It got 41.
How sure can you be? Compare the changed group's share of citations before and after. The share went from 0.478 to 0.603. The standard error of that change is √(0.478 × 0.522 ÷ 46 + 0.603 × 0.397 ÷ 68) = 0.095, so z = 0.125 ÷ 0.095 = 1.32, about a 91% chance the change helped. That's promising, not proven. Run it for four more weeks, or add prompts that are likely to cite these pages, before rolling the change out to the control group. As a second signal, check referral visits from the assistants in your analytics for the same two groups of pages.
What tends to get cited
No AI company publishes its ranking rules in full, so treat any list as a set of hypotheses to test. The claims below have public sources.
The assistant has to be able to reach you
Each assistant uses a named crawler for search, and robots.txt or firewall rules can block it:
- OpenAI says sites that opt out of OAI-SearchBot "will not be shown in ChatGPT search answers, though can still appear as navigational links" (OpenAI crawlers).
- Perplexity says PerplexityBot "is designed to surface and link websites in search results on Perplexity" (Perplexity crawlers).
- Anthropic says blocking Claude-SearchBot "may reduce your site's visibility and accuracy in user search results", and blocking Claude-User stops Claude fetching your pages when users ask (Anthropic crawler help).
Check your robots.txt and your CDN's bot settings first. A blocked crawler makes every other test pointless.
Google's AI features follow normal Search rules
Google says there are "no additional requirements to appear in AI Overviews or AI Mode", that you don't need AI text files or new markup, and that a page must be indexed and eligible to show with a snippet. Both features may use "query fan-out", running several related searches to build an answer (Google Search Central).
Ranking for the exact prompt isn't enough
Fan-out means assistants cite pages that rank for related sub-questions. Ahrefs studied 15,000 long-tail queries and found only 12% of URLs cited by AI assistants ranked in Google's top 10 for the original prompt: 8.0% for ChatGPT's in-text citations and 28.6% for Perplexity (Ahrefs, August 2025). For Google's AI Overviews, Ahrefs found 76% of citations came from top 10 pages in July 2025, falling to 38% in a 2026 update across 863,000 results pages (Ahrefs). Pages that answer the follow-up questions around a topic give assistants more to cite.
Sources, quotes and numbers helped in lab tests
The GEO paper (Aggarwal and colleagues, KDD 2024) tested content edits against a simulated AI search engine and on Perplexity. Adding citations, quotations and statistics improved their visibility measure by 30 to 40% in relative terms, while keyword stuffing gave "little to no improvement" (arXiv). The main results came from a simulated engine and 200 Perplexity queries, so test these ideas on your own pages before rewriting everything.
Common mistakes
- Checking prompts from your own logged-in account, where memory and history shape the answer.
- Rewording prompts between weeks.
- Reporting "we rank third in ChatGPT", when positions are close to random.
- Changing the whole site at once, so there's no control group.
- Ignoring wrong facts. If assistants repeat an old price, fix the pages they cite for it and track whether the error rate falls.
Questions people ask
Can I A/B test for ChatGPT like a normal A/B test?+
Not per visitor. ChatGPT, Claude and Perplexity read one version of each page, so you can't show them version A and B of the same URL. Test with page groups instead: change some pages, leave similar pages alone, and compare how often each group gets cited over the following weeks.
How many prompts do I need to track AI visibility?+
Enough that the mention rate stops jumping around. A panel of 40 prompts run 3 times in 3 assistants gives 360 answers a week, and at a 20% mention rate the margin of error is about 4 points either way. Fewer runs means only big changes will show.
Does ranking on Google mean ChatGPT will recommend me?+
It helps but it isn't enough. An Ahrefs study of 15,000 long-tail queries found only 12% of URLs cited by AI assistants ranked in Google's top 10 for the original prompt, and 8% for ChatGPT's in-text citations. Assistants run several related searches and pull from those results too.
Do I need an llms.txt file or special markup to appear in AI answers?+
Google says there are no additional requirements or special optimizations to appear in AI Overviews or AI Mode, and no need for AI text files or new markup. A page must be indexed and eligible to show with a snippet. For ChatGPT, Claude and Perplexity, the first check is that their search crawlers can reach your pages.
Read next
Let Outtest run your split tests
AI agents read your analytics and payments, find where you lose the most money, build the fix and test it. Every test is judged on revenue, not clicks. Plans from $29 a month.