Most brands that try to track AI search mentions make the same mistake: they run a question once, note whether their name showed up, and treat that single answer as the truth. It isn't. Identical prompts run against the same model can vary by 10 to 34 percentage points from one run to the next, and research on local search queries found that only 35% of domains repeat in AI answers at all between runs, meaning two-thirds disappear and reappear seemingly at random. Tracking brand mentions in AI search properly means treating it less like checking a single search result and more like running a small poll, with sample sizes, repeated runs, and a confidence interval attached to whatever number you report. This guide covers that methodology in full, along with what's actually changed about which engines matter now that ChatGPT no longer holds a majority of AI chatbot traffic on its own.
- Identical prompts show 10 to 34 percentage points of variance run to run. A single check of "am I mentioned" is closer to a coin flip than a measurement, which is why a real tracking process runs the same prompt multiple times before reporting a number.
- ChatGPT's share of AI chatbot web traffic has fallen to roughly 54%, with Gemini at nearly 28%, Claude at 9%, and DeepSeek, Grok, Perplexity, and Copilot splitting the rest. Tracking one engine alone now means missing close to half the conversation.
- Search Engine Land's recommended methodology calls for 250 to 500 high-intent prompts per tracking cycle, reported as a distribution with a confidence interval rather than a single point estimate, closer to political polling than traditional rank tracking.
- Responses are personalized by session, location, and conversation history, so two people asking the identical question can get different brand recommendations, which makes tracking methodology, beyond tracking frequency alone, the thing that determines whether your numbers mean anything.
- DeepSeek alone has around 130 million monthly users, and Microsoft Copilot spans roughly 420 million active users across its surfaces, both large enough that ignoring them in a tracking setup leaves a real, measurable blind spot.
Why a single check is closer to a coin flip
Type the same prompt into the same AI engine twice and there's a real chance you get two different answers, not because the model changed but because the underlying system samples from a distribution every time it generates a response. Researchers studying this specifically found that identical prompts run against the same model produce 10 to 34 percentage points of variance from one run to the next, purely from sampling, before personalization or timing differences even enter the picture. A separate study of local search queries found that only 35% of domains repeated in AI answers between runs, meaning roughly two-thirds of the domains that showed up in one run vanished in the next and were sometimes replaced by different ones entirely.
That single fact reframes what tracking actually requires. A one-time check of "did my brand show up" answers a question about that specific run, at that specific moment, and tells you very little about your actual, underlying mention rate. Treating one answer as a verdict is the single most common mistake in this space, and it's an easy one to make because opening a chat window and typing a question feels like it should produce a reliable result the same way a Google search does.
What "AI search" actually spans now

As of mid-2026, ChatGPT holds roughly 54% of worldwide AI chatbot web-visit share, with Gemini at close to 28%, Claude at about 9%, DeepSeek at 4%, Grok at 2%, and Perplexity and Microsoft Copilot each near 1% of that specific traffic measure. Those last few numbers understate real reach: Copilot spans an estimated 420 million active users once every surface it ships on is counted, well beyond its consumer web traffic alone, and DeepSeek's app alone carries around 130 million monthly users, concentrated heavily in markets a US-centric tracking setup can easily miss entirely.
The practical implication is that a tracking setup built around ChatGPT, Gemini, and Perplexity, the engines that get the most attention in most guides on this topic, still leaves real coverage gaps, particularly if your buyers skew toward Microsoft's ecosystem or a region where DeepSeek has meaningful share. We go deeper on the specific behavior differences between ChatGPT, Gemini, and Perplexity in our engine-by-engine guide; this piece focuses on the measurement methodology that makes any of that tracking trustworthy in the first place, regardless of which engines end up in your panel.
Building a real prompt panel
Search Engine Land's recommended approach treats this closer to political polling than traditional SEO rank tracking: 250 to 500 high-intent queries run daily or weekly, sampling a distribution of responses rather than checking one question and calling it done. That volume sounds like a lot next to the 20 or 30 queries a manual spot check typically uses, and it is, deliberately, because a panel that size actually captures the variance described above instead of getting fooled by it.
Build the panel with intentional variable control rather than however many questions come to mind in one sitting. Cover definitional questions, comparison questions, alternative-to questions, and buying-decision questions in a deliberate mix, since over-representing one intent type skews the whole panel toward whatever that intent happens to favor. Avoid a panel built entirely from your own team's guesses about what buyers ask; pull real phrasing from Reddit threads and review sites where people describe their actual research process in their own words, since that phrasing rarely matches the clean, formal questions a marketing team would draft from scratch.
Controlling for personalization
Responses are personalized by session context, location, and conversation history, which means two people typing the identical prompt can land on genuinely different brand recommendations without either answer being wrong. A tracking setup that doesn't control for this is measuring personalization noise as much as it's measuring anything about your actual brand visibility.
Run every prompt from a fresh, logged-out session where the platform allows it, since a signed-in account with history skews results toward what the model has inferred about that specific account rather than a neutral answer a new user would see. Fix the location setting for every run rather than letting it default to wherever the request happens to originate, since a category question can pull different local competitors depending on assumed geography. And run the full panel in one sitting per cycle instead of spreading it across days, since routine model updates between runs introduce a second source of variance that has nothing to do with your actual visibility and everything to do with when you happened to run the check.
Reporting a number you can actually trust
Report a range, not a single figure. The polling-style framing suggested by Search Engine Land's methodology puts it plainly: instead of stating "we appear in 42% of responses," state "we appear in 42% of responses, 95% confidence interval 38 to 46 percent." That range tells a stakeholder whether a month-over-month shift from 42% to 45% reflects a real change or normal statistical noise sitting well inside the margin either measurement already carries. Without it, teams routinely celebrate or panic over swings that a wider sample would have shown to be meaningless.
This is also where mention rate, citation rate, and category share of voice should each carry their own interval rather than getting collapsed into one blended visibility score. A brand can have a wide confidence interval on raw mention rate while still holding a narrow, reliable lead on share of voice relative to named competitors, and reporting only the blended number hides exactly the distinction a decision-maker needs to see.
Turning this into an operating cadence
A 250-to-500-prompt panel run with repeated sampling across six or seven engines is not something one person keeps up by hand indefinitely, even with the best intentions in the first month. AI visibility tracking built for this specifically automates the repeated-run sampling and the confidence-interval math, so the output lands as a trustworthy range on a schedule instead of a single anecdote someone remembers to check when they have a spare hour. Context-aware monitoring layers sentiment and citation detail on top of the raw rate, so a stable mention percentage that's quietly trending toward negative framing doesn't get reported as good news simply because the count held steady.
Agencies running this across several accounts benefit from locking the panel methodology once and reusing it, since standardized measurement makes a client's month-over-month report comparable instead of rebuilt from a different, informally sized panel every cycle.
Common mistakes
- Treating one run as the truth. With 10 to 34 points of run-to-run variance documented in research, a single check tells you almost nothing about your actual mention rate.
- Building a panel around 20 or 30 guessed questions. A panel that small, and drafted from assumption rather than real buyer language, cannot support the sample size a trustworthy measurement needs.
- Ignoring engines outside the big three. Copilot's 420 million users and DeepSeek's 130 million monthly users are large enough that skipping them isn't a rounding error, it's a real gap.
- Reporting a single percentage with no interval attached. A number with no confidence range invites a stakeholder to treat routine noise as a meaningful trend, in either direction.
- Running the panel from a signed-in, location-unfixed session. Both introduce personalization variance into a measurement that's supposed to represent a neutral buyer's experience.
Frequently asked questions
Do I really need 250 to 500 prompts, or is a smaller panel good enough?
It depends on how precise the resulting number needs to be. A smaller panel, even 50 to 100 well-constructed prompts run with repeated sampling, is a meaningful improvement over a single 20-question spot check and produces a usable directional signal. The 250-to-500 range is what's needed to report a genuinely tight confidence interval suitable for board-level reporting or a paid campaign decision. Start smaller if that's what's realistic, and treat the wider interval that comes with a smaller sample as part of the honest result rather than a flaw to hide.
How often should a tracking cycle run?
Weekly or monthly, depending on how quickly your category moves and how much budget rides on the answer.
Does this methodology replace the manual, open-three-tabs approach entirely?
Not for a first look. A quick manual check across a handful of engines is still the fastest way to get an initial read on whether your brand shows up at all, and it costs nothing but time. The methodology in this guide matters once that initial read turns into an ongoing measurement someone reports on, budgets against, or compares month over month, since that's exactly where the 10-to-34-point variance in single-run answers starts to produce misleading conclusions if nobody accounts for it.
Should DeepSeek and Grok really be part of a mainstream brand's tracking panel?
It depends on where your buyers actually are. DeepSeek's roughly 130 million monthly users skew heavily toward specific regional markets, and Grok's base concentrates around X's existing user base, so neither is a universal must-track the way ChatGPT and Gemini currently are for most Western B2B brands. Check your own audience's actual platform usage before deciding, rather than defaulting to either including or excluding them based on overall market share alone, since a brand with meaningful traction in a DeepSeek-heavy market would be making a real mistake leaving it out of the panel entirely.
AI search mention tracking done well looks a lot more like a polling operation than a search rank checker, with a real sample size, controlled variables, and a confidence interval attached to whatever number gets reported. Start with a panel of even fifty well-built prompts run three times each rather than one hundred run once, and the difference in what you can actually trust from the result will be immediate.
Track AI search mentions with a real methodology behind them
Mentient runs repeated, controlled sampling across ChatGPT, Gemini, and Perplexity, alongside news, social, and review monitoring, so what you report has a confidence interval behind it instead of a single lucky or unlucky run. Start free and see your own AI visibility today.
Start free trial


