How-To
How to Check Whether AI Chatbots Mention Your Brand
Most companies have never checked whether AI assistants name them. They track rankings weekly and have no idea what ChatGPT says when someone asks who supplies their product.
That gap matters because the assistant answer is often the whole interaction. Pew Research found people click a link inside an AI summary in about 1 percent of visits, so the mention is the impression. Across 900 million weekly ChatGPT users, being absent from those answers is a real cost that no report currently surfaces.
The measurement is free and takes about ninety minutes to set up. Here is how to do it so the numbers mean something.
The Challenges That Make This Harder Than It Sounds
- Answers vary by run. Ask twice, get two different lists. A single check tells you almost nothing.
- Your own account lies to you. Memory and personalization mean the assistant already knows who you are, so it names you when a stranger would not see you at all.
- Phrasing changes everything. Winning one question says little about the adjacent twenty.
- There is no baseline to compare against. Without your own history the first result is a number with no meaning.
Step 1: Build the Question Set
Write fifteen to twenty questions a buyer genuinely asks in the two weeks before choosing a supplier. This set is the whole method, so it is worth doing carefully.
Cover four types. Category discovery: who makes this, who supplies this in my region. Comparison: X versus Y, alternatives to the market leader. Constraint: who handles small orders, who ships internationally, who works with my industry. Direct: what does your company do, is your company any good.
Use the language your customers use, including the awkward phrasing. Buyers do not search in your marketing vocabulary, and the whole point is to see what they see.
Freeze the list once written. Changing questions between runs destroys the comparison, which is the only thing that makes this useful.
Step 2: Run It Cleanly
Use a logged out or temporary session on each assistant, with memory and personalization disabled. Paste each question exactly as written. Do not follow up, do not clarify, do not nudge. You are simulating a stranger, not conducting a conversation.
Record four things for every question: whether you were named, which competitors were named, the order they appeared in, and how you were described. That last field is the one people skip and later wish they had, because a wrong description is a separate and very fixable problem.
Run every assistant your market actually uses. For most businesses that means ChatGPT, Google's AI results, and one or two others depending on the audience.
Step 3: Score It
| Metric | How to calculate | What it tells you |
|---|---|---|
| Citation share | Questions naming you, divided by total questions | Your overall presence in the category |
| Position rate | How often you appear first among named companies | Whether you are the default or an afterthought |
| Competitor share | The same calculation for each rival | Who currently owns the category answer |
| Description accuracy | Answers describing you correctly, as a percentage | Whether the model understands what you do |
The first run is a baseline, not a verdict. Its value appears in month three when you can see direction.
What the Results Usually Tell You
Named nowhere. Either the assistants cannot crawl you, or nothing on your site makes a checkable claim worth quoting. Check crawler access first, because that failure is cheap to fix and invisible otherwise.
Named but described wrong. This is an entity problem, not a content volume problem. Your site is probably vague about what you actually do, or inconsistent across the directories and profiles that models corroborate against.
Named only on branded questions. The models know you exist but do not associate you with the category. You need topical depth on the questions buyers ask before they know your name.
Named alongside one competitor repeatedly. That company has become the category reference. Study what their pages do that yours do not, which is usually specificity rather than volume.
Keeping the Method Honest
Three habits separate a measurement you can act on from a spreadsheet that quietly misleads you.
Never change the questions mid-programme. The temptation to swap in a question you now rank for is strong and it destroys the series. If you must add questions, add them as a separate second set and keep the original intact.
Record the raw answer text, not just a yes or no. The descriptions are where the useful detail lives. A model calling you "a regional supplier" when you serve three continents is a specific, fixable problem that a binary score would have hidden completely.
Run it at roughly the same point each month. Assistants update continuously, and comparing a first-of-month run against an end-of-month run adds variance you will misread as progress.
One more caution about tools. Several vendors now sell an AI visibility score presented as though it were measured. Nobody outside the model providers has access to citation logs, so those numbers are modelled estimates at best. A spreadsheet of answers you actually observed is less impressive and considerably more true.
Pair It With Search Console
Citation share on its own is a leading indicator. Branded search volume is where it shows up commercially.
When assistants begin naming you, a share of those readers go and search your company name. That arrives in Search Console as rising branded impressions, and it typically moves before revenue does. Watching both together turns a soft metric into something you can defend in a budget conversation.
Set the baseline this month even if you plan no other changes. Six months from now the question will be what happened, and only a recorded starting point can answer it.
A final note on expectations. Citation share moves in steps rather than smoothly, because it depends on when models refresh and what they retrieve. A flat month after real content work is normal and not evidence the work failed. Judge this over a quarter, the same way you would judge any content programme.
If you want the question set built properly for your category and the first run interpreted against your competitors, tell me what you sell and I will put the baseline together.
Frequently Asked Questions (FAQs)
Do I need a paid tool to track AI mentions?
No. A spreadsheet and a fixed question set run monthly gives you the trend that matters. Paid tools mainly automate the running and add competitor tracking, which is worth paying for once the manual version proves useful.
Why do I get different answers each time I ask?
Generative responses vary by design, and personalization, memory and live retrieval all add variance. That is why a single check is meaningless and a fixed question set run repeatedly is not.
Should I log out before testing?
Yes. Use a logged out or temporary session with memory and personalization off. Otherwise you measure what the assistant has learned about you specifically, not what a prospective customer sees.
What is a good citation share?
There is no universal benchmark, because it depends entirely on how crowded your category is. Judge yourself against your own baseline and against the specific competitors who keep appearing instead of you.
How often should I run the check?
Monthly is the right cadence for most businesses. Weekly produces noise you will misread as movement, and quarterly is too slow to connect a change in results to the content work that caused it.