AI search
AI Search Optimization Services
Your buyers are no longer running one search on one engine. They are asking Google, which answers above the results. They are asking ChatGPT, which answers and sometimes cites. They are asking Perplexity, which answers and always cites. Each of those systems builds its answer differently and reaches for different sources.
Optimising for one of them and assuming the rest follow is the mistake I see most often. A brand can be well cited in Perplexity, which leans on recent indexed content, and entirely absent from ChatGPT, which leans on different retrieval behaviour and a different training base.
This engagement treats AI search as a set of surfaces with shared foundations and separate behaviours, and measures each one separately.
The surfaces, and how they differ
Google AI Overviews sit above the organic results and draw heavily on pages that already rank, which means classic SEO strength still transfers here. This is the surface where a strong existing position is most reusable.
ChatGPT answers from a mix of training data and live retrieval. Brand mentions across the wider web matter more here, and recency matters less, which makes it the slowest surface to move and the most valuable to hold.
Perplexity retrieves and cites aggressively on almost every answer. It is the most responsive surface to fresh, well-structured content, and usually the first place a restructuring effort shows up.
Gemini and Copilot each blend their own index with their model, and both reward consistency between what a site says about itself and what third parties say about it.
Why one question set matters
The discipline that makes this measurable is a fixed question set, written once and re-run unchanged. Twenty to forty questions covering the decisions your buyers actually make, run across every surface, recorded the same way each time.
Without it, AI visibility reporting collapses into anecdote. Someone asks a chatbot something, gets a good answer, and concludes the programme is working. Someone else asks a slightly different question, gets nothing, and concludes it is not. A fixed set removes the sampling problem and turns the whole thing into a number that moves.
What transfers between surfaces and what does not
Structural work transfers almost completely. A page rewritten so its claims are extractable becomes easier to quote everywhere at once, which is why it is always the first thing done.
Corroboration transfers well, because every one of these systems cross-checks against third parties.
What does not transfer is surface-specific behaviour. Ranking strength in Google helps in AI Overviews and does comparatively little for ChatGPT. Freshness helps in Perplexity and does little in a model answering from training data. This is why the programme is sequenced surface by surface after the shared work is done, rather than treating all five as one target.
Reading the report without fooling yourself
Three habits keep this measurement honest, and all three are easy to lose.
Never conclude from one run. Model output varies between identical prompts. A brand that appears in four of five runs and a brand that appears in one of five look identical if you only ask once. The report shows how many runs named you, not whether a single test did.
Separate being named from being described correctly. Appearing in an answer that describes a service you retired two years ago is not a win, and it is common on businesses that have repositioned. The report tracks accuracy as a distinct column.
Watch the citation sources, not just the mentions. If assistants keep building answers in your category from one directory you are not listed in, that single line is worth more than the rest of the report.
What you need to supply
Very little, and that is deliberate. A list of the questions your buyers actually ask, which usually comes out of one conversation with whoever answers the phone. The competitors you consider real, since the ones a model names are often not the ones you would list. And access to whoever can approve changes to page copy, because the fixes are worthless unless they ship.
No analytics access is required for the baseline itself, since it measures external systems rather than your site.
Where most people start
Almost everyone starts with the audit, because committing to a cross-surface programme before seeing a baseline is guesswork.
The baseline usually settles the priority on its own. It shows which surfaces your buyers use, which ones already name you, and which competitors are being cited instead. That is normally enough to decide whether this needs a programme or a short list of fixes.
If you already have a baseline from elsewhere, bring it and we can start from the sequencing instead.
What the engagement includes
- A cross-surface baseline. Your question set run across AI Overviews, ChatGPT, Perplexity, Gemini and Copilot, with citations and competitor mentions recorded per surface.
- Shared structural work. The page and schema changes that improve extraction on every surface at once, done first because the return is broadest.
- Surface-specific sequencing. Targeted work per surface after the shared layer, prioritised by which surface your buyers actually use.
- Competitor citation mapping. Who is being named instead of you, on which questions, and which of their sources the answer was built from.
- Monthly re-measurement. The same question set, unchanged, so the trend is comparable month to month.
How the work runs
- Write the question set. Twenty to forty real buying questions, fixed for the life of the engagement so measurements stay comparable.
- Baseline every surface. Run the set across all five surfaces and record answers, citations and competitors.
- Do the shared work. Structure, specificity and schema, which improve every surface simultaneously.
- Sequence by surface. Then work the surfaces individually, starting with the one your buyers use most.
- Re-run and report. Monthly, same questions, reported as citations by surface.
Proof
The measurement discipline described here is the same one used in the AI visibility audit, which is where most clients start.
Related
Talk about your situation
The first conversation is short and mostly questions. Get in touch and tell me what you are trying to fix. Or see how this is priced.
Frequently Asked Questions (FAQs)
What is AI search optimization?
The practice of making a business visible inside AI-generated search results across every major surface: Google AI Overviews, ChatGPT, Perplexity, Gemini and Copilot. It shares foundations with SEO but is measured on citations rather than rankings, and each surface behaves differently enough to need separate tracking.
Which AI search platform matters most for my business?
It depends on your buyers, and the baseline answers it rather than guessing. Business-to-business buyers researching technical purchases tend toward ChatGPT and Perplexity. Consumer and local queries concentrate in Google AI Overviews. The baseline shows where you are absent and where your competitors are named, which usually settles the priority.
Can you optimise for ChatGPT specifically?
Partly. ChatGPT draws on both training data and live retrieval, and training data is not something anyone can edit. What is controllable is the retrieval half and the corroboration that shapes what the model has learned: consistent facts about your business repeated across sources it reads. That moves slowly and holds well once it moves.
How many questions should the question set contain?
Twenty to forty for most businesses. Fewer than twenty and normal variance in model output swamps the signal. More than forty and the monthly re-run becomes expensive without adding much. The set should cover the full buying sequence rather than clustering on one stage.
How often should the question set be re-run?
Monthly is the right cadence for most businesses. More frequently than that and normal model variance dominates the signal, so you end up reacting to noise. Less frequently and you lose the ability to connect a change in citations to the work that caused it. The set itself stays fixed, because changing the questions destroys comparability.