
Pew Research Center tracked 68,879 real Google searches from 900 US adults in March 2025. On result pages carrying an AI summary, users clicked a traditional result link in 8% of visits. On pages without one, 15%. When an AI summary was present, users clicked the sources cited inside it in 1% of visits.
That is the commercial problem in one paragraph. Being the answer is now worth more than being the tenth blue link, and the mechanics of becoming the answer are not what most agencies are currently selling.
There is no separate lever that produces AI citations. What exists is a set of documented requirements (be crawlable, be indexed, be eligible), a set of reasonable practices supported by how retrieval systems work (clear entities, extractable structure, first-party information), and a large body of unproven tactics currently marketed as GEO.
The single highest-return technical action for most sites in 2026 is not adding a file. It is checking that you have not accidentally blocked the crawlers that feed AI answers.
SEO is optimising for retrieval and ranking in a search index.
AEO, answer engine optimisation, describes work aimed at appearing inside a generated answer rather than in a list of links.
GEO, generative engine optimisation, is used more or less interchangeably with AEO, usually with more emphasis on large language model surfaces such as ChatGPT, Claude and Perplexity.
Google's own position is worth quoting accurately because it is unambiguous. In its May 2026 guide, Google states that from Google Search's perspective, optimising for generative AI search is optimising for the search experience, and is therefore still SEO. It also points readers toward its guidance on evaluating third-party SEO advice, which is a pointed piece of signalling.
That does not make AEO meaningless. It makes it a scope question. Google's generative features and the assistant products run on different plumbing, and that difference is where the practical work sits.
Google documents two mechanisms. Retrieval-augmented generation, which it also calls grounding, uses core Search ranking systems to retrieve relevant, current pages from the index, then generates a response from the specific information on those pages with clickable supporting links. Query fan-out generates a set of concurrent related queries to gather additional results. Google's own example: for "how to fix a lawn that's full of weeds", fan-out queries might cover herbicides, chemical-free removal and prevention.
The eligibility rule is concrete. To appear in Google's generative AI features, a page must be indexed and eligible to be shown with a snippet, meeting the Search technical requirements. In addition, the site must be included in Search generative AI features in Search Console. Google also states plainly that meeting every requirement does not guarantee crawling, indexing or serving.
Note what follows from fan-out: Google explicitly warns against creating a separate page for every query variation. Doing that primarily to influence rankings or generative responses falls under its scaled content abuse spam policy. Several GEO methodologies currently recommend exactly this.
A large number of sites have blocked the wrong half of that table. The pattern is familiar: in 2024 or 2025, someone decided the company did not want its content training AI models, and either wrote a blanket robots.txt rule or enabled a one-click "block AI bots" toggle at the CDN. The result is a site that has opted out of training, which was the intention, and also out of being cited in ChatGPT and Claude answers, which was not.
This is the highest-value thirty minutes of AEO work available to most companies, and it costs nothing. Decide the training question and the citation question separately, then write robots.txt to match. A defensible default for a B2B company that wants visibility and does not want to feed training corpora is to disallow GPTBot, ClaudeBot and CCBot while allowing OAI-SearchBot, Claude-SearchBot, Claude-User, ChatGPT-User and PerplexityBot. That is a business decision, not a technical one, and it should be made by someone who understands the trade.
Two caveats. Blocking training crawlers may reduce long-term presence in model knowledge, which is a different surface from live retrieval and is not measurable from your side. And robots.txt is advisory: it governs the bots that choose to honour it.
This is the strongest documented lever, and it is the least technical. Google's guidance is unusually direct: creating content people find unique, compelling and useful will likely influence presence in generative AI search more than anything else in its guide. It contrasts commodity content, which it exemplifies as something like "7 Tips for First-Time Homebuyers", with non-commodity content that carries genuine expert or lived experience, giving as its counter-example a piece about waiving an inspection and what the sewer line revealed.
The practical translation for a B2B company: a generative system summarising common knowledge does not need you, because common knowledge is already in the model. It needs you when you hold something it cannot generate. Original numbers from your own operations, a methodology you actually run, failure cases, real pricing ranges, benchmarks from your own client base. Anonymise what you must, but publish the substance.
This is also the only AEO lever that a competitor cannot copy in an afternoon.
Retrieval systems need to resolve who you are before they can decide to cite you. That means consistency more than markup: the same legal name, the same descriptor, the same locations and the same service vocabulary across your site, your LinkedIn page, your directory listings and any press coverage. Organization schema with accurate name, url, sameAs and description helps machines connect those records. It does not create authority where none exists.
Where a brand is genuinely ambiguous (a common word, a name shared with a larger company, a recent rename), fixing that ambiguity is worth more than any on-page tactic.
A generated answer is assembled from passages. Passages that stand alone get used more easily than passages that depend on three paragraphs above them for meaning. This argues for direct answers placed immediately after the heading that poses the question, self-contained definitions, and tables where a comparison genuinely is tabular.
It does not argue for chopping content into fragments. Google states there is no requirement to break content into tiny pieces, that its systems handle multiple topics on a page, and that there is no ideal page length. Write for a reader who skims, not for an imagined parser. Article 4 in this series covers the heading and answer-first architecture in detail.
Unremarkable and still the most common blocker. Google processes JavaScript with an evergreen version of Chromium, but its own documentation notes that server-side rendering or pre-rendering remains preferable, partly because not every bot executes JavaScript. Assistant crawlers are generally less capable than Googlebot here. If your primary content only exists after client-side hydration, you are relying on the most generous crawler in the set.
Check clean HTTP status codes, a current sitemap, no accidental noindex, and that main content is present in the raw HTML response.
Retrieval systems favour current information for queries where currency matters. Accurate, visible publication and update dates, and genuine updates rather than date-bumping, are worth more than they look. This matters disproportionately for pricing, versions, regulations and anything a reader would describe as "as of".
Uncomfortable but real: a significant share of what assistants cite is not your site. Industry citation studies consistently find community platforms, reference sites and professional networks over-represented in AI answers relative to their share of the open web. Those studies are run by vendors selling AI visibility tooling, their methodologies differ, and their numbers should be read as directional rather than precise. The directional signal is nonetheless consistent enough to act on: where your category is discussed publicly matters.
Google's guidance draws a hard line here. Seeking inauthentic mentions is called out explicitly as ineffective, and its spam systems target it. Participating genuinely in the places your market talks is a legitimate strategy. Buying mentions is not.
Structured data is worth implementing. It is not an AI visibility mechanism, and one widely recommended type has just stopped doing anything in Google.
On 7 May 2026, Google deprecated FAQ rich results entirely. The documentation notice states they no longer appear in Google Search, with the FAQ search appearance, rich result report and Rich Results Test support being dropped through mid-2026. This completes a restriction that began in August 2023, when FAQ rich results were limited to well-known authoritative government and health sites.
FAQPage remains a valid Schema.org type. Unused structured data causes no harm. Other systems may parse it. What it no longer does is earn a Google search appearance for anyone. If an agency is currently selling FAQ schema as an AI visibility tactic, that claim has never been documented and, as of May 2026, its Google-side justification is gone as well.
Google's broader position in the May 2026 guide is that structured data is not required for generative AI search, there is no special schema.org markup to add, and it remains a good idea as part of overall SEO because it supports rich result eligibility.
Keep the distinction clean in client conversations. Eligibility is what valid markup buys you. Visibility is what the search engine decides to display. They are separate, and only one of them is in your control.
For most B2B sites the useful set is Organization, BlogPosting or Article, BreadcrumbList, and Person where a real author with a verifiable public profile exists. That is it.
llms.txt is a proposed convention: a plain-text or Markdown file at the root of a domain that describes the site and points to its most important pages, intended to help language models understand what a site is about.
Where it stands in September 2026:
A defensible recommendation: publish one if it is cheap, keep it accurate, and do not present it to a client as a citation strategy. It is closer to a business card than a ranking factor. If your CMS makes it a five-minute job, do it. If it requires a CDN workaround, spend that hour on crawler access instead.
Google side. The Generative AI performance report in Search Console is the only first-party measurement that exists. Use it. Google also warns directly against third-party tools that promise ranking success or claim to use internal Google metrics, noting that no third-party tool has access to its internal ranking or AI systems.
Assistant side. There is no first-party console. The honest options are server log analysis to see which AI crawlers actually fetch your pages and how often, referral traffic segmentation where assistants pass a referrer, and manual prompt testing on a fixed set of queries recorded over time. Third-party AI visibility trackers are useful for trend direction and should not be treated as measurement of a system they cannot see.
Set expectations properly. Citation share is volatile, varies enormously by engine, and is not directly controllable. A reasonable KPI set is: crawler access verified, indexation and snippet eligibility clean, first-party content published on a cadence, brand entity consistent, and a tracked panel of real prompts reviewed quarterly.
From Google's own mythbusting section, these can be dropped for Google Search purposes:
Add one that follows from the same guide: do not build a page per fan-out query. That is the scaled content abuse pattern.
Does blocking GPTBot stop my site appearing in ChatGPT?Not by itself. GPTBot is OpenAI's training crawler. ChatGPT search visibility depends on OAI-SearchBot, and user-initiated retrieval on ChatGPT-User. The controls are independent, so blocking training does not have to mean blocking citation, provided robots.txt is written precisely.
Do I need different content for AI than for Google?No. Google states that best practices for SEO continue to apply because its generative features are rooted in its core ranking and quality systems, and that you do not need to write in a specific way for generative AI search. What changes is emphasis: content that only restates common knowledge has less reason to be retrieved, because the model already contains it.
Is FAQ schema still worth adding?For Google rich results, no. FAQ rich results were deprecated on 7 May 2026. FAQPage remains valid Schema.org and causes no harm, and other systems may parse it, but no platform documents it as an AI citation factor. Write genuine questions and answers because readers use them, and treat the markup as optional.
All sources consulted on 21 September 2026.
An agency looking for a production partner, or a business with a project to launch? Tell us what you need in websites, design, AI or marketing, and receive a clear, costed proposal tailored to your context.