GEO · AEO · SEO · AI · 14 OCTOBER 2025 · 9 MIN READ
Getting found by AI assistants: structure, sources and citations
There is no markup that makes an assistant cite you. What gets quoted is a retrievable page that states a fact somebody else would have to link to you to repeat.
By being retrievable, being quotable, and being the source. Retrievable means the page is indexed, fetchable by the crawlers that feed the assistant, and readable as text without JavaScript doing the work. Quotable means a specific claim sits in a self-contained passage an extractor can lift without needing the three paragraphs above it. Being the source means the claim is yours — a number you measured, a policy you set, a comparison nobody else has written — because a summariser given five pages that all say the same thing cites whichever one it retrieved first, and that is not a game you win. Google states plainly that there are no additional technical requirements for its AI features, and no special schema.org structured data you need to add. Generative engine optimisation is therefore not a markup exercise. It is retrieval plus editorial.
IN SHORT
- Google states that to appear in its generative AI features "a page must be indexed and eligible to be shown in Google Search with a snippet, fulfilling the Search technical requirements. There are no additional technical requirements."
- Google also states that "structured data isn't required for generative AI search, and there's no special schema.org markup you need to add" — so no file or tag buys you a citation.
- Different crawlers do different jobs: OpenAI documents OAI-SearchBot for surfacing sites in ChatGPT search, GPTBot for model training, and ChatGPT-User for pages a user asks ChatGPT to visit live.
- A passage is quotable when it makes sense lifted out of the page — subject named, claim stated, qualifier attached, no pronoun pointing backwards.
- Assistants cite sources, not summaries: the page that reports a number first is the citation, the page that repeats it is training data.
- Google publishes a Generative AI performance report in Search Console; treat anything else claiming to count AI citations as an estimate, not a measurement.
- Blocking a crawler is a business decision with a cost — the pages you exclude cannot be retrieved, and there is no partial credit.
The part that is not editorial: can the thing be fetched at all
Every conversation about generative engine optimisation should start with a boring question, because the answer is often no. Can the assistant get the page?
For Google, the bar is the one that already exists. Its guidance on generative AI features says a page "must be indexed and eligible to be shown in Google Search with a snippet, fulfilling the Search technical requirements," and that "there are no additional technical requirements." The same documentation adds that you do not need to create new machine-readable files, AI text files, or markup to appear in these features. That single sentence disposes of most of what is currently being sold as GEO tooling.
For the assistants that crawl separately, access is a specific decision you have probably made by accident. OpenAI documents its crawlers individually: OAI-SearchBot is "used to surface websites in search results in ChatGPT's search features", GPTBot is "used to make our generative AI foundation models more useful and safe" through training crawls, and ChatGPT-User handles "certain user actions in ChatGPT and Custom GPTs" — a live fetch when someone asks for your page, not automatic crawling. Three different jobs behind three different tokens. A robots.txt rule that blocks the training crawler and leaves the search crawler alone is a coherent position; blocking all of them and then asking why ChatGPT does not mention you is not.
The third failure is rendering. Snippet controls apply here too — Google lists nosnippet, data-nosnippet, max-snippet and noindex as the levers that limit what can appear — and a nosnippet inherited from an old template is an unintentional opt-out. More common on Shopify: the fact a buyer needs lives inside a script-injected app widget, a tab that renders on click, or an image of a specification table. Retrieval sees HTML. If the answer is not in the HTML, the answer does not exist.
What makes a passage quotable
Assume the system lifts one paragraph and shows it beside four others. Write for that. The test is whether the passage survives being cut out of the page: does it name its subject, state the claim, and carry its own qualifier?
Most ecommerce copy fails on the first of those. "It ships in two to three days" is useless out of context — what ships, from where, to where, and what counts as a day. "Orders placed before 2pm UK time ship the same working day to UK mainland addresses" survives extraction, and it is not longer for the sake of it. Every added word is one an extractor needed.
The same discipline applies structurally. A heading that poses the question a buyer typed, followed immediately by the answer rather than three sentences of preamble, gives the extractor an unambiguous unit. Tables beat prose for anything with rows. Numbered steps beat prose for anything sequential. None of that is a trick — it is the same structure that makes a page readable by a person skimming on a phone, which is why it keeps working when the retrieval mechanics change.
- One self-contained claim per paragraph, with the subject named rather than pronouned.
- Qualifiers inside the sentence they qualify: the cut-off time, the region, the plan, the date.
- Headings phrased as the question, answered in the first line beneath them.
- Tables for comparisons and specifications; numbered lists for procedures.
- Dates on anything that changes, so a stale claim is visibly stale rather than quietly wrong.
Being the source, not the summary
This is the part no amount of structure fixes. A summariser handed six pages that say the same thing does not weigh them on quality; it cites what it retrieved. If your page is the seventh restatement of an industry statistic, there is no version of it good enough to be chosen reliably, because nothing about it is yours.
Google's own advice points the same way: "creating content that people find unique, compelling, and useful will likely influence your website's presence in generative AI search," with the emphasis on firsthand work rather than recycled summaries. For a retailer, firsthand is narrower and more available than it sounds.
You own your policies — delivery cut-offs, returns windows, warranty terms, what happens when something arrives damaged. You own your specifications, including the awkward ones: what this fits, what it does not fit, what the actual measured weight is rather than the manufacturer's. You own comparisons between the things you sell, including the ones that talk someone into the cheaper option. And you own whatever your team knows from doing the work — the sizing quirk, the failure mode, the reason the expensive one is only worth it in one specific case. That last category is the one that gets cited, because nobody else can write it.
What you do not own is the general question. A page called "what is merino wool" competes with an encyclopedia. A page called "which merino weight for UK winter commuting, and when cotton is the better buy" has one plausible source, and it is you.
Where these pages live, and why that is a production problem
The pages that earn citations are rarely product pages. They are selection guides, compatibility pages, comparison pages, policy pages written as answers rather than legal text, and the occasional genuinely opinionated piece. A catalogue does not generate any of them.
Which turns an editorial strategy into a throughput problem, and this is where most programmes stall. If publishing a comparison page means a developer ticket, a two-week queue and a deploy, the plan that called for forty of them produces six. The fix is the same one that fixes campaign landing pages: a section library — a set of composable, tested blocks a merchandiser arranges without touching code. Comparison table, spec block, FAQ, step list, citation-friendly answer block. Build the components once, then the constraint is writing, which is the constraint you actually wanted.
It also keeps the structure consistent. A hand-built page gets its FAQ markup right; the fortieth one does not. A block that emits correct markup every time removes an entire class of quiet failure — and while structured data buys you nothing directly in AI features, Google is explicit that it should still match the visible text on the page, which is exactly the rule hand-rolled pages break.
Measuring it, honestly
Google publishes a Generative AI performance report in Search Console for content appearing in its generative AI features on Search and Discover. That is a first-party number, and it is the one to start from.
Everything else is estimation. The tools that promise to tell you how often ChatGPT mentions your brand work by running prompts repeatedly and counting the results, which measures a sample of a non-deterministic system — useful as a directional signal, worthless as a metric with a target attached. Assistants answer differently to different users, in different sessions, with different retrieval. Reporting a percentage from that to a board is inventing a number, and inventing numbers is how this discipline earns its current reputation.
The measurable proxies are the ordinary ones. Are the pages indexed. Are the crawlers allowed. Do the answer passages appear in featured snippets, which draw on the same extraction behaviour. Is referral traffic from assistant domains rising. Do the questions arriving at customer service look like the ones your pages answer.
What we would not spend money on
Three things get pitched constantly and none of them survives contact with the documentation.
A file that declares your content to AI systems. Google states you do not need to create new machine-readable files or AI text files to appear in its AI features. Adding one costs an hour and does nothing; the harm is that it feels like the work, and the actual work is harder.
Schema added specifically for AI. Google says structured data is not required for generative AI search and there is no special markup to add. Keep your Product and FAQ markup because it earns rich results in ordinary search — that is a good enough reason — but no property in it is an AI lever.
Content produced at volume to blanket a topic. The failure is mechanical: more restatements of what is already on the web make you more replaceable, not less. Forty pages nobody would cite is worse than four somebody would, because the four are now harder to find inside your own site.
The honest version of generative engine optimisation is unglamorous. Let the crawlers in, put the facts in the HTML, write the passages so they survive being lifted, and spend the budget on the handful of pages where you are genuinely the source. There is no shortcut in the documentation because the vendors have said, repeatedly and in writing, that there is not one.
Questions this raises
Is there any markup that makes AI assistants cite my site?
Not according to Google. Its guidance states that structured data is not required for generative AI search and there is no special schema.org markup to add, and that no new machine-readable or AI text files are needed. Keep structured data for rich results in ordinary search, and make sure it matches the visible text — but do not buy it as an AI lever.
Should I block AI crawlers?
It is a business decision with a real cost, and the tokens do different jobs. OpenAI documents OAI-SearchBot as the crawler for ChatGPT search features and GPTBot as the training crawler, so blocking training while remaining findable in ChatGPT search is a coherent position. Blocking everything removes you from retrieval, and there is no partial credit for a page an assistant cannot fetch.
Do I need separate content for AI search and for Google?
No, and building two versions is how sites end up with duplicate thin pages. The structure that makes a passage extractable — a question as a heading, the answer immediately below it, qualifiers inside the sentence — is the same structure that wins featured snippets and reads well on a phone. One well-structured page serves all three.
How do I know whether my pages are being cited?
Search Console has a Generative AI performance report covering Google's generative AI features on Search and Discover — start there, because it is first-party. Tools that sample prompts against ChatGPT or Perplexity give a directional signal at best; they measure a non-deterministic system by repetition, so treat their percentages as estimates and never as a target.
Do product pages get cited?
Rarely, and not for the questions that matter. A product page answers "what is this", which an assistant can answer from the catalogue data it already has. The pages that get cited answer "which of these should I buy, and when should I not" — selection, compatibility and comparison pages that a catalogue never generates on its own.
Does publishing more content improve AI visibility?
Volume works against you here. A summariser choosing between pages that all say the same thing cites whatever it retrieved, so each additional restatement makes you more interchangeable. Four pages stating something only you know beat forty that restate the category — and the forty also bury the four inside your own site.
NEXT STEP
Free store audit
A senior Shopify engineer reviews your storefront, theme performance and checkout, then sends a prioritised list of fixes.
