The Evidence Base

What we work on, and how well the evidence holds up.

GEO is two years old and most of what is sold under the name has never been tested. This page lists every tactic we work on, grouped by how strongly the evidence supports it, with the primary source against each. The last group is the tactics we will not sell you, including three we used to describe more warmly than the data deserved.

Mechanically established

Measured

Directly observed in crawler logs or stated by the vendor. If one of these is wrong on your site, nothing else you do matters.

  1. 01

    No major AI crawler executes JavaScript

    Across more than 500 million crawler fetches, GPTBot, ClaudeBot, PerplexityBot and others issued a request and read the returned HTML without running a single script. They sometimes download JavaScript files and never execute them. The exceptions are Google's AI surfaces, which inherit Googlebot, and Microsoft Copilot, which inherits Bingbot. AppleBot also renders.

    If your content is assembled in the browser, ChatGPT, Claude and Perplexity see an empty shell. Disable JavaScript, reload your homepage, and look at what is left.

    Source: Vercel and MERJ, The rise of the AI crawler
  2. 02

    Blocking an answering crawler removes you from that surface entirely

    Two different kinds of bot get confused constantly. Training crawlers build background knowledge: GPTBot, ClaudeBot, Google-Extended, Applebot-Extended. Answering crawlers fetch pages so you can be cited live: OAI-SearchBot, Claude-SearchBot, PerplexityBot, Bingbot, plus the user-initiated Claude-User and Perplexity-User.

    You can block training and still appear in answers. Block the answering bots and you are gone from that surface. Google-Extended is the common trap: it does not affect AI Overviews or AI Mode eligibility. Check your CDN as well as robots.txt, because Cloudflare has blocked AI crawlers by default since 1 July 2025 and a lot of sites are blocked without anyone having chosen it.

    Source: operator documentation from OpenAI, Anthropic and Perplexity. Ours is public at /robots.txt.
  3. 03

    AI crawlers hit dead URLs at four times Googlebot's rate

    404 rates ran to 34.8 percent for ChatGPT's crawler and 34.2 percent for Claude's, against 8.2 percent for Googlebot. The gap is stale URLs and redirect chains that Google has long since re-learned and the AI crawlers have not.

    Unglamorous, cheap, and the highest-confidence fix on this page.

    Source: Vercel and MERJ, The rise of the AI crawler
  4. 04

    ChatGPT stores one short snippet per page, taken from around the H1

    The snippet is roughly 200 characters, query-independent, and frozen at index time. The same extract comes back whether the question was about your pricing or your location, and the H1 appeared in it about 84 percent of the time.

    That makes the paragraph immediately under your H1 the highest-value text on the page. It should say what the business is, where it is, and who it is for, in plain words and without a preamble.

    Source: Search Engine Land, Inside ChatGPT's retrieval stack
  5. 05

    Google names exactly two conditions, and no more

    A page must be indexed and eligible to show with a snippet, meeting the Search technical requirements. And, in Google's words, “a site must be included in Search generative AI features in Search Console to be eligible for display in generative AI features on Google Search”. That second condition is a toggle in Search Console, and it is easy to miss.

    Beyond those two, nothing. No special structured data, no new files: “Structured data isn't required for generative AI search, and there's no special schema.org markup you need to add.” Conversely, nosnippet, max-snippet:0 and noindex will remove you from those features.

    Corrected 11 September 2026. This entry previously said there was nothing to do beyond being indexed and snippet-eligible, which dropped the Search Console condition Google states in the same paragraph. We publish our corrections rather than quietly editing them.

    Source: Google Search Central, AI optimization guide

Good correlational evidence

Correlational

Large datasets pointing the same way, but none of it proves causation and the study authors say so themselves. We spend real budget here, and we tell you it is correlational while we do it.

  1. 06

    Off-site mentions track AI visibility about three times more strongly than backlinks

    Across roughly 75,000 brands, branded web mentions correlated with AI Overview visibility at 0.664. Backlinks correlated at 0.218. Separately, an analysis of more than 25 million citations across ChatGPT, Claude and Gemini put earned media at about 84 percent of AI citations and paid or advertorial placement at 0.3 percent.

    The Ahrefs authors state explicitly that correlation is not causation. We agree, and we still treat this as the highest-leverage work available, because two independent datasets point the same way and nothing points the other way. In practice it means being included in a credible roundup is worth more than publishing your own.

    Sources: Ahrefs, brand mentions and AI Overview visibility · Muck Rack, What is AI reading?
  2. 07

    Being cited and being recommended are different outcomes

    Across 3,981 domain appearances from 115 prompts on four engines, 61.7 percent of citations produced no brand mention at all. ChatGPT cited a source 87 percent of the time and named a brand 20.7 percent of the time. Gemini ran close to the inverse. Comparison-shaped queries produced 2.4 times more brand mentions than informational ones.

    If you sell something, the mention matters more than the link, and that changes what is worth writing.

    Source: Semrush and Kevin Indig, the ghost citations study
  3. 08

    The informational blog post lost the most, and it lost it to AI summaries

    A browsing-tracker study of 900 US adults across 68,879 Google queries found users clicked a traditional result on 8 percent of searches carrying an AI summary, against 15 percent without. One percent clicked a link inside the summary. Session abandonment rose from 16 percent to 26 percent. Separately, 99.2 percent of the keywords that trigger an AI Overview are informational in intent.

    Put those together and the casualty is specific: the ranked informational article that used to capture the click. Not SEO as a whole.

    Sources: Pew Research Center · Ahrefs, 300,000 AI Overview keywords

Sold widely, tested, did not hold up

Debunked

We will not bill you for any of these. Three of them appeared on this site earlier and were removed on 10 September 2026 when we checked our own copy against the sources.

  1. 09

    Schema markup does not cause AI citations

    A matched difference-in-differences study tracked 1,885 pages that added JSON-LD between August 2025 and March 2026 against roughly 4,000 control pages with comparable prior citation levels. AI Overview citations fell 4.6 percent, a statistically significant decline. AI Mode rose 2.4 percent and ChatGPT rose 2.2 percent, both indistinguishable from noise. The authors concluded there was no major uplift on any platform.

    The widely quoted stat that cited pages are more likely to carry schema is real and is a confound: sites that add schema also fix their technical SEO, build links and refresh content. Keep schema for conventional rich results and for disambiguating your business from a similarly named one. Do not budget against it as a visibility lever.

    Source: Ahrefs, does schema markup drive AI citations?
  2. 10

    llms.txt does nothing for AI visibility

    Google's AI optimization guide, updated 15 June 2026, states that you do not need to create machine readable files, AI text files, markup or Markdown to appear in Google Search including its generative AI capabilities, because Search itself does not use them. No other major provider has committed to reading one. An audit of 137,000 sites found 97 percent of published llms.txt files had never been requested by an AI crawler at all.

    We still publish one, because it costs nothing and writing it forces you to state your facts plainly. That exercise is the entire value. The file is not.

    Sources: Google Search Central · Ahrefs, 137,000 sites analysed
  3. 11

    speakable schema is not a GEO feature

    It never left a limited beta for news publishers tied to Google Assistant, and there is no evidence it influences any AI citation surface. Every article on this site emitted it until 10 September 2026, which was inconsistent with what we were telling readers. It has been removed.

    Source: Google's own speakable documentation, which still describes it as a limited news-publisher beta
  4. 12

    FAQPage markup does not win AI citations, and no longer earns a rich result either

    FAQ rich results were retired from Google Search on 7 May 2026, with the deprecation notice added to the structured data docs that day and Search Console reporting removed the following month. The markup remains valid and harmless. It is not a shortcut to being cited.

    The habit underneath it does help: put a direct answer in the first sentence under a question heading. That is an argument about writing, not about JSON-LD.

    Source: Google Search Central, FAQ structured data deprecation notice, 7 May 2026
  5. 13

    The "40 percent visibility boost" is not a client outcome

    The figure comes from the Princeton GEO paper by Aggarwal and colleagues. It is a maximum recorded on a synthetic Position-Adjusted Word Count benchmark, with a fixed context and few competing sources. It is not an average, not discoverability, and not traffic.

    It has never appeared on this site. It is listed here because it is the most repeated number in the field, and because an agency quoting it at you as a promised result has told you something useful about itself.

    Source: Aggarwal et al., GEO: Generative Engine Optimization, arXiv 2311.09735, KDD 2024

How we measure, and what we cannot tell you

The honest limits, stated once, so you can hold us to them.

  1. 14

    Nothing you can submit gets you in

    There is no submission form and no verification file, and anyone selling you a submission service is selling you nothing. The only inputs you control are whether your pages can be fetched and read, and what independent sources say about you.

    What does exist, and is worth opening: Google rolled out a Generative AI performance report in Search Console from June 2026, worldwide by 31 August 2026. It reports impressions in AI Overviews and AI Mode by page, country, device and date. It carries no clicks, no click-through rate, no position and no queries. Bing Webmaster Tools' AI Performance report is richer, showing Copilot citations and the grounding queries the model generated. Both are free.

    Corrected 11 September 2026. This entry previously said there was no Search Console for AI answers. There now is one, with the limits described above. We publish corrections rather than quietly editing them.

  2. 15

    One run of a prompt is noise, not a measurement

    Generated answers vary between identical runs. Sampling a prompt five times gives a margin of error wide enough to swamp any change you made. We sample each prompt repeatedly per period, logged out, on a fixed locale, with a frozen prompt set and a frozen methodology, and we report the uncertainty alongside the number.

    Most AI-visibility dashboards run each prompt once. That is why they move so much.

  3. 16

    We will not give you a timeline for the answer changing

    There is no index status to check and no platform publishes its retrieval cycle. We will give you dates for the implementation work, because that is ours to schedule. What happens afterwards, we sample and report. Treat any specific timeline promise, ours included, as a sales tactic.

This page changes when the evidence does. It was last revised on 11 September 2026, when every figure on this site was traced back to a primary source and the three tactics in the debunked group were removed from our own copy. If you find something here that a newer study contradicts, tell us and we will correct it. We apply the same standard to our own work in the teardowns, where the method and the raw evidence are published alongside the findings.

Questions this page gets asked

No. A matched difference-in-differences study of 1,885 pages that added JSON-LD against roughly 4,000 control pages found AI Overviews down 4.6 percent, AI Mode up 2.4 percent and ChatGPT up 2.2 percent, and concluded there was no major uplift on any platform. Google separately states that no special structured data is needed for AI Overviews or AI Mode. Schema is still worth deploying for conventional rich results and entity disambiguation.

No. Google's AI optimization guide states that you do not need to create machine readable files, AI text files, markup or Markdown to appear in Google Search including its generative AI capabilities, because Google Search does not use them. An audit of 137,000 sites found 97 percent of published llms.txt files had never been requested by an AI crawler.

Server-rendered HTML. Across an analysis of more than 500 million crawler fetches, no major AI crawler executed JavaScript. GPTBot, ClaudeBot, PerplexityBot and others fetch the page and read the returned HTML without running scripts. Only Google's AI surfaces via Googlebot and Microsoft Copilot via Bingbot render JavaScript. Content assembled in the browser is invisible to the rest.

The Princeton GEO paper by Aggarwal and colleagues. It is a maximum recorded on a synthetic Position-Adjusted Word Count benchmark, with a fixed context and few competing sources. It is not a client outcome, not an average, and not a measure of discoverability or traffic. Treat any agency quoting it as a promised result as a warning sign.

Get in touch

Want to know which of these your site actually fails?

Book a 30 minute call. We will fetch your site the way the crawlers do, read your robots.txt and CDN rules, and tell you which of the first five items is broken. If none of them are, we will say so.