I Ran 30+ GEO Experiments So You Don’t Have To: What Actually Gets You Cited in AI Search | Thomas Peham (brightonSEO San Diego 2026)

Thomas Peham's title slide at brightonSEO San Diego 2026: I Ran 30+ GEO Experiments So You Don't Have To
https://speakerdeck.com/player/689ebfa5cfca4e46b248de06b55577c1

Thomas has published the full deck on Speaker Deck. Slide references below follow it.

TL;DR

  • Google and AI search judge content completely differently. Two sites with 1,000 AI-generated blog posts each were dropped by Google within 14 and 90 days. Months later, AI search engines were citing them more than ever.
  • Every AI engine behaves differently. Schema markup lifted Google AI Overview citations by 1,500%, but ChatGPT, Gemini, and Copilot citations fell. You have to measure per platform.
  • Popularity doesn’t drive AI citations. 40.83% of YouTube videos cited by AI had fewer than 1,000 views. Written content beat video on every AI engine tested.
  • AI forgets slowly. Of more than 20 million cited URLs, 19.3% no longer returned a live page. Pages removed eight months earlier still held a 5% citation share.
  • You don’t need expensive mainstream media to get cited. News publishers with AI licensing deals get more citations on ChatGPT and Copilot, but in 9 of 10 industries studied, unlicensed publishers were cited more.

About the Speaker

Thomas Peham is the CEO and co-founder of OtterlyAI, an AI search monitoring platform. His team runs controlled experiments on generative engine optimization (GEO) and publishes the results openly, so other practitioners can learn from them and challenge them.

For context, OtterlyAI was also a headline sponsor of the event, and most of the data in this talk comes from OtterlyAI’s own platform and experiments.

The Problem: Nobody Has Data

Thomas framed the session around a familiar situation. Your boss wants an AI search strategy. LinkedIn is full of people selling one. Very few of them have data. His response was to spend a year testing the tactics the industry swears by.

He also made a point about measurement before getting into the experiments.

According to OtterlyAI’s own GA4 data on a last-touch basis, Google Search accounted for 35% of conversions, ChatGPT 7.1%, and Claude just 0.1%.

But when customers were asked directly where they first heard about OtterlyAI, the picture changed. By this self-reported measure, Claude accounted for 10.6% of conversions, not 0.1%, and Claude and ChatGPT together brought in 25% of OtterlyAI’s revenue. And the average customer who came from Google spent 15% less than a customer who came from Claude. Customers from ChatGPT also spent slightly less than those from Claude. Web analytics barely registers a channel that, in practice, brings in higher-value customers.

His one-line summary on the slide: AI search is primarily a brand channel that’s also driving revenue.

The Methodology

Every experiment follows the same loop:

  1. Form a hypothesis. Which GEO tactic or myth are we testing, and what result do we expect?
  2. Measure relentlessly. Track AI bot visits, brand mentions, and AI citation rates over a fixed time window.
  3. Publish results. Did the experiment move the needle, and what can other practitioners learn?

Feedback from the community then feeds into the next round of hypotheses.

The measurement combines four data layers: Google Search Console, Bing Webmaster Tools, AI citations tracked in OtterlyAI, and agent analytics showing which AI crawlers visit.

Experiment 1: The AI Slop Sites

This was the experiment the session abstract promised would be “stranger than you’d guess,” and it delivered.

OtterlyAI launched two brand-new websites and published 1,000 AI-generated blog posts on each, created with minimal effort. The only difference was that site B also added an author and an author box. Then they tracked what happened across Google, Bing, and AI search engines.

On Google, both sites died. Site B lost its clicks and impressions after just 14 days. Site A lasted longer, but was effectively gone after 90 days.

Google, Bing and AI citation data compared for two AI-generated test sites

On Bing and AI search, the opposite happened. Bing kept indexing both sites. And AI citations kept climbing, reaching their highest levels in August, months after Google had dropped the sites entirely.

The details:

  • Not every AI engine treated the sites equally. The effect only worked on Perplexity, Copilot, and Claude.
  • Bing, Copilot, and ChatGPT diverged. Bing still indexes the AI-generated content, and Copilot still cites it. ChatGPT rarely does.
  • Crawling isn’t citing. Just because an AI agent crawls a page doesn’t mean it will cite it.
  • Google is inconsistent with itself. AI Overviews cited the content, while the same content was fully deindexed from Google’s organic results.

Experiment 2: Video vs. Blog

OtterlyAI published the same story as a YouTube video and as a written blog post, then tracked which version AI search cited.

  • Written content was detected and cited earlier.
  • Every AI search engine cited the blog version. Not one skipped it.
  • Text is also what makes a video citable in the first place.
  • Perplexity and ChatGPT behaved so differently they were, in Thomas’s words on the slide, like different universes.
  • Don’t bet on Shorts alone. 94% of video citations went to long-form videos.

His recommendation: ship every story twice, as an article and a long-form video, on the same day.

Experiment 3: What YouTube Videos Get Cited?

OtterlyAI analyzed the YouTube videos that AI search engines cite. The findings challenge most assumptions about video SEO:

Data pointFinding
Video views40.83% of cited videos had fewer than 1,000 views
Video likes36% had fewer than 15 likes
Video length50% were shorter than 8 minutes
Title lengthAverage of 19 words
Description lengthAverage of 334 words
Hashtags50.07% of descriptions included hashtags
Channel subscribers35% of channels had fewer than 10,000 subscribers
Channel viewsThe median cited channel had about 2.2 million total views
Channel video count50% of channels had fewer than 41 videos

Two conclusions: popularity metrics like views, likes, and comments don’t matter for AI search, and active channels with more videos simply have more chances of being cited.

Case Study: A New YouTube Channel

Thomas shared results from a brand that launched a fresh YouTube channel with 25 videos. After one month, its share of voice changed as follows:

  • Google AI Mode: +114% (14 to 30)
  • ChatGPT: +83% (40 to 73)
  • Microsoft Copilot: +17% (36 to 42)
  • Perplexity: +15% (54 to 62)
  • Gemini: +3% (29 to 30)
  • Google AI Overviews: -37% (38 to 24)

19 of the 25 videos also ranked first in regular Google Search. High-quality video can clearly move brand visibility, but again, not equally everywhere.

Experiment 4: Ghost Citations

Slide: of 20,092,769 URLs cited by AI search, 19.3% do not return a live page

OtterlyAI checked 20,092,769 URLs cited by AI search engines. 80.7% returned a live page. 19.3% did not, whether because of redirects, missing pages, server errors, or unreachable domains.

Examples included Wikipedia pages that are no longer articles, PDFs (10.66% of the total were already dead), and Reddit threads removed after the platform’s crackdown.

The share of dead citations varied by engine: ChatGPT had the highest at 25.1%, followed by Google AI Mode (23.9%), Copilot (20.0%), Gemini (18.7%), Perplexity (18.0%), Claude (16.7%), and Google AI Overviews with the lowest at 12.6%.

OtterlyAI also removed pages from its own site in October 2025 and tracked their citations. Eight months later, those pages still held a 5% citation share. Thomas’s takeaways:

  • AI builds perceptions from outdated content that you, or someone else, already removed.
  • Your content footprint in AI is bigger than your actual website.
  • Think twice before publishing. AI forgets slower than you think.

Experiment 5: On-Page Tactics

The most surprising on-page result was schema. After OtterlyAI implemented schema markup across more than 2,000 pages in December 2025, its Google AI Overview citations rose by 1,500% by March. But ChatGPT, Gemini, and Copilot citations decreased, and Perplexity showed no impact. Thomas noted that growth in March was seen across the whole search landscape, so part of the lift was likely due to Google algorithm changes.

Summary table of OtterlyAI's on-page experiments for Google and AI search

His summary of on-page experiments:

TacticLearningGoogleAI search
SchemaLifted Google AIO, but ChatGPT citations droppedYesNo
URL lengthURL optimization doesn’t matter for AI searchYesNo
Page typesGuides, blogs, help pages, and home pages are most favorableYesYes
ImagesNo impact on AI search; filenames were hallucinatedYesNo
Alt tagsOnly processed by ChatGPTYesPartially
MarkdownNo measurable impactNoNo
FAQProven results on Google and AI searchYesYes

Experiment 6: Does PR With Licensed Publishers Pay Off?

Many AI companies have signed licensing deals with news publishers. OtterlyAI worked with Press Ranger to test whether those publishers are cited more.

Citations per page, with and without a licensing deal with that AI company:

  • Microsoft Copilot: 11.7 vs. 6.3 (+85%)
  • ChatGPT: 10.2 vs. 6.9 (+48%)
  • Perplexity: 12.7 vs. 13.0 (-2%)
  • Google AI Overviews: 3.8 vs. 4.1 (-9%)
  • Google AI Mode: 3.0 vs. 3.5 (-14%)
  • Gemini: 2.9 vs. 3.3 (-14%)
  • Claude: Anthropic has not publicly disclosed any publisher deals.

The content types cited most from news publishers were best-of and top lists (38.1% for licensed publishers, 22.8% for unlicensed), followed by buying guides and money advice (15.1% and 21.3%).

Chart showing which news publishers AI cites more, by industry

The industry breakdown was the most useful chart for PR teams. Licensed publishers were cited more only in consumer goods (52% vs. 48%). In every other industry studied, unlicensed publishers were cited more, with the gap widest in real estate (84.8%), transportation (84.4%), media and entertainment (81.5%), hospitality (77.7%), and education platforms (72.8%).

His conclusions for PR teams:

  • News accounts for 7.2% of all AI citations, and is the second-biggest source on Claude.
  • News publishers with OpenAI deals do get 48% more citations on ChatGPT.
  • But you don’t need a licensed publisher to get cited. Analyze your industry first.
  • Niche and trade media often beat expensive mainstream media. On a separate slide, trade media beat mainstream media in 15 of 16 industries.
  • The content types cited most are top and best-of listicles, how-to guides, problem-discovery content, and product or service reviews.

His Scorecard: What Works in AI Search Right Now

Thomas opened and closed with the same summary slide of what OtterlyAI’s experiments found.

What works:

  • Press release distribution
  • Local and English content side by side
  • On-page schema markup, for Google AI Overviews
  • YouTube video optimization
  • A presence on Wikipedia
  • Brand visibility on active Reddit communities, with replies
  • LinkedIn Pulse articles
  • External listicles
  • Industry recognition, such as awards and analyst coverage
  • Individual feature pages
  • AI-assisted content generation
  • Clean, structured URLs
  • Blog and editorial articles for top-of-funnel discovery
  • Product and category pages for evaluation and bottom-of-funnel
  • Documentation and guides for product research
  • Alternative and “vs” pages
  • AI memory injection
  • AI videos on YouTube

What doesn’t:

  • AI slop
  • An llms.txt file
  • Markdown files as page variants
  • Schema markup for ChatGPT and Gemini
  • Image alt tags and file names, for AI search
  • URL structure and length
  • Beautiful content design
  • Regular LinkedIn feed posts, which work less well than Pulse articles

What Not to Do

  • Don’t assume what works on Google works on AI search. Schema, URL optimization, and images all showed different results.
  • Don’t report “AI search” as one channel. ChatGPT, Perplexity, Copilot, and Google’s AI features respond differently to the same change.
  • Don’t chase views and likes for AI visibility. Popularity metrics didn’t predict which videos got cited.
  • Don’t rely on Shorts. 94% of video citations went to long-form content.
  • Don’t delete content and assume it’s gone. AI keeps citing removed pages for months.
  • Don’t judge AI’s value by GA4 alone. Last-touch attribution showed Claude at 0.1% of conversions, while self-reported data showed its users spend more.
  • Don’t buy mainstream coverage by default. Check which publishers AI actually cites in your industry.

Next Steps

  1. Set up all four data layers. Connect Google Search Console and Bing Webmaster Tools, and add AI citation and AI crawler tracking.
  2. Add a “where did you hear about us” question. Compare self-reported sources against GA4.
  3. Publish your next key story twice. Release a written article and a long-form video on the same day.
  4. Audit what AI says from old content. Check whether removed or outdated pages are still being cited.
  5. Test one change at a time, per platform. If you add schema or FAQs, measure results separately for each AI engine.
  6. Map cited publishers in your industry. Before a PR campaign, find which outlets AI search cites for your topics.
  7. Ask for the GEO experiment spreadsheet. Thomas offered it free to anyone who sends him a LinkedIn DM or an email.

Personal Takeaways

This was the session I found most interesting of the whole day. I caught part of Thomas’s talk at brightonSEO Brighton in April, and this was an almost entirely new set of experiments. It was the rare kind of talk that answered questions everyone is asking with actual tests rather than opinions.

What stayed with me:

  • This matches what we see in our own business. Until this year, almost all sign-ups at our Japanese language school came through Google Maps. Now almost all of them come through ChatGPT. We know this because the students tell us themselves. That shift came after we strengthened our JLPT keywords and added more JLPT content. That experience is exactly why these experiments felt so practical rather than theoretical. When I saw the slide comparing GA4 with self-reported attribution, it rang true: if we only looked at GA4, we might never have noticed what ChatGPT is doing for us.
  • The AI slop experiment is a warning, not a tactic. It shows that AI engines currently fail to filter low-quality content in ways Google does. That is a gap that will likely close, and anyone publishing AI content at scale should expect the rules to catch up. For content sites that use AI assistance for some formats, the lesson is to keep quality high enough to survive on Google too, not to exploit the gap.
  • The industry chart matters for my own work. In education platforms, unlicensed publishers were cited 72.8% of the time. For a language education business, that suggests trade media and specialist outlets are a better PR investment than chasing mainstream coverage.
  • Ghost citations change how I think about content cleanup. When clients prune or restructure content, outdated information can continue to shape AI answers for months. Redirects and corrections matter more than deletion.
  • The GA4 vs. self-reported attribution slide echoes the whole morning. Skyler Rudolfsky, Baruch Toledano, and now Thomas all showed the same thing from different data: AI influence is real, and web analytics doesn’t see most of it.

A caveat worth keeping in mind: many of these experiments are on a small number of sites, several use OtterlyAI’s own domain, and the tracking comes from OtterlyAI’s platform. The results are directional rather than definitive, which Thomas himself acknowledged by publishing the methodology and inviting others to test.

Related Resources

Written by Ayaka Uchida
CEO, A-Digital Works

This report covers Thomas Peham’s session “I Ran 30+ GEO Experiments So You Don’t Have To: What 200M AI Citations Say Actually Works” at brightonSEO San Diego, September 16, 2026.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top