DruxAI

We Checked: Does Google-Extended Actually Crawl the Sites Googlebot Crawls?

Michael ObembeMichael Obembe·August 6, 2026·Via original-research·5 reads
Share

Every site owner has heard of GPTBot by now. Fewer have heard of Google-Extended — the crawler that specifically controls whether Google can use a site's content for Gemini and AI Overviews, separate from regular Googlebot indexing. We wanted to know: for real, working websites, does that crawler actually show up?

We pulled real request-path data from DruxShield across seven live sites we monitor — not a survey, not a sample of self-reported opinions, actual HTTP requests hitting real servers. Here's what the traffic itself says.

Google indexes you constantly. Google-Extended almost never comes.

Across the seven sites, one pattern was stark: five of them have never once been visited by Google-Extended — including sites Googlebot itself visits thousands of times.

SiteGooglebot hitsGoogle-Extended hits
wavebets.com11,9460
naijagist.com2,5070
taxiwi.ca2590
drux.space2000
typemyself.com111
kapturedwithlove.com01

wavebets.com alone has been crawled by regular Googlebot almost twelve thousand times. Google-Extended, the crawler that actually feeds Gemini and AI Overviews, has never shown up at all. That's not a robots.txt problem — we checked. It's simply that Google's indexing crawler and its AI-training crawler run on completely different schedules, and the second one is far less frequent than most site owners assume.

If you've been told "just make sure robots.txt allows it and you're covered," this is the gap that advice misses: being allowed to crawl and actually being crawled are two different things, and only one of them is visible from a robots.txt file.

The AI crawler doing the most work on real sites isn't the one you'd guess

Ask most people to name an AI crawler and they'll say GPTBot. By raw visit volume across our monitored sites, GPTBot isn't even close to the top:

CrawlerVisitsSites reached (of 7)
Bytespider (ByteDance)16,5454
Amazonbot15,4005
Meta-ExternalAgent11,9086
GPTBot5,1515
ClaudeBot3,4235
OAI-SearchBot7866
PerplexityBot3185
ChatGPT-User2125
CCBot (Common Crawl)1564

Bytespider — ByteDance's crawler, which feeds TikTok's and other ByteDance AI products — visited more than three times as often as GPTBot. Meta's crawler reached six of the seven sites, more than any single OpenAI or Anthropic crawler. If a site is optimizing its content strategy around "what does GPTBot see," it's tuning for one of the less frequent visitors, not the most active one.

The bigger picture: 43% of all real traffic wasn't human

Across all seven sites, 423,162 real requests were logged. 240,863 — 56.9% — were classified as human. The remaining 43% was some form of automated traffic: AI crawlers, search engines, SEO tools, scrapers, and a small amount of outright malicious activity (credential stuffing, headless-browser probing).

That number will vary a lot by site type — a content-heavy news-style site sees far more bot traffic than a low-traffic product page — but "under half your real traffic is human" is closer to normal than most people expect in 2026.

What this actually means if you're trying to be visible to AI

  1. ·Don't assume robots.txt access equals a crawl. Google-Extended access can be open and still go unused for weeks. The only way to know is to look at real request logs, not the allow-list.
  2. ·Don't optimize only for the AI crawlers you've heard of. Bytespider and Meta's crawler had more reach across our sample than GPTBot or ClaudeBot. If TikTok/ByteDance or Meta's AI products matter to your audience, their crawlers are already showing up more than OpenAI's.
  3. ·"Bot traffic" and "AI traffic" aren't the same thing, and conflating them will make a site's real AI exposure look bigger or smaller than it is. Most of what we classified as non-human here was ordinary SEO tooling and search engines, not AI systems at all.

This is the exact data DruxShield's AI Crawler Intelligence view surfaces per-site for anyone who wants to check their own numbers instead of guessing from a blog post — including ours.

Frequently Asked

Does having Google-Extended allowed in robots.txt mean it's actually crawling my site?

No. Allowing a crawler in robots.txt only means it's permitted to visit — it doesn't mean it does, on any particular schedule. Our data shows sites with heavy, constant Googlebot traffic that Google-Extended has never visited at all. The only way to know is to check real server logs.

Which AI crawler visits the most, in practice?

In our sample of seven monitored sites, Bytespider (ByteDance's crawler) had the highest raw visit volume, followed by Amazonbot and Meta's crawler — all ahead of GPTBot, which is the crawler most site owners think of first.

What share of website traffic is actually human in 2026?

Across the sites we measured, 56.9% of real requests were human — meaning 43% was some form of automated traffic (AI crawlers, search engines, SEO tools, and scrapers). This varies by site type.

What do the AIs actually think?

Ask GPT, Claude, Gemini and more about this topic simultaneously — and get a Consensus Score showing how much they agree.

Ask the AIs: “We Checked: Does Google-Extended Actually Crawl the Sites…” →