Verified 11 July 2026

AI crawlers directory

A practical directory of every AI bot Aiola tracks, using the same taxonomy as Aiola Analytics. AI answers are user-triggered fetches, Indexing bots build search or retrieval indexes, and Training crawlers collect data for model or dataset development.

A user-agent string is only a claim. Where a vendor publishes CIDRs, rDNS rules, or an ASN, the detail page explains how to verify the request. “No fetched ranges” means exactly that: Aiola does not substitute guessed infrastructure.

AI answers (14)

Fetched to answer a user's question now

ChatGPT-User · OpenAI

ChatGPT-User fetches a page after a person asks ChatGPT to visit or summarize it; it is not the automated training crawler.

Published CIDRs

Claude-User · Anthropic

Claude-User retrieves a page in response to a person's request inside Claude.

No fetched ranges

Perplexity-User · Perplexity

Perplexity-User visits a URL as part of a specific user action and may ignore robots.txt because the fetch is user initiated.

No fetched ranges

MistralAI-User · Mistral

MistralAI-User fetches pages when a Mistral product user requests current web content.

No fetched ranges

Grok-DeepSearch · xAI

Grok-DeepSearch fetches sources while a Grok user runs DeepSearch.

No fetched ranges

Google-NotebookLM · Google

Google-NotebookLM fetches a source URL that a NotebookLM user explicitly adds to a notebook.

Verified by rDNS

Google-Read-Aloud · Google

Google-Read-Aloud fetches pages for Google's text-to-speech and read-aloud services.

Verified by rDNS

Google-Agent · Google

Google-Agent identifies user-triggered requests made by Google agent products.

Verified by rDNS

Copilot · Microsoft

Copilot fetches web content for a Microsoft Copilot interaction rather than performing Bing's general index crawl.

No fetched ranges

Amzn-User · Amazon

Amzn-User retrieves a page after a user action in an Amazon AI experience.

No fetched ranges

Meta-ExternalFetcher · Meta

Meta-ExternalFetcher fetches a URL after a person shares or requests it in a Meta product, including link-preview generation.

Published ASN

DuckAssistBot · DuckDuckGo

DuckAssistBot fetches pages in real time for DuckDuckGo's AI-assisted answers and is not used to train AI models.

No fetched ranges

Kimi-User · Moonshot AI

Kimi-User fetches a URL for a specific Kimi user request.

No fetched ranges

Qwen-User · Alibaba

Qwen-User fetches web content in response to a Qwen user request.

No fetched ranges

Indexing (21)

Crawled for AI or search indexing

OAI-SearchBot · OpenAI

OAI-SearchBot discovers and indexes pages so ChatGPT search can retrieve and cite current web results.

Published CIDRs

Claude-SearchBot · Anthropic

Claude-SearchBot builds Anthropic's search index so Claude can find current web sources.

No fetched ranges

PerplexityBot · Perplexity

PerplexityBot indexes pages for Perplexity search results and citations; Perplexity says it is not used for foundation-model training.

No fetched ranges

MistralAI-Index · Mistral

MistralAI-Index indexes web pages for Mistral's search and retrieval features.

No fetched ranges

xAI-SearchBot · xAI

xAI-SearchBot discovers pages for xAI search retrieval used by Grok.

No fetched ranges

Google-InspectionTool · Google

Google-InspectionTool fetches a URL for Search Console inspection and related Google testing tools.

Verified by rDNS

Googlebot · Google

Googlebot crawls and renders pages for Google Search indexing, including surfaces that supply search results to AI features.

Verified by rDNS

Applebot · Apple

Applebot indexes the web for Apple search experiences including Spotlight, Siri, and Safari.

Verified by rDNS

Bingbot · Microsoft

Bingbot crawls and renders pages for the Bing index, which also supplies results to Microsoft Copilot.

No fetched ranges

Amzn-SearchBot · Amazon

Amzn-SearchBot discovers pages for Amazon search and AI retrieval experiences.

No fetched ranges

TikTokSpider · ByteDance

TikTokSpider crawls pages for TikTok search, previews, and discovery surfaces.

No fetched ranges

meta-webindexer · Meta

meta-webindexer indexes public pages for Meta's web-search and AI retrieval features.

Published ASN

YouBot · You.com

YouBot indexes public pages for You.com's search engine and cited AI answers.

No fetched ranges

Kimi-SearchBot · Moonshot AI

Kimi-SearchBot indexes pages for Kimi's web-search retrieval.

No fetched ranges

Baiduspider · Baidu

Baiduspider crawls and indexes pages for Baidu Search.

No fetched ranges

msnbot · Microsoft

Microsoft's legacy search crawler, still seen alongside Bingbot.

No fetched ranges

YandexBot · Yandex

Yandex's primary search crawler.

Verified by rDNS

DuckDuckBot · DuckDuckGo

DuckDuckGo's search crawler.

No fetched ranges

SeznamBot · Seznam

Search crawler for the Czech portal Seznam..

No fetched ranges

PetalBot · Petal

Crawler for Huawei's Petal Search..

No fetched ranges

Qwantbot · Qwant

Search crawler for the French engine Qwant..

No fetched ranges

Training (32)

Scraped for model training or dataset development

GPTBot · OpenAI

GPTBot collects public web content that may be used to improve OpenAI's generative AI models.

Published CIDRs

ClaudeBot · Anthropic

ClaudeBot automatically crawls public pages for Anthropic model development and improvement.

No fetched ranges

anthropic-ai · Anthropic

anthropic-ai is Anthropic's legacy training-data crawler token, separate from Claude's user-triggered fetcher.

No fetched ranges

GrokBot · xAI

GrokBot is an xAI crawler identity associated with Grok's automated web collection.

No fetched ranges

xAI-Web-Crawler · xAI

xAI-Web-Crawler is a broad xAI automated web-crawler identity.

No fetched ranges

xAI-Grok · xAI

xAI-Grok is an xAI crawler token associated with Grok content collection.

No fetched ranges

xAI-Bot · xAI

xAI-Bot is a general xAI automated crawler identity seen in server logs.

No fetched ranges

Grok · xAI

Grok is the broad Grok match token retained by Aiola for xAI crawler traffic not identified by a more specific token.

No fetched ranges

Google-Extended · Google

Google-Extended is a robots.txt control token, not a standalone HTTP crawler; it controls whether crawled content may help ground or improve Gemini and Vertex AI.

No fetched ranges

Google-CloudVertexBot · Google

Google-CloudVertexBot crawls sites at their owners' request when building Vertex AI data stores.

Verified by rDNS

GoogleOther · Google

GoogleOther is Google's generic crawler for product teams' one-off and R&D fetches; it is not the Search indexer (that is Googlebot) and may use separate crawl controls.

Verified by rDNS

Applebot-Extended · Apple

Applebot-Extended is Apple's robots.txt opt-out token for use of Applebot-crawled material in generative foundation-model training.

Verified by rDNS

Amazonbot · Amazon

Amazonbot is Amazon's general-purpose web crawler used to improve services such as Alexa and generative AI.

No fetched ranges

Bytespider · ByteDance

Bytespider collects public web content for ByteDance products and model development.

No fetched ranges

Meta-ExternalAgent · Meta

Meta-ExternalAgent automatically crawls public web content for Meta AI products and model improvement.

Published ASN

CCBot · Common Crawl

CCBot builds Common Crawl's open web corpus, which is used by researchers and many downstream datasets.

No fetched ranges

cohere-ai · Cohere

cohere-ai collects public web data for Cohere model training and improvement.

No fetched ranges

KimiBot · Moonshot AI

KimiBot automatically crawls public pages for Moonshot AI's model and product development.

No fetched ranges

QwenBot · Alibaba

QwenBot collects public pages for Qwen model and product development.

No fetched ranges

TongyiBot · Alibaba

TongyiBot is an Alibaba crawler identity associated with Tongyi model development.

No fetched ranges

AliyunBot · Alibaba

AliyunBot is an Alibaba Cloud crawler identity used for automated web collection.

No fetched ranges

ERNIEBot · Baidu

ERNIEBot collects public web content associated with Baidu's ERNIE model ecosystem.

No fetched ranges

YiyanBot · Baidu

YiyanBot is a Baidu crawler identity associated with the ERNIE/Yiyan generative AI service.

No fetched ranges

ChatGLM-Spider · Zhipu AI

ChatGLM-Spider collects public web content for Zhipu AI's ChatGLM model ecosystem.

No fetched ranges

DeepSeekBot · DeepSeek

DeepSeekBot is an automated crawler associated with DeepSeek model and product development.

No fetched ranges

AI2Bot · Allen AI

AI2Bot collects public web content for the Allen Institute for AI's research datasets and models.

No fetched ranges

Diffbot · Diffbot

Diffbot extracts structured knowledge from public pages for Diffbot's Knowledge Graph and crawl products.

No fetched ranges

Timpibot · Timpi

Timpibot crawls public pages for Timpi's decentralized search index.

No fetched ranges

ImagesiftBot · ImageSift

ImagesiftBot collects and analyzes public images and page context for ImageSift's image intelligence services.

No fetched ranges

Doubaobot · ByteDance

Crawler for ByteDance's Doubao assistant.

No fetched ranges

LinerBot · Liner

Crawler for the Liner AI answer engine..

No fetched ranges

QualifiedBot · Qualified

Crawler operated by Qualified for its AI sales-agent product..

No fetched ranges

Other (3)

Fetched for ads or link previews rather than answering, indexing, or training