What is an AI crawler?
An AI crawler fetches web content for AI systems: training corpora, retrieval indexes, or live browsing on behalf of a user query. Different operators run different crawlers with different purposes, and they can be allowed or blocked separately. The set changes as operators launch, rename and retire them.
Why is allowing or blocking an AI crawler a decision, not a default?
Blocking a crawler means your content is not used to build answers, which also means you are far less likely to be named in them. For a publisher whose product is content that can make sense. For a local service business whose product is the job, being cited is the goal.
Most robots.txt files were never decided. They were inherited from a theme, a plugin, or a developer, years ago.
What do we find when we check a site's crawler access?
An AI crawler can be blocked in two separate places: in robots.txt, and in a CDN or firewall rule. We regularly find sites blocking crawlers their owners would never have chosen to block, and CDN or firewall rules blocking access independently of robots.txt. Both are invisible from inside a browser.
How do you check which AI crawlers you are blocking?
There are two things to check, and they are separate. One is the robots.txt file on your own domain. The other is the CDN or firewall sitting in front of it, which can block a crawler without any of it showing up in robots.txt.
| Layer | Where the rule lives | What to do |
|---|---|---|
| robots.txt | A directive in the robots.txt file on your own domain | Open yourdomain.com/robots.txt in a browser and read it. If you do not recognise a directive, you did not decide it. |
| CDN or firewall | A rule in your CDN or firewall, independent of robots.txt | Then check whether your CDN or firewall is blocking crawler traffic independently — that one is invisible in robots.txt and common. |
See where you actually stand
The free Visibility Check puts twelve real buying questions from your category to three surfaces.
Get my free Visibility Check