> ## Content Index
> Fetch the complete content index at: https://www.magicpages.co/llms.txt
> Use this file to discover other available public pages before exploring further.

# AI Crawler Controls on Magic Pages
- URL: https://www.magicpages.co/help/features/ai-crawler-controls-on-magic-pages/
- Published: 2026-05-13T14:08:52.000Z
- Updated: 2026-05-29T19:21:55.000Z
- Description: Block known AI crawlers from your Magic Pages site at the Cloudflare edge with one toggle. Here's what it blocks, what it doesn't, and the limits to be honest about.
- Author: Jannis Fedoruk-Betschki
- Tags: Help Center & Guides, Features and Benefits

AI training crawlers and AI assistants now make up a meaningful share of web traffic. Some respect the unwritten rules − they identify themselves clearly, honour `robots.txt`, and only fetch what they need. Others crawl aggressively, ignore `robots.txt`, or rotate user-agents to slip past blocks.

If you'd rather your content not be used to train large language models – or returned inside AI-generated answers without a click back to your site – Magic Pages lets you block known AI crawlers from your website with a single toggle.

## What this blocks

When you enable AI crawler blocking, we add a rule at our CDN edge that returns `403 Forbidden` for any request matching either of two lists:

- **Cloudflare's verified-bot categories** − *AI Crawler*, *AI Assistant*, and *AI Search*. These are maintained by Cloudflare and cover ChatGPT, Claude, Perplexity, Google's AI features, and many more. Cloudflare keeps this list fresh, so we don't have to.
- **A vendored list of known AI/ML user-agent strings**, sourced from the open-source [ai-robots-txt](https://github.com/ai-robots-txt/ai.robots.txt?ref=magicpages.co) project. This catches crawlers that haven't been formally verified yet. The list refreshes weekly.

Currently on the list: GPTBot, ChatGPT-User, ClaudeBot, Claude-Web, CCBot, Google-Extended, Anthropic-AI, Bytespider, FacebookBot, Meta-ExternalAgent, PerplexityBot, Applebot-Extended, Amazonbot, Cohere-AI, and around forty others.

## What this doesn't block

Regular search engines – Googlebot, Bingbot, DuckDuckBot – are ***not*** affected. Your site stays indexable and shareable. Human visitors browse normally. The block is precisely scoped to AI training and assistant bots. Everything else continues exactly as before.

RSS feeds, the Ghost Admin API, webhooks, and any integrations you have running are also untouched.

## How it works

The check happens at the Cloudflare edge, before the request ever reaches your site. That has two practical benefits:

- AI crawlers don't consume any of your site's resources (though, this is more for our peace of mind than yours) – Cloudflare absorbs the request entirely.
- The block is consistent across every page, including paid posts behind your membership, RSS-only routes, and static assets.

Changes propagate within seconds. We allow up to 60 seconds in the worst case.

## Limits and honesty

A few things worth being upfront about:

- **UA detection is honour-based.** A determined scraper can spoof a normal browser user-agent and bypass this. The well-behaved crawlers – the ones training the major AI products – identify themselves correctly, which is what this feature relies on.
- **Already-indexed content stays indexed.** Blocking AI crawlers today stops new training data, but content crawled before you enabled the toggle may already exist inside trained models.
- **The list evolves.** New AI products launch constantly. Both the vendored UA list and Cloudflare's verified-bot categories update automatically. However, it is perfectly possible that a new AI crawler emerges that slips through this for a few days.

## How to enable

Open the Customer Portal from inside your Ghost Admin, head to the **Configuration** tab, and flip the ***Block AI crawlers from accessing this site*** toggle. For a step-by-step walkthrough with screenshots, see our [how-to guide](https://www.magicpages.co/help/features/how-to-block-ai-crawlers-on-your-ghost-cms-website/).

![](https://www.magicpages.co/content/images/2026/05/image.png)

## Why we built this

Independent creators and small businesses on Magic Pages told us they wanted control over how their work shows up in AI products. Some are fine with their content being part of training data, others aren't. Both positions are valid.

This toggle gives you the choice – at the infrastructure layer, where it actually works – without you having to maintain a `robots.txt` or wrestle with Cloudflare rules yourself.