Cozmo Scan My SEO Logo

AI search and AI training: understand your website’s permissions

Start my free audit

“Allow AI” is not one decision.

You may want your public service pages to appear in search while making a different choice about model training. A visitor asking an assistant to open a page is another situation again.

Start with the purpose of the request, not just the word “AI”. Providers publish different crawler names and controls. A rule written for one purpose should not be treated as permission for all of them.

PurposeWhat to understandDocumented examples
SearchFinding pages to surface in the provider’s search experience.OAI-SearchBot; PerplexityBot.
TrainingCollecting content that may be used to train models.GPTBot.
A user-requested visitOpening a page in response to a user action, rather than automatic search crawling.ChatGPT-User; Perplexity-User.

Sources: OpenAI’s crawler documentation and Perplexity’s crawler documentation. These names describe different purposes; they are not a complete list of every provider or permission.

Choose search access and training preferences separately.

OpenAI documents independent settings for OAI-SearchBot and GPTBot. Its search crawler is not the training crawler. Perplexity describes PerplexityBot as a search crawler, not a crawler for training foundation models.

For your own website, write down the business decision first: which public pages should be discoverable, and what is your separate training preference? Keep private pages private.

Do not remove a training restriction just because a tool has labelled it as an “AI issue”. ScanMySEO separates these purposes, and a training restriction does not lower search readiness by itself.

Provider controls do not all behave identically. The Google Search generative AI control is an owner-side Search Console setting. Review it in the account rather than trying to infer its state from a public page or a different provider’s rules.

Read an access result for the page it actually covers.

A website’s robots.txt file publishes crawler rules. Those rules can vary by crawler, path and website origin. A service page and an account area can therefore have different policies.

In ScanMySEO, “Rules allow this page” means the observed policy permitted that checked path for the named crawler. It does not mean that the provider successfully visited, indexed or recommended it. A firewall, login requirement or other restriction is a separate layer.

Likewise, when Cozmo cannot open a page, that tells you about this audit’s attempt. It is not proof that every search service is blocked. Keep the exact response, path and time when asking your administrator to investigate.

When the rules cannot be read reliably, the useful result is “Could not assess these rules”. There is no reason to turn a missing result into either a pass or a sitewide failure.

A requested visit is not the same as search crawling.

OpenAI and Perplexity also document agents used for user-triggered visits. Their guidance explains that these requests may not follow robots rules in the same way as automatic crawlers.

That is another reason not to use robots.txt as protection for confidential information. Keep proper authentication and access controls in place. Ask your administrator to use the provider’s documented verification methods before trusting a request that merely claims a familiar crawler name.

Give your administrator a clear brief.

Send a short description of the intended outcome, the affected public URLs and the relevant report findings. Avoid sending a blanket instruction to “allow every AI bot”.

The administrator can check the precise rules, origin, path and security configuration. Where an exception is appropriate, it should be narrow and verified against current provider documentation, rather than a general removal of security controls.

Check the change in two stages.

First, confirm that the published website rule now matches your intended policy and that the relevant public page is accessible under the conditions being tested. A new audit can record those website changes.

Second, use the provider’s own tools and reporting for the outcomes they actually cover. A policy change and a reported appearance are different observations. Record the date of each rather than assuming the first caused the second.

For the wider picture, see what ScanMySEO’s AI search review covers and the measurement guide.

Provider guidance reviewed on 6 September 2026. Recheck the linked documentation before changing permissions; crawler names, controls and platform behaviour can change.

Choose your next website improvement.

Use Cozmo’s findings to decide what to change, then share the task with the person who can help.

Start my free audit
Hansel McKoy

Written by Hansel McKoy, founder of ScanMySEO. These guides help website owners understand their audit findings and turn them into useful improvements.

Hansel McKoy

Founder, ScanMySEO