Your website can welcome AI search tools without giving every crawler the same instructions. Practical AI crawler rules separate search discovery, training collection, and private access, so you can make choices that support your business.

We recommend starting with the purpose of each bot, then applying targeted controls instead of blocking everything. Robots.txt communicates your preferences, but authentication and server-side restrictions protect content that must stay private.

First, let’s separate what these controls can do from what they can’t promise.

What AI Crawler Rules Can Control

Robots.txt is a request, not a security measure

Think of robots.txt as instructions posted at your website’s entrance. Cooperative crawlers read those instructions before requesting pages.

The Robots Exclusion Protocol standard states that these rules aren’t access authorization. A crawler can ignore them, and someone can still open a disallowed page directly.

That distinction matters for customer records, private pricing, and unpublished documents. Listing a folder in robots.txt doesn’t protect its contents. The file is public, so its entries can reveal paths you would rather keep private.

We recommend using robots.txt for crawler preferences and real access controls for confidential information.

Crawling and indexing are separate decisions

Crawling means fetching a page. Indexing means adding information about that page to a search system. Allowing the first doesn’t guarantee the second.

According to Google’s robots.txt guidance, a blocked URL can still appear in search if Google discovers it through links.

For a public page that shouldn’t appear in Google, use a supported noindex directive and allow Googlebot to read it. Blocking the page first can prevent Google from seeing that instruction.

The same practical distinction applies to AI visibility: crawler access creates an opportunity for discovery, not a promise of inclusion or citation.

Which AI Bots Handle Search, Training, and User Requests?

Crawler-specific user agents let you give different instructions to different services. These names matter because one provider can operate several bots with different purposes.

OpenAI separates search from training

OpenAI’s crawler documentation distinguishes OAI-SearchBot, which supports ChatGPT search discovery, from GPTBot, which collects content that may support model training.

You can allow OAI-SearchBot while disallowing GPTBot. That gives you a more targeted choice than a blanket OpenAI block.

ChatGPT-User handles user-triggered visits rather than automatic crawling. OpenAI says robots.txt rules may not apply to these visits. It also doesn’t determine whether your content appears in ChatGPT Search.

Googlebot and Google-Extended have different roles

Googlebot handles Google Search crawling. Blocking it can prevent Google from accessing the content you want customers to find.

Google-Extended is a robots.txt product-use token, not a separate crawler identity in HTTP requests. It controls certain Gemini training and grounding uses.

Disallowing Google-Extended doesn’t affect inclusion in Google Search and isn’t a Search ranking signal. We recommend keeping that choice separate from your Googlebot rules.

Anthropic provides three separate controls

Anthropic’s crawler guidance identifies ClaudeBot for potential training collection, Claude-SearchBot for search, and Claude-User for user-requested access.

Anthropic says its bots honor robots.txt. Blocking Claude-SearchBot can reduce search visibility, while blocking Claude-User can prevent retrieval for user requests.

Each bot needs its own targeted instruction. A rule for ClaudeBot doesn’t automatically express your preference for the other two.

How to Set Up Robots.txt Correctly

A man wearing glasses examines a printed flowchart next to a server rack in an office.

Put the file in the right location

Publish a plain-text file named robots.txt at the root of the relevant website host. Its root-relative location is /robots.txt.

A file inside a blog folder won’t control the entire website. Subdomains also need their own rules. Your main website’s file doesn’t automatically govern a separate shop or customer portal.

For WordPress, a plugin or the platform may already generate a virtual robots.txt file. We recommend identifying the current source before editing it, so you don’t create competing configurations.

Our robots.txt SEO guide covers the surrounding crawling and indexing basics.

Understand the three main directives

User-agent identifies the crawler the following rules address. Disallow lists paths that crawler shouldn’t fetch. Allow permits a path, including an exception within a broader restriction.

For example, Disallow: / requests a complete crawling block for that group. Disallow: /wp-admin/ targets a WordPress directory instead.

The wildcard in User-agent: * addresses crawlers generally. However, a crawler-specific group doesn’t automatically inherit every rule from the wildcard group. Repeat restrictions that need to apply to that crawler.

Path matching is case-sensitive. Keep directory names accurate, and preserve access to CSS and JavaScript that search engines need to render your pages.

Practical Rules for Search Visibility and Training Choices

Allow search discovery while declining training crawls

For a business that wants ChatGPT search discovery but prefers to decline GPTBot collection, these two groups express that choice:

User-agent: OAI-SearchBot
Allow: /

User-agent: GPTBot
Disallow: /

Place each directive on its own line in the actual file. Keep a blank line between groups.

You can add separate groups using User-agent: ClaudeBot with Disallow: /, and User-agent: Google-Extended with Disallow: /, to express the corresponding preferences.

These instructions govern the identified uses. They don’t remove material already collected or prevent every possible source of AI access.

Restrict selected paths without blocking useful pages

A complete block isn’t always the right choice. You can restrict a directory for a selected crawler while leaving service pages and helpful articles accessible.

For WordPress, this group requests a restriction on the administration directory while allowing its AJAX endpoint:

User-agent: GPTBot
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php

The longer, matching allowance creates an exception to the directory restriction. This example limits that path; it doesn’t decline GPTBot collection across the whole website.

We recommend choosing rules around your actual goal. Search visibility, training preferences, and server-load management are different decisions. Combining them into one blanket block can close off access you intended to keep.

When to Use Enforceable Access Controls

A network firewall appliance in a server rack connected with blue and grey ethernet cables.

What can you do when a bot ignores robots.txt? Use controls that decide whether the server delivers the content.

Authentication requires a valid login before access. That’s the right foundation for customer portals, private documents, and staging websites. Authorization then determines which logged-in users can access particular resources.

A web application firewall, often called a WAF, can block requests based on rules. A content delivery network can also apply bot controls or rate limits before requests reach your hosting server.

User-agent names alone aren’t reliable proof of identity. A scraper can claim to be Googlebot or an AI crawler. Where available, use verified-bot features or documented network identity checks instead of trusting the name alone.

Rate limiting is useful when request volume causes performance problems. An HTTP 429 response communicates that too many requests have been made.

Keep enforcement targeted. Applying login requirements or aggressive bot blocks to public service pages can damage search access. Our guide to authentication errors and search visibility explains why those restrictions belong on the right pages.

We recommend protecting private areas first, then addressing unwanted public crawling without disrupting customer access.

How Llms.txt Differs From Access Controls

The llms.txt proposal describes a file intended to help agents use a website. It can provide context and direct them toward useful information.

For a small business, that might mean identifying your main service information, documentation, or policies. Its value depends on whether a tool retrieves and uses it.

An llms.txt file doesn’t authenticate visitors, block downloads, or prevent scraping. Writing “don’t train on this content” inside it doesn’t create an enforceable restriction.

It also doesn’t replace crawler-specific robots.txt groups. The files have different purposes: robots.txt communicates crawling preferences, while llms.txt supplies information for agents.

We recommend treating llms.txt as an optional information resource, not a requirement for AI visibility. It doesn’t guarantee citations, traffic, or better rankings.

For a small website, clear public content, working internal links, and accessible service pages deserve attention first. A machine-readable summary can’t compensate for missing information or blocked pages.

If you publish one, keep it aligned with your website. Outdated links and descriptions can send an agent toward the wrong content rather than simplify discovery.

Publish the Rules and Monitor Their Effects

Make one controlled update

Start with the existing configuration rather than replacing it with a generic bot list. A short, intentional file is easier to maintain.

We recommend this sequence:

  1. Save the existing file so you can restore it if necessary.
  2. Add only the crawler groups and path rules that match your decisions.
  3. Publish the update and confirm the root file returns plain text successfully.
  4. Test important public pages and private areas for their intended access.

If a CDN caches the file, clear the relevant cached copy after publication. Keep the change date and the reason for each restriction in your maintenance records.

Watch access, indexing, and business outcomes separately

Server logs show requested URLs, claimed user agents, response codes, and request volume. They help identify whether crawling is creating errors or unnecessary load.

For Google visibility, use Search Console’s URL Inspection tool to assess important pages. Crawling alone doesn’t guarantee indexing, even when discovery happens sooner.

IndexNow is a change-notification protocol for participating engines. It doesn’t override access restrictions or guarantee search inclusion. For a rarely updated site, content quality and confirmed indexing problems usually deserve more attention than custom notification work.

Our technical SEO audit checklist helps organize those wider checks.

Crawler identities, behavior, and provider policies can change. We recommend a monthly review, with another review after hosting, security, or website changes. Measure visibility and inquiries separately from bot activity.

Keep Crawler Choices Separate From Website Security

Effective AI crawler rules match each bot’s purpose to your business goals. You can support search discovery, express training preferences, and protect private content with different controls.

We recommend starting with a small set of targeted rules and reviewing their effects. Keep confidential information behind real access controls, regardless of what robots.txt says.

Your website should stay useful to customers and accessible to the services you choose. Clear rules support that goal without promising control they can’t deliver.

We use cookies so you can have a great experience on our website. View more
Cookies settings
Accept
Decline
Privacy & Cookie policy
Privacy & Cookies policy
Cookie name Active

Who we are

Our website address is: https://nkyseo.com.

Comments

When visitors leave comments on the site we collect the data shown in the comments form, and also the visitor’s IP address and browser user agent string to help spam detection. An anonymized string created from your email address (also called a hash) may be provided to the Gravatar service to see if you are using it. The Gravatar service privacy policy is available here: https://automattic.com/privacy/. After approval of your comment, your profile picture is visible to the public in the context of your comment.

Media

If you upload images to the website, you should avoid uploading images with embedded location data (EXIF GPS) included. Visitors to the website can download and extract any location data from images on the website.

Cookies

If you leave a comment on our site you may opt-in to saving your name, email address and website in cookies. These are for your convenience so that you do not have to fill in your details again when you leave another comment. These cookies will last for one year. If you visit our login page, we will set a temporary cookie to determine if your browser accepts cookies. This cookie contains no personal data and is discarded when you close your browser. When you log in, we will also set up several cookies to save your login information and your screen display choices. Login cookies last for two days, and screen options cookies last for a year. If you select "Remember Me", your login will persist for two weeks. If you log out of your account, the login cookies will be removed. If you edit or publish an article, an additional cookie will be saved in your browser. This cookie includes no personal data and simply indicates the post ID of the article you just edited. It expires after 1 day.

Embedded content from other websites

Articles on this site may include embedded content (e.g. videos, images, articles, etc.). Embedded content from other websites behaves in the exact same way as if the visitor has visited the other website. These websites may collect data about you, use cookies, embed additional third-party tracking, and monitor your interaction with that embedded content, including tracking your interaction with the embedded content if you have an account and are logged in to that website.

Who we share your data with

If you request a password reset, your IP address will be included in the reset email.

How long we retain your data

If you leave a comment, the comment and its metadata are retained indefinitely. This is so we can recognize and approve any follow-up comments automatically instead of holding them in a moderation queue. For users that register on our website (if any), we also store the personal information they provide in their user profile. All users can see, edit, or delete their personal information at any time (except they cannot change their username). Website administrators can also see and edit that information.

What rights you have over your data

If you have an account on this site, or have left comments, you can request to receive an exported file of the personal data we hold about you, including any data you have provided to us. You can also request that we erase any personal data we hold about you. This does not include any data we are obliged to keep for administrative, legal, or security purposes.

Where your data is sent

Visitor comments may be checked through an automated spam detection service.
Save settings
Cookies settings