Technical
AI crawler controls: a robots.txt generator workflow
Build AI crawler controls from a written policy, not a copied list of user agents. Decide which public content each crawler may fetch, distinguish search discovery from model-training preferences where providers publish separate controls, and express only those choices robots.txt can support. Generate a draft, test paths and deploy conservatively. robots.txt is publicly readable guidance rather than authentication, and allowing a crawler does not guarantee inclusion, ranking or citation in AI answers.
Create a policy matrix before writing directives
List content classes such as marketing pages, documentation, products, account areas, internal search and generated exports. For each, record whether it is public, whether crawling is desirable and what security control actually protects it. Then map current provider-documented user agents to business purposes and review dates. Crawler names and policies can change, so confirm them from primary provider documentation during implementation. Never use robots.txt to protect personal, licensed or confidential information; unauthorized users may still request those public URLs.
Draft narrow, understandable groups
- Keep a clear default User-agent group for ordinary public crawling policy.
- Add named groups only when the business has a documented reason to treat them differently.
- Use path rules based on stable URL structure rather than fragile query-string guesses.
- Avoid blocking CSS, scripts or images needed to understand allowed public pages.
- Reference the canonical sitemap when appropriate and ensure that it lists only intended URLs.
- Add owner, rationale and next review date in the internal change record, not as sensitive comments.
Use the robots.txt generator to turn the matrix into a readable draft, then run the AI readiness checker against representative URLs. Test exact matching behavior with the parser used by your crawler or an established test harness; do not assume every agent interprets obscure extensions identically. Verify case, wildcards, anchors and group precedence. Keep the file at the origin root with a successful plain-text response, and avoid environment rules that accidentally ship a staging-wide block to production.
Deploy with an access and evidence checklist
- Save the current file and crawl-policy matrix for review and rollback.
- Test allowed and disallowed examples for every rule group in a production-like environment.
- Deploy the smallest change and fetch robots.txt externally to confirm content, status and caching.
- Inspect server logs for requested paths while recognizing that user-agent strings can be spoofed.
- Verify that sensitive routes remain protected by authorization regardless of crawler behavior.
- Repeat answer and citation observations over time before attributing movement to the policy.
Run an AI visibility scan before a material policy change, then use ModelSaid to preserve recurring answer results across supported assistants. Compare citations, factual accuracy and brand presence with the deployment timeline and server retrieval evidence. A decline or increase after editing robots.txt is not proof of causation; cached content, other crawlers and model updates can complicate the timeline. This monitoring is most valuable when it exposes unintended outcomes, such as blocking current documentation while stale third-party pages remain accessible.
Plan for failure modes outside the file itself. A content delivery network may cache an obsolete robots.txt, route one hostname differently or serve an HTML error document with a successful status. Monitor the exact production origins and compare the retrieved bytes with the version in source control. During a domain migration, test both old and new hosts because redirects and host-specific rules interact. Establish an urgent rollback path for an accidental broad disallow, but remember that restoring access does not force immediate recrawling. Keep change windows, approvers and test evidence in the same operational record used for other high-impact routing changes. Include international subdomains and alternate protocols when they expose separate origin-level policies. Alert on unexpected content hashes or status changes so a routing incident does not remain invisible until traffic or citations decline.
Coordinate crawler policy with the evidence layer
Review robots.txt after site migrations, provider policy changes and new content classes. Structured data on allowed pages should still match visible page facts and is not a guaranteed AI ranking factor. Likewise, an llms.txt file cannot override robots.txt access or secure a route. Maintain one decision register covering crawler, purpose, permitted paths, business owner and verification date. The objective is deliberate access: public evidence should be retrievable under the chosen policy, private material should be protected technically, and nobody should promise that any directive controls how every AI provider trains, retrieves or answers.
Is your business visible in AI search?
Run a free check and see what ChatGPT, Claude, Gemini and Perplexity actually say about you right now.
Check your business for free