What is robots.txt
robots.txt is a plain text file at the root of your website that provides instructions to web crawlers about which parts of your site they should and shouldn't access.
Why it matters
- Express crawler preferences — publish path rules that conforming crawlers may follow
- Avoid accidental exposure assumptions — robots.txt is public and is not access control; sensitive content still needs authentication/authorisation
- Declare sitemaps — the
Sitemap:directive tells crawlers where to find your XML sitemap - Declare crawler-specific policy intent — some publishers name AI-related crawler products, but robots.txt cannot guarantee compliance, use or training outcomes
Basic format
User-agent: *
Allow: /
Disallow: /admin/
Sitemap: https://example.com/sitemap.xml