Glossary Web & HTTP

robots.txt

robots.txt is a public text file at a site's root that tells crawlers which paths they may or may not fetch.

Last reviewed:

AdSense Placeholder
Slot: header_reference_page

In more detail

Rules use User-agent, Disallow and Allow lines, plus a Sitemap line. It controls crawling, not indexing: a blocked URL can still appear in results, so use noindex to keep a page out. Never use it to hide sensitive content.

Try these tools

See also

  • Glossary Sitemap
    An XML sitemap is a file listing a site's URLs, with optional last-modified dates, to help search engines find and crawl pages.
  • Guide How SEO Meta Tags Work
    Title, description, canonical, hreflang, robots.txt.
AdSense Placeholder
Slot: footer_leaderboard