Skip to content

What is robots.txt?

robots.txt is a file at the root of your site telling crawlers which parts they may fetch. It controls crawling, not indexing — a page blocked here can still appear in results if other sites link to it.

In more detail

That distinction catches people out constantly. To keep a page out of results, use noindex and let the crawler reach the page to see it; blocking it in robots.txt means the crawler never reads the noindex. The file is also where you point crawlers at your sitemap.

Why it matters for your site

A misplaced Disallow line can remove a whole section of a site from search, and nothing on the site itself shows it. It is also where you decide whether AI assistants may read your pages — blocking their crawlers removes you from AI answers.

Does this affect your website?

We check for this and a great many other things in about a minute. No account, no card.

Audit my website free