How to Check Your robots.txt (and the Line That Deindexes a Site)
To check your robots.txt, open https://yoursite.com/robots.txt in a browser, read each User-agent group, and look for any Disallow rule that covers pages you want in search, starting with a bare Disallow: /. On a small site that takes a minute. Then confirm what Google actually fetched in the robots.txt report in Search Console, because the file you see and the file Googlebot received aren't always the same.
Most robots.txt files are a handful of lines, which is why hardly anyone reads them. They also apply to the whole host, so a single wrong line affects every URL on it.
Crawling, not indexing
robots.txt tells crawlers which URLs they may request. It doesn't decide what appears in search results. A URL blocked in robots.txt can still be indexed if other pages link to it; Google shows it without a description, and Search Console lists it as "Indexed, though blocked by robots.txt".
This matters when robots.txt is used to hide pages. Block a page and add noindex to it as well, and Google can't crawl the page, never sees the noindex, and the URL may stay in the index. To take a page out of results, leave it crawlable and serve noindex. The noindex tag, explained covers that side.
Finding the file
robots.txt always sits at the root of a host, for example https://example.com/robots.txt. Its rules apply only to the protocol, host and port it's served from, so https://shop.example.com needs its own file.
On many platforms there's no physical file to open. WordPress generates a virtual robots.txt unless you upload a real one to the web root, and SEO plugins such as Yoast and Rank Math let you edit it from the admin. Shopify serves a default file that you can customise by adding a robots.txt.liquid template to the theme. If you connect over FTP and can't find the file, that's usually the reason.
Reading it line by line
A robots.txt is made of groups. Each group opens with one or more User-agent lines followed by its rules:
User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php
Disallow: /*?s=
Disallow: /cart/
User-agent: GPTBot
Disallow: /
Sitemap: https://example.com/sitemap_index.xml
Reading this example from the top, all crawlers are kept out of /wp-admin/ apart from admin-ajax.php, and out of internal search results (?s=) and the cart. GPTBot has its own group and is blocked from everything. The Sitemap line isn't part of any group and applies to the whole file.
A few parsing rules explain most surprises. A crawler obeys only the most specific group that matches its name. Googlebot finds no User-agent: Googlebot group here, so it uses *; if there were a Googlebot group, it would ignore the * rules entirely instead of merging them. Paths are case-sensitive, so /Blog/ and /blog/ are different rules. * matches any run of characters and $ marks the end of the URL, which means Disallow: /*.pdf$ blocks PDFs and nothing else. When an Allow and a Disallow both match a URL, Google follows the longer, more specific one. An empty Disallow: blocks nothing at all.
Google ignores Crawl-delay and Noindex lines in robots.txt. Bing does honour Crawl-delay.
The two lines that block a whole site
User-agent: *
Disallow: /
That tells every crawler to stay away from every URL. It's the normal setup for a staging or development site, which is exactly how it reaches live sites: a deploy copies the staging file, or a migration brings the whole config across. WordPress has a related trap. Ticking "Discourage search engines from indexing this site" under Settings > Reading makes current versions add a noindex robots meta tag to every page, so after a launch it's worth checking both the file and the page source.
Visitors notice nothing when this happens. Pages load, forms submit, uptime checks stay green. Crawling slows down and, over the following weeks, pages drop out of results.
Checks worth running
- Open
/robots.txtdirectly and look at the HTTP status as well as the content. The Network tab in Chrome DevTools shows it. You want a 200 and the rules you expect. - Search the file for
Disallow: /on its own, and for rules on folders that hold pages you want to rank, such as/blog/,/products/or/category/. - Make sure the
Sitemap:line points to a sitemap that loads and is the current one, not a leftover from a previous plugin. - In Search Console, open Settings > robots.txt. The report lists the robots.txt files Google found for the property's top hosts, when each was last crawled, the fetch status, and any lines it couldn't parse. After fixing an urgent problem you can request a recrawl from there.
- For a single URL, use URL Inspection. If the URL is blocked, the page indexing section says "Blocked by robots.txt".
- On a bigger site, run a crawler such as Screaming Frog in its default mode, which respects robots.txt, and review the URLs it reports as blocked.
Google retired the old robots.txt Tester in 2023, so guides that send you there are out of date.
The status code matters as much as the rules. Google treats a 404 or any other 4xx except 429 as if no robots.txt exists, meaning everything may be crawled. A 5xx is handled differently: Google stops crawling for a period, then falls back on a cached copy while it keeps retrying. Google also caches the file for up to 24 hours and only reads the first 500 KiB, so rules beyond that point in an unusually long file are ignored.
Knowing when the file changes
Every check above tells you what robots.txt says today. None tells you that a plugin rewrote it on Tuesday or that a hosting move restored an old version. On one site you can reopen it after each release. Across a client portfolio where developers deploy without warning, you won't, and the first sign of a bad edit is usually a crawl drop in Search Console weeks later. The guide to robots.txt change alerts goes through which edits deserve an alert.
Deltio fetches robots.txt on each client site as part of its daily check, compares it with the previous version, and sends a Slack or email alert for that site when it has changed. The same pass checks noindex, canonicals and the URLs in the sitemap, since a staging deploy tends to break several of them together. It runs once a day, so an edit is reported the same day or the next rather than within minutes. SEO change alerts explains how those alerts are scoped, and there's a 14-day free trial if you want to add a site.
Frequently asked questions
- Does every site need a robots.txt?
- No. If the file returns a 404, Google treats the site as having no crawl restrictions. A small site with nothing to keep crawlers away from can do without one, although a Sitemap line in robots.txt is a convenient way to point every search engine to the XML sitemap.
- Should I block CSS and JavaScript files in robots.txt?
- No. Google renders pages much like a browser, and blocking the CSS or JavaScript it needs can stop it seeing the page as users do, including content loaded by scripts. Old rules such as Disallow: /wp-includes/ are worth removing for that reason.
- Can robots.txt block AI crawlers?
- You can add groups for user agents such as GPTBot, ClaudeBot or CCBot. Google-Extended controls whether content is used for Google's AI models and has no effect on Google Search crawling or rankings. Compliance is voluntary, so robots.txt only works on crawlers that choose to respect it.
- Is robots.txt a way to hide private pages?
- No. The file is public, so a Disallow rule for /private-area/ actually advertises that path to anyone who reads it. Pages that must stay private need password protection or authentication, not a crawl rule.
- How do I keep images out of Google Images with robots.txt?
- Add a group for Google's image crawler, for example User-agent: Googlebot-Image followed by Disallow: /images/private/. Regular Googlebot rules still apply to the pages themselves.