← Back to the blog

How to Catch a Noindex Before It Kills Your Rankings

How to Catch a Noindex Before It Kills Your Rankings

To catch a noindex in production, check both the HTML and the HTTP headers of the pages that matter, and set up something that compares them from one day to the next. A stray noindex produces no error: the page loads, the checkout works, and Google removes the URL from its index as it recrawls.

The tag has legitimate uses on thank-you pages, internal search results and staging environments. The damage starts when one lands on a page you need ranking and stays there for weeks before anyone looks.

Where production noindexes usually come from

Most are staging settings that travelled. Staging environments are kept out of Google with a robots meta tag or a header, and a deploy that copies configuration, environment variables or a database dump carries that setting across. On WordPress the usual culprit is the "Discourage search engines from indexing this site" checkbox under Settings > Reading, ticked on staging and restored along with the database on production.

SEO plugins come next. Yoast and Rank Math both let you set a whole post type or taxonomy to noindex from their settings screens. In Yoast it's the "Show [post type] in search results" switch, so one toggle or a settings reset affects hundreds of URLs at once.

Then there are headers. X-Robots-Tag: noindex can be added with an Nginx add_header directive, an Apache .htaccess rule or a CDN configuration, and none of these appear in the page source. A developer who hides a new section while building it and forgets the rule at launch creates exactly this kind of invisible noindex.

Checking a page by hand

When you suspect a specific URL, work through these in order. Each one catches something the previous step can miss.

  1. Fetch the headers: curl -sI https://example.com/page/ | grep -i x-robots-tag. Any line containing noindex means the header is set, whatever the HTML says.
  2. Search the served HTML: curl -s https://example.com/page/ | grep -i 'name="robots"'. Look for name="googlebot" as well, which applies to Google only.
  3. If the site renders with JavaScript, open Search Console's URL Inspection tool, click "Test live URL" and then "View tested page" to see the HTML Google gets. Google's documentation warns that when the initial HTML contains noindex, Googlebot may skip rendering, so a script that removes the tag afterwards doesn't help.
  4. In the same inspection result, "Indexing allowed?" tells you whether the block came from the robots meta tag or the X-Robots-Tag HTTP header.
  5. For a sweep of the whole site, crawl it with Screaming Frog and use the Noindex filter in the Directives tab.

Search Console's Page indexing report also groups URLs under "Excluded by 'noindex' tag". It is authoritative, but it only updates after Google recrawls, so it describes last week rather than today.

Stopping it at deploy time

If you control the deployment pipeline, a check after every production release costs very little. A CI step that runs the two curl commands above against the homepage and a few key templates, and fails the job when either finds noindex, blocks the most common staging leak.

It won't catch a plugin setting changed in wp-admin on a Tuesday afternoon, or a header rule added in the CDN dashboard. Neither goes through the pipeline.

Monitoring for the changes nobody deploys

Manual checks answer "is this page noindexed right now?". Monitoring answers "did any page become noindexed since yesterday?", which is the question that matters on client sites where changes come from people and systems you don't control.

Deltio reads the pages in each site's sitemap on a daily cycle and records the noindex state from both the meta tag and the X-Robots-Tag header. When a page that was indexable on the previous check is noindexed on the current one, it sends a Slack and email alert for that site naming the affected URLs. Being daily, it tells you within a day rather than within minutes, which in most cases is still before Google has recrawled and dropped the page. The same comparison covers robots.txt edits, canonical changes and URLs leaving the sitemap. For background on the tag itself, see the noindex tag, explained.

You can run it on your own client sites with a 14-day free trial.

Frequently asked questions

How long does it take Google to drop a page after a noindex is added?
The noindex only takes effect when Google recrawls the URL, so it depends on crawl frequency. A homepage or a busy category page can drop within days, while a rarely crawled page may stay indexed for weeks. That lag is the window in which catching the change is cheap.
Will removing the noindex bring the rankings back?
Usually, once Google recrawls the page. To speed that up for important URLs, use Request indexing in the URL Inspection tool and make sure they are listed in the sitemap. The longer a page was out of the index, the less certain it is to return to the same position.
Can robots.txt hide a noindex from Google?
Yes. If a URL is blocked with Disallow, Googlebot never fetches it, so it never reads the noindex. The URL can even stay indexed, typically without a description, because Google still knows it exists from links. A page you want deindexed has to stay crawlable.
Does a noindex rule inside robots.txt work?
No. Google never officially supported noindex lines in robots.txt and stopped honouring them completely in September 2019. Use the robots meta tag or the X-Robots-Tag header instead.
Does Deltio check pages that aren't in the sitemap?
Its daily check works from the URLs listed in each site's sitemap. If important pages are missing from the sitemap, they won't be part of that comparison, and that gap is worth fixing for SEO reasons anyway.