I finished building a site, put it on a server, pointed a real domain at it, and then made sure Google could not see a single page of it.

Not by accident. On purpose.

The site was empty. No articles, no homepage copy. Just a framework and some legal pages.

If Google had crawled it in that state, it would have found an unfinished site and made up its mind about it. First impressions are hard to undo in search.

So I locked the door. And in the process I learned that almost everything people say about hiding a site from Google is wrong.

Why you would hide a live site

This is not only for unfinished sites. There are real reasons:

Before launch. You want it on your domain and working, but not judged yet.

While redesigning. You do not want Google indexing half-finished pages.

On staging sites. A staging copy that Google indexes becomes duplicate content competing with your real site.

After killing a product. The pages are gone but the URLs still exist.

For private areas. Client portals, member pages, internal tools.

Different reasons. Same mistake made in all of them.

The mistake

Almost everyone reaches for robots.txt.

It feels right. You add two lines:

User-agent: *
Disallow: /

You check it in your browser. The file loads. You assume you are hidden.

You are not.

What robots.txt actually does

Google states it plainly, in the first paragraph of their Introduction to robots.txt:

"A robots.txt file tells search engine crawlers which URLs the crawler can access on your site. This is used mainly to avoid overloading your site with requests; it is not a mechanism for keeping a web page out of Google. To keep a web page out of Google, block indexing with noindex or password-protect the page."

Read that again. "It is not a mechanism for keeping a web page out of Google."

Robots.txt is about crawling. It is a request to a robot: please do not fetch this.

It says nothing about indexing, whether a URL may appear in search results.

Those are two different things, and that distinction is the whole article.

So what actually happens?

Google explains:

"If your web page is blocked with a robots.txt file, its URL can still appear in search results, but the search result won't have a description."

And the warning on the same page:

"Don't use a robots.txt file as a means to hide your web pages (including PDFs and other text-based formats supported by Google) from Google Search results. If other pages point to your page with descriptive text, Google could still index the URL without visiting the page."

Google will list your URL without ever reading your page.

So a visitor searching for your brand could find a result with no title, no description, just a bare link. For an unfinished site, that is worse than either being fully visible or fully hidden.

The trap: two methods that cancel each other out

Here is where most people make it worse.

They think: "Robots.txt blocks crawling, and noindex blocks indexing. If I add both, I'm double-safe."

The opposite is true.

From Google's Block Search Indexing with noindex:

"Important: For the noindex rule to be effective, the page must not be blocked by a robots.txt file, and it has to be otherwise accessible to the crawler. If the page is blocked by a robots.txt file or the crawler can't access the page, the crawler will never see the noindex rule, and the page can still appear in search results."

Follow that through:

Disallow: /   →  Google cannot fetch the page
              →  Google never reads your noindex tag
              →  the noindex does nothing

Your noindex is invisible. You have blocked the one robot that was supposed to read it.

You end up with the worst of both worlds: not crawled, and still indexable.

Google also says something people find surprising:

"Specifying the noindex rule in the robots.txt file is not supported by Google."

You cannot write Disallow in robots.txt and expect indexing to stop. It is not supported. It has never been supported.

So what actually works?

Three methods, and only one does everything.

Method 1: noindex (works, with one condition)

Put this in the <head> of every page:

<meta name="robots" content="noindex">

Or send it as an HTTP header, which also works for files like PDFs:

X-Robots-Tag: noindex

What it does: "Google will drop that page entirely from Google Search results, regardless of whether other sites link to it."

The condition: Google must be able to fetch the page to see the tag. So this requires that you do not block it in robots.txt.

Method 2: Password protection (works, and does everything)

This is what I used. The site returns 401 Unauthorized to anyone without a password.

Google recommends it explicitly, in three separate places:

"To keep a web page out of Google, block indexing with noindex or password-protect the page."

"To properly prevent your URL from appearing in Google Search results, password-protect the files on your server, use the meta tag or response header, or remove the page entirely."

"Don't use robots.txt as a way to block your page."

What it does: stops crawling and indexing. One method, both problems.

Google's Remove information page explains why:

"Password-protect your page. Limiting access to your page enables the right users to view your page, while preventing Googlebot and other web crawlers from accessing it."

Method 3: robots.txt (does not work for this)

Stops crawling. Does not stop indexing. That is its entire job.

The status code table

This is the part that makes everything else click. Google's HTTP status code documentation explains what happens with each:

Code What Google does
200 Content may be indexed. "An HTTP 2xx (success) status code doesn't guarantee indexing"
301/302 Follows the redirect, up to 10 hops
401 / 403 "Google doesn't index URLs that return a 4xx status code, and URLs that are already indexed and return a 4xx status code are removed from the index"
404 / 410 Same as above. The URL is dropped
429 Treated as a server error, not a client error
500 / 503 "already indexed URLs are preserved in the index, but eventually dropped"

That 401 row is why password protection works. Google does not index a page it cannot access, and removes it if it was already there.

Note the difference between 404 and 503:

  • 404 → removed from the index
  • 503 → preserved in the index for a while

If you take a site down with maintenance mode (503), Google keeps your pages. If you want them gone, that does not do it.

Comparing the three methods

robots.txt noindex Password (401)
Stops crawling ✅ ❌ ✅
Stops indexing ❌ ✅ ✅
URL can still appear ✅ yes ❌ ❌
Works on PDFs/images partly ✅ via header ✅
Needs the page reachable - ✅ yes ❌
Works on a brand new site ❌ ✅ ✅

If you only remember one row: robots.txt can stop crawling, but it cannot stop your URL appearing in Google.

Two timing facts nobody mentions

1. Google caches your robots.txt for up to 24 hours

From the robots.txt specification:

"Google generally caches the contents of robots.txt file for up to 24 hours, but may cache it longer in situations where refreshing the cached version isn't possible."

So when you change your robots.txt, Google may keep using yesterday's copy for a day. Do not assume your change took effect instantly.

2. A broken robots.txt means no rules at all

"Google's crawlers treat all 4xx errors, except 429, as if a valid robots.txt file didn't exist. This means that Google assumes that there are no crawl restrictions."

If your robots.txt returns a 404, say you deleted it, or misspelled the path , every rule in it stops applying. Google treats it as "no restrictions."

That is the opposite of what most people assume. A missing robots.txt does not protect you. It frees Google.

The part that catches people later

Suppose you did get indexed, and now you want out.

Adding noindex does not remove you today. Google has to come back and read it first.

"We have to crawl your page in order to see noindex tags and HTTP headers. If a page is still appearing in results, it's probably because we haven't crawled the page since you added the rule. Depending on the importance of the page on the internet, it may take months for Googlebot to revisit a page."

Months.

You can speed it up with the URL Inspection tool, which requests a recrawl. But for a large site, you are waiting.

There is also a temporary option, the Removals tool:

"Requests made in the Removals tool last for about 6 months."

Six months, then it comes back unless you have fixed it properly underneath.

Prevention is instant. Removal is not. That is why I locked my site before it had a chance to be seen.

Four myths

"I deleted the site, so Google will forget it." The URL stays. Google keeps returning to it. You get 404s in Search Console. Remove the content properly or redirect it.

"I removed my sitemap, so Google will remove my pages." A sitemap is a suggestion, not a permission slip. Google already knows the URLs.

"Disallow: / hides my whole site." It stops crawling. Your URLs can still appear, with no description.

"Disallow: / plus noindex is extra safe." It is the one combination that guarantees neither works.

What I actually did

I did not use robots.txt alone, and I did not use it with noindex.

My site returns a 401 to everyone without a password. That is one method that stops crawling and indexing, and Google's own documentation names it first.

Here is what a crawler gets from every page:

HTTP/1.1 401 Unauthorized
WWW-Authenticate: Basic realm="Private", charset="UTF-8"
X-Robots-Tag: noindex, nofollow

I also set robots.txt to Disallow: / while hidden. Not as the mechanism, as a second signal that costs nothing. But I went in knowing it was not the lock. The 401 is the lock.

How to check your own site

Check what a crawler really receives:

curl -I -A "Googlebot" https://yoursite.com/

Check a specific page:

curl -s -o /dev/null -w "%{http_code}\n" https://yoursite.com/about

Check your robots.txt:

curl -s https://yoursite.com/robots.txt

Then confirm in Google Search Console with the URL Inspection tool. It shows the exact HTML Google received, including your meta tags. That is the only way to know for certain rather than assuming.

The short version

  • robots.txt stops crawling, not indexing
  • Your URL can still appear in Google with no description
  • robots.txt + noindex cancel each other out, the noindex is never read
  • Writing noindex inside robots.txt is not supported by Google
  • Password protection (401) is the only single method that stops both
  • Google caches robots.txt for up to 24 hours
  • If your robots.txt 404s, all your rules stop applying
  • Removing a page after indexing takes months
  • Hide it before Google sees it. That is the only easy moment

Sources

Every quote above is from Google's own documentation: