
Readers now have a clear, practical lesson from a recent indexing failure: any page a tool creates on a public domain can be found and listed by search engines, so a single missing instruction in the page’s code is enough to expose private content to the world. The incident that surfaced this risk involved shared chat pages from a major AI assistant that began appearing in Google search results, exposing conversations the users thought were private.
What actually happened
An AI company published a feature that lets users generate a shareable link to a conversation. The page that the link opens lives on the company’s own domain, which is the same domain the search engine already crawls and indexes. Each share link produced a unique URL that the company intended to circulate by word of mouth, on social media, or by direct message. The unintended consequence was that search engines found these URLs on their own and added them to the public index.
Once indexed, anyone searching for the right phrase could open the page and read the conversation. Reporters and researchers found private chats that discussed medical questions, legal problems, personal relationships, and work details. The exposure was not the result of a data breach in the traditional sense. It came from a configuration choice on the publisher’s side.
The technical SEO mistake behind the leak
Search engines index a page only when three things line up: the page is reachable by the crawler, the page is allowed in the robots configuration, and the page is not marked as excluded. The shared chat pages failed on the last point. The site served them with the same default rules as every other public page, which told crawlers that indexing was welcome.
The fix that should have been applied is straightforward. A page that should not appear in search results needs an instruction that says so. The two standard tools for this are:
- A
noindexmeta tag in the page’s HTML head, which tells search engines not to include the page in their index. - A disallow rule in the site’s
robots.txtfile, which stops crawlers from requesting the page at all.
Neither was applied to the shared chat URLs. The pages had unique, hard-to-guess links, but the search engines were able to follow links to them from other indexed pages on the same domain. A long random string in a URL is not a security control. It is a guessable identifier that any crawler that lands on the linking page can read and follow.
Why “unlisted” is not the same as “private”
Many publishing tools use the word “unlisted” to describe a page that has no public navigation pointing to it. Unlisted means the page is not advertised. It does not mean the page is hidden from crawlers. A page that returns a 200 status code and has no noindex directive is treated by Google as indexable the moment any link to it is discovered.
For business owners, the same trap appears in several familiar places:
- Preview or draft versions of a page that are reachable on the live domain.
- Thank-you pages, order confirmation pages, or internal search results that print query strings in the URL.
- PDFs and documents hosted in a public folder with no protection.
- Pages generated by chatbots, quote builders, configurators, and AI assistants that create a unique URL for each session.
In each case, the pattern is the same. The page exists at a real address, the address is reachable, and the publisher did not tell search engines to leave it alone.
What site owners should do today
The checklist below covers the configuration that keeps generated, internal, or low-value pages out of Google’s index. Apply it to any URL pattern on a site that was not designed for public search traffic.
1. Add noindex to pages that should not rank
Place a <meta name="robots" content="noindex"> tag in the head section of any page that holds private, internal, or low-value content. This includes chat transcripts, draft views, internal search results, staging previews, and any AI-generated page that was designed to be shared one-to-one rather than ranked.
2. Block crawlers from sensitive paths in robots.txt
For URL patterns that should never be requested at all, add a Disallow rule. This is a stronger signal than noindex because the crawler never sees the page in the first place. Reserve it for paths that hold truly private data.
3. Require authentication before serving a private page
If a page should only be readable by a specific person, place it behind a login. A page that requires authentication before it loads is not indexable. This is the strongest control and the one that matters most for any page that holds personal information.
4. Audit what your site already exposes
Use a site search operator such as site:yourdomain.com to see which of your own pages Google has indexed. Look for patterns you did not intend to publish: chat transcripts, internal search results, staging URLs, and parameter-heavy pages. If a page appears in the index that should not, add noindex, request removal through Google’s tools, and fix the template so the next page created in that pattern inherits the right rules.
A full technical audit walks through every URL pattern a site serves and checks each one for indexability, and SEOScanPro runs that kind of audit and shows the measured result behind every check.
Lessons for anyone publishing AI-generated pages
AI chat products, AI writing tools, and AI-powered site features all generate URLs. Every one of those URLs inherits the indexability rules of the domain it lives on. If the product is built on a public domain and the team does not set noindex on the generated page, the page will end up in Google the first time any crawler finds a link to it.
Three habits make this kind of leak much harder to ship:
- Treat any dynamically generated URL as a candidate for
noindexby default, then make it indexable on purpose if a decision is made. - Audit newly launched features against a list of page templates and confirm each one has the correct robots directive before the feature ships.
- Watch the indexed URL count for a domain. Sudden growth in indexed pages after a new feature launches is a signal that something the team meant to keep private has become public.
The shared chat incident is a public reminder of a rule that has been true since the earliest search engines. A page that the server returns to anyone who asks for it is a page that the search engine can return to anyone who searches for it. The only way to keep a page out of search results is to tell search engines, in their language, that the page is not for them.
FAQ
How did private AI chat pages end up in Google search results?
Shared chat URLs were served on a public domain without a noindex directive or a robots.txt block. Search engines found links to those URLs and indexed them like any other page on the site.
What is the difference between an unlisted page and a private page?
An unlisted page has no public link pointing to it but is still reachable and still indexable. A private page is either blocked from crawlers, marked noindex, or protected behind a login so that neither users nor search engines can access it without permission.
How can a site owner keep AI-generated pages out of search engines?
Add a noindex meta tag to the template that generates each chat, quote, or session URL. Block the URL pattern in robots.txt if the page should never be requested, and require a login for any page that holds personal information. Audit the site regularly to confirm that only intended pages are being indexed.
Related coverage
- Google March 2026 Core Update: What Changed and How Sites Should Respond – BizScoreAI
- Google on AI agents reading your site, and local SEO fixes
SEOScanPro
SEOScanPro has the site audit tool runs a full technical audit of a site and shows the measured result behind every check. Open the site audit tool.
This article summarizes reporting from searchengineland.com. See our editorial disclaimer for how our articles are produced.
Run a free scan to see your AI Visibility Score, SEO rating, and local citation accuracy.
