Q1. So should we block or allow?
This article does not prescribe either. The answer turns on three facts: whether your IR domain is cited as a source in AI answers today; if it is not, what is cited instead; and how much explanation — results presentations, FAQs, business descriptions — exists only on your own site. Filed disclosures have separate retrieval routes, but explanations that exist only on your site can lose their path to machines entirely when you block. Measure those three things before deciding either way (§7).
Q2. Can we adopt a policy of "we do not want to be trained on"?
You can. What robots.txt delivers, however, is a reduction in future direct retrieval by the crawler you name. It does not reach data already retrieved, copies collected by others, or data supplied under agreement. Providers also divide purposes differently: with OpenAI you can disallow the training crawler, GPTBot, on its own, whereas Google's Google-Extended covers Gemini training and grounding under a single name. Once you set the policy, check how to express it against each provider's published documentation (§1, §6-3).
Q3. Why do Google and OpenAI require different settings?
Because they divide purposes differently. OpenAI publishes GPTBot for training and OAI-SearchBot for search as separate crawlers, and each can be addressed separately. It also states that robots.txt rules may not apply to ChatGPT-User, because those fetches are initiated by a user. Google's Google-Extended, by contrast, is not a crawler that arrives at your site but a control token used in robots.txt, and it covers Gemini training and grounding under one name. A single "block all AI" line therefore produces different results per provider (§1).
Q4. Our results materials are PDFs. What should we watch for?
HTML meta tags do not apply to PDFs, because a PDF has no HTML tags of its own. Google's documentation directs site owners to the X-Robots-Tag HTTP response header to block indexing of non-HTML resources such as PDFs. That header, however, controls indexing and serving in Google Search. How an AI provider treats the same directive has to be checked in that provider's own published documentation. Assuming that "noindex in the meta tag" covers a PDF leaves the intended control unapplied (§4).
Q5. We changed robots.txt and nothing changed. Did it not work?
That does not follow. OpenAI's documentation notes that a robots.txt change can take around 24 hours to take effect. Some fetches, such as ChatGPT-User, may not be subject to robots.txt at all, and data already retrieved is not removed. In IR, the separate disclosure routes are a further candidate explanation. Observing the answers alone cannot tell these possibilities apart, so do not conclude either that the change failed or that the information is arriving by another route (§1-1, §3-1, §7-2).
Q6. Is this legal advice?
No. This article is not legal advice. The overview in §5-2 covers the EU and Japan, where the statutory text and official material could be checked: Article 4 of Directive (EU) 2019/790, and Article 30-4 of Japan's Copyright Act. The concrete treatment depends on national law and the circumstances of each case, and other jurisdictions are not covered in this article. Neither overview assesses any particular use of your content. Where a legal assessment is required, please consult a qualified professional in the relevant jurisdiction.
Q7. Does disallowing Google-Extended affect our inclusion or ranking in Google Search?
Google's documentation states that Google-Extended does not impact a site's inclusion in Google Search and is not used as a ranking signal there. What it governs is use of content to train Gemini models that power Gemini Apps and the Vertex AI API for Gemini, and use for grounding in Gemini Apps and on Vertex AI. For IR, that means treating use in products such as Gemini Apps and inclusion in Google Search as two separate questions, each checked on its own (§1-2).
Q8. If our disclosures are on EDGAR or EDINET, does it matter if we block our IR site?
For filed disclosure documents, a machine-retrievable route exists apart from your site, so reach is not necessarily lost altogether. But a route existing is a different matter from whether a given provider's AI actually retrieves through it. And explanations that exist only on your own site — results presentations, FAQs, business descriptions — can lose their path to machines entirely when you block. Filing through the regulatory channels and being cited as the source in AI answers are separate goals. Check the chance for your domain to be cited and the reach of the information separately (§3).
Q9. Why are some sites cited even though they block?
The published observations cannot identify the reason. In the observation published in May 2026, 2.1% of the citations for which robots.txt could be retrieved came from domains blocking the answer-generation crawler. The authors themselves state that they cannot separate licensed supply, third-party indexes, or previously retrieved data. Blocked sources were also not automatically demoted to the bottom of the citation order in that sample. In IR, the existence of separate disclosure routes is one further candidate explanation, but none of these observations shows which route was actually used (§2-3, §3-2).
Q10. If we disallow a PDF's URL in robots.txt, will it disappear from search results?
Not necessarily. Google's documentation states that a page disallowed in robots.txt can still be indexed if other sites link to it, in which case the URL can appear in results without a description. It also states that meta tags and X-Robots-Tag headers are read when a URL is crawled, so on a URL disallowed in robots.txt any indexing rules are not found and are ignored. Stopping retrieval and removing a URL from search results are not achieved by the same action (§4-1).
Q11. Do the figures from these studies apply to our company?
Not directly. The four studies measure different things: the peer-reviewed study measured whether crawlers comply with robots.txt; the working paper, traffic for news publishers; the API observation, how many existing citations come from blocking domains; and the vendor observation, a two-day correlation between blocking status and citation propensity. Each is limited to its own sample and period; the working paper's causal estimate, for example, covers only the period before May 2024. The effect on your own company has to be measured for your own company (§2).
Q12. If we measure before and after a robots.txt change, will that show the effect of the change?
With the measurement conditions fixed, the before-and-after difference can be observed. That difference alone, however, does not identify the causal effect of the change, because model updates, search index updates, news coverage and new disclosures can all fall in the same period. Measure comparison questions alongside — about pages whose settings you did not change, or about other companies — and observe at several points in time, so that the change can be separated from what else happened at the same time (§7-4).