Department Use Cases

Should You Block AI Crawlers on Your IR Site? What to Check Before You Edit robots.txt

2026-09-24Reading time 17min

By Vaigate Inc. (which operates Vaipm, measuring AI-space perception through a total of 25 stateless queries across multiple AI engines)

Key point

Should you block AI crawlers on your IR site? Filings have other machine routes, but IR-only content can lose reach. Check citation and reach separately.

Summary

Should you block generative-AI crawlers on your investor relations site, or let them through? There is no single setting that answers this in one line.

There are two reasons. The first is that AI crawlers are separated by purpose, and each provider draws the lines differently. OpenAI lets you address its training crawler and its search crawler separately, but it also states that robots.txt rules may not apply to fetches initiated by a user. Google's Google-Extended governs Gemini training and grounding through a single token, and has no effect on inclusion in Google Search. "Blocking AI" as a single monolithic action does not exist.

The second reason is specific to IR. A listed company's disclosures are machine-retrievable from places other than its own IR site. EDGAR in the United States, and EDINET and TDnet in Japan, are routes built for programmatic retrieval — though not on equivalent terms.

So blocking crawlers on your IR site does not necessarily cut off access to the filed disclosures altogether. On the other hand, explanations that exist only on your own site — results presentations, FAQs, business descriptions — can lose their path to machines entirely when you block. The heart of the decision is to check two things separately: the opportunity for your own domain to be cited as the source, and whether the information itself can be reached.

This article sets out the material for that decision. It does not tell you how to configure anything. It covers what to check before you change a setting, and what a measurement taken afterwards can and cannot tell you.

Whether AI crawlers execute the JavaScript on your site — the question of reach — is covered in the article on the technical conditions for AI crawlers and IR sites. Whether the numbers inside a disclosure PDF are correctly extracted is covered in the article on reading disclosure documents. This article addresses the step before both: whether to let them fetch at all. For the wider IR picture, see the parent article on AI perception for IR.

1. Separated by purpose

1-1. OpenAI separates training from search

The major providers publish their crawlers split by purpose. OpenAI's developer documentation sets out the following.

Table 1: OpenAI's documented crawler distinctions
NamePurpose (as documented)
GPTBotCrawling content that may be used in training OpenAI's generative AI foundation models
OAI-SearchBotSurfacing websites in search results in ChatGPT's search features
ChatGPT-UserCertain user actions in ChatGPT and Custom GPTs, such as visiting a page when a user asks a question. Because these actions are initiated by a user, the documentation says robots.txt rules may not apply
OAI-AdsBotValidating the safety of landing pages submitted as ads on ChatGPT. The documentation adds that landing-page content may also be used to determine when an ad is relevant to show

What this split means is that "do not use this for training, but do show it in ChatGPT's search answers" is a setting you can actually write. In practice: disallow GPTBot, allow OAI-SearchBot. The documentation states that each setting is handled separately from the others. It also notes that sites opted out of OAI-SearchBot will not be shown in ChatGPT search answers, though they can still appear as navigational links.

ChatGPT-User is a different matter. As Table 1 shows, the same documentation states that robots.txt rules may not apply to it because its fetches are user-initiated. What can be said accurately about OpenAI, then, is that GPTBot for training and OAI-SearchBot for search can be addressed separately. It does not follow that user-initiated fetching can be controlled through robots.txt in the same way.

The same documentation notes that a robots.txt change takes time to propagate — it can take around 24 hours for the provider's systems to adjust. Measuring immediately after a change and concluding that nothing happened is too early.

1-2. Google draws the lines differently

Google does not follow the same shape as OpenAI, and this is an area where practical confusion is common.

Google's documentation says the following about Google-Extended:

Google-Extended doesn't have a separate HTTP request user agent string. Crawling is done with existing Google user agent strings; the robots.txt user-agent token is used in a control capacity.

In other words, no crawler called Google-Extended actually arrives at your site. The fetching is done by existing Google crawlers; Google-Extended is a name used in robots.txt for control purposes.

According to the same page, Google-Extended governs two things: whether content may be used to train future generations of the Gemini models that power Gemini Apps and the Vertex AI API for Gemini, and whether it may be used for grounding — providing content to the model at prompt time — in Gemini Apps and in Grounding with Google Search on Vertex AI. On Search, the page is explicit:

Google-Extended does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search.

Two consequences follow.

First, within the scope of Google-Extended, the distinction "do not train Gemini on this, but do use it to ground answers in Gemini Apps" cannot be expressed through that one token, because training and grounding sit under the same name. That is a different split from OpenAI's separation of training and search.

Second, because Google-Extended is documented as having no effect on inclusion in Google Search, it cannot be read as a way to control how you are treated inside Google Search. The search side has to be considered separately, through the rules search engines document for indexing (§4). For IR, that means checking use in products such as Gemini Apps and inclusion in Google Search as two separate questions.

If you write "disallow every AI-related token" without knowing this design difference, you will get different outcomes per provider. That is not a configuration mistake; it follows from the specifications not being aligned.

2. What happens if you block

Several studies bear on what follows from blocking. They are set out below by the kind of evidence they offer. The four studies measure different things. Every figure is an observation limited to that study's sample and period.

2-1. Peer-reviewed measurement

A peer-reviewed study was presented at the ACM Internet Measurement Conference in 2025. It analysed 24 AI crawler user agents, top-ranked site lists, and robots.txt history from October 2022 to October 2024, covering more than 40,000 robots.txt files among consistently top-ranked sites alone.

What this study mainly measured is whether crawlers comply with robots.txt. It did not measure citation in AI answers. Two findings matter here.

First, in a six-month observation of researcher-controlled sites, nine AI crawlers visited. Seven obeyed robots.txt. One fetched the robots.txt file and then did not follow its directives. The remaining one, ChatGPT-User, did not fetch robots.txt at all. The ChatGPT-User observation is consistent with OpenAI's statement that robots.txt rules may not apply to user-initiated fetches (§1-1). Either way, writing a rule in robots.txt does not make a resource technically unreachable.

Second, between August 2023 and October 2024, 484 sites removed a GPTBot restriction they had previously set. Some were publishers that had announced commercial arrangements with the provider. The authors themselves state that they have not established that the agreements caused the configuration changes. This is a correlation.

2-2. A non-peer-reviewed causal estimate

A 2026 working paper analysed how news publishers responded. It used 30 newspaper domains as the main analysis and the top 500 publishers as a supplementary one, covering November 2022 to May 2024.

What this study measured is site-wide web traffic for publishers that introduced restrictions. It did not measure citation in AI answers.

Using a staggered difference-in-differences design, it estimates a roughly 7% decline in web traffic in the six weeks after blocking. Three measurement series agreed in direction, but two of them had confidence intervals crossing zero, so statistical precision varies by series.

Three limitations matter. First, this is not a direct measurement of AI referral traffic falling by 7%. AI referrals were small at the time, and the authors discuss weakened exposure and distribution as a possible mechanism.

Second, the causal analysis is restricted to the period before May 2024, which the authors state is to avoid confounding from the introduction of AI summaries in Google Search. The estimate therefore says nothing about the period after AI search became widespread.

Third, it is a non-peer-reviewed working paper, and the estimate is not something to transfer directly onto your own company.

2-3. Cited despite blocking

An observation published in May 2026 collected 24,127 citations across 330 prompts from several providers' web-search APIs. Counting the 22,263 citations for which robots.txt could be retrieved, 7.8% came from domains that block the training-side crawler, and 2.1% from domains that block the answer-generation crawler. The remaining citations were excluded from the calculation because robots.txt could not be retrieved due to network or DNS errors. For OpenAI specifically, 1.1% of citations came from domains blocking OAI-SearchBot.

What this observation measured is how many of the citations already appearing in answers came from blocking domains. It did not measure how much citation falls when a site blocks. Blocked sources were also not automatically demoted to the bottom of the citation order in this sample.

The limitations are significant. This observes APIs, not the consumer products, and it cannot separate licensed supply, third-party indexes, or previously retrieved data. The authors themselves say this should not be read as "robots.txt is meaningless."

2-4. A vendor observation

There is also an observation by a company that sells AI-visibility tooling. It observed citations for 1,058 domains over two days, normalised against organic search appearance, and reported that domains blocking a given provider's crawler showed markedly lower citation propensity in that same provider's answers.

What this observation measured is the correspondence (correlation) between blocking status and citation propensity. It is not a randomised experiment, and the observation window is two days. Sites that block crawlers also differ in business model and scale. The observing company states that the result is directional and does not establish causation. It is an observation by a vendor of AI-visibility tooling, and should be read on those terms.

2-5. What each study supports

The four studies measure different outcomes. They cannot be added together into a single conclusion.

Table 2: What each study measured, and what this article can use
StudyWhat it measuredWhat this article can use
Peer-reviewed measurement (§2-1)Whether crawlers comply with robots.txtThat some fetching ignores robots.txt, and some never fetches it. Effect on citation not measured
Non-peer-reviewed working paper (§2-2)Site-wide traffic for publishers that introduced restrictionsAn estimate for news publishers before May 2024. Citation in answers not measured
API observation (§2-3)Share of existing citations that come from blocking domainsThat citation appears despite blocking. Neither the size of any decline nor the retrieval route was measured
Vendor observation (§2-4)Correspondence between blocking status and citation propensityA two-day correlation. Causation not shown

What can be said is this much:

A relationship between crawler blocking and citation inside AI answers has been observed. The size of the effect and whether it is causal, however, fall within different scopes in each study. Citation despite blocking has also been observed.

These observations cannot identify why citation appears despite blocking. The next section turns to a structure specific to IR that is one candidate explanation.

3. What blocking does not remove

This is the part specific to IR, and the centre of this article.

3-1. Disclosures have separate retrieval routes

A listed company's disclosures do not exist only on the issuer's IR site. Routes built for programmatic retrieval exist alongside it.

Table 3: Machine-retrieval routes for disclosures
RouteOperatorConditions of use
EDGARU.S. Securities and Exchange CommissionFree, no registration. Provides APIs and indexes intended to assist automated processing. Under its "Fair access" guidance, the current maximum is 10 requests per second, users are asked to declare a user agent in request headers and to use efficient scripting, and requests from automated tools outside the acceptable policy are managed to ensure fair access for all users
EDINETFinancial Services Agency, JapanAccount and API key required. Connections to the API are authenticated with an API key, and obtaining a key requires creating an account. There is one API for retrieving a list of filings and another for retrieving the documents themselves, with a published specification
TDnet APIJPX Research (Japan Exchange Group)Offered as a paid information service. It sits under the listed-company paid information category on the Japan Exchange Group site

The three are not open on the same terms. EDGAR can be queried without registration; EDINET requires an account and an API key; the TDnet API is a paid service. Collapsing all three into "publicly available" hides that difference.

Which organisations actually use these routes, and which AI providers the retrieved information then reaches, is not something this article has been able to establish.

What can be said is this much:

Even if you block AI crawlers on your own IR domain, statutory and timely disclosures that have been filed remain retrievable by machines through routes separate from your site. That a route exists, however, is a different matter from whether a given provider's AI actually retrieves through it.

3-2. Separate reach from the chance to be cited

This lets the decision be split into two axes.

One is whether the information itself can be reached. For filed disclosure documents, a separate route exists even if you block your own site, so reach is not necessarily lost altogether. Explanations that exist only on your own site — results presentations, FAQs, business descriptions — are a different case. Blocking can remove the path by which machines reach them at all.

The other is the opportunity for your own domain to be cited as the source. When an AI answer states the same earnings figure, whether it attributes that to your IR page, to a disclosure portal, to press coverage, to an aggregator, or to a competitor's comparison article is a separate matter.

If citations of your IR pages decline, whether other sources replace them, whether the number of citations simply falls, or whether the answer itself changes has to be measured separately. What to check is whether the composition of sources used in the answers changes.

The previous section noted observations of citation persisting despite blocking. In IR, the existence of separate retrieval routes for disclosures is one candidate explanation. The observation in §2-3, however, cannot separate licensed supply, third-party indexes, or previously retrieved data, so the actual retrieval route cannot be identified from it.

3-3. What may change depending on whose domain is cited

"The opportunity to be cited" sounds abstract, but in IR practice it can make a concrete difference.

When a disclosure portal is the source, what can be checked there is the filed primary disclosure itself. Filed documents do contain explanation of the business and results. But the scope and structure of that explanation differ from the presentation decks, FAQs and business descriptions you place on your own IR site.

Your own IR site may carry the company's supplementary account — why a figure sits where it does, what assumptions were made. Whether that gets cited can affect which explanation an AI answer builds on.

When press coverage or an aggregator is the source, the situation shifts again. Someone else's summary and assessment is layered in. It may be accurate; it may be out of date; qualifications may have been dropped. Someone else's summary may be filling in what the company has not explained.

That problem exists separately from crawler configuration. Where a company is cited but described incorrectly, or where another company's matter is told as yours, the article on AI misinformation and misattribution covers it. What this article can say is only that crawler configuration can influence the composition of that supply.

3-4. Keep the two goals apart

The same structure supports a further point.

Filing disclosures through the regulatory channels and being cited as the source inside AI answers can be considered separately. The filing and distribution routes for statutory and timely disclosures exist apart from your IR site's AI crawler settings. Whether your own domain is cited, and whether explanations that exist only on your site can be reached by machines, do depend on your site's settings.

Conflating the two into "we must allow crawlers or our disclosures won't get through" misplaces the basis for the decision. Conflating them the other way — "there is another route anyway, so we may as well block" — drops both the reach of explanations that exist only on your site and the point in §3-3 about the composition of the supply. Separate the two goals and give each its own basis.

4. The means are not equivalent

There is more than one way to block or allow, and they do not work the same way. In IR, a key question is whether it works on PDFs.

One distinction first. HTML meta tags and HTTP response header directives are documented by Google as means of controlling indexing and presentation in Google Search results. Read search-engine indexing and serving controls separately from the usage controls each AI provider defines for itself. How a given AI provider treats these directives has to be checked in that provider's own published documentation.

Table 4: Control mechanisms and their effect on PDFs
MechanismWhat it controlsPDFLimits
robots.txtWhether a given crawler may fetch a URLWorks (PDFs can be named as paths)Depends on the crawler complying. Does not remove already-retrieved data. May not apply to user-initiated fetches
HTML meta tagsSearch indexing and snippet behaviour (in Google's documentation)Does not workA PDF has no HTML tags of its own; and a crawler that cannot fetch the HTML cannot read the tag
HTTP response header directivesSearch indexing and snippet behaviour (in Google's documentation)WorksGoogle documents this as the way to block indexing of non-HTML resources such as PDFs. Whether an AI provider supports it has to be checked individually
WAF or CDN blockingThe HTTP request itselfWorksStronger enforcement than robots.txt, but blocking on a user agent string alone can be spoofed
AuthenticationAccess to the content itselfWorksRestricts access itself, but is normally unsuitable for public retail-investor IR
Licensed feedSupply to contracted parties by a separate routeWorksCan be managed separately from crawler policy

4-1. If your results materials are PDF-only

This is common in IR practice, and the following differences bite.

  • Disallowing the PDF path in robots.txt means a compliant crawler will not fetch the PDF body.
  • HTML meta tags do not apply to the PDF itself. Google's documentation directs site owners to the X-Robots-Tag HTTP response header to block indexing of non-HTML resources such as PDFs.
  • An HTTP header directive on the PDF response can specify indexing and snippet behaviour in Google Search. Whether an AI provider follows that directive has to be checked provider by provider.

Treating PDFs on the assumption that "we set noindex in the meta tag, so we're covered" leaves the intended control unapplied.

One further point: stopping retrieval and removing something from a search index are not the same operation. Google's documentation states that a page disallowed in robots.txt can still be indexed if linked to from other sites, in which case the URL can appear in results without a description. It also states that meta tags and X-Robots-Tag headers are discovered when a URL is crawled, so on a URL disallowed in robots.txt any indexing rules will not be found and will be ignored. Blocking and removing are not achieved by the same action.

5. The institutional picture

5-1. robots.txt is not access control

The robots.txt specification was standardised as RFC 9309 in September 2022, on the Standards Track. The document itself states:

The Robots Exclusion Protocol is not a substitute for valid content security measures.

It also notes that listing paths in robots.txt exposes them publicly and makes them discoverable. Where actual access control is needed, the RFC directs implementers to valid measures in the application itself, such as HTTP authentication.

robots.txt is a request that works only on those who honour it. The observation in §2-1 of a crawler that did not follow the directives is the other side of that property.

5-2. The treatment differs by jurisdiction

Viewed from the rights side, the position is not uniform. Two jurisdictions are set out here, for which the statutory text and official material could be checked.

  • In the EU, Article 4 of the Directive on copyright and related rights in the Digital Single Market (Directive (EU) 2019/790) provides an exception or limitation for reproductions and extractions for text and data mining. Article 4(3) applies that exception on condition that use has not been expressly reserved by rightholders in an appropriate manner, giving machine-readable means as the example for content made publicly available online. A directive takes effect through member states' national law, so the concrete treatment depends on each state's legislation.
  • In Japan, Article 30-4 of the Copyright Act addresses uses not aimed at enjoying the thoughts or sentiments expressed in a work, including use for information analysis. The article contains no provision under which a rightholder's stated reservation removes a use from its scope; its proviso excludes cases that would unreasonably prejudice the interests of the copyright owner. The "General Understanding on AI and Copyright" issued by the Legal Subcommittee of the Copyright Subdivision of the Council for Cultural Affairs (15 March 2024) states that it is difficult to interpret a rightholder's expressed objection, in itself, as excluding a use from the limitation. It then discusses how technical measures such as access restrictions in robots.txt may, together with other circumstances, bear on whether the proviso applies. The document states that it is not itself legally binding.

Other jurisdictions are not covered here. All of this is beyond the scope of this article. Where a legal assessment is required, please consult a qualified professional. This article is not legal advice.

5-3. Notations that providers treat differently

Notations such as noai and noimageai exist, used mainly for images. In November 2022, DeviantArt announced that it had introduced them, delivered through an HTML meta tag and an HTTP header, to communicate that content is not authorised for use in AI training.

How each AI provider treats these notations, however, has to be checked in each provider's own published documentation. They do not appear, for example, in the list of valid rules in Google's robots meta tag specification.

Ways of expressing preferences about AI use in machine-readable form are under discussion in the IETF AI Preferences working group, which is working on a vocabulary and on ways of attaching preferences through robots.txt and HTTP headers. As of 23 September 2026, no RFC published by that working group could be found.

6. Common misconceptions

Four, stated in the negative.

6-1. "Block GPTBot and ChatGPT won't cite us"

Incorrect. GPTBot is the training-side agent; retrieval at answer time runs through OAI-SearchBot and ChatGPT-User. robots.txt rules may not apply to ChatGPT-User (§1-1). Citations have also been observed from blocking domains (§2-3). In IR, the existence of separate disclosure routes (§3-1) is a further candidate explanation.

6-2. "Allow OAI-SearchBot and we'll be cited"

Not guaranteed. Allowing removes an obstacle to retrieval; no published documentation promises citation. Allowing may be necessary without being sufficient.

6-3. "Block in robots.txt and we won't be trained on"

What can be said extends only to reducing future direct retrieval by that crawler. It does not reach data already retrieved, copies collected by others, or data supplied under agreement. robots.txt is not a mechanism for removing anything retroactively.

6-4. "Ignoring robots.txt is illegal"

It cannot be stated as a general rule. As §5-2 shows, the position from the rights side differs by jurisdiction. And as §5-1 shows, robots.txt was not designed as an access control mechanism. Not honouring a technical request and the legal assessment of that conduct are settled in separate frameworks.

7. What to measure, before and after

Everything above is decision material. But "should we block" is not settled by material alone. Without measuring how your company is currently treated, neither choice has a basis.

Blocking and allowing both need to rest on measurement.

7-1. Before you block

Three things.

First, whether your IR domain is currently cited as a source inside AI answers. If it is, you can see what blocking might give up. If it is not, what that tells you is only that "no citation of our domain was observed for the questions we measured." Questions you did not measure, and use that does not surface as a citation, fall outside what that measurement can show. "We are not cited, so blocking costs nothing" does not follow.

Second, if it is cited, on which engine and for which questions. Given that specifications differ by provider (§1), results will not be uniform. What gets cited for an earnings question differs from what gets cited for a business-description question.

Third, if it is not cited, what is cited instead. This matters in practice. A disclosure portal, press coverage, an aggregator, or a competitor's comparison. The reason your company is not cited is not necessarily that you are blocking. If you are not blocking and still are not cited, several factors need to be separated: whether retrievable text exists and how the information is structured, but also relevance to the question, how search and citation candidates are selected, freshness, and language. Rewriting robots.txt alone may not change the outcome.

7-2. After you block

Fourth, how the three above changed. Anchor the comparison to the date the robots.txt change was made. As noted in §1-1, propagation can take time; measuring the next day and concluding is too early. Even if something changed, the change cannot be attributed to robots.txt alone (§7-4).

Fifth, what it means if nothing changed. This is easy to misread. No change does not mean the block had no effect. The separate retrieval routes in §3-1 are one candidate explanation. Others include propagation delay, fetches such as ChatGPT-User to which robots.txt may not apply, and data already retrieved. Observing the answers alone cannot tell these apart.

7-3. Designing the questions

Measuring does not mean asking an AI what it thinks of your company. In an IR context, you build the questions an investor would plausibly ask.

Table 5: Question types and what each reveals
Question typeExampleWhat it reveals
Naming the companyRecent results, business mix, stated policyWhat is used as the source when describing you
Without naming the companyNotable companies in the sector; companies meeting stated criteriaWhether you surface at all, and who surfaces instead
Alongside competitorsA comparison of several peers on a metricWhich sources are used in a comparison context
By your own initiative namesThe name of a plan or programme you have publishedWhether content only you have published is reflected in answers

The fourth type ties directly to crawler configuration. If questions about content only you have published produce no answer, that text may not be being retrieved — though an empty answer does not prove it was not retrieved.

Conversely, if such content comes back correctly, that is an observation that the information is in some form available to the model or to the search system. But training data, earlier retrieval, republication by others, and retrieval at answer time cannot be told apart from the answer alone. It cannot identify the retrieval route or when retrieval happened.

Phrase the questions the way an investor would. Asking in internal jargon or in-house abbreviations guarantees an empty answer, and measures nothing.

7-4. Why once is not enough

Generative AI does not return the same answer to the same question every time. Measuring once and recording "we were cited" or "we were not" does not establish a state.

Unless the measurement conditions are fixed, repeated, and disclosed, a before-and-after comparison does not hold. The conditions to fix are the number of repetitions, running without carrying over login state or history, spanning multiple engines, the language, and the design of the questions. A figure published without its conditions cannot be interpreted.

A robots.txt change is an operation with a recorded date and a recorded content. With the measurement conditions fixed, the before-and-after difference can be observed. That difference alone, however, does not identify the causal effect of the robots.txt change, because model updates, search index updates, news coverage and new disclosures can all fall in the same period.

To separate the change from what else happened at the same time, you need comparison questions — about pages whose settings you did not change, or about other companies — measured alongside, at several points in time. The difference-in-differences design in §2-2, which compares against publishers that did not introduce restrictions, is built to remove such concurrent change. Without fixed conditions, you cannot even separate the before-and-after difference from measurement variance.

7-5. How we measure

Vaipm measures AI-space perception through a total of 25 stateless queries across multiple AI engines. Stateless means measuring without carrying over login state or prior exchanges. Repeating under identical conditions is what makes it possible to read a state out of the variance.

Changes in answers, citations and source composition discussed in this article can be observed in that form: whether your IR domain is cited as a source, what is cited instead if it is not, and what moved across a robots.txt change. Identifying the retrieval route itself, or a causal link to the change, requires verification beyond observing the answers.

The design of the measurement itself — how to build the answer key, what to record, how to define the indicators — is covered in the article on designing AI answer verification tests.

8. Summary

One. There is no monolithic "block AI" operation. With OpenAI, GPTBot for training and OAI-SearchBot for search can be addressed separately, but robots.txt rules may not apply to user-initiated ChatGPT-User fetches. Google's Google-Extended governs Gemini training and grounding under one token and does not affect inclusion in Google Search. Each provider draws the lines differently.

Two. The studies measure different things. The peer-reviewed study measured whether crawlers comply with robots.txt; the working paper, publisher traffic; the API observation, how many existing citations come from blocking domains; the vendor observation, the correlation between blocking and citation propensity. A relationship between blocking and citation has been observed, but the size of the effect and its causal status fall within different scopes in each study. Citation despite blocking has also been observed.

Three. In IR, filed disclosures have retrieval routes separate from your own site. Blocking therefore does not necessarily cut off access to that information altogether. Explanations that exist only on your site, however, can lose their path entirely. Check the chance to be cited and the reach of the information separately.

Four. The mechanisms are not equivalent. HTML meta tags in particular do not work on PDFs. And what Google's documentation covers is search indexing and serving, which should be read separately from AI providers' usage controls.

Five. robots.txt is not access control. It is a request that works only on those who honour it.

And finally: whichever way you go, measure first. Without knowing whether you are cited today, and what is cited instead if you are not, a change in configuration leaves you unable to explain the outcome either way. Even with measurement, a before-and-after difference can be observed, but attributing it to the change requires comparison questions and observation at several points in time.

One note on scope: official documents from financial regulators directly stating how a listed company should configure generative-AI crawlers on its IR site were not found as of 23 September 2026, within a search of the SEC and FSA websites and the materials from those two bodies listed in the sources. Documents from securities exchanges and IR industry bodies were not checked this time. That is a statement that they were not found, not that they do not exist.

9. Frequently asked questions

Q1. So should we block or allow?

This article does not prescribe either. The answer turns on three facts: whether your IR domain is cited as a source in AI answers today; if it is not, what is cited instead; and how much explanation — results presentations, FAQs, business descriptions — exists only on your own site. Filed disclosures have separate retrieval routes, but explanations that exist only on your site can lose their path to machines entirely when you block. Measure those three things before deciding either way (§7).

Q2. Can we adopt a policy of "we do not want to be trained on"?

You can. What robots.txt delivers, however, is a reduction in future direct retrieval by the crawler you name. It does not reach data already retrieved, copies collected by others, or data supplied under agreement. Providers also divide purposes differently: with OpenAI you can disallow the training crawler, GPTBot, on its own, whereas Google's Google-Extended covers Gemini training and grounding under a single name. Once you set the policy, check how to express it against each provider's published documentation (§1, §6-3).

Q3. Why do Google and OpenAI require different settings?

Because they divide purposes differently. OpenAI publishes GPTBot for training and OAI-SearchBot for search as separate crawlers, and each can be addressed separately. It also states that robots.txt rules may not apply to ChatGPT-User, because those fetches are initiated by a user. Google's Google-Extended, by contrast, is not a crawler that arrives at your site but a control token used in robots.txt, and it covers Gemini training and grounding under one name. A single "block all AI" line therefore produces different results per provider (§1).

Q4. Our results materials are PDFs. What should we watch for?

HTML meta tags do not apply to PDFs, because a PDF has no HTML tags of its own. Google's documentation directs site owners to the X-Robots-Tag HTTP response header to block indexing of non-HTML resources such as PDFs. That header, however, controls indexing and serving in Google Search. How an AI provider treats the same directive has to be checked in that provider's own published documentation. Assuming that "noindex in the meta tag" covers a PDF leaves the intended control unapplied (§4).

Q5. We changed robots.txt and nothing changed. Did it not work?

That does not follow. OpenAI's documentation notes that a robots.txt change can take around 24 hours to take effect. Some fetches, such as ChatGPT-User, may not be subject to robots.txt at all, and data already retrieved is not removed. In IR, the separate disclosure routes are a further candidate explanation. Observing the answers alone cannot tell these possibilities apart, so do not conclude either that the change failed or that the information is arriving by another route (§1-1, §3-1, §7-2).

Q6. Is this legal advice?

No. This article is not legal advice. The overview in §5-2 covers the EU and Japan, where the statutory text and official material could be checked: Article 4 of Directive (EU) 2019/790, and Article 30-4 of Japan's Copyright Act. The concrete treatment depends on national law and the circumstances of each case, and other jurisdictions are not covered in this article. Neither overview assesses any particular use of your content. Where a legal assessment is required, please consult a qualified professional in the relevant jurisdiction.

Q7. Does disallowing Google-Extended affect our inclusion or ranking in Google Search?

Google's documentation states that Google-Extended does not impact a site's inclusion in Google Search and is not used as a ranking signal there. What it governs is use of content to train Gemini models that power Gemini Apps and the Vertex AI API for Gemini, and use for grounding in Gemini Apps and on Vertex AI. For IR, that means treating use in products such as Gemini Apps and inclusion in Google Search as two separate questions, each checked on its own (§1-2).

Q8. If our disclosures are on EDGAR or EDINET, does it matter if we block our IR site?

For filed disclosure documents, a machine-retrievable route exists apart from your site, so reach is not necessarily lost altogether. But a route existing is a different matter from whether a given provider's AI actually retrieves through it. And explanations that exist only on your own site — results presentations, FAQs, business descriptions — can lose their path to machines entirely when you block. Filing through the regulatory channels and being cited as the source in AI answers are separate goals. Check the chance for your domain to be cited and the reach of the information separately (§3).

Q9. Why are some sites cited even though they block?

The published observations cannot identify the reason. In the observation published in May 2026, 2.1% of the citations for which robots.txt could be retrieved came from domains blocking the answer-generation crawler. The authors themselves state that they cannot separate licensed supply, third-party indexes, or previously retrieved data. Blocked sources were also not automatically demoted to the bottom of the citation order in that sample. In IR, the existence of separate disclosure routes is one further candidate explanation, but none of these observations shows which route was actually used (§2-3, §3-2).

Q10. If we disallow a PDF's URL in robots.txt, will it disappear from search results?

Not necessarily. Google's documentation states that a page disallowed in robots.txt can still be indexed if other sites link to it, in which case the URL can appear in results without a description. It also states that meta tags and X-Robots-Tag headers are read when a URL is crawled, so on a URL disallowed in robots.txt any indexing rules are not found and are ignored. Stopping retrieval and removing a URL from search results are not achieved by the same action (§4-1).

Q11. Do the figures from these studies apply to our company?

Not directly. The four studies measure different things: the peer-reviewed study measured whether crawlers comply with robots.txt; the working paper, traffic for news publishers; the API observation, how many existing citations come from blocking domains; and the vendor observation, a two-day correlation between blocking status and citation propensity. Each is limited to its own sample and period; the working paper's causal estimate, for example, covers only the period before May 2024. The effect on your own company has to be measured for your own company (§2).

Q12. If we measure before and after a robots.txt change, will that show the effect of the change?

With the measurement conditions fixed, the before-and-after difference can be observed. That difference alone, however, does not identify the causal effect of the change, because model updates, search index updates, news coverage and new disclosures can all fall in the same period. Measure comparison questions alongside — about pages whose settings you did not change, or about other companies — and observe at several points in time, so that the change can be separated from what else happened at the same time (§7-4).

10. Sources

  1. Google Search Central, "List of Google's common crawlers" (last updated 14 July 2026). Statements that Google-Extended has no separate HTTP user agent string and is used in a control capacity; that it governs training of future Gemini models powering Gemini Apps and the Vertex AI API for Gemini, and grounding in Gemini Apps and Grounding with Google Search on Vertex AI; and that it does not affect inclusion or ranking in Google Search. https://developers.google.com/search/docs/crawling-indexing/google-common-crawlers
  2. OpenAI, "Bots" developer documentation. Purpose distinctions for GPTBot, OAI-SearchBot, ChatGPT-User and OAI-AdsBot; that each setting is handled separately; that robots.txt rules may not apply to user-initiated ChatGPT-User actions; that OAI-AdsBot may use landing-page content to judge ad relevance; the time required for a robots.txt change to take effect; and the note that sites opted out of OAI-SearchBot can still appear as navigational links. https://developers.openai.com/api/docs/bots
  3. RFC 9309, "Robots Exclusion Protocol" (September 2022, Standards Track). Statements that the protocol is not a substitute for valid content security measures, and that listing paths makes them discoverable. https://www.rfc-editor.org/rfc/rfc9309.html
  4. U.S. Securities and Exchange Commission, "Accessing EDGAR Data" (last reviewed or updated 26 June 2024). Under "Fair access": the current maximum of 10 requests per second, the request for efficient scripting, and the management of requests from automated tools outside the acceptable policy. Also the request to declare a user agent, the RESTful APIs on data.sec.gov and indexes for automated processing, and index update timing. https://www.sec.gov/search-filings/edgar-search-assistance/accessing-edgar-data
  5. Enze Liu, Elisa Luo, Shawn Shan, Geoffrey M. Voelker, Ben Y. Zhao, Stefan Savage, "Somesite I Used To Crawl: Awareness, Agency and Efficacy in Protecting Content Creators From AI Crawlers," ACM Internet Measurement Conference 2025 (peer-reviewed). Nine AI crawlers visited the authors' sites over six months from September 2024 to March 2025: seven respected robots.txt, one fetched robots.txt and did not respect it, and ChatGPT-User did not fetch robots.txt; 484 sites removed explicit GPTBot restrictions between August 2023 and October 2024; 40,455 sites had a robots.txt file in every Common Crawl snapshot analysed. https://arxiv.org/abs/2411.15091
  6. Hangcheng Zhao, Ron Berman, "Strategic Response of News Publishers to Generative AI" (non-peer-reviewed working paper; arXiv 2512.24968v4, 15 April 2026). Main analysis of 30 newspaper domains, extended analysis of the top 500; data from November 2022 to May 2024. Estimates over the six weeks after blocking: SimilarWeb −0.074 [−0.141, −0.007], Semrush −0.069 [−0.145, 0.007], Comscore −0.065 [−0.150, 0.021]. The causal analysis is restricted to before May 2024. https://arxiv.org/abs/2512.24968
  7. OpenAttribution, "Measuring content influence in AI assistants," published May 2026. 24,127 citations and 6,798 unique domains collected from 330 prompts. Using the 22,263 citations for which robots.txt was retrieved as the denominator, 7.8% came from domains blocking the training-side crawler and 2.1% from domains blocking the live-search bot; for OpenAI, 1.1% (39/3,576) came from domains blocking OAI-SearchBot. Mean citation rank was 4.1 for blocked sources and 4.4 for permitted sources. The authors state that this is a probabilistic audit of providers' public output, not of internal retrieval logs. https://openattribution.org/research/measuring-content-influence-in-ai-assistants
  8. cloro, "Do Sites That Block GPTBot Get Cited Less by ChatGPT?" 1,058 domains, robots.txt successfully fetched for 88%, two-day citation window. An observation by a vendor of AI-visibility tooling. The authors state that "the data can't prove causation, and domains that block crawlers may differ in other ways." https://cloro.dev/research/ai-crawler-blocks/
  9. Financial Services Agency, Japan, Corporate Disclosure Division, "EDINET API Specification (Version 2)," June 2026. That the EDINET API lets users retrieve disclosure data through programs; that connections are authenticated with an API key and obtaining a key requires creating an account; and that a document-list API and a document-retrieval API are provided. https://disclosure2dl.edinet-fsa.go.jp/guide/static/disclosure/download/ESE140206.pdf
  10. Japan Exchange Group, "TDnet API." Listed under the listed-company paid information category. https://www.jpx.co.jp/markets/paid-info-listing/tdnet/02.html
  11. Google Search Central, "Robots meta tag, data-nosnippet, and X-Robots-Tag specifications" (last updated 24 March 2026). That these rules control how a page is indexed and served in Google Search results; that X-Robots-Tag is used to block indexing of non-HTML resources such as PDFs; that indexing rules on a URL disallowed in robots.txt will not be found and will be ignored; and the list of valid rules. https://developers.google.com/search/docs/crawling-indexing/robots-meta-tag
  12. Google Search Central, "Introduction to robots.txt" (last updated 10 December 2025). That robots.txt is not a mechanism for keeping a page out of Google; that a disallowed page can still be indexed if linked to from other sites; and that its URL can then appear in results without a description. https://developers.google.com/search/docs/crawling-indexing/robots/intro
  13. Directive (EU) 2019/790 of the European Parliament and of the Council of 17 April 2019 on copyright and related rights in the Digital Single Market (EUR-Lex). Article 4(3), applying the text and data mining exception on condition that use has not been expressly reserved by rightholders in an appropriate manner, such as machine-readable means. https://eur-lex.europa.eu/eli/dir/2019/790/oj/eng
  14. Copyright Act of Japan (Act No. 48 of 1970), Article 30-4 (e-Gov Law Search). Provision on uses not aimed at enjoying the thoughts or sentiments expressed in a work, including information analysis, and the proviso for cases that would unreasonably prejudice the copyright owner's interests. https://laws.e-gov.go.jp/law/345AC0000000048
  15. Legal Subcommittee, Copyright Subdivision, Council for Cultural Affairs, "General Understanding on AI and Copyright" (Japanese original「AIと著作権に関する考え方について」, 15 March 2024, Agency for Cultural Affairs). That a rightholder's expressed objection, in itself, is difficult to interpret as excluding a use from the limitation; the discussion of technical measures such as robots.txt access restrictions; and the statement that the document itself is not legally binding. https://www.bunka.go.jp/seisaku/bunkashingikai/chosakuken/pdf/94037901_01.pdf
  16. DeviantArt Team, "UPDATE All Deviations Are Opted Out of AI Datasets" (11 November 2022). Description of the noai and noimageai directives delivered through an HTML meta tag and an HTTP header. https://www.deviantart.com/team/journal/UPDATE-All-Deviations-Are-Opted-Out-of-AI-Datasets-934500371
  17. IETF, "AI Preferences (aipref)" working group. Work on a vocabulary for AI usage preferences and on attaching those preferences to content through means including robots.txt and HTTP response header fields. https://datatracker.ietf.org/wg/aipref/about/

Of the sources above, items 1 to 9 and 11 to 17 were checked directly against the primary source on 23 September 2026. For item 9, the body of the specification PDF was read by machine on that date and the statements were confirmed. For item 14, the statutory text was retrieved on that date through the e-Gov Law Search API. For item 10, the page did not return a response to automated retrieval on that date, so it has not been directly verified by machine. The statement in item 10 rests on the categorisation used on the Japan Exchange Group site.

The Vaipm perspective

Vaipm measures AI-space perception through a total of 25 stateless queries across multiple AI engines. It asks in four ways (naming the company, asking by category without naming it, placing it alongside competitors, and asking about specific initiatives by name) and records how AI answers describe the company. Before you change robots.txt, it gives you a record of how you are described today.

Related articles