Department Use Cases

How AI Describes Your Financials: What IR Can Measure, and What It Cannot

2026-08-08Updated 2026-08-29Reading time 23min

By Vaigate Inc. (which operates Vaipm, measuring AI-space perception through a total of 25 stateless queries across multiple AI engines)

Key point

IR adopted generative AI fast, but it measures the AI it uses, not the AI investors read. Where AI gets your financials wrong, and what to measure instead.

Key conclusions

Japanese IR departments have rapidly adopted generative AI as a tool they use themselves. In the 33rd "Fact-Finding Survey on IR Activities" published by the Japan Investor Relations Association (JIRA) on May 14, 2026, 80.5% (n=780) of companies answered that they either "use it in their work" or are "trialing it in their work" when it comes to summarizing and organizing materials. At the same time, 77.2% (n=917) cited "lacking accuracy of information (so-called hallucination)" as a challenge for adoption, the most frequently selected option.

In the same survey, however, the question asking which indicators are used to measure the effectiveness of IR activities offers options built from conventional measures: shareholder composition, number of meetings, market capitalization, PBR and the like. Question Q19-① of the survey form (202605_factfinding.pdf) lists 18 indicator options together with "not conducting effectiveness measurement in particular," but within the scope of this review, we could not confirm any option relating to how AI describes the company.

In other words, the inaccuracy of the AI a company uses itself has been recognized as a challenge, while the inaccuracy of company information inside the AI answers investors read has not yet become an object of measurement. Because IR has statutory disclosure as a reference point established by institution, an error in an AI answer can be treated not as a matter of reputation but as an inconsistency with disclosure. This is the relative difference from communications and HR, where evaluation of reputation is central.

This article sets out where that asymmetry comes from, what generative AI actually gets wrong and why, how far the empirical research goes and where it stops, and what companies must not do. The procedure for actually designing the measurement is covered in the companion article on designing verification tests for AI answers.

What you will learn

  • The asymmetry between generative AI use in IR departments and the indicators used for effectiveness measurement, as shown by JIRA's 33rd survey (read accurately, with the denominators kept separate)
  • Why the correctness of an AI answer can be judged in the IR field, and what cannot be concluded from that
  • The four types of error generative AI makes with financial figures (fiscal period, accounting standard, currency/unit, consolidation scope) and why each occurs
  • What the empirical research does and does not show (checking peer-review status, reservations grounded in experimental conditions, and the distinction between benchmarks and public AI services)
  • What can and cannot be said, at this point, about the difference between HTML and PDF
  • What to measure — the idea of measuring consistency rather than exposure volume
  • The legal and disclosure-rule boundaries that must not be crossed when correcting an answer

Who this is for

IR practitioners and heads of IR at listed companies, and their executive teams. It is written for people who want to understand what AI-space perception of their company currently looks like. If you are looking for the implementation procedure for measurement, it will be faster to start from the companion article on designing verification tests for AI answers.

1. IR departments have adopted generative AI. But they are measuring in the opposite direction

1-1. What happened — the jump from the 2024 survey

On May 14, 2026, the Japan Investor Relations Association published the results of its 33rd "Fact-Finding Survey on IR Activities." The survey covered all 4,088 listed companies as of January 2026, ran from February 9 to March 25, 2026, and drew 948 responses (a response rate of 23.2%). Of those, 917 companies (96.7%) answered that they conduct IR activities.

Questions on generative AI were one focus of this survey. Among companies conducting IR, 80.3% answered that they use it more frequently than a year earlier.

Use by specific task is shown in the table below. The figures are the percentage of companies selecting either "use it in their work" or "trialing it in their work," and the denominator for this question is n=780.

IR-related task2026 (n=780)2024
Summarizing and organizing materials to gather work-related information80.5%13.8%
Preparing minutes of briefings, meetings and similar sessions65.8%8.1%
Drafting business correspondence (email and the like)65.1%8.7%
Preparing English-language disclosure materials (translation and similar work)63.7%16.0%
Preparing documents for briefings, press use and similar purposes51.3%5.5%
Handling Q&A from investors and others37.3%3.5%
Preparing reports such as annual reports and integrated reports32.9%3.2%
Preparing operating manuals26.5%2.5%
Creating images and video for briefings, press use and similar purposes19.1%1.7%

Summarizing and organizing materials rose by 66.7 points in two years. For work centered on generating text, generative AI is no longer an exceptional tool. Image and video creation, by contrast, remains at 19.1%, so the gap between use cases is clear.

1-2. What was cited as a challenge

In the same survey, the question asking about challenges for adopting generative AI in IR-related work (denominator n=917, newly added this time) produced the following order of responses.

Challenge (n=917)Share
Risk of lacking accuracy of information (so-called hallucination)77.2%
Security concerns such as information leakage60.1%
Risk of copyright infringement and ethical problems (bias and the like)45.0%
Insufficient employee skill in using it (writing prompts and the like)34.8%
Inability to handle internal or industry-specific terminology34.1%
Not yet able to substitute for highly specialized work33.3%

Accuracy is the largest challenge. Note, however, what this question is asking about: the accuracy of the output of generative AI that the company itself uses in its work. The question is phrased in terms of challenges for adopting generative AI in IR-related work, and the options sit in the context of internal use.

A note on denominators: several denominators appear in this survey. All responding companies number 948; companies conducting IR number 917; the question on use of generative AI by task has n=780; the question on challenges has n=917. These must not be added together or presented as if they shared one denominator. Note also that "use more frequent than a year earlier, 80.3%" and "summarizing and organizing materials, 80.5%" come from different questions; the figures are close enough to be easily confused.

1-3. Where the effectiveness indicators point

In the same survey, companies conducting IR were asked which indicators they use to measure effectiveness. The results were as follows.

Indicator used for effectiveness measurement (companies conducting IR)20262024
Shareholder composition86.8%84.8%
Change in the number of meetings with analysts and investors69.0%65.1%
Number of analysts and investors attending briefings51.1%47.3%
Market capitalization47.0%39.0%
Share trading volume37.1%31.6%
PBR and similar measures36.9%28.1%

Market capitalization, trading volume and PBR rose by 8.0 points, 5.5 points and 8.8 points respectively, so adoption of share-related indicators is advancing. According to the summary of results, every option outside the top six falls below 30%.

What deserves attention here is which options the question made available. Q19-① of the survey form presents the following 18 indicators, plus "not conducting effectiveness measurement in particular," as options.

  1. Change in the number of meetings with analysts and investors
  2. Number of meetings with target investors
  3. Number and quality of analyst reports
  4. Number of analysts covering the company
  5. Number of analysts and investors attending briefings
  6. Share trading volume
  7. PBR (price-to-book ratio) and similar measures
  8. Market capitalization
  9. Number of inquiries from shareholders and investors
  10. Number of visits to the website
  11. Shareholder composition
  12. Voting turnout and the percentage of votes in favor of proposals
  13. Content of articles by news organizations
  14. Results of investor questionnaires
  15. Degree of understanding of IR activities among top management and within the company
  16. Assessments by ESG rating agencies
  17. Assessment results from third-party organizations other than the above (IR excellence awards and the like)
  18. Other

Looking at this list of options, within the scope of this review, we could not confirm any item relating to AI-space perception of the company, the accuracy of AI answers, or citation by generative AI. The closest in character are "content of articles by news organizations" and "number of visits to the website," but the former concerns news organizations and the latter measures traffic reaching the company's own site. Neither measures what AI says about the company or how it says it.

This is not an assertion that Japanese IR departments do not measure AI-space perception of their companies. The absence of an item among a survey's options and the absence of the practice are two different things. The limited fact is this: in the published materials of the 33rd survey, within the scope of this review, we could not confirm any item relating to AI-space perception among the options for effectiveness indicators.

1-4. AI is classified on the "tools we use" side

The same asymmetry appears in another question. In the question on use of IR support firms, a new option was added this time: "AI-enabled services for briefings and meetings." 5.0% of companies answered that they currently use such services, and 10.0% that they would like to use them in future. Here too, AI is framed as a service the company procures and uses.

A similar pattern can be seen outside Japan. In the United States, NIRI (National Investor Relations Institute) lists the "2025 NIRI and University of Florida Research Survey on AI within IR," conducted jointly with the University of Florida, on its Research page. According to the published description, the survey's purpose is to understand how IR professionals feel about incorporating new technology into their daily work. The body of the results is members-only, and this article has not consulted its contents. We state only the fact that it is listed and the published description of its purpose. NIRI also lists a "NIRI Policy Statement on Artificial Intelligence in IR" on its Policy Statements page. Here too we record only the fact that it is listed.

One more point on companies that do not measure effectiveness. According to the summary of results, 8.9% of companies conducting IR (11.9% previously) do not measure effectiveness, and their reasons were "our IR activities have not reached the stage of measuring effectiveness" at 52.4%, "it is difficult to identify indicators for measuring effectiveness" at 42.7%, and "we do not know an effective method of measurement" at 36.6%. Designing indicators is itself recognized as difficult by practitioners.

What this article addresses is part of that gap. An indicator for company information inside AI answers can be designed as a consistency rate rather than a volume of exposure. And in IR, a consistency rate is easier to define than in other fields, because a reference point exists.

2. A reference point exists as an institution — statutory disclosure and its limits

2-1. The correct answer is machine-retrievable

IR information has a property that is not often found in other kinds of corporate information: what counts as correct can be settled by documents established by institution rather than by the company's own judgment.

  • EDGAR (U.S. SEC) provides submission histories and XBRL company facts through a JSON API. It covers Forms 10-K, 10-Q, 8-K, 20-F, 40-F, 6-K and others, and the SEC makes it available free of charge as a developer resource.
  • EDINET (Financial Services Agency) allows searching of annual securities reports and similar filings, and provides API v2 (an API key is required).
  • TDnet (Tokyo Stock Exchange) publishes timely disclosure in real time. Public viewing runs for 31 days, and search by company covers 10 years.
  • XBRL provides earnings reports, revisions to earnings and dividend forecasts, corporate governance reports and similar documents in machine-readable form.

In other words, the reference values used to check an AI answer can be taken from documents fixed outside the company by institution, rather than from someone's memory or an internal file. This is the starting point for measurement design in IR.

Distinguishing legal character: TDnet is an exchange framework, not national law itself. The timely disclosure framework rests on exchange rules. Statutory disclosure such as the annual securities report, by contrast, rests on the Financial Instruments and Exchange Act. Lumping the two together as "statutory disclosure" will lead to errors in the judgments discussed in §7.

2-2. What cannot be concluded from this

That something is machine-retrievable and that AI actually consults it are entirely different matters.

Claims such as "ChatGPT searches EDINET every time," "Gemini gives priority to the 10-K," or "Perplexity updates immediately after TDnet publication" cannot be concluded within the scope of this review. How each AI service retrieves and ranks company-specific answers is not published, and the process may combine learned knowledge, general web search, partner data, search caches, and materials uploaded by the user.

The position of this article is therefore as follows.

The reference point of statutory disclosure exists. But there is no guarantee that AI consults it and answers with the right fiscal period, currency, unit and accounting standard. That is why we measure.

2-3. The relative difference from other fields

IR is not the only field concerned with AI-space perception of companies. The same problem arises in recruiting (errors in job information), communications (persistence of past coverage) and legal affairs (correcting misinformation). For detail, see Misinformation in AI Answers and How to Address It and How Far Can AI Answers Be Corrected.

The relative difference in IR lies in the degree to which the criterion for right and wrong can be placed outside the company. With reputation or impression, it is hard to draw the line between what is an error and what is an interpretation. "Consolidated net sales for the fiscal year ended March 2025," by contrast, either matches the annual securities report or does not. That difference is why a consistency rate works as an indicator in IR.

Parts of the IR field are of course also hard to judge. Statements about business outlook, evaluation of management, or comparison with competitors carry the same interpretive range as in other fields. What this article addresses is the approach of measuring first the part where the judgment can be settled.

3. Four types of error — fiscal period, accounting standard, currency/unit, scope

Errors in financial figures within AI answers are not scattered at random. In practice they can be organized along four axes. These four types are an organizing scheme this article adopts for practical convenience; they are neither an established classification in the field nor something agreed as a standard. The purpose of the classification is to make it possible to trace a detected inconsistency back to where it originated.

3-1. Wrong fiscal period

What happens: to the question "what were sales in 2025," the answer may come back on a calendar-year basis (January to December 2025), for the fiscal year ended March 2025 (April 2024 to March 2025), on a trailing twelve months (LTM) basis, or as an annualization of quarterly results. Each of these can be called "sales in 2025," so reading the answer alone makes the error hard to notice.

Why it happens: many Japanese companies close their books in March, but a substantial number close in December. Materials for overseas investors often carry calendar-year figures alongside. Multiple period definitions coexist in both training data and search results.

The awkward part: stating the period explicitly does not necessarily resolve it. Beyond the Reported Cutoff (accepted at COLM 2025), discussed below, observed incorrect answers even when using prompts that specified company name, year and sales. It cannot be said that "specifying the fiscal period makes the answer correct."

3-2. Wrong accounting standard

What happens: figures prepared under IFRS, U.S. GAAP and Japanese GAAP become mixed. Because the concept corresponding to operating income differs by standard, the same term "operating income" returns different figures.

Why it happens: voluntary application of IFRS is not unusual among Japanese listed companies. Where a company previously used Japanese GAAP and now uses IFRS, the standard switches partway through a multi-year comparison. And when figures pass through an external data vendor that has standardized them, the original standard information drops out.

Practical consequence: the table used to check AI answers (the reference table; how to build it is covered in the companion article on designing verification tests for AI answers) must include an "accounting standard" column. Without that column, an inconsistency can be detected but its cause cannot be identified.

3-3. Wrong currency or unit

What happens: millions of yen and thousands of yen, millions of U.S. dollars and billions of U.S. dollars, local-currency amounts and converted amounts become mixed. Errors that shift the magnitude by three digits are the easiest to miss when the answer reads fluently.

Why it happens: earnings reports may be in millions of yen, annual securities reports in millions or thousands of yen, and English-language materials in millions of U.S. dollars. Exchange rates and reference dates for conversion also differ between documents.

The awkward part: standardizing units does not necessarily resolve it either. Beyond the Reported Cutoff converted sales to millions of U.S. dollars, defined an answer differing from the reference value by more than 10% as a hallucination, and then observed exactly that. Standardizing units is a necessary condition, not a sufficient one.

3-4. Wrong scope

What happens: consolidated and non-consolidated, continuing operations and company-wide, segment totals and company totals get confused. For companies that have divested businesses or reorganized, prior-year figures change depending on whether they have been retrospectively restated.

Why it happens: companies themselves disclose on multiple bases. Presentation materials on a continuing-operations basis alongside an annual securities report on a company-wide basis is an ordinary occurrence.

Practical consequence: the reference table should carry "consolidated/non-consolidated" and "continuing-operations scope" as separate columns. One of them alone is not enough.

3-5. A note on what cannot be written here

Confusion between companies with identical or similar names is generally discussed as a known risk of large language models. However, within the scope of this review, we could not confirm a published primary case specific to IR. This article therefore presents no figure for how often this type occurs.

Likewise, within the scope of this review, we could not confirm any company-by-company incident register of the form "generative AI misstated Company A's dividend and Company A issued a formal correction". What has been demonstrated extends to the results of controlled experiments and benchmarks. Named incidents and experimental results must not be mixed together.

4. What the empirical research shows (and what it does not)

4-1. Observation of numerical hallucination at scale

Beyond the Reported Cutoff: Where Large Language Models Fall Short on Financial Knowledge (Shah, Ye, Jaskowski, Xu, Chava / Georgia Institute of Technology) is a peer-reviewed paper accepted at COLM 2025 (Conference on Language Modeling). This article refers to the arXiv version (arXiv:2504.00042).

The study built 197,011 question-answer pairs from Compustat sales data. It covers 17,621 companies listed on 13 exchanges over the 43 years from 1980 to 2022. The prompt is fixed in form: "what were the sales of {company name} in {fiscal year}." Numbers were extracted from the answers and scored on three values: correct if the absolute error against the reference value was under 10%, a factual hallucination if 10% or more, and no answer where no figure was returned.

The main observations were as follows.

  • Older years are answered less well. Llama-3-70B-Chat answered accurately for 54.17% of companies for 2017, but only 6.32% for 1995 — even though financial information has been published through EDGAR in the United States since 1995.
  • The study observed a tendency for larger and newer companies to attract both more knowledge and more hallucination. That is, the higher the market capitalization, the higher the rate of correct answers, and at the same time the more incorrect answers about the same companies. This is the result of an experiment on a particular set of models, using a single question form asking about sales, and scored against Compustat values. It must not be read as a law holding for generative AI in general.
  • Search activity, institutional investor attention, access counts for SEC filings, and the readability of filings were also associated with the rate of correct answers.

What cannot be said from this study: whether a paper has been peer-reviewed and how far its results generalize are separate questions. The reservations belong not with peer-review status but with the experimental conditions.

  • Sales figures come from Compustat and are converted to millions of U.S. dollars, so differences in accounting standard are not reconciled. What the study measures is therefore consistency with the values of one standardized database, not consistency with each company's own statutory disclosure.
  • The question takes a single form: "what were the sales of {company name} in {fiscal year}." It does not represent the range of questions investors put in IR practice.
  • The API calls were made in February 2025. These are not figures for the performance of current models.
  • The subjects are U.S.-listed companies. It cannot be said that the findings apply as they stand to Japanese companies.

4-2. Adding the company name alone changes the assessment (Japanese data)

Evaluating Company-specific Biases in Financial Sentiment Analysis using Large Language Models (Nakagawa, Hirano, Fujimoto) is a peer-reviewed conference paper included in the 2024 IEEE International Conference on Big Data (BigData) (pp. 6614–6623, DOI: 10.1109/BigData62323.2024.10826008). A preprint version is also on arXiv (2411.00420).

Handling of this source: it is a peer-reviewed conference paper included in IEEE BigData 2024. However, the arXiv version published by the authors was used to check the methods and figures in the text, and any verbatim or methodological differences from the IEEE version have not been independently confirmed for this article.

The method is straightforward. For the same statement about business performance, the sentiment score returned by a large language model is compared between prompts that include the company name and prompts that do not, and the difference is defined as "company-specific bias." A positive difference means the model shifts its assessment in a favorable direction for that company. Japanese financial text data was used for the analysis.

The implication for IR practice sits in a layer separate from the correctness of figures: even for the same disclosed text, the interpretation can change once a company name is attached. This is a domain where "factually incorrect" is hard to establish, and a consistency rate will not capture it. In the indicator design of §6 of this article, and in the companion article on designing verification tests for AI answers, this layer is treated separately.

What cannot be said from this study: the authors are practitioners and researchers affiliated with Nomura Asset Management and Preferred Networks. It is a peer-reviewed conference paper and not a study conducted to sell a particular service, but it is worth noting that it was designed from the perspective of financial practice. Nor can it be said that a consistent conclusion has been reached across models about how the direction of bias corresponds to company characteristics. It must not be cited as a general rule that "large companies are assessed favorably."

4-3. "Confident and wrong"

Confidently Wrong: Detecting Hallucinations in Financial Question Answering from LLM Internal States (Richard Zhe Wang, submitted July 13, 2026) is an arXiv preprint, and the one paper among the four covered in this section that has not been peer-reviewed (arXiv:2607.11414, 8 pages). The figures below should be read on that basis.

The study uses two financial question-answering benchmarks built from actual filings, FinQA and TAT-QA. It resampled each question eight times and defined the state in which all eight answers agreed as "high confidence." The result: in FinQA, 15–23% of high-confidence answers were incorrect.

The models covered are Qwen3-8B, Llama-3.1-8B and Gemma-2-9B. The focus of the study is whether such errors can be detected by training linear probes on the model's internal states (the residual stream). The probes held an AUROC of 0.68–0.77, exceeding baselines such as token log-probabilities and the model's own self-report of correctness (at best 0.55–0.63).

The implication for IR practice: this is one of the grounds on which this article recommends repeated measurement. Asking the same question again and getting the same answer is not proof of correctness. Consistency and accuracy are different properties.

What cannot be said from this study:

  • The subjects are open-weight foundation models, not the AI search products available to the public. These are not figures for the error rate of products such as ChatGPT or Gemini that combine a model with web search.
  • FinQA and TAT-QA are benchmarks built from filings; they do not reproduce how the services are actually used.
  • The 15–23% figure is a share "of high-confidence answers," not an error rate across all answers.

4-4. What is happening on the investor side

Blankespoor, Croom and Grant, "Generative AI and Investor Processing of Financial Information" is a peer-reviewed paper published in the Journal of Accounting and Economics, Volume 82, Issue 2 (November 2026 issue) (DOI: 10.1016/j.jacceco.2026.101908), by a research team at the University of Washington.

The study uses two kinds of data: archival analysis of more than 400,000 investor queries submitted to a major brokerage's generative AI chatbot, and a survey of more than 2,000 individual investors. In the survey, close to half reported using generative AI. The uses are mainly interpreting financial information and putting market movements in context, along with screening stocks and making complex research more efficient. As users spend longer with the tools, use shifts from high-level screening toward detailed monitoring and interpretation of news about individual companies. More sophisticated individual investors are reported to have adopted the tools further.

On the institutional side there is the Brunswick 2026 US Investor Survey (published February 9, 2026). It has n=100, covering U.S. institutional investors (active equity), half long-only managers and half hedge funds. Brunswick is a firm that supports corporate IR and communications, so this is a survey by an interested party — a point worth stating plainly.

Brunswick 2026 US Investor Survey (n=100)Share
Rate generative AI as at least "moderately important" for investment research54%
Use AI as a primary tool when researching a new investment candidate in depth42%
Say AI has changed how they approach earnings calls68%
Agree they can check more information, spend time on higher-value work, and cover more names84%
Use AI to update financial models23%

On the importance of information sources, 77% rated direct human contact with management as "important" or "very important," company disclosure 66%, dialogue with IR 63%, corporate websites and IR pages 47%, and generative AI output 24%. The questions and scales differ, so a simple comparison of usage rates is not possible, but it can at least be read that AI is not replacing primary materials or human dialogue.

The same survey reports that investors are aware of AI's weaknesses. Outdated information, frequent hallucination, numerical errors and missed context are all cited. In financial forecasting, the use requiring the most caution, only 23% said they use AI to update models, and more than half expressed reservations.

The implication of these two studies: investors use generative AI not as a substitute for primary materials but as a route toward them. Precisely for that reason, the significance of an error about your company inside an AI answer is not that it sways the whole investment decision, but that the path to the primary materials becomes distorted.

4-5. Four disciplines for reading this body of research

Of the three AI studies cited in this section, two have been peer-reviewed — Beyond the Reported Cutoff (COLM 2025) and Evaluating Company-specific Biases (included in IEEE BigData 2024) — and one, Confidently Wrong, is a preprint (on the investor side, Blankespoor and colleagues is a peer-reviewed paper in the Journal of Accounting and Economics). Peer-review status, however, is not the main axis for reading them. Having been peer-reviewed does not guarantee that a result may be applied to your own company. The reading is done on these four points.

  1. Look at the experimental conditions. What was treated as correct (a standardized database such as Compustat, or the filings themselves), what form the questions were limited to, and what threshold defined an error. The 10% threshold in Beyond the Reported Cutoff is also a definition that study set.
  2. Look at the models covered. Open-weight foundation models, or public AI services integrated with web search. Benchmark results for foundation models must not be cited as the performance of public AI services.
  3. Look at where the data came from. U.S.-listed or Japanese companies, what period was covered, and when the work was run. If the API calls were made in February 2025, that is the behavior of the models at that time.
  4. Do not read correlation as causation. Causal readings such as "accurate because it is HTML" or "wrong because the company is large" are supported by none of these studies.

5. HTML and PDF — what can and cannot be said

5-1. What GenAI as a Reader observed

GenAI as a Reader: How ChatGPT & Co. use Annual Reports is a joint study by USTP (St. Pölten University of Applied Sciences), HHL Leipzig Graduate School of Management, and nexxar.

Handle with care: this document is not a peer-reviewed paper but a practitioner report (the Document Type in the HHL repository is Report). In addition, nexxar is a provider of digital IR reports and therefore an interested party. The figures below should be read as observations requiring independent replication. This article does not adopt the publisher's own assessments regarding precedence or scale.

The study has two parts.

Part 1: visibility of annual reports in LLMs

  • It covered 20 listed companies across 9 European countries and 8 industries, split into 10 companies with structured online annual reports in HTML (Group A) and 10 companies publishing annual reports mainly as PDF (Group B).
  • More than 2,500 standardized prompts were submitted to ChatGPT (GPT-4o and GPT-5). 23 academic testers divided the work between them and cross-checked results.
  • 24,662 citations were extracted and classified. Of these, 58.5% were direct links to the companies' annual reports.
  • Annual report content for Group A (HTML) was cited 3.05 times as often as for Group B (PDF).
  • Conversely, answers about PDF-centric companies carried roughly 2.7 times as many citations to external sources. Financial data aggregators such as MarketScreener, Bloomberg and Reuters appear as substitutes for the companies' own sources.
  • On the accuracy of answers, analysis of a partial sample of n=200 found Group A (HTML) correct 71% of the time and Group B (PDF) 54%.

Part 2: log analysis of digital annual reports

  • Server logs were obtained from 5 DAX constituent companies over the eight weeks from August 29 to October 25, 2025.
  • Automated accesses recorded numbered 4,838,833. After data cleansing, 759,226 bot requests became the subject of content analysis.
  • More than 175 identifiable bots were confirmed, and the top five bots accounted for 48.8% of all automated access.
  • The breakdown shows ChatGPT alone at 30.3% of bot traffic, followed by Bingbot at 6.6%, Amazonbot at 5.8%, Baiduspider at 4.6% and Googlebot at 4.6%.
  • By content, descriptions of business activities accounted for 39.3% of access and financial reporting and financial statements for 20.6%. A skew toward the most recent reporting period was also observed.

5-2. What must not be said from this

"Move to HTML and AI answers become correct" cannot be derived from this study.

There are three reasons.

  1. Even in the HTML group, 29% were inaccurate. Behind a 71% rate of correct answers, close to three in ten errors remain. Changing the format does not make errors disappear.
  2. Confounding may be mixed into the correlation. Companies that prepare HTML annual reports may also differ in their IR resourcing, update frequency and disclosure quality. This study is not a randomized experiment that isolated the factors.
  3. The subjects are 20 European companies. That is not a scale from which to generalize to Japanese companies or to other language regions.

Similarly, the claim that "PDFs are not read by generative AI" is incorrect. Google's official documentation states that text-based PDFs can enter the search index (OCR may be applied to image-based PDFs). In the GenAI as a Reader study too, PDF annual reports were cited for every company covered. What differs is the frequency of citation; this is not a binary of read or not read.

Further, a bot user agent does not prove eventual use or adoption into an answer. The fact that access identifying itself as ChatGPT accounted for 30.3% does not mean that the content was actually reflected in answers. Crawling and citation are separate steps.

5-3. On iframes, separate domains and JavaScript

IR sites often embed share price information or disclosure lists in an iframe, use a separate domain for IR, insert a cookie consent screen, or render content only through JavaScript.

As to how these affect answers in AI search, within the scope of this review, we could not confirm a primary empirical study of IR company populations that isolated the factors. This article therefore does not assert that "content in an iframe is not read by AI." Nor does it present a figure for any quantitative decline.

What can be said extends to the general requirements stated in Google's official documentation.

5-4. What Google states officially

Google Search Central's "AI features and your website" states the following in substance about AI Overviews and AI Mode (last updated December 10, 2025).

  • It states that no additional requirements or special optimization are needed to appear in AI Overviews or AI Mode. Ordinary SEO best practices apply as they are.
  • To be eligible to appear as a supporting link, a page needs to be indexed and able to appear in search results with a snippet. No other technical requirement is listed.
  • Important content should be provided as text.
  • Structured data should match the visible text of the page.
  • It states that there is no need to create new machine-readable files for AI, AI text files, or dedicated markup. Special schema.org structured data is likewise stated to be unnecessary.
  • Appearances through AI features are also counted within the "Web" search type of the Search Console performance report.

The practical conclusion that follows is unglamorous. Rather than stacking up AI-specific tactics, it works better to make the key facts retrievable as text, to state the year explicitly on prior-year pages, and to keep structured data consistent with the visible text. llms.txt and FAQ rich results should not be positioned as tactics that work in Google. llms.txt is not used by Google Search and does not affect visibility or ranking. FAQ rich results (the accordion display in search results) ended on May 7, 2026, announced on May 8. Terms such as AIO, GEO and LLMO are likewise practitioner terms rather than official standards.

Alongside this, take care not to mistake the reach of Google's statement. Google states that no dedicated SEO requirements or special markup are needed for AI Overviews or AI Mode. This does not mean, however, that the AI features involve no search, selection or generation processing of their own. The same documentation notes that AI Overviews and AI Mode may use a technique called query fan-out that issues several related searches, and that because the two may use different models and technologies, the combination of answer and links presented can differ. "No dedicated requirements" and "no dedicated processing" are two different things.

6. What to measure — consistency, not exposure volume

6-1. Confirming the premise

To be clear at the outset: recommending AI-space exposure volume as a formal IR KPI is not the position of this article.

As to whether a regulator, a professional body, or the Japan Investor Relations Association has adopted AI-space perception of companies as a formal indicator for measuring IR effectiveness, within the scope of this review, we could not confirm a primary source. As seen in §1, no corresponding item was found among the options in the effectiveness-measurement question of the 33rd survey either.

What follows is therefore a design proposal for an internal indicator (KPI/KRI) for disclosure quality management. It is not a recommendation to position this as a formal IR performance indicator reported to the board.

6-2. Why exposure volume should not be the target

Setting "volume of mentions" or "positive tone" in AI space as the target creates two problems.

  1. The means of improvement do not align with improving disclosure quality. Measures that increase exposure volume can pull in a direction separate from the accuracy of disclosure.
  2. Exposure volume rises with incorrect information too. A state in which wrong answers circulate widely looks like improvement on an exposure metric.

In IR, as seen in §2, a reference point exists as an institution. What should be measured is whether what AI says about your company is consistent with statutory disclosure.

The procedure for turning this idea into actual indicators — defining the consistency rate, choosing the denominator, reporting the no-answer rate alongside it, grading severity, and shaping the report for the executive team — is covered in the companion article on designing verification tests for AI answers.

7. Responses to avoid — the legal and disclosure-rule boundaries

This section deals with an area where mistakes create legal risk. This article is not legal advice. Individual judgments should be checked with your own legal department and outside experts.

7-1. Do not correct an AI answer using undisclosed material information

When an AI answer misstates your company's results, it is natural for someone who knows the correct figure to want to correct it. However, if that figure is undisclosed material information, it must not be used for a correction.

Corrections should be limited to re-presenting already published materials. Specifically:

  • Point to the content of already published earnings reports, annual securities reports and timely disclosure materials as they stand.
  • Do not input or present undisclosed figures, pre-publication earnings outlooks, or information ahead of its scheduled disclosure date, in any form.
  • Where a correction is warranted but no corresponding passage exists in published materials, carry out a lawful publication procedure first.

This is not a formality. Corrections to AI answers tend to be made outside the ordinary disclosure flow — in a chat, an inquiry form, or a supplementary social media post. Material information leaving the disclosure flow is precisely what the rules are aimed at.

7-2. Do not mistake the legal character of each framework

Framework or documentLegal characterScope
U.S. Regulation FDAn SEC final rule (adopted 2000)Selective disclosure of material non-public information by issuers
Japan's Fair Disclosure RuleA statutory framework under the Financial Instruments and Exchange Act (in force April 1, 2018). The FSA guidelines set out a general interpretationCommunication of material information by listed companies and others
Timely disclosure via TDnetAn exchange framework (not national law itself)Timely disclosure of information material to investment decisions
ESMA documents on AIA supervisory statement and initial guidance (not new EU legislation)Use of AI by investment firms
IOSCO documents on AIA report and a non-binding supervisory toolkit (requiring implementation in national law)Use of AI and associated risks in capital markets
FSA AI Discussion PaperA discussion paper (neither legislation nor supervisory guidelines)Organizing the issues around AI use in the financial sector

Do not lump these together as "regulation." Statutes and final rules, on one hand, and supervisory statements, discussion papers and non-binding toolkits, on the other, carry entirely different compliance obligations. Confusing them in internal materials leads to the wrong order of priorities.

7-3. AI inferring undisclosed results is not itself an FD violation

This is a point often misunderstood in practice.

Even where a third party's generative AI infers and describes your company's undisclosed results from public information, that in itself does not violate Regulation FD or Japan's Fair Disclosure Rule. What both frameworks address is, in essence, the act of an issuer communicating material non-public information to particular market participants. Inference by a third party's model is not selective communication by the issuer.

The state of AI inferring results does not, therefore, by itself immediately create any disclosure obligation for the company. What can become a problem is the company responding to it by releasing undisclosed information.

7-4. Input to external generative AI is a separate question

A third party's AI inferring something independently from public information is not communication under the FD rules. On the other hand, where the issuer inputs or provides undisclosed material information to an AI provider or similar party, separate questions arise.

  • Information leakage: input data may be used for training or operations.
  • Confidentiality obligations: where the information concerns business partners or employees, contractual or statutory obligations may be engaged.
  • Internal control: a route arises that internal information controls do not reach.
  • Consideration under the FD rules: depending on the facts, including the counterparty's legal position and any confidentiality obligation, whether this amounts to an act of communication may come into question.

The understanding that "inputting undisclosed material information to an external AI provider is automatically a violation of the FD rules" is not accurate. U.S. Regulation FD limits the recipients it covers and excludes persons owing a duty of confidentiality, among others, and Japan's Fair Disclosure Rule likewise has requirements concerning transaction-related parties and confidentiality obligations. Whether it applies depends on the facts: who the counterparty is and what obligations they owe. In any event, from the standpoint of information leakage, confidentiality and internal control, avoiding such input is the practical default.

As seen in §1, 60.1% (n=917) of companies in the 33rd survey cited "security concerns such as information leakage" as a challenge. In the same survey, 55.2% of companies (32.1% previously) now have company-wide guidelines on the use of generative AI, and "no guidelines established" fell to 18.9% (53.3% previously). Guidelines are being put in place rapidly.

This boundary needs to hold in verification-test practice as well. Only published figures go into the reference table. Undisclosed figures must not be sent to an external service as the "correct answer."

7-5. Do not treat AI output as a substitute channel for explanation

A state in which AI describes your company in detail must not be treated as "part of the explaining being done."

  • Do not treat AI output as a substitute channel for selective explanation. Presenting an AI answer to a particular investor in place of an explanation amounts to communicating information outside the disclosure framework.
  • Do not rewrite the language of statutory disclosure into a form AI might prefer. The purpose of a disclosure document is to give investors accurate information, not to optimize machine readability. Supplementary provision on the IR site (also presenting key facts as HTML text, stating the year explicitly on prior-year pages, and so on) must be distinguished from changing what the statutory disclosure itself says.

7-6. What could not be confirmed about correction mechanisms

What can a company do when it finds an error in an AI answer?

At this point, within the scope of this review, we could not confirm a correction mechanism for companies common to the major AI services, or any procedure guaranteeing the outcome of a correction. Individual services may offer feedback functions, but it cannot be said that these are established as a mechanism with processing deadlines or guaranteed outcomes.

The realistic response therefore looks like this. "Correcting an AI answer" does not mean requesting deletion. Treat it as a matter of disclosure quality management.

  1. Save reproducible prompts together with the date and time.
  2. Identify exactly where the answer contradicts published primary materials.
  3. Fix gaps on your own IR site: missing content, outdated pages, unit display, and entity identification (variations in how the company name is written, and separation from identically named companies).
  4. After an interval, re-verify across multiple AI engines.

Rather than demanding deletion, put correct information into a retrievable state and measure the effect. That is the response this article recommends. Manipulating word of mouth, reverse-SEO techniques, and leaving the matter alone on the assumption that "it will disappear with time" are all approaches we do not recommend.

8. FAQ

Q1. Why can the correctness of an AI answer be judged in the IR field?

Because a reference point exists as an institution. Annual securities reports, earnings reports and timely disclosure materials can be retrieved mechanically through EDGAR, EDINET, TDnet and XBRL. For something like "consolidated net sales for the fiscal year ended March 2026," matching the AI answer against published materials settles whether it is consistent or not. With reputation or impression it is hard to draw the line between error and interpretation, and that is the relative difference. Parts of the IR field are also hard to judge, however. Statements about business outlook, evaluation of management, or comparison with competitors carry the same interpretive range as in other fields. Measuring first the part where the judgment can be settled is the practical order of work.

Q2. If we move our IR site to HTML, will AI answers become correct?

That cannot be said at this point. In "GenAI as a Reader," the joint study by USTP, HHL and nexxar, companies with HTML annual reports were cited by ChatGPT 3.05 times as often as PDF-centric companies, and the rate of correct answers in a partial sample (n=200) was also higher, at 71% against 54%. Even in the HTML group, however, 29% were inaccurate. The subjects were 20 European companies, and because there was no random assignment — the two groups are simply different sets of companies — differences in company size and IR resourcing may be confounded. It is also worth noting that this document is a practitioner report rather than a peer-reviewed paper, and that nexxar, a provider of digital IR reports, took part in it. It is at a stage requiring independent replication.

Q3. Is it true that PDFs are not read by AI?

That is not accurate. Google's official documentation states that text-based PDFs can enter the search index (OCR may be applied to image-based PDFs). In the study mentioned above, PDF annual reports were cited for every company covered. The difference is the frequency of citation, not a binary of read or not read. That said, tables spanning pages, footnotes, two-column layouts, and tables whose units sit away from the main body all leave room for extraction errors.

Q4. AI is describing our undisclosed results by inference. Is this a fair disclosure violation?

Where a third party's generative AI has simply inferred and described this independently from public information, that in itself does not, as a rule, amount to communication under U.S. Regulation FD or Japan's Fair Disclosure Rule. What both frameworks address is, in essence, the act of an issuer communicating material non-public information to particular market participants. What can become a problem is the company responding to it by releasing undisclosed information. Note also that where the issuer inputs or provides undisclosed material information to an AI provider or similar party, a separate question arises, and consideration under the FD rules may be needed depending on the facts, including the counterparty's legal position and any confidentiality obligation. This article is not legal advice. Please check individual judgments with your legal department and outside experts.

Q5. If an AI answer is wrong, may we correct it by giving the right figure?

If you correct it, limit the correction to re-presenting already published materials. Undisclosed figures, pre-publication earnings outlooks, and information ahead of its scheduled disclosure date must not be presented in any form. Where no corresponding passage exists in published materials, a lawful publication procedure has to come first. Corrections to AI answers tend to be made outside the ordinary disclosure flow, in a chat or an inquiry form, and that is where the greatest care is required.

Q6. Should AI-space exposure volume be an indicator for measuring the effectiveness of IR activities?

This article does not recommend it. As to whether a regulator or professional body has adopted AI-space perception of companies as a formal indicator for measuring IR effectiveness, within the scope of this review, we could not confirm a primary source. Nor was any corresponding item found among the options in the effectiveness-measurement question of JIRA's 33rd survey. What we do recommend is measuring, as an internal indicator for disclosure quality management, whether what AI says about your company is consistent with statutory disclosure. If exposure volume becomes the target, a state in which incorrect information is spreading will also look like improvement. Specific definitions of the indicators are covered in the companion article on designing verification tests for AI answers.

Q7. Do Japanese IR departments not measure AI-space perception of their companies?

That cannot be asserted. What can be said is that, in the published materials of JIRA's 33rd "Fact-Finding Survey on IR Activities," within the scope of this review, we could not confirm any item relating to AI-space perception among the options in the question asking which indicators are used to measure the effectiveness of IR activities. The absence of an item among a survey's options and the absence of the practice are two different things. In the same survey, however, use of generative AI in IR-related work rose sharply (80.5% for summarizing and organizing materials, n=780), and hallucination was the most frequently cited challenge at 77.2% (n=917). The asymmetry is visible: "the accuracy of the AI we use" has become a challenge, while "the accuracy of the AI answers investors read" has not yet been made into a question.

Q8. What practical value is there in separating errors into types?

It lets you identify what to fix. Knowing that the overall consistency rate is 80% does not determine what to do next. But if the breakdown of inconsistencies shows that "60% are wrong fiscal periods," the likely suspects are prior-year pages on the IR site without the year stated, or several period definitions coexisting. If "currency and unit errors are frequent," you check unit notation in English-language disclosure. Mapping error types one-to-one onto the columns of the table used for checking makes detection and cause identification a single continuous process. Table design is covered in the companion article on designing verification tests for AI answers.

Q9. Are investors really using AI to research companies?

What the surveys show is use confined to particular purposes. A peer-reviewed paper by a research team at the University of Washington (Journal of Accounting and Economics, 2026, Volume 82, Issue 2) analyzed more than 400,000 queries to a major brokerage's generative AI chatbot and a survey of more than 2,000 individual investors, and reports that close to half use generative AI. The uses are mainly interpretation and contextualization of information. On the institutional side, Brunswick's 2026 survey (n=100, a survey by an interested party, an IR support firm) put the importance of information sources at 77% for contact with management, 66% for company disclosure and 63% for dialogue with IR, against 24% for generative AI output. AI is not replacing primary materials or human dialogue.

Q10. Should we change how we write disclosure documents to make them easier for AI to read?

We do not recommend rewriting the statutory disclosure itself into a form AI might prefer. The purpose of a disclosure document is to give investors accurate information. What should change is how the IR site presents things: also providing key facts as text, stating the year explicitly on prior-year pages, and keeping structured data consistent with the visible text. According to Google's official documentation, no additional requirements or special optimization are needed to appear in AI Overviews or AI Mode, and there is no need to create new machine-readable files for AI or dedicated markup. This does not mean, however, that the AI features involve no processing of their own. Note too that terms such as AIO, GEO and LLMO are practitioner terms rather than official standards.

Q11. Are the four types of error a classification treated as standard in the industry?

No. The four types — fiscal period, accounting standard, currency/unit and scope — are an organizing scheme this article adopts for practical convenience. They are neither an established classification in the field nor something agreed as a standard. The purpose of setting them out is to make it possible to trace a detected inconsistency back to where it originated. Depending on what your company discloses, a different set of axes may work better. For a company with a high overseas sales ratio, making the reference date for currency conversion an independent axis is one option; for a company with frequent business reorganizations, making the presence of retrospective restatement an independent axis is another.

9. Summary and next actions

9-1. Key points

  1. IR departments have adopted generative AI, but they are measuring in the direction of "the AI we use." In JIRA's 33rd survey, use for summarizing and organizing materials reached 80.5% (n=780), and hallucination was the most frequently cited challenge at 77.2% (n=917). Meanwhile, within the scope of this review, we could not confirm any item relating to AI-space perception among the options in the question on effectiveness indicators.
  2. IR has statutory disclosure as a reference point established by institution. Reference values can be obtained from EDGAR, EDINET, TDnet and XBRL. However, it cannot be concluded within the scope of this review that AI always consults them.
  3. Errors can be organized into four types. Fiscal period, accounting standard, currency/unit, and scope. This is this article's organizing scheme, however, not an industry-standard classification.
  4. Read empirical research by its experimental conditions, not its peer-review status. The companies covered, the question format, the models covered, and when the work was run. Having been peer-reviewed does not guarantee that a result may be applied to your own company.
  5. HTML is not a cure-all. Even in the HTML group 29% were inaccurate, and the study had no random assignment — the two groups are simply different sets of companies.
  6. Measure consistency, not exposure volume. If exposure volume becomes the target, a state in which incorrect information is spreading will also look like improvement.
  7. Limit corrections to re-presenting already published materials. AI answers must not be corrected using undisclosed material information.

9-2. Next actions

This week: ask the major generative AI services for your company's most recent full-year net sales, operating income and dividend per share. Do this without carrying over history, and put the same question several times. Do not judge from a single answer.

This month: match the answers you get against the corresponding passages in the annual securities report and earnings report. Where there is an inconsistency, classify whether it stems from fiscal period, accounting standard, currency/unit or scope.

At your next results announcement: put the same question on the day of the announcement and again two weeks later, and record how the answers change. This is the right timing for observing when and how AI answers shift once disclosure has been updated.

The procedure for turning this into a continuing system — how to build the reference table, what to record, how to keep measurement conditions aligned, how to define the indicators, and how to report to the executive team — is covered in the companion article on designing verification tests for AI answers.

9-3. On sustaining measurement

Everything above can be done in-house. But making repeated, stateless measurement across several AI services a routine, and keeping records under aligned conditions, carries real operating cost. All the more so if the aim is to follow each disclosure rather than run a one-off study each quarter.

Vaipm measures AI-space perception through a total of 25 stateless queries across multiple AI engines. This is not a tool for AIO tactics; it is a mechanism for continuously managing perception in AI space (AI Perception Management). On the field of measuring AI-space perception itself, see What Is AIPM; on how long information persists in AI space, see How Long Information Stays in AI Answers. General measures against misinformation and its correction are covered in Misinformation in AI Answers and How to Address It and How Far Can AI Answers Be Corrected.

Sources

Surveys and statistics

  1. Japan Investor Relations Association (JIRA), "Results of the 33rd Fact-Finding Survey on IR Activities" (published May 14, 2026) — news release https://www.jira.or.jp/download/survey/202605_newsrelease.pdf / Subjects: all 4,088 listed companies; 948 responses (response rate 23.2%); 917 companies conducting IR. The generative AI usage question with n=780 appears on page 6/9 of the same PDF, "Graph 2. Use of generative AI in IR-related work (n=780)," and the challenges question with n=917 on page 7/9, "Graph 3. Challenges for adopting generative AI in IR-related work (n=917)"
  1. Same, "Summary of Results of the 33rd Fact-Finding Survey on IR Activities" https://www.jira.or.jp/download/survey/202605_summary.pdf
  1. Same, "Fact-Finding Survey on IR Activities" survey form (February 2026) — includes the options for the effectiveness-measurement question (Q19-①) https://www.jira.or.jp/download/survey/202605_factfinding.pdf
  1. Same, survey and research index https://www.jira.or.jp/activity/research.html
  1. NIRI (National Investor Relations Institute) Research page — confirmed listing of the "2025 NIRI and University of Florida Research Survey on AI within IR." The body of the results is members-only and this article has not consulted its contents https://www.niri.org/publications/research/
  1. NIRI Policy Statements page — confirmed listing of the "NIRI Policy Statement on Artificial Intelligence in IR." The text is members-only and this article has not consulted its contents https://www.niri.org/resources/policy-statements/
  1. Brunswick Group, "Brunswick's 2026 US Investor Survey" (published February 9, 2026) — n=100, U.S. institutional investors (active equity). A survey by an interested party, a firm supporting IR and corporate communications https://review.brunswickgroup.com/article/investor-survey-2026/

Academic research

  1. Agam Shah, Liqin Ye, Sebastian Jaskowski, Wei Xu, Sudheer Chava, "Beyond the Reported Cutoff: Where Large Language Models Fall Short on Financial Knowledge" (Georgia Institute of Technology) — accepted at COLM 2025 (Conference on Language Modeling). The text referred to here is the arXiv version (arXiv:2504.00042). 197,011 QA pairs, 17,621 companies, 1980–2022, Compustat-derived and converted to millions of U.S. dollars. Single question format. API calls made in February 2025 https://arxiv.org/html/2504.00042v2
  1. Kei Nakagawa, Masanori Hirano, Yugo Fujimoto, "Evaluating Company-specific Biases in Financial Sentiment Analysis using Large Language Models" — 2024 IEEE International Conference on Big Data (BigData), pp. 6614–6623, a peer-reviewed conference paper. DOI: 10.1109/BigData62323.2024.10826008. Preprint version: arXiv:2411.00420. The arXiv version published by the authors was used to check the methods and figures in the text, and any verbatim or methodological differences from the IEEE version have not been independently confirmed for this article. The authors are affiliated with Nomura Asset Management and Preferred Networks https://arxiv.org/abs/2411.00420
  1. Richard Zhe Wang, "Confidently Wrong: Detecting Hallucinations in Financial Question Answering from LLM Internal States" (submitted July 13, 2026) — arXiv preprint, 8 pages. FinQA and TAT-QA benchmarks. Models covered: Qwen3-8B, Llama-3.1-8B, Gemma-2-9B. A benchmark of foundation models, not an evaluation of the AI search products available to the public https://arxiv.org/abs/2607.11414
  1. Elizabeth Blankespoor, Joe Croom, Stephanie M. Grant, "Generative AI and Investor Processing of Financial Information" — a peer-reviewed paper published in the Journal of Accounting and Economics, Volume 82, Issue 2 (November 2026 issue). DOI: 10.1016/j.jacceco.2026.101908. University of Washington. More than 400,000 chatbot queries; a survey of more than 2,000 individual investors https://doi.org/10.1016/j.jacceco.2026.101908
  1. Monika Kovarova-Simecek, Henning Zülch, Leon Kirschbaum, Konstantin Klammer, Eloy Barrantes, Alexandra Horváthová, Christina Schilling, "GenAI as a Reader: How ChatGPT & Co. use Annual Reports" (USTP, HHL Leipzig, nexxar) — not a peer-reviewed paper but a practitioner report (Document Type in the HHL repository: Report). nexxar is a provider of digital IR reports and therefore an interested party. This article does not adopt the publisher's own assessments regarding precedence or scale. Part 1: 20 European companies (10 HTML / 10 PDF), more than 2,500 prompts, GPT-4o/GPT-5, 24,662 citations; the rate of correct answers comes from a partial sample of n=200. Part 2: 5 DAX companies, August 29 to October 25, 2025, 4,838,833 automated accesses, 759,226 analyzed https://digital-investor-relations.com/_assets/downloads/DIR_GenAI-as-Reader.pdf?h=E0wP6LL9

Primary and technical materials

  1. U.S. Securities and Exchange Commission, "EDGAR Application Programming Interfaces" https://www.sec.gov/search-filings/edgar-application-programming-interfaces
  1. Financial Services Agency, EDINET https://disclosure2.edinet-fsa.go.jp/week0020.aspx
  1. Japan Exchange Group, "TDnet (Timely Disclosure network)" — an exchange framework, not national law itself. Public viewing 31 days; search by company 10 years https://www.jpx.co.jp/english/equities/listing/disclosure/tdnet/
  1. Japan Exchange Group, "XBRL" https://www.jpx.co.jp/english/equities/listing/disclosure/xbrl/03.html
  1. Google Search Central, "AI features and your website" (last updated December 10, 2025) — no additional technical requirements for AI Overviews / AI Mode; no dedicated markup needed; structured data should match the visible text https://developers.google.com/search/docs/appearance/ai-features
  1. Google Search Central Blog, "PDFs in Google search results" — text-based PDFs can be indexed https://developers.google.com/search/blog/2011/09/pdfs-in-google-search-results

Regulation and institutional frameworks (with legal character noted)

  1. U.S. SEC, "Selective Disclosure and Insider Trading (Regulation FD)" — an SEC final rule (adopted 2000) https://www.sec.gov/rule-release/33-7881
  1. Financial Services Agency, "Publication of the Cabinet Office Ordinance and related documents on the Fair Disclosure Rule" — Japan's Fair Disclosure Rule is a statutory framework under the Financial Instruments and Exchange Act (in force April 1, 2018). The FSA guidelines set out a general interpretation https://www.fsa.go.jp/news/29/syouken/20180206.html
  1. ESMA, "ESMA provides guidance to firms using artificial intelligence in investment services" — a supervisory statement and initial guidance (not new EU legislation) https://www.esma.europa.eu/press-news/esma-news/esma-provides-guidance-firms-using-artificial-intelligence-investment-services
  1. IOSCO, "Artificial Intelligence in Capital Markets" — a report (no direct legal force) https://www.iosco.org/library/pubdocs/pdf/IOSCOPD788.pdf
  1. IOSCO, "AI Supervisory Toolkit" — a non-binding supervisory toolkit (requiring implementation in national law) https://www.iosco.org/library/pubdocs/pdf/IOSCOPD823.pdf
  1. Financial Services Agency, "AI Discussion Paper" — a discussion paper, neither legislation nor supervisory guidelines https://www.fsa.go.jp/news/r7/sonota/20260303/aidp.html

Disclaimer: This article is practical information on disclosure quality management and the organization of institutional frameworks, and is not legal advice. It also offers no recommendation on investment decisions regarding any particular security, no share price forecast, and no investment advice. Please consult your own legal department and professionals such as attorneys on individual legal judgments.

The Vaipm perspective

Vaipm measures AI-space perception through a total of 25 stateless queries across multiple AI engines. This is not a tool for AIO tactics; it is a mechanism for continuously managing perception in AI space (AI Perception Management). In IR, a reference point exists as an institution, so what should be measured is whether what AI says about the company is consistent with statutory disclosure — not the volume of exposure.

Related articles

Department Use Cases

Do AI Crawlers Read Your IR Site's JavaScript? | AIO & LLMO for IR (Technical)

Are the JavaScript parts of your IR site reaching AI crawlers? A 41-day controlled experiment and large-scale log observation, plus how to check your own site.

IROnsiteCrawlersFetchabilityAIPM
Read more
Department Use Cases

AIO & LLMO for IR | Which Version of Your Disclosure Is AI Describing? — Forecasts, Actuals, and Corrective Disclosure

AI can quote a figure you published and still be wrong about its status. How IR adds status and version columns, sets precedence, and re-measures on disclosure.

IRDisclosureReference TableVersion ManagementAIPM
Read more
Department Use Cases

What AI Cites When It Describes Your Company — The Reputation Supply Chain for PR

What AI cites about your company. Japan and global data: McKinsey estimates owned sites at 5-10% of AI sources; 37.9% of citations in the first 10 SERP blocks.

PR & CommunicationsAI perception managementAIPMCitationsAMEC
Read more
Practical Guides

When the Language Changes, So Does the AI's Answer — Cross-Border AI Perception for Companies Expanding Abroad

Ask in English and you get a different answer than you get in Japanese. That prompt language changes what a model outputs is demonstrated in several peer-reviewed studies. This guide for companies operating abroad sets out why translation is a variable rather than a switch, the path by which Japanese-language primary information is less likely to be picked up by an English question, how citation sources and regulation differ market by market, and which received ideas do not hold — all organized from primary sources and peer-reviewed research.

Cross-BorderMultilingualAI Perception ManagementAIPMGlobal Expansion
Read more
Risks & Issues

AI Misinformation and Misattribution: Detect, Correct, Prevent

AI misinformation and misattribution management is the practice of continuously managing, across three layers of detection, correction, and prevention, the risk that generative AI or answer engines describe, attribute, or summarize your company wrongly. Separate from the problem of "not being cited by AI" (absence), there is the problem of "being cited, but with the content wrong" (false presence). AI citation is not a matter of careful operation but is structurally incomplete (Tow Center, ALCE, CiteFix), and the presence of a source link does not guarantee accuracy. In Japan, 87.3% of corporate staff have witnessed false presence and 76.7% say they "are measuring," yet false presence has not stopped. The problem is not the absence of measurement but that the way of measuring does not prove a state. How to deal with false presence divides into three questions: can it be challenged legally, can it be removed, and how is it measured. This article is the entry point; each question is explored in depth in a separate article.

AI misinformationMisattributionReputationAIPMRisk
Read more