Department Use Cases

How AI Describes Your Financials: An IR Guide to AI Accuracy

2026-08-08Reading time 26min

By Vaipm (which measures AI-space perception through a total of 25 stateless queries across multiple AI engines)

Key point

When AI describes your earnings or dividends incorrectly, IR faces something other than a reputation problem: a conflict with published disclosure. Drawing on an evaluation of 17,621 listed companies, peer-reviewed research using Japanese financial results summaries, and the Japan Investor Relations Association's 2026 survey, this article sets out how IR should measure and manage corporate perception in AI answers, and what it should not measure, on the basis of primary sources.

The conclusion of this article

IR information has a property that other categories of corporate information do not. Revenue, operating profit and declared dividends are examples: many such items can be checked objectively against published disclosure once the period, scope and accounting standard are specified. Annual securities reports, financial results summaries, timely disclosures and Form 10-K filings can all be retrieved by machine through EDGAR, EDINET, TDnet and XBRL. Even so, there is no guarantee that a generative AI system consults them every time and answers with the right fiscal period, currency, unit and accounting standard. An evaluation covering 17,621 US-listed companies and 197,011 questions observed numeric hallucinations even when the company name and the year were stated explicitly, and a financial QA benchmark found that 15-23% of answers were wrong even in the "confident" state where all eight re-answers agreed.

When AI describes your financial information incorrectly, that is not merely a reputation problem. It is a conflict with published disclosure. The same error carries different weight here than it does in recruiting or corporate communications.

Within the scope of this review, however, we could not confirm any primary source from a regulator or professional body that recommends AI visibility as a formal IR KPI. That is precisely why this article argues that what should be measured is agreement with disclosure, the source of citations, recency and the accuracy of units and currency, rather than volume of exposure. This is the article's own reasoning, not a public standard.

What you will learn

  • Why an AI error in IR means something different from an error in other functions (the existence of disclosure to check against)
  • What has been demonstrated, and how far, about AI misstating corporate financials
  • The error types specific to IR, the influence of an IR site's HTML and PDF structure, and the limits of that evidence
  • How institutional investors, individual investors and IR functions actually use AI, including primary data from Japan
  • The precise relationship with fair disclosure rules (an area where errors are serious)
  • What IR should measure about perception in AI answers, and what it must not measure

Who this is for

IR and disclosure staff at listed companies, CFOs and corporate planning functions. Secondarily, institutional investors and analysts, and individual investors.

0. The evidence base of this article - what is established and what is not

Before the main discussion, we state what this article rests on. The gap between expectation and reality is wide in AI perception for IR, and confident assertions circulate easily.

What is established comes from controlled experiments and benchmarks: the accuracy rate when a model is asked for a company's revenue; the difference in evaluation between the same disclosure text with and without the company name attached; whether HTML and PDF are cited differently. Each was designed as an experiment and the figures have been published.

What is not established is an accumulation of real-world incidents. A named-company incident register of the form "a given AI answered company A's dividend incorrectly and the company issued a formal correction" barely exists as of August 2026. This article names no incidents and gives no incident counts. Nor is it public how each AI service retrieves and ranks company-level answers. This article keeps experimental results and live operation separate.

The territory covered here is domain-specific and narrower than AI misinformation in general: errors that conflict with statutory and timely disclosure. For the wider picture of misinformation see AI misinformation and misattribution, for legal liability see the legal liability of AI misinformation, and for correction and removal see correcting and removing AI misinformation.

1. Why an AI error in IR means something different from other functions

If AI misstates your average salary in a recruiting context, candidate decisions are distorted. If your business is summarised incorrectly in a communications context, brand perception is distorted. Both matter, but there is not necessarily an external standard that settles what is correct.

IR is different.

Revenue, operating profit, declared dividends and segment information at a listed company: many of these can be checked objectively against published disclosure once the period, scope and accounting standard are specified. In Japan, the annual securities report is statutory disclosure under the Financial Instruments and Exchange Act, and the financial results summary rests on the exchange's timely disclosure framework. Dividends of course include forecasts and policies, shareholder composition varies with the record date and the method of identification, and adjusted (non-GAAP) measures are managed separately. Even so, IR differs from other functions in that checkable items exist across a wide range.

When AI describes your prior-year revenue with a 10% error, what the company faces is not "this looks bad" but a divergence between the content of published disclosure and the information circulating in the market. This article treats that as a monitoring issue for the external information environment, adjacent to disclosure quality control rather than a statutory disclosure-control obligation in itself.

On that basis, this article takes the following position. Corporate perception in AI answers should be handled as an extension of disclosure quality control, not as an extension of communications-style reputation management (a practical proposal specific to this article, not a legal requirement). The question is not who is saying what, but whether what is being said agrees with published disclosure. That distinction is the foundation of everything that follows.

The supply structure behind citations belongs to the communications lane and is not re-explained here; see what AI cites about your company. Differences by language and country are covered in cross-border AI perception, and the recruiting context in candidates research you with AI.

2. The correct answer is machine-retrievable. That is not a guarantee that it is consulted.

We start by confirming how open to machines the disclosure that serves as the reference point actually is.

Data routes for statutory and timely disclosure

RouteWhat it providesAccess conditions
EDGAR API (US SEC)Filing history and XBRL company facts through a JSON API. Forms 10-K, 10-Q, 8-K, 20-F, 40-F, 6-K and othersPublic and free; no API usage fee
EDINET (Financial Services Agency, Japan)Search of annual securities reports and related filings. API v2 availableAPI key required
TDnet (Tokyo Stock Exchange)Timely disclosure published in real time (31 days of open viewing, 10 years of company-level search). Financial results summaries and revisions to earnings and dividend forecasts are provided in XBRLViewing is free. Historical data and API access are paid

What this establishes is only that the statutory and timely disclosure serving as the reference point is machine-retrievable. Note also that TDnet is an exchange framework operated by the Tokyo Stock Exchange, and that the financial results summary is timely disclosure under exchange rules. Its basis differs from statutory disclosure such as the annual securities report. To judge whether an AI answer is right or wrong, you first have to decide which document, under which framework, is the reference point.

The distance between machine-retrievability and actually being consulted

That an API is public does not mean AI is calling it. Claims such as "ChatGPT searches EDINET every time" or "Gemini gives 10-K filings top priority" cannot be concluded from public information.

Google has published a detailed explanation of its own generative AI features. The condition for appearing as a linked source in AI Overviews or AI Mode is that the page is indexed and eligible to be shown with a snippet, with no additional technical requirement. Machine-readable files such as llms.txt and special markup are unnecessary, and Google states explicitly that Search does not use them, and that no special schema.org markup exists for generative AI features.

Two points matter here. Google's AI features are built on the same index and technical requirements as ordinary Search, and no AI-specific schema.org markup is required. That said, whether page selection and generation inside AI answers is identical to ordinary ranking cannot be said from official materials. And this is Google's account of its own products, not a specification for other AI services. How ChatGPT, Perplexity, Claude and others retrieve company information is not published.

Google also states that text PDFs can be indexed in Search and that OCR may be applied to image-based PDFs. The widely circulated claim that "PDFs are not read by generative AI" is an overgeneralisation.

3. What is actually happening - the established range

This section sets out how far published empirical research goes. All of it consists of controlled experiments, and none of it directly shows what is happening to your company inside a live AI service.

3-1. An evaluation of 17,621 listed companies and 197,011 questions - knowledge and overconfidence together

A research team at the Georgia Institute of Technology built 197,011 questions about the revenue of US-listed companies and had multiple large language models answer them. The scope covered 43 years from 1980 to 2022 and 17,621 companies. The data came from Compustat Capital-IQ, with revenue standardised in millions of US dollars. Prompts were a fixed form stating the company name and the year, modelled on how a general user would ask.

Defining a numeric hallucination as an absolute error above 10% against the correct value, the following results were obtained.

  • Accuracy is higher for more recent years. Llama-3-70B-Chat answered correctly for 54.17% of companies for 2017, against 6.32% for 1995. The paper analyses 1995 as a turning point in the availability of disclosure information through EDGAR (the SEC's electronic filing mandate was phased in from 1993 and completed in May 1996).
  • Accuracy is higher for companies with larger market capitalisation. For the same model, the log odds of a correct answer rose by 1.0091 per tenfold increase in market capitalisation. The same tendency appeared for institutional investor attention.
  • Among the cases they did not answer correctly, however, larger companies were also more likely to return a numeric hallucination. For the same model the log odds rose by 0.1914 per tenfold increase in market capitalisation. The denominator of this ratio is the set of cases the model did not answer correctly, not the total number of questions.

The research team frames this as knowledge and overconfidence existing side by side. The implication for IR is that the better known a company is, the more confidently AI answers, and the larger the miss when it misses. This is a correlation with firm size and related factors, not evidence that increasing disclosure causes more errors. The converse intuition, that increasing disclosure automatically makes you safer, is equally unsupported.

Note that because the data comes from Compustat, differences in accounting standards are not harmonised. The study was accepted at the international conference COLM 2025.

3-2. Wrong even when "confident" - consistency does not prove correctness

Research released in July 2026 reports a finding that bears directly on measurement design for IR. Using financial QA benchmarks built from actual filings (FinQA and TAT-QA), the same question was answered eight times, and the case where all eight converged on the same answer was defined as the "confident" state. On FinQA, 15-23% of those confident answers were wrong.

Important limits apply. The models evaluated were Qwen3-8B, Llama-3.1-8B and Gemma-2-9B, three open models in the 8-9B class, and not the publicly available AI search products themselves. It is also a preprint. The 15-23% figure cannot be read across as the probability that a particular AI service is wrong about your company.

The practical implication is nonetheless clear. At least in this experiment, agreement across re-answers did not guarantee correctness. In monitoring commercial AI as well, the sound approach is to treat consistency as no substitute for a truth signal and to prioritise checking against published figures.

3-3. In Japanese financial results summaries, the presence of the company name changes the evaluation

The study that bears most directly on IR in Japan is peer-reviewed work by researchers at Nomura Asset Management and Preferred Networks (published in IEEE BigData 2024).

The performance narrative sections of financial results summaries from January 2019 to December 2023, retrieved from the Tokyo Stock Exchange's TDnet, were given to models under two prompts. One stated the company name; the other withheld it. Everything else was identical. Models scored sentiment from 1 (bad) to 5 (good), and the difference between the two scores for the same text was defined as company-specific bias.

Seven models were evaluated: GPT-4o, GPT-3.5-turbo, Gemini 1.5 Pro and Flash, Claude 3.5 Sonnet, Claude 3 Haiku and Qwen2-7B. Scores changed in roughly 10-20% of cases for each model; for GPT-4o the figure was about 17.8% of 10,249 valid observations.

More important still, the direction of the bias did not agree across models. Analysed against the 20 factors of the MSCI Barra Japan Equity Model, GPT-4o showed a negative spread on the value factor (-0.328) while GPT-3.5 showed a positive spread (+0.230), and the signs were also reversed on momentum.

The meaning from an IR standpoint is plain. Even when the text of the financial results summary you published does not change, attaching the context of a company name can change how AI evaluates it. The way it changes differs by model, and there is a layer that cannot be controlled by redrafting the disclosure. The same study also analyses the relationship with share prices, but it does not establish causation.

3-4. Grounding answers in statutory disclosure is itself an unsolved technical problem

A benchmark study released in March 2026 built 755 annotated examples from 300 pages of SEC Form 10-K filings and evaluated methods for judging whether an AI answer is grounded in the source document. The 755 figure is the size of the evaluation dataset, not an error rate. What it shows is that mechanically verifying which statement in statutory disclosure supports an AI answer is an independent technical problem in its own right (preprint). In other words, the fact that AI has shown a source does not mean that source supports the whole answer. This has practical implications for IR staff and institutional investors alike.

4. Error types specific to IR

Building on the previous section, we classify the errors that can arise with IR information. No reliable public statistics on frequency exist, so what follows is a classification of types, not of frequency.

TypeWhat happensWhat IR should settle in advance
Wrong fiscal period"Revenue for 2024" can be read as the calendar year, the fiscal year, the trailing twelve months or an annualised quarter. A fiscal year ending March 2024 and "fiscal 2024" diverging is not unusualState the definition of the fiscal period. Note that the 17,621-company evaluation above observed errors even with the year stated, so specifying the period is not sufficient on its own
Wrong currency or unitMillions of yen, thousands of yen, millions of US dollars, billions of US dollars, and local-currency versus converted figures become mixedPlace the unit and currency next to the figure. Note that the study above observed errors above 10% even with units standardised
Mixed accounting standards"Operating profit" is defined differently under Japanese GAAP, IFRS and US GAAP, and companies that adopted IFRS voluntarily cannot be compared across years without adjustmentMake it possible to tell from the answer key whether a comparison crosses a change of standard
Stale information persistingDeparted officers, divested businesses and superseded dividend policies are described in the present tenseState the fiscal year as text on prior-year pages. This kind of mixing can originate with your own site
Confusion with same-name or similar-name companiesA listed company and an unlisted company of the same name, or a domestic entity and a similarly named overseas entity, are conflatedPlace entity-identification cues in disclosure and on the site. No IR-specific public primary example has been confirmed, so no frequency figure is given here

5. How is your IR site actually being read?

The most concrete question for an IR function is probably whether the way you build your site changes how AI reads it. One experiment sets out to answer it head-on. It needs to be handled with considerable care, so we set out the conditions.

5-1. A comparison of HTML and PDF annual reports

This is 2026 joint research by St. Pölten University of Applied Sciences (USTP), HHL Leipzig Graduate School of Management and nexxar, a digital IR reporting vendor. It consists of two studies.

Study 1 (citation patterns): 20 listed companies across nine European countries and eight sectors were split into peer groups of 10 publishing structured HTML annual reports and 10 publishing mainly in PDF. More than 2,500 fixed prompts were run through ChatGPT (GPT-4o and GPT-5). Twenty-three testers coded the sources cited in the answers, extracting 24,662 citations.

FindingFigureConditions
Direct links to each company's annual report, out of all citations58.5%Based on 24,662 citations
Citation frequency of annual report content in the HTML group3.05x the PDF group20 companies (10 against 10)
Reliance of the PDF group on external information sourcesAbout 2.7xBloomberg, Reuters, MarketScreener and others
Accuracy of answersHTML group 71% / PDF group 54%Accuracy is from a subsample of n=200

Study 2 (server logs): eight weeks of access logs, from 29 August to 25 October 2025, were analysed for the digital annual reports of five DAX-listed companies. Automated accesses totalled 4,838,833, with more than 175 identifiable bot types. The top five bots accounted for 48.8%, and accesses identifying themselves as ChatGPT made up 30.3% of bot traffic (followed by Bingbot at 6.6%, Amazonbot at 5.8% and Googlebot at 4.6%). Among the 759,226 requests remaining after cleaning, business review content accounted for 39.3% and financial reporting and financial statements for 20.6%.

5-2. What this experiment can and cannot support

Start with what it cannot support.

  • "Move to HTML and AI answers become correct" is unproven. Even in the HTML group, 29% of answers were judged inaccurate. Companies that build HTML annual reports may also have stronger IR resourcing and update more frequently, and this was not an experiment that separated causation.
  • nexxar is a digital IR reporting vendor and therefore an interested party, and the peer-review status is not stated. Independent replication is needed.
  • The samples are 20 companies and five companies, and they are European rather than Japanese.
  • Study 1 counted only the sources explicitly displayed inside answers. "Not cited" cannot be read as "not read".
  • The user agent is self-declared, and an access identifying itself as ChatGPT does not prove that the content was adopted into an answer.

There are also things it can support. Even in the PDF-centred group, 2,385 citations of annual report PDFs were recorded, so a flat assertion that PDFs are not read is wrong. The more measured reading is this. PDFs are not a binary of "unreadable" against "read accurately". Tables that span pages, footnotes, financial statements rendered as images, unit labels placed away from the main table: structures like these create room for extraction errors. And in PDF-centred configurations, AI answers were observed to lean more on external financial information vendors than on the company's own primary materials. That is a loss of control, not a matter of correctness in itself.

For iframes, dedicated IR subdomains, cookie consent screens and content rendered only through JavaScript, no primary empirical study isolating the effect on accuracy could be confirmed. This article does not assert that content in an iframe is not read by AI.

6. Who is using AI, and how?

Next, who is actually using AI. Populations and question wording differ across surveys, so the figures cannot be compared directly. Note also how many are surveys by interested parties.

6-1. Institutional investors - AI has not replaced primary sources and human dialogue

In a 2026 survey of 100 US institutional investors conducted by an IR advisory firm, 54% rated generative AI at least moderately important for investment research, and 42% said they use it as a main tool for detailed work on new investment candidates. The most telling result concerns the importance of information sources: contact with management 77%, company disclosure 66%, dialogue with the IR function 63% and the IR site 47%, against generative AI answers at 24% (an interested-party survey, n=100; question wording and scales differ by item, so these cannot be lined up for direct comparison).

At least in this survey, generative AI ranks below executives, company disclosure and IR dialogue in importance, so it cannot be said to have replaced primary sources and human dialogue (note also that this survey did not measure whether respondents always return to primary sources for final checks). The risk for IR is not that AI makes the final decision, but that questions get built on a mistaken premise before the divergence is detected.

6-2. Implementation at asset managers - a cross-sectional review by Japan's supervisor

A more reliable domestic data point is the Financial Services Agency's Progress Report 2026 on the sophistication of asset management services, published in July 2026. This is a supervisory progress report based on monitoring, not law or rules.

Surveying 13 major asset managers (10 Japanese and three foreign-affiliated) as of the end of November 2025, generative AI had been adopted or was being prepared by 11 companies in research work, 11 in reporting work and 10 in stewardship activity. One case cites a reduction of 40 hours per analyst per month in earnings analysis (self-reported by a single company; a measure of working hours, not of investment outcomes).

What matters for IR is where AI is being used. The report states that in proxy voting and engagement, practices are spreading in which AI supports the collection and analysis of financial and non-financial information from investee companies' published materials and the initial judgment on proxy voting. In other words, the disclosure materials you publish are being read and summarised by investor-side AI, and are starting to feed into the initial judgment on proxy voting. Note that 10 of the 13 companies cite database design and construction, and nine cite securing explainability, as issues, and that the practice is one in which humans verify and decide.

6-3. Individual investors - "having concerns" and "always verifying" are different things

For individual investors, peer-reviewed research is the most reliable source. A study by a research team at the University of Washington (forthcoming in the November 2026 issue of the Journal of Accounting and Economics) combines more than 400,000 actual queries to a major brokerage's generative AI chatbot with a survey of more than 2,000 people, and reports that roughly half of individual investors use generative AI. The main uses are interpreting and contextualising financial information and market trends, and stock screening is among them. The most common concern is reliability and accuracy (54%). But having concerns is not the same as always verifying against primary sources. The study also reports that users with greater financial knowledge and investment experience tend to use AI for more complex purposes, so the picture of "young beginners who believe AI uncritically" is not supported.

6-4. IR functions themselves - what Japan's primary data shows

The key data on IR functions in Japan was published in May 2026: the 33rd Survey on IR Activities by the Japan Investor Relations Association, a non-profit body that promotes IR. It covered all 4,088 listed companies as of January 2026, ran from 9 February to 25 March 2026, and drew 948 responses (a 23.2% response rate).

The results are unambiguous. 80.3% said their use of generative AI in IR activity had increased in frequency over the past year. For usage in IR-related work (the sum of "in use" and "trialling", n=780), summarising and organising materials stood at 80.5% (13.8% in the 2024 survey), minute-taking at 65.8%, preparing English-language disclosure materials at 63.7% (16.0% previously) and preparing briefing and press documents at 51.3%. Establishing guidelines has also advanced, to 55.2% (from 32.1%).

And the top issue in adoption (n=917) is decisive. Lack of accuracy in information, that is, hallucination, at 77.2%. It is followed by security concerns at 60.1%, copyright and ethical risk at 45.0%, insufficient employee skills at 34.8% and no support for internal terminology at 34.1%.

Close to 80% of responding companies cite generative AI hallucination as their largest issue. The population needs to be discounted for: 948 responses from 4,088 companies contacted, a 23.2% response rate. Even so, it is readable that AI inaccuracy is becoming a shared practical concern in IR.

Yet the indicators used to measure effectiveness in the same survey do not line up with that.

Effectiveness indicator2026Previous
Shareholder composition86.8%84.8%
Change in the number of meetings with analysts and investors69.0%65.1%
Attendance of analysts and investors at briefings51.1%47.3%
Market capitalisation47.0%39.0%
Trading volume / PBR and similar37.1% / 36.9%31.6% / 28.1%

At least in the effectiveness-measurement question of that survey, AI perception is not identifiable as a measurement item. Whether each company measures it separately cannot be determined from this survey.

Here is the asymmetry this article wants to point out. IR functions cite hallucination as the largest issue with the AI they use themselves. But how the AI that investors use describes their company is not, at least, inside the framework of IR effectiveness measurement. This is so even though the same models can generate the same inaccuracy about their own earnings.

The picture abroad is similar. In a survey of roughly 700 IR practitioners conducted in the fourth quarter of 2025 by an exchange-affiliated IR solutions vendor, 51% said they had already built AI into their work, against 30% in 2024 (an interested-party survey). Both are data on IR functions using AI, not data on measuring corporate perception in AI answers.

6-5. The measurement gap in Japan

The Japan Securities Dealers Association's Survey on Individual Investors' Attitudes Toward Securities Investment 2026 is a large study with 5,000 valid responses, showing websites at 52.9% as an information source among securities holders, and Instagram, YouTube, TikTok and similar at 43.4% among those in their thirties and below.

That survey, however, contains no question asking directly about AI or ChatGPT. This does not mean that Japanese individual investors are not using AI; it means it has not been measured. No public statistic showing how far Japanese individual investors use generative AI for stock research could be confirmed as of August 2026.

7. The relationship with fair disclosure rules - getting this wrong is serious

This section answers the question IR functions are most likely to raise. "AI is describing our unpublished results by inference. Is that not a fair disclosure violation?" The answer, as a general matter, is that it is not a violation.

7-1. What both frameworks regulate is the conduct of the issuer

Regulation FD in the United States is a final rule, in effect since October 2000. Japan's fair disclosure rule is a statute, set out in Article 27-36 and following of the Financial Instruments and Exchange Act and in force from 1 April 2018. Regulation FD does not apply to foreign private issuers. Many Japanese companies listed in the US fall into this category and are not direct addressees of Reg FD. What applies directly to Japanese companies is the FD rule under the FIEA.

What both frameworks address is, fundamentally, conduct by the issuer in transmitting material non-public information to specified market participants and others. Under Japan's rule, when a listed company transmits material non-public information to securities firms, investors and others in the course of its business, publication on a website or by similar means is required simultaneously if the transmission was intentional, and promptly if it was not. Accordingly, a third-party AI inferring unpublished results from public information is not selective disclosure by the issuer. The understanding that an AI inference automatically constitutes an FD violation is legally inaccurate.

Note that the Financial Services Agency states at the head of its Fair Disclosure Rule Guidelines that they set out the agency's general interpretation of the law as of that time and do not answer whether the law applies to any individual case. Individual cases are judged on their facts. Always confirm specific judgments with your own legal department and outside professionals.

7-2. Other issues do arise, however

That something is not an FD violation does not mean there is no issue. Where a company inputs material non-public information into an external generative AI service, a separate set of problems arises: information leakage, confidentiality obligations, internal control and, depending on the circumstances, selective transmission.

Even so, input to an external AI does not automatically constitute a Reg FD violation. Reg FD defines the categories of recipients it covers, and exceptions exist for persons who owe a duty of confidentiality or who have agreed to keep the information confidential. The legal assessment differs according to whether the AI provider falls within the covered recipients under FD and whether a confidentiality obligation or non-disclosure agreement is in place. What is serious here is the information management risk; it should not be converted into a flat assertion of legal violation.

The 60.1% of respondents citing security concerns such as information leakage in the Japan Investor Relations Association survey above are pointing at this issue. Having AI draft anticipated Q&A before an earnings announcement, or summarise pre-disclosure materials, is natural enough as an efficiency measure, but it requires checking where that AI sends the data and how it retains it. The Financial Services Agency's AI discussion paper (version 1.1, which is neither law nor supervisory guidelines) likewise treats input of confidential and material information into external AI, and verification of outputs, as issues.

There is professional-body support for the point as well. NIRI, the US IR professional body, states in the policy statement on AI approved by its board in September 2024 that evaluating pre-release earnings materials with public AI tools could unintentionally trigger an SEC Form 8-K disclosure obligation. It is an example of an IR professional body itself identifying input of material non-public information into external AI as a distinct practical issue.

7-3. The most dangerous error - trying to correct AI with unpublished information

Using material non-public information to correct an AI answer because it is wrong is extremely dangerous. If AI is describing your current-year outlook as weaker than reality, and you produce figures you have not published, that can amount to transmission of material non-public information by the issuer. The motive of correcting an AI error does not justify selective disclosure.

The conservative practice this article recommends is therefore to make corrections to an individual AI using published information only. The law does not provide that only published information may be used when correcting AI. The substance is the sequence: if new information has to be conveyed, carry out appropriate broad public disclosure first, and explain afterwards. In practice, point at published facts, in the form "in the financial results summary we disclosed on a given date, the figure is X", and add nothing beyond that. Using exposure in AI answers as a substitute channel for formal disclosure should also be avoided.

Note that correction and update duties where an issuer's prior disclosure has become misleading, anti-fraud liability, insider trading, and confidentiality obligations are separate issues from FD. The Regulation FD final rule itself states that anti-fraud liability under the federal securities laws survives independently of FD.

How to pursue a legal claim against an AI provider, and how correction and removal procedures are designed, fall outside the scope of this article (see legal liability and correction and removal).

8. What IR should measure

On the basis of everything above, we come to this article's conclusion.

8-1. What to establish first - no official KPI is confirmed

It is not that professional bodies have no formal documents on AI. NIRI in the United States has a formal policy statement on the use of AI, approved by its board on 18 September 2024. Its content, however, concerns governance and ethics for the IR function's own use of AI, and it does not set out KPIs for measuring AI visibility.

CFA Institute also published "Artificial Intelligence and the Future of Finance: A Framework for Structural Change" on 20 July 2026. What that report addresses is governance of AI use on the investor and asset management side; it does not ask issuers to measure their own visibility in AI answers. In the same body's 2024 survey of 200 investment professionals, 85% said the investment industry needs common standards and 82% said their absence is an obstacle to adoption, confirming that standards are not yet settled. ESMA's supervisory statement, IOSCO's report and non-binding supervisory toolkit, and the Financial Services Agency's discussion paper likewise make no reference to AI visibility KPIs for issuers.

Within the scope of this review, we could not confirm any primary source from a regulator or professional body that recommends AI visibility as a formal IR KPI, as of August 2026.

What follows is therefore this article's reasoning. It is not a public standard.

8-2. Measure agreement with disclosure, not volume of exposure

In communications, indicators such as mention volume, citation rate and share of voice inside answers are meaningful, because brand perception has a volume dimension. In IR, increasing volume does not resolve a conflict with disclosure. Chasing exposure alone can amplify errors. We propose the following as the operating indicators IR should hold.

IndicatorDefinitionWhy it is specific to IR
Agreement with statutory and timely disclosureThe share of AI answers to a core question set that agree with the annual securities report, the financial results summary and timely disclosureDefinable only in a domain where the reference point exists externally
Primary-source citation rateThe share of an answer's sources accounted for by the company's own disclosure materials and IR siteMeasures the degree of dependence on external vendors
Latest-year adoption rateThe share of answers using figures from the most recently disclosed fiscal yearDetects the persistence of stale information
Error rate on units, currency and accounting standardThe share of answers where the figure is right but the unit, currency or standard is wrongAn error mode specific to financial figures
Persistence rate of former officers and divested businessesThe share of answers describing departed officers or divested businesses in the present tensePrior-year pages can be the cause
Time from detection of a serious error to observed improvementThe period from detecting an error to observing improvement through re-measurementThis includes update lag on the third-party AI side, so it does not by itself indicate the quality of the company's own controls

All of them share the property that correctness can be judged objectively. That is what makes them usable in IR, and what makes them non-transferable to other functions.

8-3. What not to measure

Equally important is what not to set as a target.

  • Do not make share of answers or positive sentiment alone a management target. AI describing you favourably does not mean enterprise value has improved, and there is no evidence of correlation either.
  • Do not report exposure in AI answers to management as an outcome measure of IR activity as it stands. There is no public standard and its validity is not established.
  • Do not claim causation from AI errors to share price or cost of capital. No causal estimate has been confirmed.

9. Operating design for the IR function

Next, operations.

9-1. Build the answer key

Judging whether an AI answer is right requires a standard. The figures themselves exist in many places, but they are usually not organised into a form usable for judgment. Give the answer key the following attributes.

  1. Fiscal period (the fiscal year ending March 2026, fiscal 2026 and so on, stated with its definition)
  2. Consolidated or parent-only
  3. Accounting standard (Japanese GAAP, IFRS or US GAAP)
  4. Currency and unit (millions of yen, thousands of yen, millions of US dollars and so on)
  5. Scope: continuing operations or the whole company
  6. Announcement date and name of the source document
  7. Whether the figure is an actual result or a forecast or guidance
  8. Whether it is the originally reported value or the value after a correction or restatement
  9. Whether it is a statutory measure or an adjusted (non-GAAP) measure

Only when these are in place can you say objectively that an AI answer is wrong. Until they are, even internal debate about correctness fails to converge.

9-2. Design a regular test of core questions

Next, what to ask. Revenue, operating profit and net profit for the most recent period; dividends (dividend per share and dividend policy); the names of the representative and principal officers; segment composition and core businesses; disclosed targets such as cost of capital and ROE; and the main KPIs of the medium-term management plan. Starting from this question set is realistic.

Ask them regularly, with the question wording and conditions held fixed, and check against the answer key. Around earnings announcements, after the annual securities report is filed, after officer changes are announced: designing the process to re-measure immediately after disclosure events makes the persistence of stale information easier to detect.

9-3. What the IR site can do with confidence

As set out in section 5, the evidence that would let you assert how implementation affects AI answers is limited. Even so, measures with a clear basis and no side effects do exist.

  • Provide the key facts in HTML text as well. You do not need to abolish PDFs. Publish the main figures in a form readable as text alongside them.
  • State the fiscal year on prior-year pages. Write it as text within the page, in the form "fiscal 2023 financial results summary". Avoid a state where the year appears only in the file name.
  • Place units and currency near the figures. A layout where the unit appears only in a table margin or a footnote creates room for extraction errors.
  • Do not obstruct crawling. Check for blocking in robots.txt, at the CDN and in the hosting layer, and make sure IR information is reachable through internal links. These are the basic requirements Google sets out (and as noted in section 2, no special markup is required for generative AI features).

On the other hand, we do not recommend budgeting for llms.txt or AI-specific files as measures that work in Google. Google states that Search does not use them and that they neither help nor harm visibility or ranking. Organising content in FAQ form is not objectionable in itself, but it cannot be recommended as a way of producing display treatment in Google (FAQ rich results were retired in May 2026).

9-4. What not to do

  1. Do not change the wording of statutory disclosure to something AI might prefer. Disclosure documents are prepared under law and exchange rules, and neither their content nor their wording should be distorted for AI readability.
  2. Do not use unpublished information to correct AI (section 7).
  3. Do not treat a single model's answer to a single prompt as a stable state of perception.
  4. Do not transcribe AI answers into investor materials as they stand. Even when source links are attached, there is no guarantee that the whole answer is supported (section 3-4).

And "correcting an AI answer" does not mean only a removal request. Save the reproducible prompt and timestamp; identify where the answer conflicts with published primary material; fix gaps, stale pages, unit labelling and entity identification on your own IR site; and re-verify across multiple models after an interval. The substance is designing this flow as a disclosure quality control process. Note, however, that no corporate correction procedure common to the major AI services, and no process guaranteeing the outcome of a correction, could be confirmed. What the company can control extends to the quality and structure of the information it publishes, and to the speed of its measurement and detection.

10. What is still unestablished

We set out the items that could not be confirmed. These are things this article cannot assert, and if another source states them definitively, check the basis.

Unconfirmed itemHow this article handles it
A comprehensive list of named cases where AI answered incorrectly and a company issued a formal correctionNo counts and no company names are given
The share of Japanese individual investors using generative AI for stock researchNo public statistic could be confirmed
A public benchmark directly measuring the error rate on Japanese companies' financial figuresCould not be confirmed
The effect of iframes, separate domains and JavaScript-only rendering on accuracyNo quantitative value (no study isolating the factor could be confirmed)
Guidance from a regulator or body recommending AI perception as a formal IR KPINot an official KPI
A correction and appeal procedure for companies common across AI providers, with processing deadlinesNo such procedure could be confirmed
A causal estimate that AI errors alone moved share price or cost of capitalCausation unconfirmed
A general law that moving to HTML necessarily raises AI accuracy on IR informationLimited evidence only (not independently replicated, not randomised)
The route by which financial vendors' closed data enters public AI answersUnconfirmed, as it is not disclosed

11. So how do you measure it?

The constraints above determine the requirements for measurement in IR.

First, repeat. Given that 15-23% can be wrong even when eight runs agree, a single run is one observation, but not enough to estimate a stable state. Second, hold conditions fixed. A question asked with prior conversation history in place produces an artefact of that conversation rather than a stable state of perception. Going stateless removes conversation history as a confounder and improves comparability across conditions (it does not remove the influence of model updates or the time of retrieval). Third, span multiple AI engines. As section 3-3 shows, the direction of bias does not agree across models. Fourth, measure content. Not whether a citation appeared, but whether the figures described agree with published disclosure. That is the heart of it.

Vaipm measures AI-space perception through a total of 25 stateless queries across multiple AI engines. Note that the body of research cited in this article establishes only that repetition is required and that conditions must be held fixed; this article does not demonstrate the statistical validity of any particular number of runs. For the wider picture of AI Perception Management see what AIPM is, and for the relationship with AI search optimisation see what AIO is and what LLMO is.

For IR, none of this is new work. The existing job of assuring the accuracy of disclosure has simply extended to a new kind of reader.

Frequently asked questions (FAQ)

Q1. AI got our financial figures wrong. Is that a fair disclosure violation?

No. Both Regulation FD and Japan's FD rule govern conduct by the issuer in transmitting material non-public information to specified market participants and others. A third-party AI inferring from public information is not selective disclosure by the issuer. (Note that Regulation FD does not apply to foreign private issuers.) Individual cases are judged on their facts, so confirm the position with your own legal department.

Q2. May we give AI accurate figures we have not published yet, in order to correct it?

Avoid it. Transmitting material non-public information can itself raise an FD issue. The conservative practice this article recommends is to limit corrections to information already published and, if necessary, to carry out lawful public disclosure procedures first.

Q3. If we rebuild our IR site around HTML, will AI answers become correct? Are PDFs not read?

Neither can be asserted. In the comparison across 20 European companies, accuracy was 71% for the HTML group and 54% for the PDF group, but 29% of the HTML group's answers were still inaccurate. The sample is small and differences in company size and IR resourcing are mixed in. Google also states that text PDFs can be indexed, and PDFs were cited in that same experiment. Providing HTML is reasonable, but it does not guarantee accuracy.

Q4. Will installing llms.txt give us an advantage in AI search?

For Google Search, the company denies it explicitly. It states that llms.txt and AI-specific markup are unnecessary and that installing them neither helps nor harms visibility or ranking. You are free to prepare files for other services, but there is no basis for promising an effect.

Q5. We asked several times and got the same answer. Does that mean the AI understands us correctly?

No. In a financial QA benchmark, 15-23% of answers were wrong even in the "confident" state where all eight re-answers agreed (three open models in the 8-9B class, preprint). At least in that experiment, agreement did not guarantee correctness, and checking against an answer key is required.

Q6. Can we use volume of exposure in AI answers as an IR KPI?

We do not recommend it. Within the scope of this review, no primary source from a regulator or professional body recommending AI visibility as a formal IR KPI could be confirmed. Chasing exposure can amplify errors. Put agreement with disclosure at the centre.

Q7. Is there evidence of company-specific bias when AI reads disclosure from Japanese companies?

Yes. In peer-reviewed research using Japanese financial results summaries from 2019 to 2023, sentiment scores for the same performance text changed depending on whether the company name was shown, in roughly 10-20% of cases for each model, and the direction of the bias did not agree across models. This is bias in evaluation, however, not a rate of numerical error in revenue or dividends. No public benchmark directly measuring how often AI misstates Japanese companies' financial figures was found within the scope of this review.

Q8. Are institutional investors really researching companies with AI?

They use it, but at least in this survey no substitution is visible. In a survey of 100 US institutional investors, importance as an information source was 77% for contact with management and 66% for company disclosure, against 24% for generative AI answers. Meanwhile, among 13 major Japanese asset managers, adoption is advancing across research, reporting and stewardship work.

Q9. What should an IR function watch most carefully when using generative AI?

Input of unpublished information into external AI. In the Japan Investor Relations Association's 2026 survey, 60.1% of responding companies cited security concerns and 77.2% cited hallucination as issues. NIRI also notes that evaluating pre-release earnings materials with public AI tools could unintentionally trigger an SEC Form 8-K disclosure obligation.

Q10. Can we ask the AI provider to correct an answer?

No corporate correction procedure common to the major AI services, and no process guaranteeing the outcome of a correction, could be confirmed. In practice, record the reproducible prompt and timestamp, identify where the answer conflicts with published material, fix gaps, stale pages and unit labelling on your own site, and confirm through re-measurement.

Q11. Where should we start?

Building the answer key: fiscal period, consolidated or parent-only, accounting standard, currency and unit, the scope of continuing operations, the announcement date, the name of the source document, whether the figure is an actual result or a forecast, whether it is the value after a correction, and whether it is an adjusted measure. Only when these are in place can you say objectively that an AI answer is wrong. Measurement can come after that.

Sources

All sources in this article were re-retrieved and re-checked at the listed URLs on 8 August 2026. Because they differ in the character of their reliability, they are presented in three tiers.

A. Primary and official sources (law, rules, regulator documents, operators' official explanations)

#SourceCharacterURL
A1US SEC, "EDGAR Application Programming Interfaces"Official technical documentation from the regulatorhttps://www.sec.gov/search-filings/edgar-application-programming-interfaces
A2US SEC, "Developer Resources"As above (public and free; no API usage fee)https://www.sec.gov/about/developer-resources
A3Financial Services Agency (Japan), "EDINET"Disclosure viewing system operated by the regulatorhttps://disclosure2.edinet-fsa.go.jp/week0020.aspx
A4Japan Exchange Group, "TDnet"An exchange framework (not national law in itself)https://www.jpx.co.jp/english/equities/listing/disclosure/tdnet/
A5Japan Exchange Group, "XBRL (TDnet)"Official explanation from the exchangehttps://www.jpx.co.jp/english/equities/listing/disclosure/xbrl/03.html
A6Google Search Central, "AI features and your website"Official explanation from the provider itselfhttps://developers.google.com/search/docs/appearance/ai-features
A7Google Search Central, "PDFs in Google search results"As abovehttps://developers.google.com/search/blog/2011/09/pdfs-in-google-search-results
A8US SEC, "Regulation FD" (Release 33-7881)Final rule (in effect since October 2000)https://www.sec.gov/rule-release/33-7881
A9Financial Services Agency (Japan), "Fair Disclosure Rule Guidelines"The regulator's general interpretation of a statute (FIEA Article 27-36 and following, in force 1 April 2018)https://www.fsa.go.jp/news/29/syouken/20180206.html
A10ESMA, guidance on the use of AIA supervisory statement and initial guidance (not new EU legislation)https://www.esma.europa.eu/press-news/esma-news/esma-provides-guidance-firms-using-artificial-intelligence-investment-services
A11IOSCO (2025), report on AIA report and consultation paper (no direct legal force)https://www.iosco.org/library/pubdocs/pdf/IOSCOPD788.pdf
A12IOSCO (2026), supervisory toolkitA non-binding supervisory toolkithttps://www.iosco.org/library/pubdocs/pdf/IOSCOPD823.pdf
A13Financial Services Agency (Japan), "AI Discussion Paper" (version 1.1)A discussion paper (not law or supervisory guidelines)https://www.fsa.go.jp/news/r7/sonota/20260303/aidp_version1.1.pdf
A14Financial Services Agency (Japan), Progress Report 2026 on the sophistication of asset management servicesA supervisory progress report based on monitoring (not law or rules). 13 companies, as of the end of November 2025https://www.fsa.go.jp/policy/pjlamc/20260724/01.pdf
A15Japan Investor Relations Association, 33rd Survey on IR Activities (published May 2026)A primary survey by a non-profit body. All 4,088 listed companies contacted, 948 responses (23.2% response rate)https://www.jira.or.jp/news/detail?id=281&category=2 / https://www.jira.or.jp/file/topics_file1_281.pdf
A16Japan Securities Dealers Association, Survey on Individual Investors' Attitudes Toward Securities Investment 2026A primary survey by an industry body (5,000 valid responses). No direct question on AI or ChatGPThttps://www.jsda.or.jp/shiryoshitsu/toukei/2026kozintoushika.pdf
A17CFA Institute, "Artificial Intelligence and the Future of Finance" (20 July 2026)A professional body report (not a regulatory document). Addresses governance on the investor and asset management sidehttps://rpc.cfainstitute.org/research/reports/2026/artificial-intelligence-future-of-finance
A18CFA Institute, "Survey on AI in the Investment Sector" (2024)Professional body survey (n=200)https://www.cfainstitute.org/about/press-room/2024/ai-in-investment-sector-survey
A19NIRI, "Policy Statement - Artificial Intelligence in IR" (board approved 18 September 2024)A formal policy of an IR professional body. It addresses governance and ethics for the IR function's own use of AI and is not a KPI standard for AI visibility. It states that evaluating pre-release earnings materials with public AI tools could trigger an SEC Form 8-K disclosure obligationhttps://community.niri.org/HigherLogic/System/DownloadDocumentFile.ashx?DocumentFileKey=57c9423d-bb43-fd8e-75fa-c9f7de08fab5

B. Peer-reviewed research

#SourceWhere peer review sitsURL
B1Blankespoor and others (University of Washington and others) / individual investors' use of generative AIPeer-reviewed paper. Journal of Accounting and Economics (forthcoming, November 2026 issue). More than 400,000 actual queries plus more than 2,000 survey respondentshttps://doi.org/10.1016/j.jacceco.2026.101908
B2Nakagawa, Hirano, Fujimoto, "Evaluating Company-specific Biases in Financial Sentiment Analysis using Large Language Models"Peer-reviewed research published in IEEE BigData 2024. TDnet financial results summaries (January 2019 to December 2023), seven modelshttps://arxiv.org/abs/2411.00420
B3"Beyond the Reported Cutoff" (Georgia Institute of Technology and others) / 17,621 listed companies, 197,011 questionsAccepted at COLM 2025 (the arXiv version is referenced in the body). Data from Compustat Capital-IQ, standardised in millions of US dollarshttps://arxiv.org/abs/2504.00042

C. Preprints and interested-party surveys (reference values)

#SourceHandling notesURL
C1High-confidence errors in financial QA (FinQA / TAT-QA)Preprint. The models evaluated are Qwen3-8B, Llama-3.1-8B and Gemma-2-9B, three open models in the 8-9B class, not public AI search productshttps://arxiv.org/abs/2607.11414
C2A groundedness benchmark using Form 10-KPreprint. The 755 figure is the size of the evaluation dataset, not an error ratehttps://arxiv.org/abs/2603.20252
C3FinLFQA (joint evaluation for long-form financial QA)Preprinthttps://arxiv.org/abs/2510.06426
C4"GenAI as a Reader" (USTP, HHL and nexxar, 2026)nexxar is a digital IR reporting vendor and therefore an interested party. Peer-review status unclear. Samples are 20 European companies and five DAX companies. Accuracy is from a subsample of n=200. The citation tally covers only sources explicitly displayed inside answers. The user agent is self-declaredhttps://digital-investor-relations.com/_assets/downloads/DIR_GenAI-as-Reader.pdf?h=E0wP6LL9
C5US institutional investor survey (2026, n=100)An interested-party survey by an IR advisory firm. Question wording and scales differ by item, so items cannot be lined up for direct comparisonhttps://review.brunswickgroup.com/article/investor-survey-2026/
C6Exchange-affiliated IR solutions vendor, "Global Issuer Pulse" (Q4 2025, approximately 700 IR practitioners)An interested-party survey by a company that sells IR platformshttps://www.nasdaq.com/solutions/ir-intelligence/resources/trends/global-issuer-pulse

The status of the cited URLs: these are reference links supporting the statements in this article. They do not represent the full set of models and data that each AI service referenced internally.

About this article

By Vaipm (which measures AI-space perception through a total of 25 stateless queries across multiple AI engines)

Sources verified: 8 August 2026

This article is written for IR functions, CFOs and corporate planning teams at listed companies, and sets out a way of treating corporate perception in generative AI answers as an extension of disclosure quality control. It is not investment advice and makes no share price forecasts or stock recommendations. Because the application of law, including fair disclosure rules, is judged on the facts of each individual case, always confirm specific judgments with your own legal department and outside professionals.

The proposed indicators in section 8 for how IR should measure corporate perception in AI answers are this article's reasoning, not official standards adopted by regulators or professional bodies.

The Vaipm perspective

Corporate perception in AI answers should be handled as an extension of disclosure quality control rather than as an extension of communications-style reputation management (this is a practical proposal specific to this article, not a legal requirement). Vaipm measures AI-space perception through a total of 25 stateless queries across multiple AI engines. Note that the body of research cited here establishes only that repetition is required and that conditions must be held fixed; this article does not demonstrate the statistical validity of any particular number of runs. For IR, none of this is new work. The existing job of assuring the accuracy of disclosure has simply extended to a new kind of reader.

Related articles

Department Use Cases

What AI Cites When It Describes Your Company — The Reputation Supply Chain for PR

What AI cites about your company. Japan and global data: McKinsey estimates owned sites at 5-10% of AI sources; 37.9% of citations in the first 10 SERP blocks.

PR & CommunicationsAI perception managementAIPMCitationsAMEC
Read more
Practical Guides

When the Language Changes, So Does the AI's Answer — Cross-Border AI Perception for Companies Expanding Abroad

Ask in English and you get a different answer than you get in Japanese. That prompt language changes what a model outputs is demonstrated in several peer-reviewed studies. This guide for companies operating abroad sets out why translation is a variable rather than a switch, the path by which Japanese-language primary information is less likely to be picked up by an English question, how citation sources and regulation differ market by market, and which received ideas do not hold — all organized from primary sources and peer-reviewed research.

Cross-BorderMultilingualAI Perception ManagementAIPMGlobal Expansion
Read more
Risks & Issues

AI Misinformation and Misattribution: Detect, Correct, Prevent

AI misinformation and misattribution management is the practice of continuously managing, across three layers of detection, correction, and prevention, the risk that generative AI or answer engines describe, attribute, or summarize your company wrongly. Separate from the problem of "not being cited by AI" (absence), there is the problem of "being cited, but with the content wrong" (false presence). AI citation is not a matter of careful operation but is structurally incomplete (Tow Center, ALCE, CiteFix), and the presence of a source link does not guarantee accuracy. In Japan, 87.3% of corporate staff have witnessed false presence and 76.7% say they "are measuring," yet false presence has not stopped. The problem is not the absence of measurement but that the way of measuring does not prove a state. How to deal with false presence divides into three questions: can it be challenged legally, can it be removed, and how is it measured. This article is the entry point; each question is explored in depth in a separate article.

AI misinformationMisattributionReputationAIPMRisk
Read more