4-1. Observation of numerical hallucination at scale
Beyond the Reported Cutoff: Where Large Language Models Fall Short on Financial Knowledge (Shah, Ye, Jaskowski, Xu, Chava / Georgia Institute of Technology) is a peer-reviewed paper accepted at COLM 2025 (Conference on Language Modeling). This article refers to the arXiv version (arXiv:2504.00042).
The study built 197,011 question-answer pairs from Compustat sales data. It covers 17,621 companies listed on 13 exchanges over the 43 years from 1980 to 2022. The prompt is fixed in form: "what were the sales of {company name} in {fiscal year}." Numbers were extracted from the answers and scored on three values: correct if the absolute error against the reference value was under 10%, a factual hallucination if 10% or more, and no answer where no figure was returned.
The main observations were as follows.
- Older years are answered less well. Llama-3-70B-Chat answered accurately for 54.17% of companies for 2017, but only 6.32% for 1995 — even though financial information has been published through EDGAR in the United States since 1995.
- The study observed a tendency for larger and newer companies to attract both more knowledge and more hallucination. That is, the higher the market capitalization, the higher the rate of correct answers, and at the same time the more incorrect answers about the same companies. This is the result of an experiment on a particular set of models, using a single question form asking about sales, and scored against Compustat values. It must not be read as a law holding for generative AI in general.
- Search activity, institutional investor attention, access counts for SEC filings, and the readability of filings were also associated with the rate of correct answers.
What cannot be said from this study: whether a paper has been peer-reviewed and how far its results generalize are separate questions. The reservations belong not with peer-review status but with the experimental conditions.
- Sales figures come from Compustat and are converted to millions of U.S. dollars, so differences in accounting standard are not reconciled. What the study measures is therefore consistency with the values of one standardized database, not consistency with each company's own statutory disclosure.
- The question takes a single form: "what were the sales of {company name} in {fiscal year}." It does not represent the range of questions investors put in IR practice.
- The API calls were made in February 2025. These are not figures for the performance of current models.
- The subjects are U.S.-listed companies. It cannot be said that the findings apply as they stand to Japanese companies.
4-2. Adding the company name alone changes the assessment (Japanese data)
Evaluating Company-specific Biases in Financial Sentiment Analysis using Large Language Models (Nakagawa, Hirano, Fujimoto) is a peer-reviewed conference paper included in the 2024 IEEE International Conference on Big Data (BigData) (pp. 6614–6623, DOI: 10.1109/BigData62323.2024.10826008). A preprint version is also on arXiv (2411.00420).
Handling of this source: it is a peer-reviewed conference paper included in IEEE BigData 2024. However, the arXiv version published by the authors was used to check the methods and figures in the text, and any verbatim or methodological differences from the IEEE version have not been independently confirmed for this article.
The method is straightforward. For the same statement about business performance, the sentiment score returned by a large language model is compared between prompts that include the company name and prompts that do not, and the difference is defined as "company-specific bias." A positive difference means the model shifts its assessment in a favorable direction for that company. Japanese financial text data was used for the analysis.
The implication for IR practice sits in a layer separate from the correctness of figures: even for the same disclosed text, the interpretation can change once a company name is attached. This is a domain where "factually incorrect" is hard to establish, and a consistency rate will not capture it. In the indicator design of §6 of this article, and in the companion article on designing verification tests for AI answers, this layer is treated separately.
What cannot be said from this study: the authors are practitioners and researchers affiliated with Nomura Asset Management and Preferred Networks. It is a peer-reviewed conference paper and not a study conducted to sell a particular service, but it is worth noting that it was designed from the perspective of financial practice. Nor can it be said that a consistent conclusion has been reached across models about how the direction of bias corresponds to company characteristics. It must not be cited as a general rule that "large companies are assessed favorably."
4-3. "Confident and wrong"
Confidently Wrong: Detecting Hallucinations in Financial Question Answering from LLM Internal States (Richard Zhe Wang, submitted July 13, 2026) is an arXiv preprint, and the one paper among the four covered in this section that has not been peer-reviewed (arXiv:2607.11414, 8 pages). The figures below should be read on that basis.
The study uses two financial question-answering benchmarks built from actual filings, FinQA and TAT-QA. It resampled each question eight times and defined the state in which all eight answers agreed as "high confidence." The result: in FinQA, 15–23% of high-confidence answers were incorrect.
The models covered are Qwen3-8B, Llama-3.1-8B and Gemma-2-9B. The focus of the study is whether such errors can be detected by training linear probes on the model's internal states (the residual stream). The probes held an AUROC of 0.68–0.77, exceeding baselines such as token log-probabilities and the model's own self-report of correctness (at best 0.55–0.63).
The implication for IR practice: this is one of the grounds on which this article recommends repeated measurement. Asking the same question again and getting the same answer is not proof of correctness. Consistency and accuracy are different properties.
What cannot be said from this study:
- The subjects are open-weight foundation models, not the AI search products available to the public. These are not figures for the error rate of products such as ChatGPT or Gemini that combine a model with web search.
- FinQA and TAT-QA are benchmarks built from filings; they do not reproduce how the services are actually used.
- The 15–23% figure is a share "of high-confidence answers," not an error rate across all answers.
4-4. What is happening on the investor side
Blankespoor, Croom and Grant, "Generative AI and Investor Processing of Financial Information" is a peer-reviewed paper published in the Journal of Accounting and Economics, Volume 82, Issue 2 (November 2026 issue) (DOI: 10.1016/j.jacceco.2026.101908), by a research team at the University of Washington.
The study uses two kinds of data: archival analysis of more than 400,000 investor queries submitted to a major brokerage's generative AI chatbot, and a survey of more than 2,000 individual investors. In the survey, close to half reported using generative AI. The uses are mainly interpreting financial information and putting market movements in context, along with screening stocks and making complex research more efficient. As users spend longer with the tools, use shifts from high-level screening toward detailed monitoring and interpretation of news about individual companies. More sophisticated individual investors are reported to have adopted the tools further.
On the institutional side there is the Brunswick 2026 US Investor Survey (published February 9, 2026). It has n=100, covering U.S. institutional investors (active equity), half long-only managers and half hedge funds. Brunswick is a firm that supports corporate IR and communications, so this is a survey by an interested party — a point worth stating plainly.
| Brunswick 2026 US Investor Survey (n=100) | Share |
| Rate generative AI as at least "moderately important" for investment research | 54% |
| Use AI as a primary tool when researching a new investment candidate in depth | 42% |
| Say AI has changed how they approach earnings calls | 68% |
| Agree they can check more information, spend time on higher-value work, and cover more names | 84% |
| Use AI to update financial models | 23% |
On the importance of information sources, 77% rated direct human contact with management as "important" or "very important," company disclosure 66%, dialogue with IR 63%, corporate websites and IR pages 47%, and generative AI output 24%. The questions and scales differ, so a simple comparison of usage rates is not possible, but it can at least be read that AI is not replacing primary materials or human dialogue.
The same survey reports that investors are aware of AI's weaknesses. Outdated information, frequent hallucination, numerical errors and missed context are all cited. In financial forecasting, the use requiring the most caution, only 23% said they use AI to update models, and more than half expressed reservations.
The implication of these two studies: investors use generative AI not as a substitute for primary materials but as a route toward them. Precisely for that reason, the significance of an error about your company inside an AI answer is not that it sways the whole investment decision, but that the path to the primary materials becomes distorted.
4-5. Four disciplines for reading this body of research
Of the three AI studies cited in this section, two have been peer-reviewed — Beyond the Reported Cutoff (COLM 2025) and Evaluating Company-specific Biases (included in IEEE BigData 2024) — and one, Confidently Wrong, is a preprint (on the investor side, Blankespoor and colleagues is a peer-reviewed paper in the Journal of Accounting and Economics). Peer-review status, however, is not the main axis for reading them. Having been peer-reviewed does not guarantee that a result may be applied to your own company. The reading is done on these four points.
- Look at the experimental conditions. What was treated as correct (a standardized database such as Compustat, or the filings themselves), what form the questions were limited to, and what threshold defined an error. The 10% threshold in Beyond the Reported Cutoff is also a definition that study set.
- Look at the models covered. Open-weight foundation models, or public AI services integrated with web search. Benchmark results for foundation models must not be cited as the performance of public AI services.
- Look at where the data came from. U.S.-listed or Japanese companies, what period was covered, and when the work was run. If the API calls were made in February 2025, that is the behavior of the models at that time.
- Do not read correlation as causation. Causal readings such as "accurate because it is HTML" or "wrong because the company is large" are supported by none of these studies.