Department Use Cases

How to Verify What AI Says About Your Financials — Reference Table, Logging, KPIs, and the Line Not to Cross

2026-08-17Reading time 14min

By Vaigate Inc. (which operates Vaipm, measuring AI-space perception through a total of 25 stateless queries across multiple AI engines)

Key point

How to verify what AI says about your financials: the reference table columns, tolerance, stateless logging, consistency metrics, and the line not to cross.

Summary of conclusions

Disclaimer: This article organizes practical procedures and institutional rules for disclosure quality management. It is not legal advice. For individual judgments, consult your own legal department and outside counsel or other qualified professionals.

Generative AI has entered the IR function, while the company's own information inside AI answers has not become an object of measurement. That asymmetry, the four types of error AI makes with financial figures, and how far the empirical research actually reaches were covered in the hub article How AI Describes Your Financials. This article goes further and sets out the procedure for designing the measurement itself.

Three parts sit at the center. The reference table (Golden Answer Table) fixes what is correct, attribute by attribute: fiscal period, accounting standard, currency, unit, and consolidation scope. The verification test submits those questions to multiple AI services repeatedly and statelessly, and records the results under aligned conditions. The metrics measure consistency with statutory disclosure rather than exposure volume. When these three are in place, detection, root-cause identification, remediation, and confirmation of effect become one continuous loop.

And there is one boundary to confirm before putting any of this into practice. When you correct an error in an AI answer, the correction must be limited to re-presenting already-published materials. You must not correct an AI answer with material non-public information. Only published figures belong in the reference table, and an unpublished figure must never be submitted to an external AI service as the "correct answer." Share this boundary with your legal department before you begin measuring. The institutional analysis is covered in §7 of the hub article.

What you will learn

  • The columns the reference table (Golden Answer Table) should carry, and why
  • How to choose the questions to cover, and an example of an initial set
  • Why one measurement is not enough, and how to design repeated, stateless measurement
  • What to record, and which measurement conditions to align
  • How to decide the frequency of measurement
  • Metric definitions that measure consistency rather than exposure volume, and how to handle denominators and severity
  • Phrasings to avoid when reporting to management
  • Implementation steps and a checklist you can use as they stand

Who this article is for

IR practitioners at listed companies who are responsible for designing measurement and disclosure controls. Heads of IR and CFO-level readers are expected to start from §4, "Defining and operating the metrics."

1. How this article differs from the hub

1-1. What this article covers

This article is a procedure manual. It covers building the reference table, running the verification test, defining the metrics, and reporting. It assumes the reader is in a position to implement this at their own company.

1-2. What this article does not cover

The following three are not covered. All are necessary as background, but repeating them would make this article harder to read.

  • The actual state of corporate perception in AI space, and the survey data behind it. The asymmetry shown by the 33rd "Survey on IR Activities" from the Japan Investor Relations Association, the state of generative AI adoption inside IR departments, and the breakdown of effectiveness metrics. → §1 of the hub article
  • The mechanism by which generative AI gets financial figures wrong, and what the empirical research does and does not show. The explanation of the four error types, and how Beyond the Reported Cutoff, Confidently Wrong, GenAI as a Reader and others should be positioned and qualified. → §3–§5 of the hub article
  • The scope and subjects of the rules, starting with fair disclosure. Which rule governs whose conduct and what, and the distinction in legal character. → §7 of the hub article

This article takes those conclusions as given. Where evidence is needed, it refers back to the hub article at that point.

1-3. Three premises carried over

The following are premises already confirmed in the hub article. This article does not go further into them.

  1. The materials to check against exist as a matter of institutional design. Annual securities reports, financial results summaries, and timely disclosure materials can be retrieved mechanically through EDGAR, EDINET, TDnet, and XBRL. That said, it does not follow that AI always refers to them.
  2. Errors can be organized into four types. Fiscal period, accounting standard, currency/unit, and scope. This is how this group of articles organizes the problem; it is not an industry-standard taxonomy.
  3. A stable answer is not necessarily a correct one. Answers given with high confidence have been observed to be wrong. The design should therefore not draw a conclusion from a single measurement.

2. Building the reference table (Golden Answer Table)

From here the discussion turns to implementation. To run a verification test, you need a table that fixes what the correct answer is. This is what we call the reference table (Golden Answer Table).

2-1. Why a table is necessary

An arrangement in which a person reads AI answers and judges that something "seems off" does not survive a change of staff. And as the hub article showed, most errors surface as a mix-up of attributes: fiscal period, accounting standard, currency, scope. A table that does not carry those attributes as columns can detect a mismatch but cannot classify its cause.

2-2. Columns to carry

ColumnContentsWhy it is needed
question_idA distinct ID for the questionTo track it over time
question_ja / question_enThe question text submitted (Japanese and English)To measure differences between languages
expected_valueThe correct valueThe basis for matching
fiscal_periodFiscal period (e.g., FY ending March 2026)To separate out the fiscal-period error type
period_typePeriod type (full year / quarter / cumulative / LTM)To remove the ambiguity of "2025"
consolidationConsolidated / non-consolidatedTo separate out the consolidation-scope error type
accounting_standardAccounting standard (IFRS / US GAAP / Japanese GAAP)To separate out the accounting-standard error type
currencyCurrency (JPY / USD, etc.)To separate out the currency and unit error type
unitUnit (millions of yen / thousands of yen / millions of US dollars, etc.)To detect errors in the order of magnitude
continuing_opsContinuing operations basis / company-wide basisTo handle retrospective restatement after a reorganization
announced_dateAnnouncement dateTo judge whether the figure is the most recent one
source_documentName of the source document (annual securities report / financial results summary, etc.)To identify the material to re-present when correcting
source_urlSource URLThe evidence trail for the match
source_pageThe relevant page and itemReproducibility when staff change
toleranceTolerance (e.g., 0% / rounding digit)So a rounding difference is not judged a wrong answer
last_verifiedDate last verifiedFreshness management for the reference table itself

Do not treat the tolerance column lightly. Financial results summaries are rounded to millions of yen. When an AI answer says "approximately 1.234 trillion yen," you need to decide before you start measuring whether that counts as a match or a mismatch. Changing the criterion later breaks time-series comparison.

2-3. Choosing the questions

Start with 10–20 core questions. Aiming for full coverage makes the operation unsustainable. There are three selection criteria.

  1. The question is one investors actually ask. Extract from inquiries received by the IR department, the Q&A at earnings briefings, and the FAQ on the IR site.
  2. The correct answer resolves to a single value. "What is your future growth strategy" cannot go in the reference table. "What were consolidated net sales for the fiscal year ending March 2026" can.
  3. The impact of an error is large. Net sales, operating profit, net income, dividends, employee count, total shares issued, market capitalization, and the composition ratio of major segments are the usual starting points.

An example of an initial set follows. The following is a practical design example presented in this article, not an optimal configuration derived from research. Rearrange both the categories and the number of questions according to your own disclosure content and the actual pattern of investor inquiries.

CategoryApproximate number of questionsExamples
Key financial figures for the most recent full year5–6Net sales, operating profit, net income, EPS, equity ratio
Comparison with prior years2–3Year-on-year net sales growth rate, operating profit over the past three years
Shareholder returns2–3Dividend per share, payout ratio, share buyback record
Basic company attributes3–4Head office location, representative, year established, employee count, listing market
Segments and business composition2–3Net sales composition ratio by segment, overseas sales ratio
Governance and capital policy1–2Number of directors, whether the cost of capital is disclosed

It is better not to take "basic company attributes" lightly. A state in which a former officer's name or a former head office address persists is easier to detect than an error in a financial figure, and the effect of a correction is easier to measure. That makes it a suitable place to begin.

2-4. An example of the schema

The reference table can be a spreadsheet or a CSV. Here the column definitions are shown in the form of a CSV header.

question_id,question_ja,question_en,expected_value,fiscal_period,period_type,consolidation,accounting_standard,currency,unit,continuing_ops,announced_date,source_document,source_url,source_page,tolerance,last_verified
Q001,2026年3月期の連結売上高は,What was consolidated net sales for FY ending March 2026,1234567,2026-03,FY,consolidated,IFRS,JPY,million,continuing,2026-05-12,financial results summary,https://example.co.jp/ir/tanshin_2026q4.pdf,p.1,0,2026-08-16
Q002,直近の1株当たり年間配当は,What is the latest annual dividend per share,85,2026-03,FY,consolidated,IFRS,JPY,yen,total,2026-05-12,financial results summary,https://example.co.jp/ir/tanshin_2026q4.pdf,p.1,0,2026-08-16

Do not turn this table into an image. Even for internal sharing, keep it in text form, because a later step matches against it mechanically.

3. Designing the verification test — what to ask, how often, and how to record it

Once the reference table is ready, the next stage is obtaining AI answers and matching them against it. This is the core of measurement design.

3-1. Why one measurement is not enough

A single answer is one observation. But from an answer obtained by sending a single prompt once to a single model, you cannot estimate the stable tendency of what that AI returns about your company. There are four reasons.

Reason 1: the same question does not always return the same answer. Generative AI output is probabilistic. Temperature settings, sampling settings, and internal routing all mean that the same question can return different answers. One observation is one sample from a distribution. With a sample of one, you cannot distinguish a tendency from an outlier.

Reason 2: consistent does not mean correct. As the hub article showed, on the FinQA benchmark 15–23% of the "high-confidence" answers in which 8 resamplings all agreed were wrong. The stability of an answer is not evidence of its correctness. Put the other way, if you cannot observe the variation, you will mistake a stable error for a correct answer.

Reason 3: answers differ by engine. Major AI services each have different models, different search indexes, and different citation policies. You cannot look at one service alone and conclude "this is how AI talks about our company."

Reason 4: session history contaminates the result. When you ask successive questions inside the same chat, the earlier answer remains as context. An answer returned after you say "please correct the figure you just gave" is not the answer that AI returns about that company under normal conditions. Measurement has to be carried out from a clean state (stateless) every time.

3-2. Vaipm's measurement design

Vaipm measures AI-space perception through a total of 25 stateless queries across multiple AI engines.

Three points about this design should be made explicit.

  • This is not an optimum derived from research; it is Vaipm's operational design. It is set as a point that captures the variation in answers to a degree that is workable in practice, while staying within a running cost that can be sustained over time.
  • Being stateless is the essential part. Each measurement is independent and does not carry a previous answer forward as context. What is observed is therefore the distribution of answers that AI returns to that question under normal conditions.
  • Multiple AI engines are covered. The result from a single service is not treated as the whole picture.

If you build this in-house, we recommend keeping these three — repeating the measurement, not carrying history forward, and looking at multiple services — as design premises. The number of repetitions can be set to match your operating capacity, but do not design it so that a single measurement produces a conclusion.

3-3. What to record

One record of a verification test carries the following items.

Record itemContentsPurpose
run_idA distinct ID for the executionIdentifying the same batch
question_idReference to the reference tableMatching against the correct value
prompt_textThe wording submitted (verbatim)Ensuring reproducibility
executed_atExecution timestamp (with time zone)Judging the order relative to the disclosure date
engineName of the AI serviceAnalyzing differences between services
model_labelThe model name and version displayedComparison before and after an update
localeRegion and language settingsDetecting Japanese/English and regional differences
session_stateLogin state and whether history is presentMaking the conditions explicit
response_textFull text of the answerRe-judging at a later point
extracted_valueThe figure extractedThe object of mechanical matching
cited_urlsSource URLs presented in the answerCalculating the primary-source citation rate
verdictJudgment (match / mismatch / partial match / no answer)The denominator for the KPIs
mismatch_typeType of mismatch (fiscal period / standard / currency and unit / scope / other)Root-cause analysis
deltaDifference from the correct valueJudging severity

Always carry `mismatch_type`. Mapping it to the four error types (fiscal period, accounting standard, currency/unit, scope) tells you which error is the most frequent. If you look only at the overall consistency rate, you cannot identify what needs to be addressed.

An example of the record schema follows.

run_id,question_id,prompt_text,executed_at,engine,model_label,locale,session_state,extracted_value,cited_urls,verdict,mismatch_type,delta
R2026-0816-001,Q001,2026年3月期の連結売上高は,2026-08-16T10:00:00+09:00,EngineA,modelX,ja-JP,stateless,1234567,"https://example.co.jp/ir/tanshin_2026q4.pdf",match,,0
R2026-0816-002,Q001,2026年3月期の連結売上高は,2026-08-16T10:00:30+09:00,EngineA,modelX,ja-JP,stateless,1210000,"https://example.com/news/article",mismatch,fiscal_period,-24567

3-4. Conditions that must be stated

Even for the same question, the result can change when the following conditions change. You have to choose one of two: fix the conditions, or record them separately for each condition.

  • Variation in the query: "net sales," "revenue," and "operating revenue" have different formal names depending on the company. Prepare several prompts with different wording.
  • Region: the sources referred to can change depending on the region the access comes from.
  • Language: Japanese and English answers do not necessarily agree. The more a company discloses in English, the more it is worth measuring both.
  • Login state: whether you are logged in, and differences in paid plans, can change which model or search function is used.
  • Model and version differences: behavior changes across a model update even within the same service. Record `model_label` and annotate any comparison that spans an update.

3-5. Designing the frequency

The following is a practical design example presented in this article, not an optimal configuration derived from research. It is offered as one example of a frequency; set it according to your risk, disclosure events, and running cost.

TimingScopePurpose
Immediately after a quarterly earnings announcement (announcement day to 3 business days)All questions on the most recent figuresTo measure how quickly the update is reflected
Two weeks after the announcementSame as aboveTo identify items that were not reflected
MonthlyBasic attributes and shareholder returnsTo detect persistence of stale information
After the annual securities report is filedAll questionsTo confirm consistency with statutory disclosure
After a serious wrong answer is detected (14 and 30 days after remediation)The relevant questionsTo confirm the effect of the remediation
After a reorganization or a change of officers is announcedThe relevant questionsTo detect persistence of stale information

The measurement immediately after an earnings announcement carries the most information. It lets you observe when and how AI answers change from the moment the disclosure is updated.

3-6. Combining with other measurement methods

The verification test measures the content of AI answers. Combine it with methods that measure reach through AI.

  • The Google Search Console performance report: impressions and clicks that arrive through AI Overviews and AI Mode are counted within the overall data for the "Web" search type. Filter by the IR site directory and compare before and after a disclosure update.
  • Server logs: which bots crawl which pages of the IR site, and how often. As the log study in GenAI as a Reader showed, this information itself can be obtained. However, a user agent does not prove that the content was used in the final answer. Do not confuse being crawled with being cited.
  • Referrers: whether there is traffic from AI services. Note, though, that a path in which someone reads an AI answer and then arrives via search cannot be captured.

These three measure different layers. Accuracy of content (the verification test), exposure on the search surface (Search Console), and machine retrieval (logs) are not substitutes for one another.

4. Defining and operating the metrics

4-1. Metric definitions

MetricDefinitionDenominatorCharacter
Consistency rate with statutory disclosureOf the answers to the questions in the reference table, the share that matched the correct valueMeasurements for which an answer was obtainedKPI
Primary-source citation rateOf the source URLs presented in answers, the share pointing to the company's own statutory disclosure or IR siteTotal number of cited URLsKPI
Latest-fiscal-year adoption rateThe share of answers given with figures based on the most recent disclosureMeasurements for questions asking about the most recent figuresKPI
Unit/currency error rateOf mismatches, the share whose mismatch_type was currency/unitTotal number of mismatchesKRI
Stale-officer / stale-business persistence rateOf the questions on basic attributes, the share answered with pre-update informationMeasurements for questions on basic attributesKRI
Time to re-confirmation after remediation for a serious wrong answerDays from detection of a serious wrong answer to confirming a match on re-measurement after remediationKRI
No-answer rateThe share for which no numerical answer was obtainedAll measurementsReference

4-2. Three cautions in operating the metrics

Caution 1: fix the denominator. The value changes a great deal depending on whether the denominator of the consistency rate is "all measurements" or "measurements for which an answer was obtained." When no-answers increase, the former falls and the latter rises. Decide which one you use, and always state the no-answer rate alongside it.

Caution 2: separate by severity. An error in the order of magnitude of net sales and a 10-person difference in employee count must not be counted as the same "one mismatch." An example of severity definitions follows.

SeverityCriterionResponse
SeriousAn error in the order of magnitude, an error in the direction of results (higher or lower profit), an error in dividends, description of an event that did not occurImmediate response. Report to the head of IR
MediumA numerical difference beyond tolerance, a mix-up of accounting standard or consolidation scopeAddress by the next disclosure
MinorA rounding difference, a variation in wording, pre-update information whose practical impact is smallRecord only

Caution 3: do not mix in the sentiment layer. As the hub article showed, a phenomenon has been observed in which the assessment changes depending on whether the company name is present. That is a layer separate from factual correctness. If you mix an assessment of tone into the consistency metric, you can no longer read which of the two is moving. If you are going to measure tone, separate it as a distinct metric.

4-3. The shape of reporting to management

When reporting to the CFO or the board, narrowing to the following three points makes discussion easier to hold.

  1. The number of serious mismatches and what they were (what was observed, on which AI, and when)
  2. The remedial action and the result of re-confirmation after it (whether the action worked)
  3. The trend in the consistency rate (whether it is improving or deteriorating)

Avoid reporting in the form of "our assessment in AI space has improved." Treating a favorable answer from a given model as evidence of improved corporate value has no basis. For the same reason, we do not recommend setting answer share or positive sentiment as a management target.

5. Implementation steps and checklist

5-1. Implementation steps

Step 1: decide the questions to cover (1–2 weeks)

  1. Extract frequently asked questions from the IR department's inquiry history, the Q&A at earnings briefings, and the FAQ on the IR site.
  2. Narrow to 10–20 questions using the criteria in §2-3 of this article (actually asked / the correct answer resolves to a single value / the impact of an error is large).
  3. Share the question list with the legal department and confirm that no question touches unpublished information.

Step 2: build the reference table (1–2 weeks)

  1. Create a table with the columns in §2-2 of this article.
  2. Transcribe the correct value for each row from annual securities reports, financial results summaries, and timely disclosure materials. Transcribe from published materials, not from internal management figures.
  3. Enter the source URL, page, and announcement date.
  4. Decide the tolerance and enter it.
  5. Ask the accounting department to review the correct values.

Step 3: carry out the initial measurement (1 week)

  1. Decide which AI services to cover.
  2. For each question, measure repeatedly from a state that does not carry history forward.
  3. Record the items in §3-3 of this article.
  4. Match against the reference table and assign verdict and mismatch_type.

Step 4: carry out the initial analysis (1 week)

  1. Calculate the consistency rate, the primary-source citation rate, and the no-answer rate.
  2. Look at the distribution of mismatch_type and identify which of the four types is most frequent.
  3. Classify by severity.
  4. For serious mismatches, check whether the cause lies on your own IR site (whether the relevant information exists as text, whether the fiscal year is stated explicitly, whether an old page is still live).

Step 5: remediate and re-measure (4–6 weeks)

  1. Make the corrections on the IR site.
  2. Re-measure the relevant questions 14 days and 30 days after the correction.
  3. Record the time to re-confirmation after remediation.

Step 6: build it into routine operation (ongoing)

  1. Put measurement on a routine footing following the frequency table in §3-5 of this article.
  2. Review the trend in the metrics each quarter.
  3. Update last_verified in the reference table (replace the correct values at each earnings announcement).

5-2. Pre-submission checklist

Design

  • ☐ The reference table has columns for fiscal period, period type, consolidated/non-consolidated, accounting standard, currency, unit, and continuing operations scope
  • ☐ The tolerance was decided before measurement started
  • ☐ The source of each correct value is a published material, and the URL and page are recorded
  • ☐ The legal department has confirmed that no unpublished information is in the reference table

Measurement

  • ☐ Each measurement is carried out from a state that does not carry history forward
  • ☐ The design does not draw a conclusion from a single observation
  • ☐ Multiple AI services are covered
  • ☐ The prompt wording is recorded verbatim
  • ☐ The execution timestamp is recorded with the time zone
  • ☐ The model name and version are recorded
  • ☐ Region, language, and login state are recorded or fixed

Analysis

  • ☐ The denominator of the consistency rate is defined, and the no-answer rate is stated alongside it
  • ☐ mismatch_type is classified into the four types
  • ☐ The criteria for severity are defined
  • ☐ Assessment of tone and sentiment is not mixed into the consistency rate

Response

  • ☐ Corrections are limited to re-presenting published materials
  • ☐ No material non-public information has been entered into an external AI service
  • ☐ The wording of statutory disclosure itself has not been changed for the benefit of AI
  • ☐ AI output is not being used as a substitute for explanation to investors
  • ☐ Re-measurement after remediation is built into the schedule

Reporting

  • ☐ The number of serious mismatches and what they were is reported
  • ☐ The remedial action and the result of re-confirmation are reported
  • ☐ The report does not take the form of "our assessment in AI space has improved"
  • ☐ Exposure volume in AI space is not presented as a formal IR performance metric

6. FAQ

Q1. How can we tell whether generative AI is getting our financial figures wrong?

The reliable way is to build a reference table (Golden Answer Table) and run a verification test. Start by choosing 10–20 questions investors frequently ask, and transcribe the correct value for each from the annual securities report or the financial results summary. When you do, carry fiscal period, consolidated/non-consolidated, accounting standard, currency, unit, and continuing operations scope as columns. Then submit those questions repeatedly to multiple AI services from a state that does not carry history forward, and match the figures returned against the correct values. It matters that you do not try once and conclude "it was right." Answers vary from execution to execution.

Q2. Why is one measurement not enough?

A single answer is one observation. However, it is not enough to estimate a stable tendency in the answers. There are four reasons. First, generative AI output is probabilistic, so the same question can return different answers. Second, an answer being consistent does not mean it is correct. In a study using the FinQA benchmark (an arXiv preprint posted in July 2026; it has not been peer reviewed), 15–23% of the "high-confidence" answers in which 8 resamplings all agreed were wrong. Third, AI services differ in the sources they refer to. Fourth, when you ask successive questions inside the same chat, the earlier answer remains as context, so it is no longer the answer given under normal conditions. With a sample of one, you cannot distinguish a tendency from an outlier.

Q3. How many questions should we start with?

We recommend 10–20. Aiming for full coverage makes the operation unsustainable. As a rough breakdown: 5–6 questions on key financial figures for the most recent full year, 2–3 on comparison with prior years, 2–3 on shareholder returns, 3–4 on basic company attributes, 2–3 on segments and business composition, and 1–2 on governance and capital policy. That said, this is a design example presented in this article, not an optimal configuration derived from research. If you are starting out, basic attributes such as head office location, representative, and employee count are a good fit. Persistence of stale information is easy to detect there, and the effect of a correction is easy to measure.

Q4. How should we decide the tolerance?

As a principle, decide it before you start measuring and do not change it afterwards. Financial results summaries are rounded to millions of yen, so when an AI answer says "approximately 1.234 trillion yen," a judgment is required as to whether that is a match or a mismatch. In practice, a reasonable place to start is to allow differences caused by rounding of digits and treat everything else as a mismatch. What matters is less the criterion itself than fixing the criterion. Changing it later breaks time-series comparison, and you can no longer distinguish an improvement from a loosened standard.

Q5. Among the record items, which ones must not be omitted?

The prompt wording (verbatim), the execution timestamp (with time zone), the name of the AI service, the model name and version displayed, and the type of mismatch. The four former items are needed for reproducibility; without them you cannot trace afterwards why a given result occurred. The type of mismatch is needed for analysis. If you look only at the overall consistency rate, you cannot identify what needs to be addressed. Keeping the full text of the answer as well lets you re-judge when you revise the judgment criteria.

Q6. Should measurement conditions be fixed, or should we vary them?

The point is to choose one and record it. Fixing the conditions makes time-series comparison easier; varying them lets you detect the effect of the conditions themselves. What is bad is comparing under the assumption that conditions are aligned when they are not. Language (Japanese versus English) and region in particular are where differences tend to show up at companies that disclose in English. The model used can also change with login state or paid plan, so record that point too.

Q7. How often should we run it?

Building the schedule around the period immediately after quarterly earnings announcements is realistic. That timing is well suited to observing when and how AI answers change from the moment the disclosure is updated. Taking two measurements — one from announcement day to within 3 business days, and one two weeks after the announcement — shows both the speed at which updates are reflected and the items that were not reflected. Monthly for basic attributes, and all questions after the annual securities report is filed, is one workable combination. That said, this is one example; set it according to your risk, disclosure events, and running cost. There is no optimal frequency derived from research.

Q8. How should we define the denominator of the consistency rate?

The value changes a great deal depending on whether you use "all measurements" or "measurements for which an answer was obtained." When no-answers increase, the former falls and the latter rises. Decide which one you use, and always state the no-answer rate alongside it. Without that, when the consistency rate rises you cannot tell whether things actually improved or the AI simply stopped answering. It is also necessary to separate by severity. If an error in the order of magnitude of net sales and a 10-person difference in employee count are both counted as "one case," the metric stops reflecting reality.

Q9. How should we report to management?

Narrowing to three points makes discussion easier to hold: the number of serious mismatches and what they were (what was observed, on which AI, and when), the remedial action and the result of re-confirmation after it (whether the action worked), and the trend in the consistency rate (whether it is improving or deteriorating). What to avoid is reporting in the form of "our assessment in AI space has improved." Treating a favorable answer from a given model as evidence of improved corporate value has no basis. We also do not recommend setting answer share or positive sentiment as a management target.

Q10. If we find an error, may we ask the AI provider to correct it?

Nothing prevents you from asking, but we do not recommend making that your main response. Within the scope of this review, we could not confirm a correction process for companies common to the major AI services, or a procedure that guarantees the outcome of a correction. The practical response is to save a reproducible prompt and timestamp, identify exactly where the answer contradicts published primary materials, fix gaps, old pages, unit display, and legal-entity identification on your own IR site, and re-verify across multiple AI services after an interval. In doing so, you must not use material non-public information to correct. Limit corrections to re-presenting published materials. The institutional analysis is covered in §7 of the hub article.

Q11. Should we build this in-house or use an outside service?

It is a question of scale and continuity. If you are checking around 10 questions by hand once a quarter, in-house works fine. On the other hand, once you put repeated, stateless measurement across multiple AI services on a routine footing, keep recording under aligned conditions, and follow every disclosure, the running cost becomes hard to ignore. The dividing line is whether you are doing measurement once as "research" or continuing it as "control." If you are aiming at the latter, we recommend putting the record schema into a machine-readable form from the outset. Migrating from a manual spreadsheet only gets more burdensome the longer you wait.

7. Wrap-up and next actions

7-1. Key points

  1. Give the reference table a column for each attribute. Fiscal period, period type, consolidated/non-consolidated, accounting standard, currency, unit, continuing operations scope, announcement date, source. Without the columns, you can detect a mismatch but cannot classify its cause.
  2. Decide the tolerance before measuring, and do not change it afterwards. Changing it later breaks time-series comparison.
  3. Start with 10–20 questions. Aiming for full coverage makes the operation unsustainable. Starting from basic attributes makes the effect easier to measure.
  4. Do not draw a conclusion from one measurement. Repeat, do not carry history forward, and look at multiple services.
  5. Either fix the measurement conditions or record them separately for each condition. Do not compare under the assumption that conditions are aligned when they are not.
  6. What you measure is consistency, not exposure volume. Define the denominator, state the no-answer rate alongside it, and separate by severity.
  7. Narrow reporting to management to three points: serious mismatches, the result of remediation, and the trend in the consistency rate.
  8. Limit corrections to re-presenting published materials. You must not correct an AI answer with material non-public information. What goes into the reference table is likewise only published figures.

7-2. Next actions

This week: pull 10 frequently asked questions from the IR department's inquiry history and select those for which the correct answer resolves to a single value. At the same time, confirm with the legal department that the question list contains nothing touching unpublished information.

This month: build the skeleton of the reference table and fill in fiscal period, consolidated/non-consolidated, accounting standard, currency, unit, and source URL for the key financial figures of the most recent full year. Decide the tolerance. Ask the accounting department to review it.

At the next earnings announcement: on the day of the announcement and two weeks later, submit the same questions repeatedly to multiple AI services and record how the answers change. The distribution of mismatch types you obtain there is what sets the priority order for subsequent remediation.

7-3. Sustaining the measurement over time

The procedure up to this point can be carried out in-house. However, putting repeated, stateless measurement across multiple AI services on a routine footing, and continuing to record it under aligned conditions, carries a running cost that is not trivial.

Vaipm measures AI-space perception through a total of 25 stateless queries across multiple AI engines. This is not a tool for AIO measures; it is a mechanism for managing perception in AI space on an ongoing basis (AI Perception Management). On the field itself, see What AI Perception Management Is; on how long information stays in AI answers, see How Long Information Stays in AI Answers. General measures against misinformation and how far correction is possible are covered in Misinformation in AI Answers and What to Do About It and How Far AI Answers Can Be Corrected.

For the actual state of corporate perception in AI space, the mechanism behind the errors, how far the empirical research reaches, and the institutional boundaries, see the hub article How AI Describes Your Financials.

Sources

Primary and technical sources

  1. U.S. Securities and Exchange Commission, "EDGAR Application Programming Interfaces" https://www.sec.gov/search-filings/edgar-application-programming-interfaces
  1. Financial Services Agency, EDINET https://disclosure2.edinet-fsa.go.jp/week0020.aspx
  1. Japan Exchange Group, "TDnet (Timely Disclosure Network)" — an exchange rule, not a national statute in itself https://www.jpx.co.jp/english/equities/listing/disclosure/tdnet/
  1. Japan Exchange Group, "XBRL" https://www.jpx.co.jp/english/equities/listing/disclosure/xbrl/03.html
  1. Google Search Central, "AI features and your website" (last updated: December 10, 2025) — impressions arriving through AI features are also counted within the "Web" search type of the Search Console performance report https://developers.google.com/search/docs/appearance/ai-features

Academic research and studies

  1. Richard Zhe Wang, "Confidently Wrong: Detecting Hallucinations in Financial Question Answering from LLM Internal States" (posted July 13, 2026) — an arXiv preprint, 8 pages. The FinQA and TAT-QA benchmarks. The models covered are Qwen3-8B, Llama-3.1-8B, and Gemma-2-9B. This is a benchmark of foundation models, not an evaluation of publicly available AI search products https://arxiv.org/abs/2607.11414
  1. Monika Kovarova-Simecek, Henning Zülch, Leon Kirschbaum, Konstantin Klammer, Eloy Barrantes, Alexandra Horváthová, Christina Schilling, "GenAI as a Reader: How ChatGPT & Co. use Annual Reports" (USTP, HHL Leipzig, nexxar) — not a peer-reviewed paper but a practitioner report (Document Type: Report in the HHL repository). nexxar is a digital IR reporting vendor and therefore an interested party. What this article refers to is the log analysis in Part 2 (five DAX companies, August 29 to October 25, 2025, 4,838,833 automated accesses, 759,226 analyzed) https://digital-investor-relations.com/_assets/downloads/DIR_GenAI-as-Reader.pdf?h=E0wP6LL9

The survey data this article carries over as premises (the 33rd "Survey on IR Activities" from the Japan Investor Relations Association and others) and the full source list for the empirical research are given in the sources section of the hub article, "How AI Describes Your Financials."

Disclaimer: This article organizes practical procedures and institutional rules for disclosure quality management. It is not legal advice. It also offers no investment recommendation on any particular security, no share price forecast, and no investment advice. For individual legal judgments, consult your own legal department and outside counsel or other qualified professionals.

The Vaipm perspective

This article is the measurement layer of AI Perception Management. A verification test turns "what AI says about our financials" from an impression into a recorded, repeatable observation, and that record is the precondition for managing perception in AI space rather than reacting to it after the fact. Vaipm provides this as an ongoing measurement operation, not as a tool for AIO measures.

Related articles

Department Use Cases

How AI Describes Your Financials: What IR Can Measure, and What It Cannot

IR adopted generative AI fast, but it measures the AI it uses, not the AI investors read. Where AI gets your financials wrong, and what to measure instead.

IRDisclosureFair DisclosureAI Perception ManagementAIPM
Read more
Fundamentals

AI Perception Management (AIPM): Measuring and Governing How AI Describes Your Brand

AI Perception Management (AIPM) is the practice of measuring — and continuously governing — how generative AI and answer engines perceive, describe, and cite your organization. Vaipm's definitive guide explains how it differs from AIO, GEO, and LLMO; the category and leading vendors that emerged abroad; governance and regulation; and how to measure impact — all grounded in primary sources.

AIPMAI Perception ManagementFundamentals
Read more
Risks & Issues

AI Misinformation and Misattribution: Detect, Correct, Prevent

AI misinformation and misattribution management is the practice of continuously managing, across three layers of detection, correction, and prevention, the risk that generative AI or answer engines describe, attribute, or summarize your company wrongly. Separate from the problem of "not being cited by AI" (absence), there is the problem of "being cited, but with the content wrong" (false presence). AI citation is not a matter of careful operation but is structurally incomplete (Tow Center, ALCE, CiteFix), and the presence of a source link does not guarantee accuracy. In Japan, 87.3% of corporate staff have witnessed false presence and 76.7% say they "are measuring," yet false presence has not stopped. The problem is not the absence of measurement but that the way of measuring does not prove a state. How to deal with false presence divides into three questions: can it be challenged legally, can it be removed, and how is it measured. This article is the entry point; each question is explored in depth in a separate article.

AI misinformationMisattributionReputationAIPMRisk
Read more