This is the core of the article. The general design of verification testing that matches AI answers against published materials — the columns of a reference table, what to log, how to define the indicators — belongs to the article on how to implement verification testing. This section is confined to the checks specific to whether what was said has become text.
8-1. Establishing the starting point
In the Japan Investor Relations Association's 33rd survey, 51.1% of companies named "the number of analysts and investors attending briefings" as an indicator for measuring the effect of IR activity (previously 47.3%; base: companies that conduct IR activities). The survey notes that every option outside the top six fell below 30%, throwing into relief how difficult it is to measure contribution.
Briefings are measured by attendance. That is a natural indicator, but it does not measure what reached those who were not there. How far what was said became referenceable afterward cannot be read from attendance.
8-2. Building questions in three layers
This is the core procedure of the section. Divide what was said at the briefing into three layers according to where the text lives.
| Layer | Example content | Where the text lives | What this layer tells you |
| Layer A | What is written as text on the slides | The deck PDF (already published) | Whether published text is being read |
| Layer B | What was said but not written as text in the deck (the background to a figure, assumptions, the basis for an outlook) | Video only, or a transcript | Whether what was said is arriving |
| Layer C | Points raised during the Q&A session | Nowhere, if there is no record | How content that existed only in the room is handled |
Build questions from each layer: three to five from Layer A, three to five from Layer B, two to three from Layer C. Phrase the questions the way an investor who wanted that information would actually phrase them. With internal jargon or in-house abbreviations, getting no answer is a foregone conclusion, and you are not measuring anything.
For Layer C, decide in advance how you will interpret an answer if one comes back. If an answer consistent with content that should exist nowhere does come back, first check that it is not a wrong answer or a guess, and then check the cited URLs (§8-6). If a record did exist outside your company, that is a discovery, not an achievement.
8-3. Reading which layer the answers reach
Sort the answers into three kinds.
- Cannot answer (an answer stating that it does not hold the relevant information, or one that stays at the level of generalities)
- Answers within Layer A (states only what is within the material written in the deck)
- Goes into Layer B or Layer C (touches on content outside the deck)
Always separate "cannot answer" from "answers wrongly". The first is a state in which text does not exist or is not arriving; the second is a state in which wrong content has been taken in from another source, and the remedies differ. The first is about producing text; the second is about matching against published materials (the province of the parent article).
Stopping at Layer A is not itself a failure. It means content at the Layer A level is appearing in the answers (whether the deck itself was referenced is checked in §8-6). What becomes a problem is where Layer B and Layer C content matters to investment decisions and yet no text for it exists anywhere.
8-4. Building an inventory table of "what was said"
Alongside the measurement, check your own stock. Build it by content, not by event.
| Column | What to record |
| Event | Which briefing, and when |
| Content | Each individual thing said (broken down to a unit that fits on one line) |
| Layer | A / B / C (the division in §8-2) |
| Format | Slide body text / figure or table inside a slide / audio or video / text |
| Location | A URL on your own site / a video platform / internal only / nowhere |
| Text present | Yes / inside an image / no |
| Published | Published / internal only / not yet produced |
The purpose of this table is to make the location of the gaps visible. Rows where location is "nowhere" and the layer is B or above are the gaps this article has been pointing to. Whether to fill them is a management decision informed by how much the content matters and by the questions in §6.
Recall §2-4. At companies that already produce minutes, rows reading "published: internal only" may line up. In that case, the next question is not the production process but the decision to publish.
8-5. Whether the figures and tables in the slides are images
Even where you publish a deck as a PDF, the key points may be embedded in images. Checking is simple. Look at it by figure, not by page.
Open the PDF you publish in a browser or viewer, and try selecting and copying the numbers inside a chart or table. If you cannot select them, then at least in that viewer they cannot be confirmed as selectable text. Segment results, progress against the mid-term management plan, charts on cost of capital and return on capital — the parts investors care most about are often the ones turned into images for the sake of the layout.
As noted in §2-6, the medium most often chosen for describing how a company is responding to TSE's request is the earnings call deck. It is worth checking whether that account is embedded as an image.
Note that text existing does not guarantee it will be retrieved. Retrieval routes are the province of the technical article in the IR lane. What is being checked here is only the prior step: whether the thing to be retrieved exists.
8-6. Where the URLs cited in an answer point
When an AI answer shows source URLs, record where they lead.
| Destination | What can be read from it |
| Your transcript or summary page | Text corresponding to what was said is being referenced |
| Your deck PDF | Being referenced within the range of Layer A |
| Your video page | The page is being referenced, but the content is not necessarily text |
| Another page on your site (a prior-year page, for instance) | The intended version may not be the one being referenced |
| Third-party articles only | Either no corresponding text exists on your side, or it is not arriving |
A run of "third-party articles only" is the state this article has been dealing with. As set out in §4-3, where what was said has not become text on your side, what gets referenced is the account a third party summarized.
That said, the question of how much each type of source is cited overall — the composition of citations — is not addressed here. It is the province of What AI cites about your company. What is being looked at here is your own side: whether text corresponding to what your company said is on the reference route at all.
8-7. Tie the timing of re-measurement to the event
A single measurement tells you nothing. Repeat the same questions under the same conditions, at fixed points.
- Before the earnings call (how the previous period's content is being described)
- Immediately after the earnings call
- After the deck PDF is published
- After the video is published
- After the transcript or summary is published (where you publish one)
- And again after a set interval
It cannot be written that "publishing a transcript increases citation by X%". Within this review, we could not confirm published data showing such an effect (§9). What can be written goes only as far as the procedure: observe before and after publication, and record what changed. Whether an observed change was caused by publication cannot be separated from other factors (the results themselves, the volume of press coverage, market conditions).
8-8. How many times, and against what, to measure
The answer to a single question put once to a single model cannot be treated as a standing perception. Answers vary from run to run even for the same question, and tendencies differ by model. You need to measure several times, against several AI engines, in a state where no prior conversation carries over.
Vaipm measures AI-space perception through a total of 25 stateless queries across multiple AI engines.
This is not an optimum derived from research; it is Vaipm's operational design.
How you design the number of runs and what you run them against depends on what you want to measure. What matters is fixing the conditions and repeating, not any particular number in itself.
8-9. Things to watch when you measure
| Do | Do not |
| Build questions separately for Layers A, B, and C | Ask "about our company" without separating the layers |
| Record "cannot answer" and "answers wrongly" separately | Lump the two together as "low accuracy" |
| Record where the cited URLs lead, and re-measure on a schedule tied to the event | Count only whether a citation appeared, and measure whenever it occurs to you |
| Observe before and after publication, and record the change | Explain the change as the effect of publication |
| Use the inventory table to make the location of the gaps visible | Set exposure volume as a target |
On the last row: within the scope of this review, we could not confirm primary material from a regulator or professional body recommending exposure volume in AI space as a formal IR indicator. Rather than exposure volume, looking at which layer of what you said the answers reach, and at what those answers cite, is better suited to practice — that is this article's inference.