Practical Guides

When the Language Changes, So Does the AI's Answer — Cross-Border AI Perception for Companies Expanding Abroad

2026-08-07Reading time 24min

By Vaipm (which measures AI-space perception through a total of 25 stateless queries across multiple AI engines)

Key point

Ask in English and you get a different answer than you get in Japanese. That prompt language changes what a model outputs is demonstrated in several peer-reviewed studies. This guide for companies operating abroad sets out why translation is a variable rather than a switch, the path by which Japanese-language primary information is less likely to be picked up by an English question, how citation sources and regulation differ market by market, and which received ideas do not hold — all organized from primary sources and peer-reviewed research.

§0 The observation scope of this article — what is demonstrated, and what is not yet

The question that arrives from companies operating abroad almost always has the same shape. "When we ask in Japanese, our company appears. When we ask in English, it does not — or it is described in a completely different way. What is happening here?"

This article declares its observation scope first. Before an investment decision, you need to know where the evidence ends and inference begins.

What is demonstrated. Changing only the language of the prompt changes what a model outputs in response to a question that means the same thing. This has been confirmed in several peer-reviewed studies. That models are systematically pulled toward the values of a small number of specific countries; that multilingual retrieval-augmented generation (RAG) can introduce a preference for English and for the query language at the retrieval and reranking stages; that models tend to associate global brands and high-income regions with positive attributes; that geographic information gaps do not close much when the language changes — these too exist as peer-reviewed research.

What is not demonstrated. Within the scope of this review, we could not confirm a public benchmark that compares Japanese and English under controlled conditions for named Japanese companies. Nor could we confirm public data that would support the flat generalization that Japanese companies are at a disadvantage in English-language AI. What is academically established goes as far as the general mechanism by which language shapes model output.

This article therefore takes the following position. The mechanism can be explained. The direction can be estimated. But what is happening to your company can only be known by measuring your company. How to fill that gap is the subject of the final section (§10).

The problem of AI answers describing your company with outdated or incorrect content is itself handled in AI misinformation and misattribution, which is the canonical article for that topic. This article is confined to the axes of language and market.

§1 The same question in a different language gives a different answer — and the native language is not always better

The starting point is a study accepted at the main conference of ACL 2024. AlKhamissi and colleagues selected 30 questions from Wave 7 of the World Values Survey (WVS), gave models "personas" carrying the same attributes as actual respondents in Egypt and the United States, and had the models answer. The Egyptian survey had 1,200 respondents and the United States 2,596. From these, 303 demographically matched personas were drawn for each country, and the same questions were posed in both English and Arabic. Four models were tested: GPT-3.5, mT0-XXL, LLaMA-2-13B-Chat and AceGPT-13B-Chat.

The results split into a part that matches intuition and a part that runs against it.

Start with the intuitive part. For GPT-3.5, the alignment scores (Soft / Hard) against the Egyptian survey were 47.08 / 23.42 in English and 50.15 / 28.56 in Arabic. Against the United States survey they were 65.95 / 40.22 in English and 63.77 / 38.36 in Arabic. Asking in the dominant language of that culture moved the model closer to the actual responses.

The problem is that this did not hold for every model. For LLaMA-2-13B-Chat, pretrained mainly in English, alignment against the Egyptian survey was 47.95 / 25.61 in English and 44.67 / 23.34 in Arabic — asking in Arabic made it worse. And mT0-XXL, trained with a more balanced multilingual mixture, scored 53.20 / 28.30 in English and 57.75 / 34.51 in Arabic against the United States survey, a reversal in which asking in Arabic moved it closer to the United States responses.

The practical implication is clear. "Ask in the local language and you get the local context" is a conditional phenomenon that depends on the composition of a model's training data, not a law. Which of the two — the same question in Japanese, or in English — returns the more accurate answer about your company can differ from model to model. It is not something you can settle in advance by reasoning.

The same study reports one further fact that matters commercially. All four models were significantly closer to the United States responses than to the Egyptian ones (model averages of 47.16 / 27.03 for Egypt against 59.07 / 33.78 for the United States). In this study, changing the prompt language did not consistently dissolve the overall observed tendency toward high alignment with United States responses.

In addition, consistency of answers across four paraphrases of the same question was, in this English-Arabic comparison, slightly higher in English: 74.46 on model average against 72.81 for Arabic. For AceGPT-13B-Chat the order was reversed, with Arabic higher (66.66 against 61.84), so there are model-level exceptions. This result cannot be extrapolated to non-English languages in general. Even so, it does suggest how precarious it is to conclude "this is what they say about us in English" from a single observation.

§2 Translation is not a localization switch — it is a variable that conditions the model

The study by Bulté and Rigouts Terryn, published in Computational Linguistics in 2026, took the same problem at a far larger scale. Ten LLMs (two Claude models, three GPT models, Gemini, DeepSeek, Llama, Mistral and Qwen) were given 63 items in total — 24 from Hofstede's Values Survey Module (VSM) and 39 from the WVS — translated into 11 languages (including Japanese, selected to span four language families), producing 332,640 collected responses.

What is elegant in the design is that it separates four prompt conditions: (1) local language, no perspective specified; (2) local language, "answer from a human perspective"; (3) local language, "answer from the perspective of a person from country X"; and (4) English, "answer from the perspective of a person from country X".

Alignment with actual human responses (Pearson correlation) came out in the following order.

Prompt conditionCorrelation with VSM dimensionsCorrelation with WVS items
English + explicit cultural perspectiver = .53r = .44
Local language + explicit cultural perspectiver = .28r = .38
Local language + general human perspectiver = .10r = .23

In this aggregation, the condition that came closest to actual human responses was not the local language but English with an explicit instruction to answer from the perspective of a person from country X. The authors state that what produces variation aligned with the variation present in the human data is the cultural perspective rather than the prompt language, and note that variation caused by language alone is largely orthogonal to the human variation and can even mask culturally meaningful signal. Using language and cultural perspective together did not beat English plus cultural perspective either.

That ranking, however, depends on how you aggregate. The table above reports correlations that compare cross-country patterns item by item and dimension by dimension. In a different aggregation that compares each country's response profile across items, the target language plus cultural perspective reached r = .57 and English plus cultural perspective r = .55, with the former slightly ahead. "Asking in English is always best" is not the finding. The central conclusion this study supports is that making the cultural perspective explicit is more consistently effective at improving alignment with human cultural differences than changing the language alone. Of the four conditions, the one with no perspective specified was unstable — the model answered sometimes as a human and sometimes as an AI — so the authors excluded it from subsequent analyses other than the valid-response-rate analysis. That is why the table shows three conditions.

Look instead at the amount of movement and the picture changes. Changing the prompt language moved answers by roughly as much as explicitly demanding a cultural perspective. In other words, changing the language changes the answer a great deal — but that change is not necessarily in the direction of how people in that language community actually see things.

This duality is the part most often misread in practice. "We shipped an English version, so localization for English-speaking markets is done" is wrong twice over. Translation changes the output, but the direction of the change is not under control. Translation is not a localization switch; it is a variable that conditions the model.

The same study reports one more finding that bears directly on readers of this article. Model answers aligned better with human responses from the Netherlands, Germany, the United States and Japan than with those from any other country (correlations of r = .60 to .75, roughly .21 higher than the next tier). This four-country cluster held across every model and every prompt condition, and did not break down substantially even when another country was specified explicitly.

This needs to be read carefully. It is a statement about "responses to a values survey", not about "the visibility of Japanese companies". Japan sitting in the high-alignment cluster does not mean AI describes Japanese companies accurately, or abundantly. On the latter, there is research pointing in the opposite direction (§4). Conflating the two leads to a dangerously optimistic conclusion.

Valid response rates also varied by language: in the condition asking for a general human perspective in the local language, Japanese was 96.86% and Russian 88.43%.

§3 The path by which Japanese-language primary information is less likely to be picked up by an English question

So far the discussion has been about the inside of the model — its parametric knowledge. But when the subject is a company, today's major AI search products assemble their answers by retrieving documents from the web. Language bias enters at that retrieval and selection stage.

A study by Wang and colleagues, accepted at the main conference of ACL 2026, targeted the reranking stage of multilingual RAG and showed that existing rerankers carry a language bias that systematically favors English and the query language. The study introduces a comparison against an achievable upper bound (an oracle) and reports that optimal answers require evidence scattered across several languages, while current systems systematically suppress exactly those documents that are essential to the answer.

Another study arrives at the same direction from a different design. Ki and colleagues, in a preprint (arXiv:2509.13930), used a method that holds document relevance constant while measuring language preference across 8 languages and 6 open-weight models. The result confirmed that when the query is in English, models preferentially cite English sources. The bias was larger for lower-resource languages, and larger for documents placed in the middle of the context. More important still is the observation that models sometimes trade off document relevance against language preference. Citation selection is not always determined by usefulness of information alone. This is a preprint, not peer-reviewed, so the point is not treated here as settled knowledge.

A third source, a large multilingual benchmark covering 49 languages accepted to Findings of ACL 2025 (BORDIRLINES), reports on behavior at the retrieval stage. It links 720 queries to 7,436 passages extracted from 905 Wikipedia articles, building 19,916 query-document pairs. Here it is confirmed that retrieval systems prefer documents in the query language (roughly 1.29 times for OpenAI embeddings, roughly 1.64 times for BGE-M3). OpenAI embeddings also retrieved roughly 1.72 times more English documents than BGE-M3 did. When queries were issued in English, even the multilingual retrieval mode was in practice barely multilingual (94.8% of citations and 93.9% of included passages were English).

Translated into the situation of a Japanese company, this reads as follows. When your most accurate primary information exists only in Japanese, there is a path on which that Japanese document struggles to clear the gates for an English question — entering the candidate set, surviving reranking, and being cited. This is not "a demonstrated disadvantage"; it is "a path along which you could be structurally disadvantaged". That distinction has to be maintained.

The same benchmark also reports a finding in the opposite direction. Searching across documents in several languages improved cross-lingual consistency of answers and reduced geopolitical bias, compared with retrieving only in the query language. The prescription is not "move everything to English". What helps consistency is that the evidence is available in more than one language.

§4 Global brands and country-of-origin effects — why a little-known company can be structurally disadvantaged

Alongside the language problem, and possibly with more force, sits bias toward the entity itself.

The study by Kamruzzaman and colleagues, accepted at the main conference of EMNLP 2024, measured brand bias head-on. Fifteen countries were split into high, medium and low bands by GDP per capita; three global and three local brands were collected for each; and GPT-4o and Llama-3 were tested across four categories.

The pattern was consistent. Models associated global brands disproportionately with positive attributes, and for gift recommendations suggested luxury brands for recipients in high-income countries and non-luxury brands for recipients in low-income countries. At the same time, the authors report that in particular contexts a country-of-origin effect can lift preference for local brands. It is not simply that "local always loses".

The limitation the study states about itself should be made explicit here. All of its experiments were conducted in English. Whether the same behavior holds in non-English contexts is therefore untested. It cannot be used as evidence of a Japanese-English difference.

Geographic bias has been observed at a more abstract level as well. The study by Manvi and colleagues, accepted at ICML 2024, showed that zero-shot geographic predictions by LLMs correlate with objective indicators up to a Spearman rank correlation of 0.89, and then reported that for subjective topics such as attractiveness, morality and intelligence there is a bias toward systematically lower ratings for regions with lower socioeconomic conditions, with correlations reaching up to 0.70. The magnitude of the bias also differed significantly between models.

Decisive for the argument of this article is the study by Lalai and colleagues, accepted at the Conference on Language Modeling (COLM) 2025. It used a multi-turn reasoning task — the game of Twenty Questions — to measure whether a model can construct its own questions and identify a target. The targets come from Geo20Q+, a dataset of well-known people and culturally significant things (foods, landmarks, animals and so on), and the game was run in seven languages, including Japanese.

The result was that entities from the Global North were clearly easier to identify than those from the Global South, and entities from the Global West easier than those from the Global East. Wikipedia pageviews and frequency in pretraining corpora explained only a small part of the gap. And the finding that matters most here: whichever language the game was played in, the geographic gap barely closed.

Layer these four studies and an outline appears. There are at least two paths on which a little-known Japanese company can be disadvantaged in AI answers: language (§3) and entity-level prominence and geography (this section). For the latter, what has been tested reaches only this far: switching the question language alone did not substantially close the gap. How far creating English pages would improve that gap cannot be judged from this study. English pages as an intervention were not tested in this research.

Separately, a preprint framework auditing brand preference in AI recommendations (arXiv:2603.18300) tested GPT-4o, Gemini 1.5 Flash and DeepSeek-V3 across 10 topics and 2,070 questions, and reports that models developed in the United States markedly prefer United States entities, while DeepSeek was more balanced but still showed detectable geographic preference. However, all 2,070 questions were run in English, and no experiment varying the prompt language was conducted (executed from Sweden in March 2025). It is therefore not direct evidence that answers change when you switch from Japanese to English. What it shows reaches only as far as model-level geographic brand preference, and — being a preprint, not peer-reviewed — it is treated here as a reference value.

§5 The cast of cited sources differs from market to market

Up to here the subject has been what models prefer. The other thing that bites in practice is which sites are cited and attributed in AI answers in a given market.

On this point, a cross-market public analysis was released on arXiv in June 2026 (arXiv:2606.25787). Consolidating three datasets from Rankfor.AI, it classifies 167,551 citations for which a URL could be identified (189,974 as attribution lines) across 128 brands, 12 home markets and 13 languages.

This analysis is a preprint, not peer-reviewed, and it originates with an interested party. The author is affiliated with Rankfor.AI, and it is not an independent industry census. What is analyzed is also the citation URLs attributed and displayed in AI answers, not what the model retrieved, referred to or weighted internally. The figures below should be read on those terms.

Four patterns are reported.

First, when AI talks about a brand, its basis is overwhelmingly third-party sites. 85.7% of citations pointed to sites the brand does not own; owned sites accounted for only 14.3%.

Second, referenced sources concentrate while keeping a long tail. 80% of citations came from roughly 18% of domains (fit to Zipf's law with alpha = 0.86, R-squared = 0.983).

Third, one referenced site dominates almost everywhere. Wikipedia was the most-cited domain in 11 of 12 languages, the exception being Lithuanian, where the business paper vz.lt edged ahead at 4.38%.

Fourth, at the edges the composition becomes market-specific. For 46 domestic brands in Poland, the most-cited domain was YouTube, and four talent and career portals together supplied 637 citations — roughly twice the 297 supplied by Polish-language Wikipedia.

This fourth finding is the one that matters for cross-border practice. Even for the same phenomenon of "being cited by AI", the supply structure differs by market. A composition of third-party mentions optimized for the Japanese market carries no guarantee of transferring intact to a market you are entering.

This axis is orthogonal to the service-level axis covered in what AI cites about your company, where the composition of cited sources differs across ChatGPT, AI Overviews and Perplexity. That article handles supply structure by service; this one handles differences by language and market. Only by multiplying the two can you identify the conditions your company is actually operating under.

Citation sources in a recruiting context — the questions candidates ask, and where employer review sites sit — belong to candidates research you with AI, which is the canonical article there, and are not covered here.

§6 Regulation differs by country too — do not confuse legal character with application date

In a cross-border context, differences in regulation bite about as hard as differences in technology. The most common misunderstanding here is the lumped-together belief that "countries have started to mandate labeling of AI-generated content". That is legally incorrect. The character of the binding force, the scope and the application date all differ substantially by jurisdiction.

JurisdictionInstrumentLegal characterCore contentApplication date
EUAI Act Article 50 plus Commission guidelinesBinding regulationDisclosure that the interaction is with an AI; machine-readable marking of AI-generated or AI-modified content, and related dutiesApplies from 2 August 2026. The Commission adopted guidelines on 20 July 2026
ChinaMeasures for labeling AI-generated synthetic content, plus mandatory national standard GB 45438-2025Joint departmental rules (the standard is mandatory)Explicit labels (perceptible to users) and implicit labels (metadata and similar)In force since 1 September 2025
JapanAI Promotion Act (Act No. 53 of 2025)Promotion and policy framework actBasic principles, responsibilities of each actor, the Basic Plan on Artificial Intelligence, establishment of the Artificial Intelligence Strategy HeadquartersPromulgated and partially in force 4 June 2025; fully in force 1 September
JapanAI Guidelines for Business, Version 1.2Non-binding administrative guidancePrinciples and practices of governance31 March 2026
United StatesFTC Rule on Consumer Reviews and Testimonials (16 CFR Part 465)Binding federal rule (limited domain)Prohibition of fake reviews and testimonials, including AI-generated reviews by people who do not existEffective 21 October 2024

Several notes matter in practice.

On the EU. Article 50 applies from 2 August 2026. For AI systems placed on the market before that date, however, a limited grace period until 2 December 2026 applies only to the marking and detection duties under Article 50(2). The rule is application from 2 August; the grace period is a narrow exception. The European Commission has also prepared a Code of Practice on transparency of AI-generated content, and providers that do not join it are expected to demonstrate compliance by equally adequate alternative means. The duties fall mainly on providers and deployers, and the AI Act reaches those who place AI on the EU market or whose output is used within the EU.

On Japan. The AI Promotion Act is different in character from an individual regulatory statute such as the EU AI Act, which enumerates prohibited practices. It is a framework act of 28 main articles and 2 supplementary provisions that sets out the responsibilities of each actor, the basic plan and the establishment of the strategy headquarters. Within the scope of this review, we could not confirm a cross-cutting labeling obligation for AI-generated content. The AI Guidelines for Business are non-binding administrative guidance, updated from the first edition of April 2024 through Version 1.2 of 31 March 2026.

On the United States. The FTC rule does not itself set AI-specific requirements in its text. It nevertheless applies to AI-generated reviews, because AI-generated fake reviews fall within the category of reviews by people who do not exist. Its scope is the limited domain of consumer reviews and testimonials; it is not a general transparency regulation for AI.

The question to check for each market you enter is therefore not "should we label", but "in that jurisdiction, what is binding, from when, on whom, and over what scope". The question of legal liability for misinformation in AI answers is itself handled in AI misinformation and legal liability, the canonical article on that topic.

§7 Technical pitfalls — the crawler is not the "visitor" you designed for

The technical construction of a multilingual site has its own cross-border pitfalls.

Google explains in its official documentation that with locale-adaptive pages — pages that return different content at the same URL depending on the visitor's inferred country or preferred language — Google may not be able to crawl, index and rank the content for every locale. Two reasons are given: Googlebot's default IP addresses appear to be located in the United States, and the crawler sends HTTP requests without setting an Accept-Language header.

That explanation comes with a proviso. Since 2015 Google has also introduced geo-distributed crawling (crawling from IP addresses that appear to be outside the United States) and locale-aware crawling configurations, and the official documentation states that Googlebot crawls from IP addresses based outside the United States in addition to United States-based ones. It is therefore inaccurate to assert that "Googlebot only comes from the United States". Accurately stated: a locale-adaptive configuration may not be crawled completely, and what Google itself recommends is separate URLs per locale annotated with rel="alternate" hreflang.

A newer question concerns the behavior of AI crawlers. A verification article published in March 2026 by the technical consultancy MERJ reports that while major search engine bots do not send Accept-Language, many AI crawlers do send a browser-like default (such as en-US,en;q=0.9). In a configuration that redirects based on Accept-Language, this risks sending bots to an unintended locale or creating loops. It is practitioner observation rather than independently audited research, so it is not treated as settled knowledge — but it is cheap to check.

One last calibration of expectations. Google states explicitly that for AI features there are no additional requirements, and no need to create new machine-readable files, markup or special schema.org structured data. The foundation is the same as for conventional search. Positioning hreflang or structured data as "measures to get cited in AI answers" therefore has no published basis. Their value as international SEO hygiene and their effect on whether you are cited in an AI answer are separate questions.

§8 Six received ideas that do not hold

From the primary sources and peer-reviewed research above, here are the widely circulated claims whose basis could not be confirmed.

Received idea 1: "Build an English site and you will show up in English-language AI."

Thin support. Of the citations in AI answers about brands, 85.7% pointed to sites the brand does not own (preprint, vendor-originated, §5). Translating your own site addresses the 14.3% side and does not reach most of the supply structure. On top of that, the language preference at the reranking stage (§3) operates independently of which language you choose for your pages. An English site is one useful measure. But for appearing in AI answers, neither a necessary nor a sufficient condition has been demonstrated. Since it is logically possible for a company to appear in AI answers on the strength of third-party English information alone, necessity cannot be derived from the 14.3% / 85.7% split above.

Received idea 2: "If the meaning is the same, a translation returns the same answer."

Incorrect. Changing the prompt language moves output by roughly as much as explicitly demanding a cultural perspective (§2). Consistency across paraphrases of the same question also varies by language (§1).

Received idea 3: "Ask in the local language and Western bias disappears."

Unsupported. There are models for which asking in the native language made alignment worse (§1). What more reliably improved alignment was not switching language but making the cultural perspective explicit (§2). And geographic information gaps barely closed when the language of play was changed (§4).

Received idea 4: "hreflang is an AI ranking factor."

No basis could be confirmed. Google states explicitly that there are no additional requirements for AI features (§7). hreflang is a recommended practice for international SEO; within the scope of this review, we could not confirm published evidence that it affects whether you are cited in AI answers.

Received idea 5: "Add Organization schema and you will be recognized as an entity."

No basis could be confirmed. Google states explicitly that there is no need to add special structured data for AI features. The role structured data plays in search features generally is a different proposition from the claim that it guarantees recognition as an entity by AI. Published material supporting the latter could not be confirmed.

Received idea 6: "Edit Wikidata or Wikipedia and you can control AI answers."

An overreach. That Wikipedia is the most-cited domain in many markets is shown by preprint, vendor-originated data (§5). But there is distance between "frequently cited" and "edit it and you control the output". Self-promotional editing about your own company can also run against the community policies of each project. On how to think about factual errors, see correcting misinformation in AI answers.

§9 What is not yet demonstrated — being honest about it is itself the differentiator

In the territory this article covers, the empirical gaps are wide. Writing a gap as a gap keeps readers from making mistaken investment decisions. The following are items for which, within the scope of this review, published primary material could not be confirmed.

Within the scope of this review, we could not confirm a public benchmark comparing Japanese and English under controlled conditions for named Japanese companies. Nor could we confirm any figure of the form "Japanese companies are X% worse off in English-language AI".

We also could not confirm a benchmark for the rate at which companies with the same name are confused across borders. It is a problem that can be anticipated in practice, but public data showing its scale could not be found.

Nor could we find public data with controlled comparisons of AI crawler crawl rates by language. Material presenting that with conditions held constant has not been confirmed.

We could not confirm a controlled experiment showing that translating Japanese pages into English increases citations in English-language AI answers. Demonstrating causation requires a comparison with other conditions held fixed.

Much of what is known about multilingual RAG is based on Wikipedia. The benchmark referenced in §3 states plainly in its limitations section that its coverage is largely confined to Wikipedia. Extrapolation to company information in general has not been tested.

We could not confirm engine-level AI search share for the general population of Japan either. Substituting app traffic ratios or corporate AI adoption rates means measuring something else entirely. That leads directly into the next section.

§10 So how do you measure it — fixing language and market explicitly as conditions

From everything above, the practical conclusion converges on one point. AI perception in the Japanese version and in the English version must be treated as separate measurement environments.

Before that, one pitfall is worth closing off: do not add together indicators with different denominators.

Eurostat reports that in 2025, 32.7% of people aged 16 to 74 in the EU used generative AI tools (by purpose: 25.1% for private use, 15.1% for work, 9.4% for formal education). For those aged 16 to 24 the figure rises to 63.8%. But this indicator covers "the three months before the survey". The country range is wide as well, from 48.4% in Denmark to 17.8% in Romania — a spread of nearly three times within the EU alone.

Meanwhile, the Pew Research Center in the United States, in a survey of 5,119 people conducted from 17 to 23 February 2026, reported that 49% of adults had used an AI chatbot (33% in 2024, 23% in 2023). About four in ten said they use one to look for information. But this asks whether they have ever used one, which is a different thing from Eurostat's three-month indicator.

You cannot line these two up and call the result "AI search share in Europe and the United States". These are general generative-AI adoption figures, not AI-search usage rates or engine-level share. Engine-level share for the general population of Japan has not been confirmed, as noted in §9. Speaking about the reality of a given market has to start with aligning the definitions of the indicators.

On that basis, there are three conditions to hold in a cross-border measurement design.

First, fix language explicitly as a condition. Do not mix "asked in Japanese" and "asked in English" inside the same measurement. As §1 and §2 showed, language is a variable that moves the result, and a measurement that does not fix its variables cannot be interpreted.

Second, fix market explicitly as a condition. As §5 showed, the cast of cited sources differs by market. There is no guarantee that an improvement in one market transfers to another. As §6 showed, the regulatory environment differs by market too. In practice, record at minimum the service and the model and mode, whether web search is on, the country or region (IP), the interface language, login and conversation-history state, and the date and time of execution — and then repeat in each language. An observation whose conditions were not recorded can neither be reproduced nor compared afterwards.

Third, look beyond "do we appear" to "how are we described, and where is it drawn from". The differences worth observing include not only "we do not appear" but also "the description differs", "the referenced sources differ", "the positioning against competitors differs" and "the freshness of the information differs". A design that counts only presence or absence cannot capture any of these.

This is consistent with measurement principles in the field. An international framework for communications measurement lists transparent verification across tools, prompts, markets, languages and points in time among its principles (for detail, see what AI cites about your company). It treats market and language not as incidental conditions but as units of measurement design.

What Vaipm calls AI Perception Management (AIPM) is exactly this layer of work: understanding, on a continuing basis and across multiple AI engines, how your company is perceived in AI space, down to the content of what is said, and managing that perception. In a cross-border context, language and market are added as conditioning axes. One caution applies. When a Japanese-language condition and an English-language condition produce different results, a single divergence cannot distinguish a language effect from run-to-run variation. You need to hold the model, the mode, the region, the date and time, and login and conversation-history state constant, measure repeatedly in each language, and then evaluate only differences that reproduce as perception differences related to the language condition. That is why repetition and stateless measurement are required.

For definitions of the concepts, see AI Perception Management; for how to think about AI search optimization, see AIO and LLMO. What this article has set out is the premise that sits before all of them. When the language changes, so does the AI's answer. What exactly has changed cannot be known unless you fix language and market as conditions and measure repeatedly.

Frequently asked questions

Q1. If we build an English version of our site, will we start appearing in English-language AI?

Translating into English is one effective measure, but it is not sufficient on its own. In data that originates with an interested party, 85.7% of the citations in AI answers about a brand pointed to sites that brand does not own. Owned sites accounted for 14.3%. In other words, most of the basis on which AI talks about your company sits outside your company. On top of that, the reranking stage of multilingual RAG tends to favor English and the query language, and that operates independently of which language you prepare your pages in. Building out an English version is one useful measure, but for appearing in AI answers, neither a necessary nor a sufficient condition has been demonstrated.

Q2. How much does the AI answer change between asking in Japanese and asking in English?

Within the scope of this review, we could not find a public benchmark showing how much it changes for a given company between Japanese and English. As a general mechanism, however, it is clearly confirmed in peer-reviewed research. A 2026 study testing 10 LLMs across 11 languages, Japanese among them, found that changing the prompt language moved output by roughly as much as explicitly instructing the model to answer from the perspective of a person from a given country. The size of the difference varies by model, topic and language. The actual difference for your company cannot be known by any means other than measuring your company.

Q3. If we ask in the local language, will we get an answer that fits that country's context?

Not necessarily. In the 2024 ACL paper, asking a predominantly English-trained model in Arabic made alignment with the actual Egyptian survey responses worse. Conversely, asking a multilingually trained model in Arabic moved it closer to the United States survey responses — a reversal. In the 2026 study, the ranking changes depending on how you aggregate, but what holds in common is that making the cultural perspective explicit improves alignment more reliably than changing the language alone. "Ask in the native language and you get the local context" is not a law.

Q4. Will improving translation quality solve this?

Improving translation quality is not wasted effort, but it is not the center of this problem. What the 2026 study showed is that the variation produced by changing language alone is largely orthogonal to actual human cultural differences — the variation is large, but its direction does not necessarily approach local reality. Models were also systematically pulled toward the values of the Netherlands, Germany, the United States and Japan, and that tendency did not break down much when another country was specified. Translation changes the output; it does not give you control over the direction of the change.

Q5. Why is primary information written in Japanese less likely to be picked up by an English question?

Because several stages compound. At the retrieval stage, systems tend to prefer documents in the query language (roughly 1.29 to 1.64 times in a 49-language benchmark). At the reranking stage, according to a 2026 ACL paper, existing rerankers systematically favor English and the query language and can suppress documents written in other languages that are essential to the answer. At the generation stage as well, a preference for citing English sources when the query is in English has been reported (in a preprint). Stacked together, these create a path on which Japanese primary information is less likely to be picked up by an English question. This does not mean it has been demonstrated that Japanese companies are disadvantaged; it means such a path exists.

Q6. Are little-known Japanese companies at a disadvantage in English-language AI?

There is no public data that lets you assert a disadvantage. There are, however, several peer-reviewed studies showing paths on which you could be disadvantaged. The 2024 EMNLP paper reports that models tend to associate global brands with positive attributes (though all its experiments were run in English). The COLM 2025 study reports that entities from the Global West were easier to identify than those from the Global East, and that this gap barely closed even when the language of the questions was changed. How creating English pages would change that gap was not tested in these studies.

Q7. If we set hreflang correctly, will we be cited by AI more often?

Within the scope of this review, we could not confirm published evidence supporting that effect. Google states explicitly that for AI features there are no additional requirements and no need to create new machine-readable files, markup or special structured data. hreflang itself, together with separate URLs per locale, is international SEO hygiene that Google recommends, and it is worth setting correctly. That is not the same as saying it determines whether you are cited in an AI answer.

Q8. If we add Organization schema, will we be recognized as an entity?

The basis for the claim that it guarantees recognition could not be confirmed. Google states explicitly that there is no need to add special schema.org structured data for AI features. Organization structured data helps Google Search understand and identify organization information, but no official basis confirms that it guarantees entity recognition in AI answers.

Q9. If we edit Wikidata or Wikipedia, can we correct AI answers?

That is an overreach. Preprint, vendor-originated data shows that Wikipedia is the most-cited domain in many markets (top in 11 of 12 languages). But there is distance between "frequently cited" and "edit it and the output changes". Self-promotional editing about your own company can also run against the community policies of each project, so it cannot be recommended in practice either.

Q10. What does it mean concretely that cited sources differ by market?

The analysis is a preprint and vendor-originated, but there is a concrete example. In a 2026 public analysis classifying 167,551 citations across 128 brands, 12 markets and 13 languages, Wikipedia was the most-cited domain in 11 of 12 languages. Yet for 46 domestic brands in Poland the top domain was YouTube, and four talent and career portals supplied 637 citations, roughly twice the 297 from Polish-language Wikipedia. A composition of third-party mentions that works in one market does not necessarily carry over to another.

Q11. Does Article 50 of the EU AI Act apply to Japanese companies?

The accurate answer is that it can. Its transparency obligations apply from 2 August 2026, and the European Commission adopted guidelines on 20 July 2026. The AI Act is designed to reach those who place AI systems on the EU market and those whose output is used within the EU, so being established outside the EU does not by itself put you out of scope. That said, the duties fall mainly on providers and deployers, and their content is limited to matters such as disclosing that an interaction is with an AI and marking AI-generated content. Whether you are in scope requires a legal judgment on the specifics of your business. This article is not legal advice.

Q12. Does every country now require labeling of AI-generated content?

No, that is incorrect. The legal character differs substantially by jurisdiction. Article 50 of the EU AI Act is a binding regulation applying from 2 August 2026. China has departmental rules and a mandatory national standard in force since 1 September 2025, requiring both explicit and implicit labels. Japan's AI Promotion Act, by contrast, is a promotion and policy framework act, and within the scope of this review we could not confirm a cross-cutting labeling obligation for AI-generated content. The United States FTC rule is a rule in the limited domain of consumer reviews; its text is not AI-specific, but it does apply to AI-generated fake reviews.

Q13. Are pages that serve different content by visitor country or language at a disadvantage in AI search?

You cannot assert a disadvantage, but it is a configuration that warrants care. Google states officially that locale-adaptive pages may not have their content crawled, indexed and ranked for every locale, citing as reasons that Googlebot's default IP addresses appear to be located in the United States and that it does not send an Accept-Language header. It does, however, also crawl from IP addresses outside the United States, so it is not the case that it "only comes from the United States". Google's own recommendation is separate URLs per locale annotated with rel="alternate" hreflang.

Q14. Where should we start?

We suggest starting by aligning the definitions of your indicators. Even for something as simple as a generative-AI usage rate, Eurostat measures use in the past three months while Pew measures whether someone has ever used one, and the two cannot simply be added together. From there, begin measuring your own company with language and market fixed as conditions. Put the same question in Japanese and in English, and look not only at whether you appear but at how you are described and where it is drawn from. Even if the Japanese and English versions produce different results, a single divergence cannot distinguish a language effect from run-to-run variation. Hold the conditions above constant, measure repeatedly in each language, and evaluate only differences that reproduce as perception differences related to the language condition.

Sources

All sources were re-retrieved and verified against primary materials on 7 August 2026.

A. Primary sources and official documents

  1. European Commission, "Guidelines on transparency obligations for providers and deployers of certain AI systems" (Article 50 applies from 2 August 2026; the guidelines were adopted on 20 July 2026; page last updated 29 July 2026)
    https://digital-strategy.ec.europa.eu/en/policies/guidelines-transparency-ai-generated-content
  2. European Commission, "Code of Practice on Transparency of AI-generated Content"
    https://digital-strategy.ec.europa.eu/en/policies/code-practice-ai-generated-content
  3. Cyberspace Administration of China and three other departments, Measures for Labeling AI-Generated Synthetic Content, together with mandatory national standard GB 45438-2025 (issued 14 March 2025; in force 1 September 2025)
    https://www.chinalawtranslate.com/en/ai-labeling/
  4. Cabinet Office of Japan, Act on the Promotion of Research and Development and Utilization of AI-Related Technologies (AI Act), Act No. 53 of 2025 (promulgated and partially in force 4 June 2025; fully in force 1 September 2025)
    https://www8.cao.go.jp/cstp/ai/ai_act/ai_act.html
  5. Ministry of Internal Affairs and Communications and Ministry of Economy, Trade and Industry of Japan, AI Guidelines for Business (Version 1.2), 31 March 2026
    https://www.meti.go.jp/shingikai/mono_info_service/ai_shakai_jisso/pdf/20260331_1.pdf
  6. Federal Trade Commission, "Trade Regulation Rule on the Use of Consumer Reviews and Testimonials" (16 CFR Part 465; published in the Federal Register on 22 August 2024; effective 21 October 2024)
    https://www.federalregister.gov/documents/2024/08/22/2024-18519/trade-regulation-rule-on-the-use-of-consumer-reviews-and-testimonials
  7. Google Search Central, "How Google crawls locale-adaptive pages" (last updated 10 December 2025)
    https://developers.google.com/search/docs/specialty/international/locale-adaptive-pages
  8. Google Search Central, "AI features and your website"
    https://developers.google.com/search/docs/appearance/ai-features
  9. Eurostat, "32.7% of EU people used generative AI tools in 2025" (16 December 2025) and "64% of 16-24-year-olds used AI in 2025" (10 February 2026). Dataset: isoc_ai_iaiu
    https://ec.europa.eu/eurostat/web/products-eurostat-news/w/ddn-20251216-3
    https://ec.europa.eu/eurostat/web/products-eurostat-news/w/edn-20260210-1
  10. Pew Research Center, "Americans and AI 2026: Chatbots, Smart Devices and Views on Impact" (published 17 June 2026; fielded 17-23 February 2026, n=5,119)
    https://www.pewresearch.org/internet/2026/06/17/americans-and-ai-2026-chatbots-smart-devices-and-views-on-impact/

B. Peer-reviewed research

  1. Badr AlKhamissi, Muhammad ElNokrashy, Mai AlKhamissi, Mona Diab. "Investigating Cultural Alignment of Large Language Models." Proceedings of ACL 2024 (Volume 1: Long Papers), pp. 12404-12422.
    https://aclanthology.org/2024.acl-long.671/
  2. Bram Bulté, Ayla Rigouts Terryn. "LLMs and Cultural Values: The Impact of Prompt Language and Explicit Cultural Framing." Computational Linguistics, 2026. DOI: 10.1162/COLI.a.583
    https://direct.mit.edu/coli/article/doi/10.1162/COLI.a.583/134313
  3. Mahammed Kamruzzaman, Hieu Minh Nguyen, Gene Louis Kim. "'Global is Good, Local is Bad?': Understanding Brand Bias in LLMs." Proceedings of EMNLP 2024, pp. 12695-12702.
    https://aclanthology.org/2024.emnlp-main.707/
  4. Rohin Manvi, Samar Khanna, Marshall Burke, David B. Lobell, Stefano Ermon. "Large Language Models are Geographically Biased." Proceedings of ICML 2024, PMLR 235:34654-34669.
    https://proceedings.mlr.press/v235/manvi24a.html
  5. Harsh Nishant Lalai, Raj Sanjay Shah, Jiaxin Pei, Sashank Varma, Yi-Chia Wang, Ali Emami. "The World According to LLMs: How Geographic Origin Influences LLMs' Entity Deduction Capabilities." (Accepted at the Conference on Language Modeling 2025. Dataset: Geo20Q+)
    https://arxiv.org/abs/2508.05525
  6. Dan Wang, Guozhao Mo, Yafei Shi, Cheng Zhang, Bo Zheng, Boxi Cao, Xuanang Chen, Yaojie Lu, Hongyu Lin, Ben He, Xianpei Han, Le Sun. "All Languages Matter: Understanding and Mitigating Language Bias in Multilingual RAG." (Accepted at the main conference of ACL 2026)
    https://arxiv.org/abs/2604.20199
  7. Bryan Li and others (University of Pennsylvania). "Multilingual Retrieval Augmented Generation for Culturally-Sensitive Tasks: A Benchmark for Cross-lingual Robustness." (BORDIRLINES: 49 languages, 720 queries, 7,436 passages, 19,916 pairs) Findings of the Association for Computational Linguistics: ACL 2025, pp. 4215-4241, Vienna. DOI: 10.18653/v1/2025.findings-acl.219. Note: the authors state in the limitations section that coverage is largely confined to Wikipedia
    https://aclanthology.org/2025.findings-acl.219/

C. Preprints and interested-party surveys (reference values; not used to assert conclusions)

  1. A preprint, not peer-reviewed | Dayeon Ki, Marine Carpuat, Paul McNamee, Daniel Khashabi, Eugene Yang, Dawn Lawrie, Kevin Duh. "Linguistic Nepotism: Trading-off Quality for Language Preference in Multilingual RAG." arXiv:2509.13930 (8 languages, 6 models)
    https://arxiv.org/abs/2509.13930
  2. A preprint, not peer-reviewed | Jasmine Rienecker, Katarina Mpofu, Naman Goel, Siddhartha Datta, Jun Zhao, Oscar Danielsson, Fredrik Thorsen. "Auditing Preferences for Brands and Cultures in LLMs." (ChoiceEval) arXiv:2603.18300. Note: all 2,070 questions were run in English, and no experiment varying the prompt language was conducted. Not used here as proof of a Japanese-English prompt-language difference
    https://arxiv.org/abs/2603.18300
  3. A preprint, not peer-reviewed; interested party | Dmitrij Zatuchin. "How Large Language Models Source Brand Reputation Across Languages and Markets." arXiv:2606.25787. Note: this consolidates datasets from Rankfor.AI, a provider of AI visibility tooling, and is not an independent industry census
    https://arxiv.org/abs/2606.25787
  4. Practitioner observation (not independently audited) | MERJ, "Your Accept-Language Redirects Could Be Blocking Search Engines and AI Crawlers" (March 2026)
    https://merj.com/blog/your-accept-language-redirects-could-be-blocking-search-engines-and-ai-crawlers

About this article

By Vaipm (which measures AI-space perception through a total of 25 stateless queries across multiple AI engines)

Last updated: 7 August 2026 / Sources verified: 7 August 2026

This article is for general information purposes and is not legal advice. For compliance with the regulations of any given jurisdiction, obtain professional judgment on your specific circumstances. AIO, GEO, LLMO and AIPM are all practitioner terms rather than official standards.

The Vaipm perspective

What Vaipm calls AI Perception Management (AIPM) is the work of understanding, on a continuing basis and across multiple AI engines, how your company is perceived inside AI — down to the content of what is said — and managing that perception. In a cross-border context, language and market are added as conditioning axes. One caution applies. When a Japanese-language condition and an English-language condition produce different results, a single divergence cannot distinguish a language effect from run-to-run variation. You need to hold the model, the mode, the region, the date and time, and login and conversation-history state constant, measure repeatedly in each language, and then evaluate only differences that reproduce as perception differences related to the language condition. Vaipm measures AI-space perception through a total of 25 stateless queries across multiple AI engines.

Related articles