Q1. If we build an English version of our site, will we start appearing in English-language AI?
Translating into English is one effective measure, but it is not sufficient on its own. In data that originates with an interested party, 85.7% of the citations in AI answers about a brand pointed to sites that brand does not own. Owned sites accounted for 14.3%. In other words, most of the basis on which AI talks about your company sits outside your company. On top of that, the reranking stage of multilingual RAG tends to favor English and the query language, and that operates independently of which language you prepare your pages in. Building out an English version is one useful measure, but for appearing in AI answers, neither a necessary nor a sufficient condition has been demonstrated.
Q2. How much does the AI answer change between asking in Japanese and asking in English?
Within the scope of this review, we could not find a public benchmark showing how much it changes for a given company between Japanese and English. As a general mechanism, however, it is clearly confirmed in peer-reviewed research. A 2026 study testing 10 LLMs across 11 languages, Japanese among them, found that changing the prompt language moved output by roughly as much as explicitly instructing the model to answer from the perspective of a person from a given country. The size of the difference varies by model, topic and language. The actual difference for your company cannot be known by any means other than measuring your company.
Q3. If we ask in the local language, will we get an answer that fits that country's context?
Not necessarily. In the 2024 ACL paper, asking a predominantly English-trained model in Arabic made alignment with the actual Egyptian survey responses worse. Conversely, asking a multilingually trained model in Arabic moved it closer to the United States survey responses — a reversal. In the 2026 study, the ranking changes depending on how you aggregate, but what holds in common is that making the cultural perspective explicit improves alignment more reliably than changing the language alone. "Ask in the native language and you get the local context" is not a law.
Q4. Will improving translation quality solve this?
Improving translation quality is not wasted effort, but it is not the center of this problem. What the 2026 study showed is that the variation produced by changing language alone is largely orthogonal to actual human cultural differences — the variation is large, but its direction does not necessarily approach local reality. Models were also systematically pulled toward the values of the Netherlands, Germany, the United States and Japan, and that tendency did not break down much when another country was specified. Translation changes the output; it does not give you control over the direction of the change.
Q5. Why is primary information written in Japanese less likely to be picked up by an English question?
Because several stages compound. At the retrieval stage, systems tend to prefer documents in the query language (roughly 1.29 to 1.64 times in a 49-language benchmark). At the reranking stage, according to a 2026 ACL paper, existing rerankers systematically favor English and the query language and can suppress documents written in other languages that are essential to the answer. At the generation stage as well, a preference for citing English sources when the query is in English has been reported (in a preprint). Stacked together, these create a path on which Japanese primary information is less likely to be picked up by an English question. This does not mean it has been demonstrated that Japanese companies are disadvantaged; it means such a path exists.
Q6. Are little-known Japanese companies at a disadvantage in English-language AI?
There is no public data that lets you assert a disadvantage. There are, however, several peer-reviewed studies showing paths on which you could be disadvantaged. The 2024 EMNLP paper reports that models tend to associate global brands with positive attributes (though all its experiments were run in English). The COLM 2025 study reports that entities from the Global West were easier to identify than those from the Global East, and that this gap barely closed even when the language of the questions was changed. How creating English pages would change that gap was not tested in these studies.
Q7. If we set hreflang correctly, will we be cited by AI more often?
Within the scope of this review, we could not confirm published evidence supporting that effect. Google states explicitly that for AI features there are no additional requirements and no need to create new machine-readable files, markup or special structured data. hreflang itself, together with separate URLs per locale, is international SEO hygiene that Google recommends, and it is worth setting correctly. That is not the same as saying it determines whether you are cited in an AI answer.
Q8. If we add Organization schema, will we be recognized as an entity?
The basis for the claim that it guarantees recognition could not be confirmed. Google states explicitly that there is no need to add special schema.org structured data for AI features. Organization structured data helps Google Search understand and identify organization information, but no official basis confirms that it guarantees entity recognition in AI answers.
Q9. If we edit Wikidata or Wikipedia, can we correct AI answers?
That is an overreach. Preprint, vendor-originated data shows that Wikipedia is the most-cited domain in many markets (top in 11 of 12 languages). But there is distance between "frequently cited" and "edit it and the output changes". Self-promotional editing about your own company can also run against the community policies of each project, so it cannot be recommended in practice either.
Q10. What does it mean concretely that cited sources differ by market?
The analysis is a preprint and vendor-originated, but there is a concrete example. In a 2026 public analysis classifying 167,551 citations across 128 brands, 12 markets and 13 languages, Wikipedia was the most-cited domain in 11 of 12 languages. Yet for 46 domestic brands in Poland the top domain was YouTube, and four talent and career portals supplied 637 citations, roughly twice the 297 from Polish-language Wikipedia. A composition of third-party mentions that works in one market does not necessarily carry over to another.
Q11. Does Article 50 of the EU AI Act apply to Japanese companies?
The accurate answer is that it can. Its transparency obligations apply from 2 August 2026, and the European Commission adopted guidelines on 20 July 2026. The AI Act is designed to reach those who place AI systems on the EU market and those whose output is used within the EU, so being established outside the EU does not by itself put you out of scope. That said, the duties fall mainly on providers and deployers, and their content is limited to matters such as disclosing that an interaction is with an AI and marking AI-generated content. Whether you are in scope requires a legal judgment on the specifics of your business. This article is not legal advice.
Q12. Does every country now require labeling of AI-generated content?
No, that is incorrect. The legal character differs substantially by jurisdiction. Article 50 of the EU AI Act is a binding regulation applying from 2 August 2026. China has departmental rules and a mandatory national standard in force since 1 September 2025, requiring both explicit and implicit labels. Japan's AI Promotion Act, by contrast, is a promotion and policy framework act, and within the scope of this review we could not confirm a cross-cutting labeling obligation for AI-generated content. The United States FTC rule is a rule in the limited domain of consumer reviews; its text is not AI-specific, but it does apply to AI-generated fake reviews.
Q13. Are pages that serve different content by visitor country or language at a disadvantage in AI search?
You cannot assert a disadvantage, but it is a configuration that warrants care. Google states officially that locale-adaptive pages may not have their content crawled, indexed and ranked for every locale, citing as reasons that Googlebot's default IP addresses appear to be located in the United States and that it does not send an Accept-Language header. It does, however, also crawl from IP addresses outside the United States, so it is not the case that it "only comes from the United States". Google's own recommendation is separate URLs per locale annotated with rel="alternate" hreflang.
Q14. Where should we start?
We suggest starting by aligning the definitions of your indicators. Even for something as simple as a generative-AI usage rate, Eurostat measures use in the past three months while Pew measures whether someone has ever used one, and the two cannot simply be added together. From there, begin measuring your own company with language and market fixed as conditions. Put the same question in Japanese and in English, and look not only at whether you appear but at how you are described and where it is drawn from. Even if the Japanese and English versions produce different results, a single divergence cannot distinguish a language effect from run-to-run variation. Hold the conditions above constant, measure repeatedly in each language, and evaluate only differences that reproduce as perception differences related to the language condition.