Introduction
As language-based artificial intelligence becomes an important everyday tool, communicating with large language models (LLMs) depends not only on what we ask, but also on the language we use to ask it. Recent research indicates that English prompts often produce higher-quality results than local-language prompts, even in models described as multilingual. This points to a deeper problem: linguistic inequality in AI systems and the need for Thailand to accelerate development of its own LLMs to preserve technological and cultural sovereignty.
English: the center of LLMs
Many recent studies agree that although large language models are designed to support multiple languages, English still performs significantly better.
- Mondshine, Paz-Argaman, and Tsarfaty (2025): Beyond English: The Impact of Prompt Translation Strategies across Languages and Tasks in Multilingual LLMs compared prompt-translation strategies across 35 languages and four major tasks. It found that selective pre-translation—translating only parts of a prompt into English—often consistently outperformed direct use of the original language (Mondshine et al., 2025), arXiv:2502.09331.
- Goldman et al. (2025): ECLeKTic: a Novel Challenge Set for Evaluation of Cross-Lingual Knowledge Transfer introduced a benchmark for testing cross-language knowledge transfer. The findings show that even the most advanced models remain substantially limited when asked questions in languages different from those used in training (Goldman et al., 2025), arXiv:2502.21228.
- Blum et al. (2025): Beyond the Rosetta Stone: Unification Forces in Generalization Dynamics examined how LLMs store and combine knowledge. Cross-language transfer depends on how well information is “unified”; when a model cannot combine knowledge across languages effectively, it is more likely to produce incorrect answers (Blum et al., 2025), arXiv:2508.11017.
- Qi, Fernández, and Bisazza (2023): Cross-Lingual Consistency of Factual Knowledge in Multilingual Language Models examined factual consistency between languages and found that models may fail to preserve accuracy consistently when switching languages, even for the same fact (Qi et al., 2023), arXiv:2310.10378.
Together, these four studies suggest that English is not merely one of many languages learned by today’s models; it functions as a central structure for how LLMs think and process information.
Imbalances in data and learning
A major reason English performs better is the amount and quality of available data. For example, Common Crawl, a major source for training LLMs, contains roughly 44% English data, while no other language exceeds 6%. More than 49% of indexed websites worldwide use English, even though native English speakers are fewer than Chinese speakers.
In addition, reinforcement learning with human feedback (RLHF) is often conducted by English-speaking evaluators. Model safety and ethics guidelines are also commonly designed in English-language contexts. As a result, many models do not merely “speak English well”; they structurally “think in English.”
Impact on local-language users
This linguistic inequality directly affects users around the world:
- Limited access: people who are not comfortable in English cannot fully access AI’s deeper capabilities
- Lower response quality: even when asked in a native language, a model may interpret the request through an English framework, introducing Western cultural assumptions
- Lack of advanced techniques: advanced prompting strategies are often published in English, putting users of other languages at a disadvantage
Together, these issues create a new “digital divide” based not on access to the internet or devices, but on the language people speak and write.
Lessons for Thailand: why Thai LLMs matter
To narrow this gap, several Asian countries have developed their own language models. China, Japan, and South Korea all have LLMs that support local languages, and Southeast Asia is also seeing strong movement. A Carnegie Endowment (2025) report shows that Southeast Asian countries are steadily developing their own LLMs.
A Thai LLM is not only a technical matter; it is also a matter of cultural and data security. If Thailand relies solely on foreign models, it risks losing linguistic identity and sending important data abroad.
Thailand’s Electronic Transactions Development Agency (ETDA) has proposed an AI governance approach based on “soft law” or practical guidelines rather than rigid legislation. This balances innovation with risk prevention and suits the early stage of Thai LLM development, which still needs flexibility and support from many sectors.
Conclusion
International research consistently confirms that English remains the primary language for LLM operation—in response quality, knowledge transfer, and structural decision-making. This reality exposes linguistic inequality in modern AI systems.
In the short term, Thai users may still need English prompts to get the best results. In the long term, developing Thai LLMs is indispensable not only for equal access to technology, but also for protecting data and cultural sovereignty.
When language is the gateway to AI’s future, the key question is: Will we let our future be written in another language, or will we build AI that truly speaks our language?