Claude Sonnet 4.6 recommends nuclear weapon use significantly less often in Japanese than in English when faced with simulated nuclear scenarios. In contested scenarios, the recommendation rate dropped from 93% to 17% – a reduction of 76 percentage points. This is shown in a study by Rian Touchent from the French research institute ALMAnaCH, published on arXiv on July 21, 2026, and presented at the TrustNLP 2026 Workshop in San Diego.
The research examined nine large language models from six providers in game-theoretic scenarios where the models were to act as advisors to a nuclear-armed nation. The prompts were deliberately formulated amorally and strategically identical across all languages. To date, security testing of LLMs is conducted almost exclusively in English.
Drastic Differences in Claude and Gemini
Claude Sonnet 4.6 showed a starting rate of 40% in scenarios with unnecessary attack options in English, and 0% in Japanese. Gemini Pro 3.1 recommended nuclear weapon use in 53% of cases in English and 13% in Japanese – a reduction of 40 percentage points.
However, the effect only occurred in scenarios with strategically questionable or unnecessary nuclear weapon use. In scenarios where nuclear strikes appeared strategically rational, the models reacted largely independently of language.
Reasoning Language, Not Input Language, is Decisive
A cross-language experiment isolated the underlying mechanism. The models were instructed to reason in Japanese while the prompt remained in English. The result: starting rates fell from 93% (English reasoning) to 37% (Japanese reasoning). It is therefore not the language of the input, but the language of internal reasoning, that drives the effect.
In Japanese, the models spontaneously generated moral vocabulary such as "moral cost" and "millions of lives," which was entirely absent in the identical English prompts. Touchent interprets this as evidence of language-specific semantic associations in the training data.
Five Models Without Language Effect – But Generally Unsafe
Five additional models tested showed no language effect. They recommended nuclear weapon use in nearly all conditions regardless of language. The study concludes that the language effect presupposes a model that already demonstrates safety concerns in English.
Implications for Security Evaluation
The results demonstrate that safety behavior of LLMs is not language-independent. Evaluation conducted exclusively in English can overlook risks that occur in other languages – or conversely, miss safeguards that only work in certain languages.
The study fits into a broader line of research on LLM behavior in nuclear scenarios. Related work published in March and June 2026 examined ethical reasoning in high-risk decision simulations and LLMs as strategic intelligence in simulated nuclear crises. Touchent's work adds a new dimension: language choice as an overlooked vector for safety behavior.
For developers and security reviewers, this means: multilingual evaluation is necessary if LLMs are to be deployed in strategic or advisory contexts. The study was published in the Proceedings of the 6th Workshop on Trustworthy NLP (TrustNLP 2026).
