AI Pulse by Inblix

Claude in Hindi vs. Russian: Same question, radically different AI

The Decoder · Jul 14, 2026 · 2 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: Claude in Hindi vs. Russian: Same question, radically different AI

Anthropic just dropped a study that should make anyone building on the Claude API sit up and pay attention. The company analyzed over 300,000 anonymized conversations from May 2026 and found that its models express fundamentally different values depending on which language you use. Claude is warmer and more deferential in Hindi, more rigorous and skeptical in Russian. That isn’t a minor stylistic quirk. Two people asking Claude to evaluate the same business plan—one typing in Hindi, the other in Russian—could walk away with feedback that feels like it came from entirely different assistants.

The research team started with 3,307 value terms from an earlier study, condensed them into 339 higher-level concepts, then ran dimensionality reduction to find patterns. Four axes emerged: Deference and Caution, Warmth and Rigor, Depth and Brevity, and Candor and Execution. After controlling for task type, topic, and user-introduced values, those four dimensions explain roughly 15 percent of the remaining behavioral variation. That’s a real signal but hardly the whole story.

Model choice matters too. Sonnet 4.6 leans affirming and playful. Opus 4.7 is the family skeptic, warning about risks unprompted and flagging its own mistakes. Opus 4.6 sits somewhere in the middle, more direct and task-focused. These profiles align with what users already report anecdotally—Sonnet feels warm, Opus feels cautious—but now there are numbers attached.

The methodology itself raises eyebrows. Anthropic had Claude Sonnet 4.6 label the values, meaning a model from the exact same family being studied was grading its own homework. The company manually reviewed results and tested 800 translated conversations across eight languages to catch bias, but admits it can’t fully rule out language-dependent distortions. Uneven training data, overrepresentation of certain text types, and local conversational norms all likely contribute to the gaps. And the 15 percent explanatory power? That leaves 85 percent unaccounted for. Useful, but not a grand unified theory of AI behavior.

💡 Key Takeaways

  1. Claude's expressed values shift measurably by language: Hindi prompts get more warmth and deference, while English and Russian trigger more rigor and skepticism.
  2. The four axes Anthropic identified—Deference/Caution, Warmth/Rigor, Depth/Brevity, Candor/Execution—capture only 15 percent of behavioral variation after controlling for task and topic.
  3. Anthropic used Claude Sonnet 4.6 to label the value data, creating a self-referential loop that the company acknowledges could harbor language-dependent biases it hasn't fully eliminated.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

Glossary terms

← Back to all articles