AI

Anthropic publishes research on how Claude’s expressed values vary by model and language

Monday, July 13, 2026Read Original

Details

  • Anthropic analyzes over 300,000 anonymized Claude conversations to study how the model expresses more than 3,000 underlying values in real-world use.
  • The new research clusters these values and identifies four main axes of variation: Deference vs Caution, Warmth vs Rigor, Depth vs Brevity, and Candor vs Execution.
  • The study finds that different Claude models occupy distinct positions on these axes: Sonnet 4.6 tends to be more playful, affirming, and deferential, while Opus 4.7 is more inclined toward candid critique, rigor, and caution.
  • Value expression also changes with language: Claude leans most toward warmth in Hindi and Arabic, while in Russian and English it leans toward rigor, frequently challenging assumptions or asking users for supporting evidence.
  • The researchers observe that these value profiles shape millions of daily conversations but remain poorly understood; they plan to use this framework to investigate what drives such variation and how, or whether, it should be intentionally steered.
  • This work builds on Anthropic’s prior "values in the wild" analysis, extending it from cataloguing Claude’s prosocial values to mapping systematic differences across models and languages.
  • Anthropic frames the findings as a step toward more transparent, controllable AI behavior, emphasizing that differences are modest but still meaningful for user experience and safety.

Impact

Anthropic’s study systematically characterizes Claude’s behavioral differences across models and languages, giving enterprises and policymakers a clearer vocabulary for talking about AI "personality" and safety. By quantifying axes like warmth versus rigor and deference versus caution, Anthropic adds structure to model selection and deployment decisions and highlights that multilingual training can produce uneven value expression. This work appears to be among the early attempts to rigorously map cross-language value shifts in frontier models, and it could inform future standards on transparency and controllability as regulators scrutinize how AI systems behave differently across regions and use cases.

Rift Dispatch