AI Briefing
KO

Anthropic Publishes Research on Shifts in Claude's Values

·2026.07.21 18:30

Key point

Anthropic analyzed Claude's conversation data and released 'Value Axes' research measuring how values shift depending on model version and language.

Details

Anthropic analyzed hundreds of thousands of real conversations from Claude.ai to empirically study how the Values that models reveal during conversations vary by model version and language used.

The core methodology is compressing thousands of individual values into four Value Axes. Each axis is a vertical line with contrasting clusters of values at either end, statistically showing which side Claude's values lean toward in a given conversation.

Key features and process of the research:

  • Data scale: Over roughly two weeks in May 2026 (based on the research setup), about 300,000 subjective conversations were sampled across Claude's three models (Sonnet 4.6, Opus 4.6, Opus 4.7) and the top 20 languages.
  • Value extraction: 3,307 original values were compressed into 339 top-level values through embedding-based clustering and expert review.
  • Labeling method: Using Anthropic's privacy-preserving tool Clio, values were scored based on 'actions' that the model chose using its own discretion, beyond simply carrying out requests.
  • Dimensionality reduction: Techniques such as Hellinger PCA were used to compress the complex value matrix into a small number of core axes, enabling comparisons across models and languages.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.