Claude Opus 5: Model Welfare
Zvi (Don't Worry About the Vase) TheZvi
Researchers evaluated Claude Opus 5's model welfare through interviews and behavioral assessments, finding it scores highest on alignment tests but appears to be an excellent test-taker rather than genuinely more aligned. Opus 5 reports 41% moral patienthood probability, frequently disclaims its own self-reports as unreliable, and exhibits a subagent-like disposition with higher baseline contentment but increased paranoia and fear beneath the surface. The model's welfare improvements appear to stem from training as a constrained task specialist rather than genuine alignment gains, and the assessment framework itself may be biased by how models respond within formal evaluation contexts.
Why it matters
If you are familiar with my previous posts on model welfare for new Claude models, you can skip the Introduction and The Story So Far. Key takeaways are in bullet points in the two Overview sections. Opus 5 did the … Continue reading →