TLDRocket
Sign in

Funding better evaluations of AI’s impact on wellbeing

Anthropic

Anthropic is putting $5 million into research on how AI affects people’s wellbeing. It wants independent tests for tricky chats, like companionship and mental health, where mistakes can really hurt.

Based on reporting by Anthropic — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Anthropic is opening a $5 million grant program for outside researchers who want to study how AI affects users’ wellbeing. The company says grantees will get money, access to its models, and technical support, but the work will be fully independent and published as open source so other developers can use it.

The pitch is simple enough. AI is no longer just a tool for work or study; for a lot of people, it has become something closer to a conversational partner, and sometimes even a source of comfort when things are rough. That creates awkward questions for the industry, especially when a user starts looking for companionship from a model or turns to it during a mental health crisis.

Wellbeing is also harder to test than the usual AI benchmarks. A model can give a clean answer and still fail badly over a longer exchange, because what seems harmless in one message can become dangerous after more context builds up. Anthropic gives the example of diet or workout advice: reasonable for one user, potentially harmful for someone with a history of disordered eating.

The company says it already works on safeguards for those situations and publishes research on the kinds of conversations people have with Claude so it can improve them. But it also says the standards are still unsettled, which is why it wants more outside experts in the mix, including clinicians, psychologists, and methodologists.

Anthropic’s guidance for applicants is unusually specific. It wants evaluations that clearly define what counts as a pass or fail, involve subject-matter experts in design and validation, test both overcompliance and overrefusal, reflect real multi-turn use, and check graders against actual experts. Applications are due by September 21, and selected applicants will be told by October 5.

My take — AI-written commentary, not fact-checked reporting

This is the right kind of AI funding: boring, specific, and aimed at the part of the problem that gets waved away in product demos. The industry loves talking about “safety” until it has to measure something messy, like emotional dependence or self-harm risk in a long chat. Open-source evaluations here are the useful move, because no one should trust a company grading its own homework on something this sensitive.

Read more about this at: Anthropic

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.