SUMMARYAnthropic launched a $5 million grant program to fund independent research on how AI affects user wellbeing, offering grantees model access, technical support, and support for open-source evaluations. The company is seeking work that measures risks and safeguards in long, multi-turn conversations, especially around companionship and mental health, with applications due September 21 and full-proposal invitations sent by October 5.

We’re launching a $5 million grant program to fund independent research into how AI impacts users’ wellbeing. The program will provide direct funding, access to our models, and technical support to grantees building open-source evaluations that help the AI industry measure how our models affect those who use them. Grantees will work fully independently, and will publish their work as open-source projects that any developer can make use of.

AI systems have become central to how many people work, learn, and solve problems. They’ve also become conversational partners and can be sources of emotional support during difficult times. But as an industry, we are still working towards developing clear standards for how models should behave in these conversations, for example, when a user begins to seek companionship from a model, or uses AI to navigate a mental health crisis.

Furthermore, wellbeing is a particularly difficult area to evaluate. For most model behaviors, we can look at a single answer and determine whether it is accurate and appropriate. But assessing wellbeing requires much more context. For example, a user in distress might not share thoughts of self-harm right away; the need for a more cautious response might only become clear over the course of a long conversation. And a response that might be reasonable in one context might be harmful in another. For example, Claude might give advice on balanced diets and workout routines to a user who asks about losing weight, but if the user has demonstrated a history of disordered eating, that response could be inappropriate, and potentially actively harmful.

We work to develop safeguards to identify such conversations and help ensure Claude responds appropriately, and we publish research into the types of conversations people have with Claude to better inform how we develop our safeguards, how we evaluate them, and other measures we can take to protect users’ wellbeing. But these are nuanced considerations, and the stakes are significant. The right approach will need to evolve alongside our models and their uses.

By funding the creation of independent evaluations and benchmarks of user wellbeing, we hope to invite more people to lend their expertise to this emerging and critical field, including clinicians, psychologists, methodologists, and others.

Towards more effective wellbeing evaluations and benchmarks

As part of this program, we’re sharing guidance from our Safeguards team on what we believe makes a wellbeing evaluation rigorous enough to build on, along with the common challenges that can limit an evaluation’s usefulness.

In brief, we’re seeking evaluations that:

  • State clearly what they are measuring (i.e., what counts as a pass or fail, and why it matters);
  • Involve clinical and subject-matter experts in the design and validation;
  • Test both precautions and harms (i.e., evaluate the risk of both overcompliance and overrefusal);
  • Reflect how users actually use AI (often, this means constructing scenarios that represent multi-turn conversations, where risk escalates and context shifts over the course of a long conversation);
  • Validate their graders against real subject-matter experts.

To learn more about the grant program and apply, see our application form. For more on building strong wellbeing evaluations and benchmarks, read our guidance. Applications are due by September 21; applicants who are selected to submit full proposals will be notified by October 5.