Anthropic has launched a $5 million grant program to support independent research into how AI systems affect users’ wellbeing. The initiative will provide direct funding, access to models, and technical support to grantees developing open-source evaluations that help the AI industry assess the impact of models on users. Grantees will work independently and publish their findings as open-source projects accessible to all developers. Source: anthropic

The program aims to create rigorous benchmarks for evaluating wellbeing by addressing the complexities of assessing emotional and mental health outcomes. Unlike straightforward accuracy checks, wellbeing evaluation requires nuanced context, such as understanding the progression of a conversation involving a user in distress or recognizing when a response might be harmful despite appearing reasonable in a different context. For instance, advice on diet and exercise could be inappropriate for someone with a history of disordered eating. Source: anthropic

Anthropic emphasized the importance of involving clinical and subject-matter experts in designing evaluations and validating their effectiveness. The company also highlighted the need for evaluations to reflect real-world AI usage, including multi-turn conversations where context and risk evolve over time. Additionally, grantees must ensure their grading systems align with expert assessments to maintain accuracy and reliability. Source: anthropic

Source: anthropic