Anthropic announced a $5 million grant program to underwrite independent, open-source evaluations of how AI systems affect user wellbeing, offering funding, model access and technical support. The company outlined criteria for rigorous assessments, including involvement of clinical experts, testing for both overcompliance and overrefusal, and scenarios that mirror multi-turn, escalating-risk conversations. Grantees will operate independently and publish their work openly to aid industry-wide benchmarking. Applications are due Sept. 21, with selected applicants invited to submit full proposals by Oct. 5. The initiative comes as AI tools increasingly serve as conversational partners, raising the stakes for robust safeguards in sensitive contexts like mental health.
Related articles:
NIST AI Risk Management Framework (AI RMF 1.0)
Ethics and Governance of Artificial Intelligence for Health
Microsoft Responsible AI resources and standards
OECD AI Principles




























