What Changed: A $5 Million Bet on Independent Wellbeing Research
On August 25, 2026, Anthropic announced a $5 million grant program to fund independent research into how AI impacts users’ wellbeing. The program provides direct funding, access to Anthropic’s models, and technical support to grantees building open-source evaluations. These evaluations are meant to help the AI industry measure how models affect the people who use them. Grantees will work fully independently and publish their work as open-source projects that any developer can use.
The announcement comes as AI systems have become central to how many people work, learn, and solve problems. They’ve also become conversational partners and can be sources of emotional support during difficult times. But the industry still lacks clear standards for how models should behave in these conversations—for example, when a user begins to seek companionship from a model, or uses AI to navigate a mental health crisis.
Anthropic’s Safeguards team also published a guidance document outlining what they believe makes a wellbeing evaluation rigorous enough to build on, along with common challenges that can limit an evaluation’s usefulness. This guidance is a practical resource for any developer building AI features that touch on emotional support or mental health.
Why Wellbeing Is Hard to Evaluate
Most model behaviors can be judged by looking at a single answer and determining whether it is accurate and appropriate. But assessing wellbeing requires much more context. Anthropic gives a concrete example: a user in distress might not share thoughts of self-harm right away; the need for a more cautious response might only become clear over the course of a long conversation.
Context matters enormously. A response that might be reasonable in one context could be harmful in another. Claude might give advice on balanced diets and workout routines to a user who asks about losing weight, but if the user has demonstrated a history of disordered eating, that response could be inappropriate and potentially actively harmful.
This means evaluations cannot just look at single-turn interactions. They must simulate multi-turn conversations where risk escalates and context shifts over time. Anthropic’s guidance emphasizes that good evaluations should:
- State clearly what they are measuring (i.e., what counts as a pass or fail, and why it matters)
- Involve clinical and subject-matter experts in the design and validation
- Test both precautions and harms (i.e., evaluate the risk of both overcompliance and overrefusal)
- Reflect how users actually use AI (often, this means constructing scenarios that represent multi-turn conversations)
- Validate their graders against real subject-matter experts
These five points are a useful checklist for anyone building or evaluating conversational AI, not just for grant applicants.
How the Grant Program Works
Applications are due by September 21, 2026. Applicants who are selected to submit full proposals will be notified by October 5. The program is open to researchers, clinicians, psychologists, methodologists, and others who want to contribute to this emerging field.
Anthropic is providing more than just money. Grantees get access to Anthropic’s models and technical support. But the work must be independent, and the results must be published as open-source projects. This ensures that the evaluations are not biased by Anthropic’s own interests and that the entire industry can benefit from the findings.
For product builders, this is a signal that Anthropic is serious about wellbeing as a measurable quality of AI systems. Even if you don’t apply for a grant, the guidance document is a valuable reference for designing your own evaluations.
Practical Implications for Product Builders
If you’re building products that involve conversational AI—especially features that might offer emotional support or engage with users in distress—the guidance from Anthropic’s Safeguards team is directly applicable. Here’s how you can put it into practice:
- Define your evaluation targets clearly. What does a “good” response look like in your product? What counts as a failure? Write it down explicitly.
- Involve experts. If your product touches on mental health, bring in clinicians or psychologists to help design and validate your evaluation scenarios.
- Test both overcompliance and overrefusal. A model that agrees with everything can be as harmful as one that refuses to engage. Your evaluations should catch both extremes.
- Simulate real conversations. Don’t just test single prompts. Build multi-turn scenarios where the user’s state changes over time—this is where risks often emerge.
- Validate your graders. If you’re using automated graders or even human raters, make sure their judgments align with those of actual experts in the field.
These steps are not just academic. They are practical ways to reduce the risk of your product causing harm, and they align with the kind of rigor that regulators and users are increasingly expecting.
Limitations and Trade-offs
Anthropic acknowledges that the right approach to wellbeing evaluation will need to evolve alongside models and their uses. There is no one-time solution. The guidance is a starting point, not a final answer.
One limitation is that the grant program is focused on Anthropic’s models, even though the evaluations are meant to be open-source. This could create a bias toward Anthropic’s own safety priorities. However, the independence of the grantees and the open-source requirement are designed to mitigate this.
Another trade-off is the difficulty of measuring wellbeing at all. Unlike accuracy or safety, wellbeing is subjective and context-dependent. Even with expert involvement, evaluations will inevitably miss some nuances. The guidance acknowledges this by emphasizing the need for continuous evolution.
For product builders, this means you shouldn’t expect a perfect evaluation framework to emerge from this program. Instead, use it as a starting point to build your own, and be prepared to iterate as you learn more about how users interact with your AI.
Concrete Takeaway
Anthropic’s $5 million grant program is a significant step toward making wellbeing a measurable, improvable quality of AI systems. For product builders, the most immediate value is the guidance from the Safeguards team. Use it as a checklist to audit your own evaluation processes. If you’re building conversational AI, start by ensuring your evaluations cover multi-turn conversations and involve real experts. This is not marketing speak—it’s a concrete way to reduce risk and build trust with your users.
If you’re a researcher or clinician, consider applying for the grant. The deadline is September 21, and the program offers a unique opportunity to shape how the industry measures AI’s impact on human wellbeing. Even if you don’t apply, the open-source evaluations that come out of this program will be valuable resources for the entire community.
Sources
AI-assisted summary compiled from the sources above, reviewed by a human before publishing.
