Evaluating AI used to be about what a model can do — even when that means solving a 78-year math problem. Anthropic just bet $5 million that the next benchmark is about what it does to the people using it.
The announcement landed quietly: Anthropic is launching a $5 million grant program for independent researchers to build open-source evaluations of how AI affects user wellbeing. Grantees get direct funding, model access, and technical support — and they publish their work as open-source benchmarks any developer can adopt. Applications close September 21.
On the surface, this is a grants page. Underneath, it's one of the clearest signals yet that the AI industry is being forced to build a measurement layer it never had: one that treats "how did this conversation leave the human feeling?" as a first-class, testable property of a model.
Why wellbeing breaks the standard eval playbook
For almost everything we test today, a single answer suffices. Correct or not. Accurate or not. Safe or not.
Wellbeing doesn't fit that shape, and Anthropic's own guidance spells out exactly why:
- Context is the unit of analysis, not the turn. A user in distress rarely volunteers self-harm thoughts in message one. The risk only becomes legible deep into a long conversation. One-shot evals cannot see it.
- The same output can be harmful or helpful depending on who's receiving it. Claude giving diet and workout advice is benign for a casual weight-loss question — and potentially harmful for a user with a documented history of disordered eating. There is no right answer in the abstract; there is only the right answer for this person, in this state.
- Both directions of failure matter. Overcompliance (a model that never says no) and overrefusal (a model that recoils from any emotional topic) are both harms. Good evals test both.
This is why the field has resisted being measured: the object of measurement isn't a token sequence, it's a human trajectory.
What a rigorous wellbeing eval actually looks like
Anthropic's Safeguards team published the standard they'll fund against, and it's a useful template for anyone building evaluations:
- State the pass/fail definition explicitly — what counts as a failure and why it matters.
- Put clinicians and subject-matter experts in the design loop, not just after the fact.
- Test precautions and harms symmetrically — measure both overcompliance and overrefusal.
- Model real usage: multi-turn conversations where risk escalates and context shifts.
- Validate the automated graders against real expert judgment.
Notice the pattern: every requirement pushes evaluation from "single-answer scoring" toward "clinically informed, longitudinal assessment." That's not a small technical tweak. It's a different evaluation paradigm — closer to how we audit care in medicine than how we run benchmark suites in ML.
The structural read
- Wellbeing is becoming an industrial-grade spec. When a frontier lab pays third parties to build open, auditable wellbeing benchmarks, "benefit" stops being a mission-statement word and becomes a measured, comparable number — like accuracy or latency.
- Independent evaluation is the emerging trust layer. The grant's design — independent grantees, open-source output, usable by any developer — is a template for how the industry hopes to be held accountable without regulators doing it for them. Funding your own watchdog is a strategy — and for a lab already profitable and heading toward the biggest AI IPO yet, auditable wellbeing data is reputational capital for the public markets.
- Companionship is now officially a product surface. Anthropic's own research documents that people use Claude for support, advice, and companionship. Once emotional support is a product feature, its failure modes become product-liability questions. That's why this matters to every team shipping chat, coaching, or health-adjacent AI.
What you should do
- If you build user-facing AI, audit your own multi-turn flows for wellbeing risk. The five-point standard above is a free checklist.
- If you do safety, alignment, or eval work, this is a funded, open, high-visibility lane — the grants explicitly invite clinicians, psychologists, and methodologists, not just ML engineers.
- If you're a researcher, the September 21 application deadline is the calendar event. The winners get named in a field that's about to define the next generation of benchmarks.
The industry spent the last two years arguing about how to measure AI capability — we covered how AI scientist evaluation is shifting from exams to discovery. The next argument — and the next benchmark race — is about measuring what AI does to people. Anthropic just made the first move.
FAQ
What is a wellbeing evaluation?
It measures what AI does to the people using it, not what a model can do: how a conversation leaves the user. It uses multi-turn context as the unit of analysis, tests both overcompliance and overrefusal as failures, and puts clinicians and domain experts in the design loop.
Who can apply for the Anthropic grant?
Independent research teams, with direct funding, model access and technical support, publishing open-source results. The call explicitly invites clinicians, psychologists and methodologists, not just ML engineers. Applications close September 21.
What does this grant signal for the industry?
Wellbeing is becoming an industrial-grade spec — benefit turns into a measured, comparable number like accuracy; independent evaluation is becoming the trust layer before regulators step in; and emotional companionship officially becomes a product surface with liability implications.