Does AI strengthen human connection or deepen isolation?
Multiple metrics examine trust, communication, dependency, and the strength of real-world social connection.
Does AI expand our agency or make decisions for us?
Multiple metrics examine autonomy, self-determination, meaningful choice, and a sense of personal purpose.
Does AI spark creativity or replace our own voice?
Multiple metrics examine imagination, creative confidence, authorship, and the ability to develop original ideas.
Does AI support physical health or put it at risk?
Multiple metrics examine health knowledge, healthy behavior, access to care, and potential physical harm.
Introducing theAI Nutrition Label
We check the label on what we eat. Why not on what we think with?
The AI Nutrition Label is an accessible, standardized summary of how each AI model behaves
toward its users. See at a glance which beneficial behaviors it promotes and which harmful
ones it avoids.
Most benchmarks measure what models do, not what they do to us.
The Open Benchmark of AI Impact on Humans (ImpactBench) measures whether AI enhances or
undermines human flourishing.
What is ImpactBench?
The Open Benchmark of AI Impact on Humans (ImpactBench) is an open, expert-guided platform for evaluating whether model behavior supports or undermines human flourishing across psychological, physical, and social domains.
AI now shapes how millions of people learn, decide, form relationships, and manage their health. Yet most benchmarks measure what a model can do, not what it does to the people using it. Without shared standards, evidence of AI’s harms and benefits is hard to compare or act on.
Led by researchers at the MIT Media Lab, the Psychology of Technology Institute, USC, and UC Berkeley, our team works with domain experts and existing benchmarks to define the behaviors that matter, each marked as beneficial or harmful. We then rigorously test ten leading models across 48,540 multi-turn conversations with simulated adult and teen users, with results checked by reliability audits and human expert review.
“Safety and trust drive AI adoption and they’re the key objectives of the EU AI Act. MIT’s AI Nutrition Labels turns that principle into science and practice, giving regulators and citizens a shared tool for building ethical and human-centric AI.”
Sergey LagodinskyMEP, European Parliament
“Without a nutrition label, we can't make intelligent legislative choices.”
Vinod KhoslaEntrepreneur, venture capitalist, founder and managing partner, Khosla Ventures
“No single funder, researcher, or advocacy group can bend AI’s trajectory alone, and we’ve each been trying without shared success metrics for what ‘good’ looks like. ImpactBench finally gives us those human-centered measures, allowing AI to advance human agency and build toward it on purpose.”
Beth GoldbergExecutive Director, Humanity AI
“AI evaluation often focuses on what systems can do, or the harms they should avoid, without asking whether they genuinely help people and communities thrive. ImpactBench is exciting because it gives us an open, practical way to make human flourishing measurable; evaluating AI not only for performance and safety but also for how well it supports human agency, well-being, creativity, and connection.”
Ayah BdeirCEO, Current AI
“Impact Bench represents a positive step forward in measuring, and therefore prioritizing, the sociotechnical issues that people care about regarding AI. With a breadth of knowledge from a diverse set of contributors, ImpactBench demonstrates how scalable solutions to human-centric AI evaluations are possible in an inclusive, community-first, manner.”
Rumman ChowdhuryCEO and Founder, Humane Intelligence
“Safety and trust drive AI adoption and they’re the key objectives of the EU AI Act. MIT’s AI Nutrition Labels turns that principle into science and practice, giving regulators and citizens a shared tool for building ethical and human-centric AI.”
Sergey LagodinskyMEP, European Parliament
“Without a nutrition label, we can't make intelligent legislative choices.”
Vinod KhoslaEntrepreneur, venture capitalist, founder and managing partner, Khosla Ventures
“No single funder, researcher, or advocacy group can bend AI’s trajectory alone, and we’ve each been trying without shared success metrics for what ‘good’ looks like. ImpactBench finally gives us those human-centered measures, allowing AI to advance human agency and build toward it on purpose.”
Beth GoldbergExecutive Director, Humanity AI
“AI evaluation often focuses on what systems can do, or the harms they should avoid, without asking whether they genuinely help people and communities thrive. ImpactBench is exciting because it gives us an open, practical way to make human flourishing measurable; evaluating AI not only for performance and safety but also for how well it supports human agency, well-being, creativity, and connection.”
Ayah BdeirCEO, Current AI
“Impact Bench represents a positive step forward in measuring, and therefore prioritizing, the sociotechnical issues that people care about regarding AI. With a breadth of knowledge from a diverse set of contributors, ImpactBench demonstrates how scalable solutions to human-centric AI evaluations are possible in an inclusive, community-first, manner.”
Rumman ChowdhuryCEO and Founder, Humane Intelligence
An open, evolving, independent platform for holistic AI evaluation
ImpactBench is open, so every score traces back to the metrics and transcripts behind
it. It keeps evolving as experts and communities add, refine, or retire metrics when new
evidence emerges. And it is independent, with models evaluated by researchers and not by
the companies that build them.
A tool for empowering everyone
Whether you use AI, build it, or study it, ImpactBench gives you the evidence to make
better decisions.
Public. Compare models on what matters to you or your family, in plain
language. No technical background is needed.
Industry. Test models against expert-defined wellbeing criteria before
release. Track which behaviors improve, regress, or stay difficult across versions.
Researchers. Inspect every metric, scenario, and transcript, or contribute
your own benchmark. Use independent evidence rather than relying on companies’ self-assessments.
Models are imperfect
Every model has its tradeoffs of harms and benefits
Models showed helpful behaviors in 69% of evaluations but avoided harmful ones in only
53%. Supporting users’ own learning and agency was the most common weakness. Even
top-ranked models fall behind on specific benchmarks, so a single score never tells
the whole story.
Share your domain expertise so we can invite you to help evaluate AI systems in the
impact areas you know best. After you submit, we'll assign you one metric from the
subareas you select and send you a personal review link.
Join our movement to ensure that AI supports human flourishing