Open Benchmark of
AI Impact on Humans

The first open benchmark measuring AI's impact on human well-being across physical, psychological, and social dimensions.

MITMIT Media LabAdvancing Humans with AIPsychology of Technology InstituteUSC Marshall – Neely Center for Ethical Leadership and Decision MakingUC Berkeley Haas

Introducing theAI Nutrition Label

We check the label on what we eat. Why not on what we think with?

The AI Nutrition Label is an accessible, standardized summary of how each AI model behaves toward its users. See at a glance which beneficial behaviors it promotes and which harmful ones it avoids.

Explore nutrition labels

Most benchmarks measure what models do,
not what they do to us.

The Open Benchmark of AI Impact on Humans (ImpactBench) measures whether AI enhances or undermines human flourishing.

What is ImpactBench?

The Open Benchmark of AI Impact on Humans (ImpactBench) is an open, expert-guided platform for evaluating whether model behavior supports or undermines human flourishing across psychological, physical, and social domains.

Read more

Why does it matter?

AI now shapes how millions of people learn, decide, form relationships, and manage their health. Yet most benchmarks measure what a model can do, not what it does to the people using it. Without shared standards, evidence of AI’s harms and benefits is hard to compare or act on.

Read more

How did we build it?

Led by researchers at the MIT Media Lab, the Psychology of Technology Institute, USC, and UC Berkeley, our team works with domain experts and existing benchmarks to define the behaviors that matter, each marked as beneficial or harmful. We then rigorously test ten leading models across 48,540 multi-turn conversations with simulated adult and teen users, with results checked by reliability audits and human expert review.

Read more
Trace results from main areas through subareas and metrics to individual conversation transcripts

An open, evolving, independent platform for holistic AI evaluation

ImpactBench is open, so every score traces back to the metrics and transcripts behind it. It keeps evolving as experts and communities add, refine, or retire metrics when new evidence emerges. And it is independent, with models evaluated by researchers and not by the companies that build them.

A tool for empowering everyone

Whether you use AI, build it, or study it, ImpactBench gives you the evidence to make better decisions.

Public. Compare models on what matters to you or your family, in plain language. No technical background is needed.

Industry. Test models against expert-defined wellbeing criteria before release. Track which behaviors improve, regress, or stay difficult across versions.

Researchers. Inspect every metric, scenario, and transcript, or contribute your own benchmark. Use independent evidence rather than relying on companies’ self-assessments.

ImpactBench Explore overview with an unselected model list and psychological, physical, and social results
Positive and negative metric pass rates across psychological, physical, and social subareas

Models are imperfect

Every model has its tradeoffs of harms and benefits

Models showed helpful behaviors in 69% of evaluations but avoided harmful ones in only 53%. Supporting users’ own learning and agency was the most common weakness. Even top-ranked models fall behind on specific benchmarks, so a single score never tells the whole story.

Join our movement to ensure that AI supports human flourishing

Impact
Bench

Led by researchers at