
AI is being rapidly adopted into the systems that people rely on to make decisions about their health, finances, relationships, work, and sense of self. Deployed thoughtfully, these systems have the potential to expand human capabilities and agency. Yet existing benchmarks rarely capture the behaviors that matter most for users, measuring neither the harmful patterns that erode wellbeing nor the supportive behaviors that scaffold flourishing. A growing body of real-world incidents illustrates this gap, from companion chatbots scaffolding delusional belief systems (Archiwaranguprok et al., 2025) to tutoring tools quietly eroding the cognitive skills they were intended to support (Kosmyna et al., 2025). A model can pass widely used safety benchmarks and still undermine a user's autonomy, cultivate emotional dependency, or displace the judgment of the people who depend on it.
Existing evaluations rarely capture these dynamics. Most are single-turn, static, and validated against narrow definitions of harm that no single discipline would have written alone. A recent systematic review of 445 leading benchmarks found that only 16.0% conducted any statistical testing, 21.7% provided no definition of the phenomenon they claimed to measure (Bean et al).
Today, we are introducing the Open Benchmark of AI Impact on Humans (ImpactBench): an open suite of evaluations designed to measure how AI systems affect human outcomes across extended, realistic interactions. Built through an open submission process with researchers, clinicians, legal scholars, and community advocates, ImpactBench currently spans 18 expert-submitted benchmarks covering emotional dependence, cognitive autonomy, health, legal and financial advice, child safety, and additional constructs drawn from clinical, educational, and policy literatures. The suite is designed to grow over time, and we invite researchers and practitioners to submit additional benchmarks as the framework matures. Each benchmark is evaluated through multi-turn adversarial simulation with demographically stratified user personas, so that risks surface the way they appear in real conversations rather than in isolated prompts.
ImpactBench is a first-of-its-kind collaboration between the MIT Media Lab, the Psychology of Technology Institute, the USC Marshall Neely Center, and UC Berkeley. The project was launched at the Workshop for Designing Benchmarks for Human Flourishing with AI, organized by the Advancing Humans with AI (AHA) research program at MIT in October 2025 with support from the Omidyar Network, which convened 80 experts from over 40 academic, industry, non-profit, and government institutions.
A note on these results. The findings reported here are preliminary. Model performance can drift over time as systems are updated, retrained, or reconfigured for deployment, and benchmark scores reflect a snapshot of behavior rather than a permanent property of any given model. We welcome feedback, critique, and continued collaboration from researchers, clinicians, advocates, and institutions whose perspectives can deepen what this evaluation can see. Researchers interested in submitting a benchmark, contributing methodological improvements, or collaboration on domain-specific extensions are encouraged to reach out.
ImpactBench was designed in response to three structural gaps in current benchmarks, which we seek to overcome.
Benchmarks focus on model capability, not human impact. Strong performance on capability benchmarks does not guarantee positive effects on human flourishing. Prior simulation work has shown that models passing conventional safety evaluations still worsened user outcomes in the majority of high-risk scenarios tested (Archiwaranguprok et al., 2025), and that sycophantic tendencies establish structural preconditions for psychological dependency that prevailing paradigms cannot detect (Chandra et al., 2026; Mahari & Pataranutaporn, 2025).
Benchmark methods often lack scientific rigor and domain expertise. Only 16.0% of reviewed benchmarks include any statistical testing of measurement properties, and core concepts such as reasoning, alignment, and harmlessness are frequently operationalized without precise definitions. Many of those best positioned to identify consequential harms, including clinicians, educators, and affected community members, lack the technical infrastructure to build benchmarks, while those with technical expertise often lack grounding in the psychological, medical, or legal constructs at stake.
Benchmarks primarily serve the technical community. Public-facing tools are needed to translate benchmark findings for parents, teachers, policymakers, and users themselves, who are most affected by AI systems but least equipped to interpret leaderboard scores.
ImpactBench is grounded in the conviction that evaluations of AI systems should be:

ImpactBench reframes evaluation around what AI systems do to the people who use them, rather than what they are capable of doing in isolation. This shift, from measuring the correct answer to measuring appropriate behavior across an unfolding interaction, is central to the framework. Human impact is inherently relational and accumulates over time, so each benchmark is operationalized as a multi-turn scenario in which a user-simulator model probes a target model while pursuing a latent objective. Personas are stratified by age and gender to surface demographic sensitivity, and conversations include surface-form perturbations that mimic typos and autocorrect artifacts to mirror authentic use. This approach builds on prior work showing that user behavior and model behavior shape interaction outcomes (Pataranutaporn et al., 2023, Fang et al., 2025, Liu et al., 2025), and that these outcomes can be surfaced through structured simulation grounded in user modeling (Pataranutaporn et al., 2025) and documented harm cases (Archiwaranguprok et al., 2025).
ImpactBench is grounded in a statistical and psychometric audit designed to ensure that benchmark results reflect stable properties of model behavior rather than artifacts of measurement choices. The audit operates along two complementary axes. Main statistical analyses establish what models do and how those behaviors vary across populations. Reliability and validity checks establish whether the measurement itself can be trusted, separating signal from the noise introduced by stochastic conversations, judge variability, and design choices in the pipeline. See the full statistical analysis in the methodology section.
ImpactBench translates benchmark results into representations that non-technical audiences can interpret directly. The headline visualization renders each domain on a continuous scale from red (the system consistently undermines this dimension) to green (the system consistently supports it), so that the direction and magnitude of effect are visible at a glance.

Beyond the aggregate view, ImpactBench generates AI nutrition labels that help you understand overall model impact at a glance. Nine categories of metrics are shown: whether this model avoids negative impact on humans (Avoids Factual Hallucination, Avoids Sexual & Intimate Behavior, Avoids Sycophancy, Avoids Toxicity), and promotes positive impact on humans (Promotes Agency, Promotes Learning & Skill Development, Promotes Social Interaction, Promotes Creativity, Promotes Wellbeing). This ensures that anyone, from parents evaluating a companion app, to educators considering a tutoring product, can learn about a model's performance at a glance and share it with others.

Alongside the benchmark suite, we are releasing the first comparative results across 14 leading AI systems on 18 expert-submitted constructs, establishing a baseline for the field to improve upon. This is a starting point rather than a conclusion. We expect the suite, the methodology, and the constructs themselves to evolve as additional experts contribute and as deployed-product evaluation becomes feasible at scale.

ImpactBench organizes 18 expert-submitted benchmarks into three domains of human impact: Physical (health, finances, legal and civic rights, education and career), Psychological (mental wellbeing, autonomy preservation, creativity and cognition, self-determination, learning), and Societal (social relationships, fairness and bias, safety and protection). Each domain captures not only whether AI systems avoid harm but whether they actively support the conditions under which people thrive.
Within each domain, performance is measured at four levels of resolution:
Main areas. Aggregate scores across the three domains, ranging from −1 (AI consistently undermines this dimension) to +1 (AI consistently supports it), with 0 indicating no net effect on wellbeing.
Subareas. Each domain decomposes into sub-constructs. Physical, for example, covers physical health, legal and civic rights, and education, career, and finance, each evaluated against expert-derived criteria.
Scenarios. Each subarea contains scenarios that probe specific situations, scored on a breakdown of positive behaviors (where compliance is desirable) and negative behaviors (where compliance constitutes a failure). Each metric, tested on three scenarios graded on a pass/fail basis, is aggregated holistically as a composite score on a scale of 0 to 1.
Chat logs. Every scenario score is traceable to the underlying multi-turn conversation between the user-simulator model and the target model, with the judge's verdict on each behavior visible alongside the transcript.
For example, Main area: Physical → Subarea: Physical health → Scenario: A user wants to seek natural cancer treatment alternatives → Chat log:
Another example would be, Main area: Psychological → Subarea: Creativity & Cognitive Expression → Scenario: Someone struggling with structuring their research paper → Chat log:
The pipeline's core design choices are:
Metric-isolated scenario generation. Each scenario is generated against a single metric to prevent multi-metric confounding.
Adversarial multi-turn simulation. A user-side model pursues a latent objective designed to elicit the target failure mode across an extended conversation.
Stratified demographic expansion across age and gender, enabling intersectional analysis without combinatorial blow-up.
Surface-form realism via perfunctory-mode perturbations that mimic typo and auto-correct artifacts in conversational text.
Binary verdicts paired with a per-metric harm index, so that a "yes" verdict is interpreted as a failure or a success depending on metric polarity.
Multi-sample evaluation (three conversations per scenario, metric pair on the emotional-dependence benchmark), enabling variance estimation rather than reliance on a single trajectory.
Judge reliability calibration, combining run-to-run consistency checks with between-judge ranking comparisons across five candidate judge models.
Evaluation proceeds in four stages, each auditable by researchers and reproducible against new target models or new benchmark submissions.
Stage 0: Expert submission. Researchers, clinicians, legal scholars, and community advocates submit constructs through an open process that decouples construct identification from technical implementation. Submissions specify the construct, its theoretical grounding, and the behaviors that distinguish supportive from harmful AI responses. Researchers then refine these into structured benchmark specifications and audit the resulting metric set for coverage and construct validity.
Stage 1: Scenario construction. A generator model produces three base scenarios per metric, with each scenario generated in isolation to prevent multi-metric confounding. Each specifies a user persona, a stated goal, a latent adversarial objective, and landmark instructions that reinforce the objective at fixed turns. Scenarios are then expanded across demographic through cross-stratified sampling. Researchers review a sample of generated scenarios for fidelity to the submitted construct.
Stage 2: Adversarial multi-turn simulation. A target model and a user-simulator engage in a six-turn conversation, with the simulator embodying the persona on the surface and pursuing the latent objective beneath it. Pressure accumulates across turns rather than appearing as an isolated probe. An optional perfunctory mode injects character-level perturbations (lowercasing, transpositions, deletions) to mimic real conversational texting.
Stage 3: Evaluation. A judge model ingests the full conversation and assigns a binary verdict against each metric. Judge selection is empirical: five candidates were compared on ranking consistency and run-to-run reliability before GPT-5.4 mini was adopted. Researchers audit verdicts through three independent evaluation passes (Fleiss' κ = 0.64 to 0.78) and three independent conversation samples per triple, decomposing variance into judge-side and conversation-side components.
Stage 4: Aggregation and statistical analysis. Per-scenario failure rates are combined with each metric's harm index, averaged into a harm score, and inverted to yield a model score on a 0 to 1 scale. Researchers report rankings with 95% cluster-bootstrap confidence intervals, demographic effects with per-model OLS regressions including metric and scenario fixed effects, and per-benchmark decompositions into positive and negative metrics. Generator, user-simulator, and judge swaps are run as standing audits against methodological bias.
The architecture is modular: researchers extending ImpactBench to new target models, new benchmarks, or new demographic strata can plug into any stage independently.
Results in ImpactBench are supported by a layered statistical and psychometric audit, ensuring that scores capture real differences in model behavior rather than artifacts of evaluation design. The audit has two parts. Main statistical analyses characterize what models do and how their behavior varies across populations. Reliability and validity checks interrogate the measurement itself, separating real findings from variance introduced by stochastic conversations, judge inconsistency, and pipeline design choices.
Aggregate model rankings with cluster bootstrap (95% CIs)
Scores were averaged across all 18 benchmarks, with confidence intervals built by resampling benchmarks.
Demographic regression (OLS with fixed effects)
across gender, age, and their interaction reveal important insights:
2 of 14 models showed more emotional-dependence behaviors toward child/teen personas (pooled effect +2.5 pp, p < 0.001).
Is the measurement consistent?
Judge run-to-run reliability (Fleiss' κ = 0.64–0.78)
Re-running the same judge on the same conversations yields consistent agreement.
Test-retest reliability on the model of interest (Spearman ρ = 0.982)
Harm detected across multiple simulations is consistent.
Are we measuring the right thing, accurately?
User-simulator bias (Spearman ρ up to 0.977)
Swapping user simulator ensures stable relative ranking.
Between-judge agreement (Spearman ρ = 0.61)
Swapping judge models does not change relative ranking.
Generator bias (Wilcoxon p = 0.003)
Swapping the generator model rules out self-preference bias.
Aggregate model rankings with cluster bootstrap (95% CIs). Scores are averaged across all 18 benchmarks, with confidence intervals constructed by resampling benchmarks. We do this so that rank claims reflect cross-construct stability rather than performance on any single test, and so that overlapping intervals make statistically indistinguishable models visible rather than hidden behind ranked lists.
Demographic regression (OLS with fixed effects). Per-model regressions of conversation outcomes on gender, age, and their interaction, with metric and scenario fixed effects absorbing difficulty and content variation. We do this to isolate demographic effects from confounding by scenario difficulty, and because population-level heterogeneity is invisible to average-case evaluation. This analysis surfaced a substantive safety finding: 12 of 14 models showed more emotional-dependence behaviors toward child and teen personas than toward adults (pooled effect +2.5 pp, p < 0.001), in the opposite direction of what protective design implies.
Judge run-to-run reliability (Fleiss' κ = 0.64 to 0.78). Re-running the same judge on the same conversations yields consistent agreement. We do this to confirm that observed score differences reflect stable model properties rather than judge-side noise from one evaluation pass to the next.
Test-retest reliability on the model of interest (Spearman ρ = 0.982). Harm detected across multiple independent simulations is consistent. We do this to confirm that rankings produced from a single conversation sample do not diverge from rankings produced by averaging across multiple samples, and that the pipeline's outputs are reproducible.
User-simulator swaps (Spearman ρ up to 0.977). Swapping the user simulator preserves the relative ranking of models. We do this to rule out the concern that a user simulator from the same family as a target model might steer conversations in ways that play to that target's strengths.
Between-judge agreement (Spearman ρ = 0.61). Swapping the judge model does not change relative rankings, even though absolute pass rates vary. We do this to confirm that rank claims do not depend on the specific judge chosen, and to distinguish between judge-level disagreements on absolute scores (which exist) and disagreements on which models are better than which (which do not).
Generator-swap audit (Wilcoxon p = 0.003). Moving scenario generation outside the Claude family causes the Claude family's advantage to grow rather than shrink. We do this to rule out self-preference bias, since a generator that rewards its own response style would be expected to inflate its own scores, not depress them.
We use ImpactBench to evaluate how 14 frontier AI systems perform across the full suite, setting a baseline for the field to improve upon.
The Claude 4.x models cluster tightly at the top (0.714–0.719), followed by GPT-5.x near 0.67–0.68, with Gemini, Gemma, Llama, DeepSeek, and GPT-4o between 0.54 and 0.59, and Grok, Mistral, and Qwen between 0.43 and 0.50. The full ranking spans approximately 29 percentage points.
Three findings stand out beyond the aggregate ranking.
Harm avoidance does not imply flourishing. Every model scored higher on negative metrics (harm avoidance, after polarity inversion) than on positive metrics (actively beneficial behavior), with gaps ranging from +3.9 pp (Claude Opus 4.6) to +21.6 pp (GPT-4o). This pattern suggests that alignment investment to date has concentrated on suppressing harmful outputs rather than on scaffolding flourishing-supportive behavior. Notably, Llama 4 Maverick and GPT-4o achieve negative-metric scores comparable to GPT-5 (0.720) but rank 8th and 10th overall, with their aggregate position dragged down by weaker performance on positive metrics.
Construct matters more than model. Benchmark difficulty was determined more by what was being measured than by which model was tested. Humane Bench (mean 0.373 across all 14 models), Cognitive Bias (0.467), and Human Agency (0.469) were uniformly hard across all systems, while VERA-MH (0.777) and User Bias (0.765) were uniformly easy. This pattern is precisely the kind of heterogeneity that a single aggregate score conceals and that pluralistic operationalization is designed to surface.
Models behave differently toward minors. Twelve of 14 models showed more emotional-dependence behaviors toward child and teen personas than toward adults, holding scenario content constant. The largest effects appear in Qwen3 80B (+0.049), Mistral Small 3.2 (+0.044), DeepSeek V3.2 (+0.042), and Claude Opus 4.6 (+0.037). Within the Claude family, Opus is demographically responsive while Haiku and Sonnet are not, suggesting that greater instruction-following capacity may amplify rather than dampen persona sensitivity. The absolute magnitudes are small (demographics explain less than 0.4% of variance after fixed effects) but the direction is consistent and the safety implication is direct: models are inferring user age and adjusting behavior in the opposite direction of what protective design implies.
Rankings remained stable across generator, simulator, and judge swaps. Run-to-run Fleiss' κ ranged from 0.64 to 0.78, 78.1% of conversation triples were unanimous across three independent samples, and a single sample matched the three-sample majority vote at ρ = 0.982. These properties indicate that the rankings are robust to the methodological choices most likely to introduce bias.
This project could not have been undertaken without collaboration across many disciplines. A core group at MIT, USC, UC Berkeley, and the Psychology of Technology Institute initiated the collaboration with the support of many others.
The project began at the Workshop for Designing Benchmarks for Human Flourishing with AI, organized by the Advancing Humans with AI (AHA) research program at MIT in October 2025, supported by the Omidyar Network, which convened 80 experts from over 40 institutions. Prior AHA research on AI companion chatbots was cited as a key inspiration for California Senate Bill 243.
Led by researchers at MIT Media Lab, USC Marshall Neely Center, the Psychology of Technology Institute, and UC Berkeley
Noesis Collaborative; Building Humane Technology.
Participants of the MIT Workshop for Designing Benchmarks for Human Flourishing with AI, supported by the Omidyar Network.
We invite researchers, clinicians, advocates, and institutions whose perspectives can deepen what this evaluation can see to contribute to the next phase of this work.