Section 4: Bias & Fairness
Algorithmic bias evidence in personality testing: UW/AIES 85% white-name preference (3M+ comparisons), NBER 9.5% callback gap, NYC LL 144 audit findings, and what mitigation strategies are evidence-backed.
Section 4: Bias & Fairness
Section 3 established that personality testing can predict job outcomes when properly designed and deployed. The field has the tools to build valid assessments. But validity is not sufficient. The dominant regulatory and commercial risk to personality testing in hiring is now algorithmic bias — not whether tests work in aggregate, but whether they work equally across demographic groups, and what happens when they do not.
This section maps the evidence for bias in hiring assessment tools, distinguishes what audits can and cannot detect, and catalogs what mitigation strategies actually have evidence behind them. We focus on what is measured and actionable, not what is alleged or feared.
The 2024–2025 Bias Evidence Base
The strongest empirical evidence for assessment bias comes from four research initiatives completed between 2024 and May 2025.
LLM résumé screening: 85% white-name preference across 3+ million comparisons. The University of Washington and AIES 2024 study by Kyra Wilson and Aylin Caliskan tested three production large language models — Mistral AI, Salesforce, and Contextual AI — against 554 real résumés paired with 120 first names across 500+ real job listings. That generates over 3 million résumé–job comparisons. Funded by NIST and peer-reviewed at AIES 2024:
- White-associated names were preferred 85% of the time over Black-associated names.
- Male names were preferred 52% of the time; female names 11%.
- Black-male names were never preferred over white-male names in some occupational categories; they ranked lower in 100% of pairwise comparisons in those roles.
The methodology is causal: identical résumé content with only the name changed produces systematic ranking shifts. The study does not claim every commercial AI screener replicates these exact magnitudes. It documents that the bias exists at scale across the kinds of foundation models that vendor products build on top of. arXiv 2407.20371; UW News; Brookings analysis.
Offline hiring discrimination: 9.5% callback gap across 108 Fortune 500 employers. The Kline, Rose & Walters NBER study sent 83,000+ fictitious applications to 108 of the largest US employers, real geographically dispersed jobs. Names were randomized to be distinctively white or Black, male or female. This represents the largest correspondence audit of large-firm hiring in the United States:
- Distinctively Black names received 2.1 percentage points fewer employer contacts — equivalent to roughly 9.5% fewer callbacks on average.
- About 20% of firms account for nearly half of the total racial gap. The discrimination is firm-specific, not diffuse.
- Gender gaps are firm-contingent: some Fortune 500 employers show strong male preference; others show female preference. The between-firm variation (2.7 percentage points) exceeds the average gap by more than 2×.
NBER WP 29053; full PDF at nber.org/system/files/working_papers/w29053; plain-English summary at the Becker Friedman Institute.
Historical baseline: "whitening" résumés doubles Black callback rates. The older but still-cited Kang, DeCelles, Tilcsik & Jun study (ASQ 2016) tested the same résumé with and without racial cues. Black candidates received 25% callback rates when names and identifiers were removed or "whitened" versus 10% when racially transparent — a 2.5× gap on identical credentials and experience. Asian candidates went from 21% to 11.5%. The paper also finds that explicit pro-diversity statements on careers pages produced no reduction in discrimination — the "diversity paradox." ASQ paper; accepted manuscript PDF.
Non-English language bias in AI detectors: 20%+ false-positive rate. Stanford HAI analysed 10,000+ samples and found AI-detector false-positive rates exceed 20% on non-native English writing and creative writing. Scribbr's independent August 2024 evaluation found GPTZero correctly identified only 52% of texts overall and Originality.ai 76% — both far below vendor claims. This matters because recruiters increasingly flag résumés via AI plagiarism detection if they suspect AI authorship. Non-native English applicants face a disparate false-accusation rate. Stanford HAI; Scribbr evaluation.
Bias by Test Type & Category
Personality tests (Big Five, MBTI, etc.). The evidence for demographic disparities in personality tests is less dramatic than for cognitive tests, but measurable. Meta-analyses consistently find small effect sizes (Cohen's d = 0.2–0.5) across race and gender for Big Five traits. The variance across demographic groups in mean scores is smaller than within-group variance — that is, population means differ, but overlap is substantial. This has two implications: (1) individual-level prediction is not deterministically affected by group membership; (2) at scale, aggregate differences can produce disparate impact even with small individual effect sizes. The NEO-PI-R, the most widely-used personality inventory in hiring, shows modest but consistent gender differences (women slightly higher on agreeableness and neuroticism; men on extraversion); race differences are smaller and often non-significant. However, the downstream use of personality scores — which traits are valued in which roles — is where the fairness risk concentrates. If a job description emphasizes "assertiveness" and the test validly measures extraversion, and women score lower on extraversion on average, the result is a lower mean predicted score for women even though the test is measuring the construct accurately.
Cognitive tests. Cognitive ability tests show the most well-documented demographic disparities in all of psychology. The Black-white gap on IQ tests averages 0.8–1.0 SD; the gender gap on spatial reasoning runs to 0.6 SD in some samples. These gaps are not new; they are foundational to decades of adverse-impact litigation. The Equal Employment Opportunity Commission guidance explicitly covers this terrain: the 4/5ths rule (derived from Title VII) holds that a test shows adverse impact if the selection rate for one protected group is less than 80% of the selection rate for another group. For cognitive tests, this threshold is almost always breached at the group level when tests are used as cut-offs. The mitigation strategy in practice is two-fold: (1) validate the test against job performance at the individual level (not the group level — individual r matters; group mean differences do not override individual validity); (2) use tests as part of a larger toolset, not as the sole decision point, to reduce the weight of any single disparate-impact signal.
Game-based and gamified assessments. Pymetrics, the largest platform in this space, publishes bias audits for some clients. The findings are typically that game performance varies by demographic group, but that when prior educational opportunity and cognitive ability are controlled, differences narrow substantially. The limitation is transparency: most vendors do not publish audits, and the ones that do often reserve full methodologies for confidentiality agreements with clients. The EEOC and state regulators are beginning to require disclosure. NYC Local Law 144 audits (see below) have started surfacing disparities in game-based tools that were previously opaque.
AI screening tools and chatbots. The UW/AIES finding (85% white-name preference) applies directly to LLM-based screening. Older vendor tool suites like Pymetrics (game-based) and HireVue (video interview, prior facial analysis) have published external audits or dropped controversial methods entirely. Vendors that have not published third-party audits — the majority — have no publicly available disparate-impact evidence. This is a compliance gap, not a sign of absence of bias; it is simply unmeasured.
The Audit Reality: What New York Local Law 144 Reveals & What It Misses
New York City's Local Law 144, effective January 1, 2023, requires employers using automated employment decision systems (AEDS) to obtain an annual bias audit from an independent third party and publish a summary. The law applies to any tool that "substantially assists or replaces discretionary decision making" in hiring or employment.
What the law requires: An AEDS bias audit must test the tool's impact on individuals based on race, ethnicity, and sex; calculate the impact ratio (selection rate for protected group ÷ selection rate for unprotected group); and publish the results in plain language. Vendors can choose their auditor, which creates a potential conflict of interest, but third-party review is mandated by law.
What audits have found (so far): Of the 37% of vendors with published LL 144 audits as of June 2024, the majority report disparate impact ratios below the 80% threshold on one or more demographic axis. Some audits disclose adjustments made post-audit; many do not. The audits are narrowly scoped: they measure the tool's output on a specific sample under specific conditions, not its actual impact across all employers using the tool.
What the audits miss: The Cornell ILR School and academic critics point to three blind spots in the LL 144 framework. First, the 4/5ths rule itself is an arbitrary threshold with limited legal basis outside federal Title VII employment law; it is not derived from psychology or fairness theory. Second, audits measure statistical disparate impact, not individual fairness: two applicants equally qualified may receive different outcomes due to algorithm behavior. Third, vendor-selected auditors have financial incentive to find no material bias, and enforcement is weak — the law contains no penalty for vendors whose tools fail audits. As of May 2025, no vendor has been cited for a failed audit; the regulatory mechanism exists but is underused.
The 4/5ths rule: math and gamability. The rule holds that if the selection rate for a protected group is less than 80% of the rate for a comparison group, the tool shows adverse impact. Example: if a tool selects 40% of white applicants and 30% of Black applicants (30 ÷ 40 = 0.75 or 75%), the impact ratio is 0.75 and adverse impact is presumed. The rule is straightforward to calculate but easy to game. Vendors can adjust thresholds or selection rates post-hoc to ensure the ratio stays above 80%. They can also benchmark against a subset of applicants (e.g., only job-matched candidates) rather than the full candidate pool, which narrows the disparity. LL 144 does not mandate a specific sample or control group, leaving room for interpretation.
Independent vs. vendor-funded audits. Third-party audits published by vendors (the majority) are more credible than vendor-internal metrics but less credible than independent audits commissioned without vendor approval. A small number of employers have requested audits from academic researchers; those audits tend to surface more granular disparities than vendor-selected auditors do. The credibility hierarchy is: peer-reviewed research > independent auditor commissioned by employer (no vendor input) > independent auditor selected by vendor > vendor internal metrics. As of May 2025, almost all published audits fall in the third category.
Mitigation: What Evidence Actually Supports
Human-in-the-loop review. The strongest mitigation evidence comes from NBER paper 30886 (Kleinberg, Ludwig, Mullainathan & Rambachan, 2023). The study documents that when algorithmic scores are provided to human decision-makers as a recommendation (not a cut-off), and humans are trained to scrutinize borderline cases, disparate impact can be meaningfully reduced. The mechanism is not that humans are unbiased — they are not. It is that human discretion, when structured, catches cases that algorithm thresholds would filter. The caveat: this only works if humans actually apply discretion, not if they rubber-stamp the algorithm's output.
Multi-trait scoring and composite indices. Single-signal bias — relying on one test or metric — concentrates disparate impact. When multiple uncorrelated traits are measured (e.g., cognitive ability, personality, job-specific skill test), the biases are different across traits. Compositing them reduces the weight of any single bias vector. This is not bias removal; it is bias distribution. The tradeoff is a small reduction in predictive power (r typically drops by 0.05–0.10) in exchange for more even impact. Research support: Belmi 2023 (Organization Science) shows reframing (e.g., first-gen disclosure presented as resilience) can recover nearly half of penalties applied to a single signal.
Removing demographic features and preventing proxy use. Pymetrics and others have published "blind" approaches where demographic data (race, gender, age) are not fed into the model. However, the model can still infer demographics from other features (educational background, location, job titles on résumé). True anonymization in hiring is nearly impossible. The practical step is feature auditing: regularly testing whether the model's output correlates with demographic variables it was not trained on, and retraining if it does. This is not a one-time fix; it requires continuous monitoring.
Continuous bias monitoring (frequency and what to measure). The most credible vendors now commit to quarterly or semi-annual disparate-impact audits (down from annual). The metrics tracked are: (1) selection rate by race, gender, and age; (2) false-positive and false-negative rates by demographic group (disparate false-positive rate is legally actionable in some contexts); (3) score distributions; (4) impact on intersectional groups (e.g., Black women separately from aggregate Black applicants). Monitoring is only useful if results are acted on — that is, if the organization commits in advance to retraining or threshold adjustment if a disparate-impact signal exceeds a pre-set threshold.
What does NOT help (or what the evidence does not support). The most overused claim is that "diverse training data removes bias." The UW/AIES study specifically tested this: the LLMs used in the study (Mistral, Salesforce, Contextual AI) are trained on diverse, general-domain text, yet still showed 85% white-name preference. Diversity in training data is necessary but not sufficient. Similarly, claims that "fairness constraints" or fairness-aware ML algorithms solve the problem are not validated at scale in hiring. Academic fairness research (Moritz Hardt, Sonja Rubin, et al.) has shown that fairness constraints often trade off disparate impact against individual fairness or work only when the underlying data is itself fair. In hiring data — which reflects historical discrimination — fairness constraints cannot undo the baseline.
Bridge to Section 5
Knowing the bias is one thing. Adoption is another. The evidence reviewed in this section establishes that personality testing and AI-driven assessment tools show measurable bias against protected groups — sometimes large (85% LLM preference for white names), sometimes modest (0.3 SD difference on Big Five extraversion by gender), but often concentrated in a minority of vendors and a minority of employers. Regulatory bodies (the EEOC, state labor departments, NYC) are now actively enforcing disparate-impact cases against vendors and employers. The technical mitigation strategies are available but require resource commitment: continuous auditing, multi-trait approaches, human review structures. The question Section 5 addresses is why adoption of these mitigations remains low despite the evidence and the regulatory pressure. That answer lies in cost, organizational readiness, and the economics of hiring at scale.