Chapter 10

Human Resources and AI

Bias at scale or fairness at scale?

CSUN · David Nazarian College of Business and Economics · Version 2026-08-22 · Download as Word (.docx)

Part 1: The Hook — One Hundred Rejections and Eight Hundred Hires

Opening Scene

Derek Mobley applied for more than one hundred jobs through employer portals powered by Workday, the software platform that manages hiring for thousands of companies. He was rejected from every one — sometimes within an hour of applying, sometimes by emails that arrived in the middle of the night, timestamps suggesting no human being had ever seen his application.34 In 2023 he sued — not the employers, but Workday itself, arguing that its AI screening tools had a disparate impact on applicants like him: over 40, Black, and living with anxiety and depression. What followed has become the defining legal test of algorithmic hiring. In July 2024 a federal court let the claims proceed on a novel theory: Workday could be liable as the agent of its client-employers. In May 2025, Judge Rita Lin certified a nationwide collective of applicants 40 and older — potentially millions of people — and Workday disclosed that during the relevant period its tools had rejected applications "numbering in the billions."124 In March 2026 the court rejected Workday's last major dismissal argument, and by June 2026 claims spanning age, race, sex, and disability were moving toward discovery.45 Every employer that had switched on the screening features found itself pulled into a lawsuit it never chose to join.5

A decade earlier, Unilever had confronted the arithmetic that makes such tools irresistible. Its entry-level program drew roughly 250,000 applications a year for about 800 positions; screening took up to six months, and human reviewers — tired, inconsistent, unconsciously biased — were the bottleneck.7 In 2016 Unilever rebuilt the funnel around AI: neuroscience-based games (Pymetrics) assessing cognitive and behavioral attributes, then AI-scored video interviews (HireVue), with humans entering only at the final assessment-center stage. The reported results became HR legend: time-to-hire cut by roughly 90 percent, some 50,000 hours of interviewing eliminated, over £1 million in annual savings, candidate completion rates near 96 percent — and a 16 percent increase in the diversity of hires, as the algorithm surfaced candidates human screeners had been passing over.67

The Unresolved Question

The same category of technology thus carries two opposite headlines: bias industrialized to the scale of billions of rejections, and bias reduced below what human screeners could achieve. Both stories are plausible because both are possible — and that is the puzzle. An algorithm has no unconscious to be biased by; it also has no conscience to catch itself. It applies whatever it learned, identically, to everyone, forever, at scale. Whether that consistency is fairness or discrimination depends entirely on what it learned and who is checking. So the real questions for the manager you are becoming: what does the evidence say actually predicts job performance? What does the law require? And when a machine makes a people decision, who is accountable for it?

Chapter Roadmap

The central question: when AI enters people management, what actually changes — the accuracy, the bias, or the accountability? The tools: what selection is for, and a century of evidence on what actually predicts performance; the structured-judgment revolution that algorithms inherit; algorithmic hiring's real mechanics — where bias enters and where it can be audited out; the legal frame, from Griggs to Mobley; performance management and people analytics; and the parts of HR that remain stubbornly human. The two cases then return as the field's live boundary markers.

Part 2: Core Concepts

2.1 What Selection Is For: A Century of Validity Evidence

Definition and Origin

Human resource management is the system through which organizations acquire, develop, evaluate, and retain people; the evidence that HR systems matter strategically is robust — Mark Huselid's landmark study found that firms with high-performance work practices (rigorous selection, training, performance-linked pay, participation) showed substantially lower turnover and higher productivity and financial performance.13 Selection is the system's front door, and its quality is measured as predictive validity: the correlation between a hiring method's scores and later job performance.

What the Research Shows

Frank Schmidt and John Hunter's synthesis of 85 years of research remains the field's reference point: general mental ability tests and structured interviews sit at the top of the validity table; work samples and integrity tests perform well; and the methods organizations love most — unstructured interviews, years of experience, and (near the bottom) education credentials and reference checks — predict weakly.9 A major 2022 re-analysis by Sackett and colleagues, correcting earlier statistical assumptions, revised most validities downward but promoted structured interviews to the top of the table — reinforcing rather than overturning the practical hierarchy.10 The scandal of selection research has always been this gap between evidence and practice: the free-flowing conversational interview, in which every manager trusts their gut, is among the worst instruments ever validated, dominated by first impressions, similarity attraction, and the Week 2 perception biases — yet it persists because it feels diagnostic to the interviewer.

Why It Matters and Where It Breaks Down

Validity compounds: across thousands of hires, small correlational advantages become large workforce differences, which is the business case for structure and, ultimately, for algorithms. The limits: validity is not the only criterion (adverse impact, cost, candidate reactions, and legality all constrain), top predictors can conflict with diversity goals if used naively, and validity evidence is mostly about typical performance in existing jobs — it says less about potential, reinvention, or the jobs AI is currently rewriting.

2.2 The Structure Revolution: Mechanical Beats Holistic

Definition, Evidence, and Limits

The deepest finding beneath the validity table is about how judgments are combined. A structured interview asks every candidate the same job-analyzed questions, scores answers against anchored rubrics, and combines scores mechanically; an unstructured one lets impressions form and merge holistically. Seventy years of research on clinical-versus-mechanical prediction — reinforced by the Week 7 noise framework — reaches an uncomfortable conclusion: simple, consistent combination rules match or beat expert holistic judgment across domains, largely because they eliminate noise, the unwanted variability between judges and occasions that silently degrades every people process.11 Structure is therefore the true ancestor of algorithmic hiring: an AI screener is a structured evaluation taken to its limit — perfectly consistent, infinitely scalable, and utterly faithful to whatever criteria it embodies. That inheritance cuts both ways and frames the whole chapter: structure removes the *idiosyncratic* biases of individual judges while hard-coding whatever *systematic* patterns live in the criteria or the training data. The practical discipline for any evaluator, human or machine: decide the criteria before meeting the candidates, score components independently, delay the holistic verdict — and audit the outputs by group, because consistency without auditing is just bias with better handwriting.

2.3 Algorithmic Hiring: Where Bias Enters, Where It Can Leave

Definition, Evidence, and Limits

AI screening tools rank, score, or reject applicants using models trained on historical data — past applicants, past hires, past performance ratings. Three entry points admit bias. Training data: if past decisions were biased, the model learns the bias as signal — the canonical case is Amazon's experimental recruiting engine, trained on a decade of male-dominated tech résumés, which taught itself to penalize the word "women's" and graduates of women's colleges; Amazon scrapped it in 2018.8 Proxy variables: models barred from using age or race find correlated stand-ins — graduation years, zip codes, gaps in employment, résumé phrasing — which is precisely the mechanism the Mobley plaintiffs allege operates against applicants over 40.23 Outcome definition: a model optimizing "similar to our current top performers" replicates the current workforce, whatever its composition. The same architecture, honestly built, runs the machinery in reverse: because an algorithm's criteria are inspectable and its outputs are countable, it can be audited in ways a thousand interviewers' intuitions never can — selection rates by group can be measured continuously, features can be tested and removed, and Unilever's reported diversity gains illustrate the upside of replacing noisy, biased human screening with validated, monitored assessment.67 The honest limit: audit is a choice, vendor claims are marketing until independently verified, and the research community has raised sustained concerns about the scientific basis of some assessment types — video-interview analysis prominent among them.6 The technology guarantees consistency; only governance decides what gets consistently done.

2.4 The Legal Frame: From Griggs to Mobley

Definition, Evidence, and Limits

U.S. employment law recognizes two theories of discrimination. Disparate treatment is intentional; disparate impact, established in Griggs v. Duke Power (1971), requires no intent: a facially neutral practice that disproportionately excludes a protected group is unlawful unless the employer proves it is job-related and consistent with business necessity.12 Disparate impact is the theory built for algorithms — no one need prove the model *meant* anything, only that its outputs skew and its criteria cannot be justified. Mobley's contributions to the frame are why the case matters beyond its parties: the court accepted that a software vendor can be liable as the employers' agent (dissolving the "we just make the software" defense and the "that's the vendor's problem" defense simultaneously), certified a collective defined by exposure to an algorithm rather than employment by any one company, and in 2026 confirmed that applicants — not just employees — can bring age disparate-impact claims.1345 Regulators are converging on the same logic: New York City's Local Law 144 requires annual independent bias audits and candidate notice for automated employment decision tools, and similar regimes are spreading.2 The managerial translation is blunt: adopting an AI tool outsources the work but not the liability — demand validation evidence, audit results by protected group, and contractual accountability from any vendor, because the court will.

2.5 Performance Management and People Analytics

Definition, Evidence, and Limits

The same structure-versus-judgment logic runs through the rest of the employee lifecycle. Performance evaluation is judgment under noise: ratings vary more with the rater than the ratee, which is why the field has moved from annual ranking rituals toward frequent, specific feedback tied to observable behavior (Week 3's goal-setting and Week 5's psychological safety supply the mechanics), and why forced-ranking systems — which convert rater noise into career consequences — have largely collapsed under their justice costs.11 People analytics extends measurement across the lifecycle: attrition prediction, engagement sensing, skills mapping, and — its frontier — algorithmic management, where scheduling, task allocation, and even discipline are automated, as in gig platforms. The evidence-based promises are real (Google's Project Oxygen and Aristotle, Week 5, are people analytics), and so are the failure modes: surveillance that destroys the trust it measures, metrics that invite Week 3's Goals-Gone-Wild gaming, and opaque automated decisions that violate every element of Week 3's procedural justice — voice, consistency, explanation. The design rule that survives the evidence: automate the measurement, keep humans accountable for the consequences, and make every algorithmic decision explainable to the person it lands on.

2.6 What Stays Human

Definition, Evidence, and Limits

Two findings anchor the boundary. First, applicant reactions matter economically: candidates who experience a process as unfair — no explanation, no human contact, opaque rejection — withdraw, disparage the brand, and litigate; procedural justice research (Week 3) applies with full force to selection, and a 1:50 a.m. auto-rejection is a justice violation with a timestamp.3 Second, the activities that build rather than sort talent — coaching, development, trust repair, hard conversations, the LMX relationships of Week 6 — run on exactly the individualized human attention that automation cannot supply and, done well, frees time for: the defensible promise of HR automation is not replacing judgment but reallocating human hours from screening (where structure beats intuition anyway) to development (where nothing substitutes for a person). The synthesis the two cases will test: AI changes the cost of consistency and the location of accountability; it does not change what predicts performance, what the law forbids, or what employees experience as fair.

Part 3: Comparative Case Study — Mobley v. Workday and Unilever

Setup

The cases bracket the design space. Both deploy algorithmic screening at massive scale; they differ in governance, transparency, and — so far — legal fate. If the technology itself determined outcomes, the cases should look alike. They do not, which localizes the causal variable where Sections 2.3 and 2.4 put it: in validation, auditing, and accountability — the human decisions wrapped around the machine.

Case A: Mobley v. Workday — Consistency Without Visible Accountability

The facts: Workday's screening tools operate across thousands of employers; Mobley's 100-plus rejections — some within an hour, some at 1:50 a.m. — became the lead allegations in a 2023 suit claiming disparate impact by age, race, and disability.34 The procedural milestones each carved new law: July 2024, claims survive dismissal on the agent theory; May 2025, nationwide ADEA collective conditionally certified, with Workday disclosing billions of rejected applications in the period; July 2025, Workday ordered to identify every customer using its HiredScore AI features — pulling client employers into the litigation's orbit; March 2026, the court holds the ADEA's disparate-impact protections cover applicants; June 2026, claims confirmed across protected groups, heading into discovery.1245 Nothing has been proven discriminatory yet — that is what discovery will test — and Workday denies the allegations. As a theory test, the case's significance is independent of its verdict: it establishes that algorithmic screening is legible to disparate-impact law exactly as Section 2.4 predicted (outputs countable, criteria discoverable); that liability follows the decision, not the org chart — vendor as agent, employers as principals, "the algorithm did it" available to no one; and that scale, the technology's selling point, is also its legal exposure — a biased interviewer harms dozens, a biased model certified into a collective implicates millions.15

Case B: Unilever — Structure, Audited

The facts: facing 250,000 applications for 800 places, Unilever's 2016 redesign staged the funnel — application, gamified assessment, AI-scored video interview, human assessment center — with the machine narrowing and humans deciding.67 The reported outcomes: roughly 90 percent reduction in time-to-hire (months to weeks), about 50,000 interviewing hours and over £1 million saved annually, 96 percent candidate completion (versus roughly half before — a candidate-experience gain, not just an efficiency one), and a reported 16 percent increase in hire diversity, consistent with the Section 2.2 logic that structured, blind-to-demographics assessment out-fairs tired human screeners reading names and universities off résumés.67 As a theory test, Unilever illustrates the defensible architecture: algorithms positioned where structure demonstrably beats intuition (high-volume screening), humans retained where judgment and justice demand them (final selection), and outcomes tracked by group. The honest caveats are substantial and belong in the record: the headline numbers originate with the company and its vendors, not independent audit;67 video-interview analysis specifically has drawn sustained scientific criticism, and the assessment industry subsequently retreated from its most contested features;6 "diversity increased" is an outcome claim, not proof the instrument was valid for every group; and a process this automated stands one unaudited model update away from becoming Case A. The difference between the cases is not virtue; it is governance — and governance is revocable.

The Comparison

Mobley v. Workday Unilever (2016– )
Scale Thousands of employers; billions of applications processed ~250,000 applications/year for ~800 roles
Human role in decision Allegedly none visible — auto-rejections within minutes, at 1:50 a.m. Machine narrows; humans decide at assessment centers
Auditing / transparency Contested — the subject of discovery Company-reported outcome tracking, incl. diversity; not independently verified
Legal posture Defining vendor-as-agent liability; ADEA covers applicants; collective of millions No comparable litigation; operates under spreading audit-mandate regimes
Chapter concept made visible Disparate impact built for algorithms; accountability follows the decision Structure + staging + monitoring — consistency governed

Analytical Interpretation

Three conclusions, with the caution both cases demand. First, the pairing dissolves the question "is AI hiring biased?" into the answerable one: is this system validated, audited, staged, and owned? The same technology sits on both sides of the table; the governance differs.68 Second, Mobley's doctrinal moves — agent liability, applicant coverage, algorithm-defined collectives — mean the accountability question is now settled in the direction managers should have assumed all along: you own what your tools decide, and "the vendor's model" is your model in the eyes of the law.15 Third, symmetry in skepticism: Unilever's glowing numbers are self-reported and its most contested assessment technology aged badly, while Workday's tools stand accused, not convicted — outcome-grading (Week 7) tempts in both directions, and the disciplined reading is that Case B shows what defensibility looks like, not that it achieved it, while Case A shows what exposure looks like, not that discrimination occurred. The verdicts belong to the auditors and the court; the design lessons are already yours.

Part 4: Synthesis — Owning the Machine's Decisions

Return to the central question: when AI enters people management, what changes — the accuracy, the bias, or the accountability? Four conclusions hold.

The evidence hierarchy hasn't changed; the delivery mechanism has. Structured, job-related, mechanically combined assessment beat intuition long before software existed;91011 AI is that finding at scale. Deploy it where structure was already winning — screening — and keep humans where development and justice live.

Consistency is not fairness; it is amplification. An algorithm scales whatever it learned — Amazon's penalized résumés and Unilever's reported diversity gains are the same mechanism pointed at different training regimes.68 The fairness is in the validation, the feature choices, and the group-level audits — governance, not code.

Accountability follows the decision. Griggs made intent irrelevant; Mobley is making the org chart irrelevant — vendor, employer, and manager share exposure for what the tool decides, and audit regimes like Local Law 144 are converting best practice into legal floor.1212 Before adopting any people-deciding tool, demand the validity evidence and the bias audit you would need to defend in court — because you may.

Guard the human core. Candidates experience your process as your culture (Week 8) and judge it by procedural justice (Week 3); a rejection with no human, no reason, and a 1:50 a.m. timestamp teaches everyone watching what your organization is.3 Spend the hours automation returns on the work only people can do — developing, coaching, and deciding with names attached.

Sources and Further Reading

  1. Holland & Knight LLP. (2025, May 27). Federal court allows collective action lawsuit over alleged AI hiring bias. https://www.hklaw.com/en/insights/publications/2025/05/federal-court-allows-collective-action-lawsuit-over-alleged
  2. University of Miami Law Review. (2026, February 6). Help wanted, screened by algorithms: Mobley v. Workday and the legal limits of AI hiring. https://lawreview.law.miami.edu/help-wanted-screened-by-algorithms-mobley-v-workday-and-the-legal-limits-of-ai-hiring/
  3. Maynard Nexsen. (2026, April 23). Emerging liability for AI-driven hiring tools: Key developments in Mobley v. Workday, Inc. https://www.maynardnexsen.com/publication-emerging-liability-for-ai-driven-hiring-tools-key-developments-in-mobley-v-workday-inc
  4. Callaham, S. (2026, May 29). A federal judge, a 1967 law and a billion rejected job applications. Forbes. https://www.forbes.com/sites/sheilacallaham/2026/05/29/a-federal-judge-a-1967-law-and-a-billion-rejected-job-applications/
  5. Kress Inc. (2026, July). Mobley v. Workday: What the AI hiring rulings mean for you. https://kressinc.com/blog/mobley-v-workday-ai-hiring-liability-10-steps/
  6. Zhang, Y. (2023). Unilever's practice on AI-based recruitment. Highlights in Business, Economics and Management. https://pdfs.semanticscholar.org/c4bd/5d281b6d06afad5dd15f50b4de3e31746d3a.pdf
  7. Best Practice AI. (n.d.). AI case study: Unilever saved over 50,000 hours in candidate interview time [vendor-reported outcomes]. https://www.bestpractice.ai/ai-case-study-best-practice/unilever_saved_over_50,000_hours_in_candidate_interview_time_and_delivered_over_%C2%A31m_annual_savings_and_improved_candidate_diversity_with_machine_analysis_of_video-based_interviewing.
  8. Dastin, J. (2018, October 10). Amazon scraps secret AI recruiting tool that showed bias against women. Reuters. https://www.reuters.com/article/us-amazon-com-jobs-automation-insight-idUSKCN1MK08G
  9. Schmidt, F. L., & Hunter, J. E. (1998). The validity and utility of selection methods in personnel psychology: Practical and theoretical implications of 85 years of research findings. Psychological Bulletin, 124(2), 262–274. https://doi.org/10.1037/0033-2909.124.2.262
  10. Sackett, P. R., Zhang, C., Berry, C. M., & Lievens, F. (2022). Revisiting meta-analytic estimates of validity in personnel selection. Journal of Applied Psychology, 107(11), 2040–2068. https://doi.org/10.1037/apl0000994
  11. Kahneman, D., Sibony, O., & Sunstein, C. R. (2021). Noise: A flaw in human judgment. Little, Brown Spark.
  12. Griggs v. Duke Power Co., 401 U.S. 424 (1971).
  13. Huselid, M. A. (1995). The impact of human resource management practices on turnover, productivity, and corporate financial performance. Academy of Management Journal, 38(3), 635–672. https://doi.org/10.5465/256741

Notes appear as superscript numbers in the text and correspond to the numbered sources above. DOIs are provided where available; classic books and cases are cited to their original publishers and reporters.

← Chapter 9: Strategy and StructureChapter 11: Ethics and Social Responsibility →