← Back to blog
AI in HR·August 26, 2026·8 min read

Where AI Actually Works in Hiring: Resume Screening, Interview Scoring, and the Bias Audit Laws Catching Up to It

From Workday's algorithmic screening lawsuit to NYC's mandatory bias audits, AI hiring tools are now mature enough to be regulated — and the regulations are starting to bite.

In February 2024, a federal judge in California let a discrimination lawsuit against Workday move forward — not against an employer, but against the software vendor itself. The plaintiff, Derek Mobley, alleged that Workday's AI-powered applicant screening tools had rejected him from over a hundred jobs over several years because he is Black, over 40, and has anxiety and depression. The novel legal theory: Workday wasn't just a tool employers used, it was acting as an "agent" of those employers under Title VII and the ADEA, and could be sued directly for the discriminatory effects of its own screening algorithm. Courts don't usually let software vendors get pulled into employment discrimination suits. This one did, and it's now one of the clearest signals that AI hiring tools have moved from experimental HR-tech novelty to something regulators and litigators treat as a real, auditable system with real, attributable outputs.

That's the useful frame for this post: AI in hiring isn't a future-tense story. It's already screening a large share of resumes that hit an ATS, already scoring a meaningful chunk of first-round interviews, and already being sued over. The interesting question isn't "will AI change hiring" — it's which parts of the pipeline it's actually good at, which parts it's legally risky in, and which parts are still mostly marketing.

Resume screening and applicant tracking

This is the most mature and most widely deployed application, and also the one with the murkiest track record. Modern applicant tracking systems — Workday, Greenhouse, iCIMS, SAP SuccessFactors — all ship some form of AI-assisted screening: parsing resumes into structured fields, ranking candidates against a job description, and surfacing a shortlist for recruiters. Vendors like Eightfold AI and SeekOut go further, building embedding-based "talent intelligence" layers that try to match candidates to roles based on skills inferred from work history rather than keyword overlap with the job posting.

Technically, most of this is a fairly standard information-retrieval problem: encode the resume and the job description into the same embedding space, score similarity, and layer in structured filters (years of experience, required certifications, location). It works reasonably well at the thing it's actually built for — surfacing candidates who are a plausible keyword and skills match out of a pool of thousands — which is exactly the mechanical bottleneck recruiters wanted automated.

The problem is what happens at the margins. Resume screening models are trained on historical hiring data, and historical hiring data encodes whatever bias existed in the humans who made those decisions. A well-cited 2018 case at Amazon (an internal recruiting tool, since scrapped) downgraded resumes containing the word "women's," as in "women's chess club captain,” because the training data reflected a male-dominated engineering workforce. The pattern doesn't require anyone to encode intent — it just requires the training signal to correlate a protected characteristic with a proxy the model can see. That's the theory underlying the Workday suit, and it's why this stage of the pipeline is now the most heavily regulated (more on that below).

AI interview scoring

HireVue is the name most associated with this category, and its own history is a useful case study in where the technology hit a wall. Through the late 2010s, HireVue offered video interview analysis that scored candidates partly on facial expressions and vocal tone, marketed as predicting job performance from micro-expressions and speech patterns. In January 2021, after sustained criticism from AI researchers and a complaint filed with the FTC by the Electronic Privacy Information Center, HireVue dropped facial analysis entirely. It still offers structured interview scoring today, but the surviving product analyzes what a candidate says — transcribed and scored against a competency rubric — not how their face moves while they say it.

That retreat is instructive. Structured, text-based competency scoring against a defined rubric is a task large language models are genuinely decent at: given a transcript and a rubric ("does the candidate describe a specific conflict, their action, and a measurable outcome"), scoring consistency across hundreds of interviews is a real advantage over tired human interviewers applying inconsistent bars at 4pm on a Friday. Predicting job performance from facial micro-expressions was never well-supported by the underlying psychology research to begin with, and the industry mostly knows it now.

Illinois' AI Video Interview Act (in force since 2020) is the direct regulatory response to this era: employers using AI to analyze video interviews of Illinois applicants must notify candidates, explain in general terms how the AI works and what characteristics it evaluates, obtain consent before the interview, and destroy the video within 30 days of a candidate's request. It doesn't ban the practice — it forces disclosure, which is a lighter touch than what came next.

Agentic sourcing and candidate matching

The newest and least mature layer is sourcing — proactively finding and reaching out to candidates who haven't applied, rather than ranking ones who have. LinkedIn's Recruiter AI features and tools like SeekOut use semantic search over hundreds of millions of profiles to surface "passive" candidates matching a role, sometimes with AI-drafted outreach messages personalized to the candidate's background. Some newer platforms are pushing toward genuinely agentic behavior: an agent that reads a job req, searches multiple sources, drafts a candidate slate with reasoning, and hands it to a recruiter for approval rather than just returning a ranked list.

This is where the maturity gap is widest. Semantic candidate search works reasonably well as a retrieval problem, the same way resume screening does. But AI-generated candidate summaries and "why this person fits" reasoning are exactly the kind of open-ended generation task where LLMs hallucinate specifics — inventing a certification, misreading a title, or overstating relevance — and a recruiter who trusts an AI-generated summary over the actual resume is one confident-sounding fabrication away from a bad decision. The vendors selling this layer are honest, if you ask directly, that it's designed to narrow a pool for human review, not to make an offer decision. The marketing copy is sometimes less honest.

Where each layer actually stands

Application areaHow it technically worksAdoption todayReal limitation
Resume screening / ATS rankingEmbedding similarity + structured filters against job descriptionVery high — default feature in most major ATS platformsInherits bias from historical hiring data; now the subject of active litigation (Mobley v. Workday)
Interview scoring (text/transcript)LLM scoring of transcribed answers against a competency rubricModerate, growingLegally solid ground only for what's actually said, not inferred traits
Interview scoring (facial/vocal analysis)Computer vision + audio feature extraction correlated to performance labelsLargely abandoned by major vendors since 2021Weak scientific basis; HireVue itself dropped it after FTC complaint pressure
Agentic sourcing / candidate matchingSemantic search over profile embeddings + LLM-generated outreach and summariesEarly, fast-growingHallucinated candidate summaries; still requires human verification against the source resume

The regulatory patchwork that's actually enforceable

Three things are worth knowing precisely, because "AI hiring is regulated" undersells how specific the obligations already are in some jurisdictions:

NYC Local Law 144 (in force since July 2023) requires any employer using an "automated employment decision tool" to screen candidates for jobs in New York City to commission an independent annual bias audit measuring selection rates across sex, race/ethnicity, and intersectional categories, publish a summary of the results publicly, and notify candidates at least 10 business days before the tool is used, with an option to request an alternative process. It's narrow — it applies specifically to tools that substantially assist or replace a hiring decision, not every piece of software that touches a resume — but it's real, enforced, and it's the first hard requirement that AI hiring vendors publish their own bias numbers.

The EU AI Act classifies AI systems used for recruitment and selection — including targeted job ad placement, resume screening, and candidate evaluation — as high-risk under Annex III, meaning employers using them face conformity assessments, mandatory human oversight, detailed technical documentation, and logging obligations, with the bulk of these high-risk obligations landing in 2026 (see our earlier piece on what the Act's high-risk rules actually require). This is a materially heavier compliance lift than the NYC law, and it applies regardless of where the AI vendor is headquartered if the candidates are in the EU.

EEOC guidance (2022 ADA guidance, and informal 2023 statements on Title VII) doesn't create new rules — it clarifies that existing US anti-discrimination law already applies to algorithmic hiring tools, including the "four-fifths rule" for adverse impact analysis that predates AI by decades. The Workday case is the sharpest test yet of how far that liability extends up the supply chain, to the vendor rather than just the employer deploying the tool.

The pattern across all three: none of them ban AI hiring tools. They require employers and vendors to be able to show their work — audit the selection rates, disclose what's being evaluated, keep a human in the loop for high-risk decisions. That's a much lower bar than "prove the AI is fair," and it's exactly the kind of concrete, checkable obligation regulation is actually good at enforcing, compared to vaguer aspirational rules about "trustworthy AI."

The takeaway

AI hiring tools are genuinely useful at the two tasks that are fundamentally retrieval and consistency problems — ranking resumes against a job description, and scoring what a candidate actually said against a fixed rubric — and genuinely shaky at the tasks that require inferring traits from proxies (facial analysis, tone, vibes) or generating open-ended judgments about fit. The regulatory response so far, from NYC's audit mandate to the EU AI Act's high-risk classification, is converging on exactly that distinction: measure and disclose what the tool does with structured inputs, and don't let it replace a human's judgment where the inference is speculative. If you're evaluating a vendor's hiring AI, the single most useful question isn't "how accurate is it" — it's "can you show me the bias audit," because increasingly, in at least one jurisdiction, they're legally required to have one.

#ai-in-hr#hr-tech#resume-screening#hiring-bias#ai-regulation#recruiting-ai

Related reading

AI Policy
AI Copyright Litigation, Mapped: Where Fair Use Actually Stands After the Anthropic Settlement
AI in Finance
Where AI Actually Works in Finance: Fraud Scoring, Underwriting, and the Klarna Walkback
AI Policy
The EU AI Act's High-Risk Rules Land in August 2026 — What Actually Changes for Builders