AI Is Screening Your CV — and It Has a Bias Problem
Automated screening promises to strip bias out of hiring. The evidence is that it more often launders bias in — learned from old decisions, hidden in a score, and now sorting millions of applications before a human sees them.
Somewhere between clicking “submit” on a job application and hearing anything back, a decision is increasingly made without a person in the room. A model reads your CV, scores it, ranks it against others, and passes a shortlist up the chain — or quietly filters you out before a recruiter ever sees your name. It is sold as the fair option: tireless, consistent, free of the gut-feel prejudice that dogs human hiring. The evidence, gathered over nearly a decade now, points the other way. Automated hiring tools do not reliably strip bias out. More often, they launder it in — learned from old decisions, hidden inside a score, and applied to far more people than any single biased human ever could reach.
This is not a story about a rogue algorithm or a single bad vendor. It is a structural point about how these systems work, and it is worth understanding whether you are applying for jobs, building hiring software, or just trying to judge the gap between what “AI-driven recruitment” promises and what it delivers. The short version: a model that learns from who a company hired before tends to recommend more of the same, and “the same” is exactly the thing anti-discrimination law was written to interrupt.
We have argued before that algorithmic bias is not a glitch but a predictable output of the training process. Hiring is where that argument stops being abstract, because the outcome is not a slightly worse search result — it is whether you get the interview.
Bias in, bias out: how it actually happens
Machine-learning systems are pattern-matchers. Show one enough examples of a thing and it will learn to predict more of that thing. A hiring model is typically trained on a company’s own history: the applications it received, and the candidates it chose to advance. The system’s job is to find the statistical signature of “the kind of person we hire” and apply it to new applicants. If the historical record reflects a workforce skewed by gender, race, age or class — whether through overt discrimination or the quieter accretion of structural advantage — the model treats that skew as the target, not the problem.
That is the mechanism, and it is not hypothetical. The most cited example is Amazon’s. As Reuters reported in 2018, the company had built an experimental tool to score CVs, trained on ten years of applications submitted to a heavily male engineering workforce. The model duly taught itself that men were preferable. It reportedly downgraded CVs that contained the word “women’s” — as in “women’s chess club captain” — and penalised graduates of at least two all-women’s colleges. Amazon tried to neutralise those specific signals, then abandoned the project altogether after concluding it could not be confident the system would not simply find other proxies for the same bias. That last part is the tell: they could not guarantee it was clean, so they stopped. Most of the industry did not stop.
The evidence got newer, and it got worse
You might hope that eight years and a leap to large language models would have fixed this. The most rigorous recent study says otherwise. In 2024, researchers Kyra Wilson and Aylin Caliskan at the University of Washington tested how leading language models rank real CVs when used for résumé screening. They paired 554 genuine CVs with names statistically associated with different races and genders, ran them against more than 500 real job listings, and generated over three million comparisons.
The results are hard to wave away. Across those millions of comparisons, the models preferred CVs with white-associated names 85% of the time, and female-associated names only around 11% of the time. Intersection made it starker still: Black men fared worst of all, with the models preferring other candidates in close to every single test. Same CV, different name, systematically different outcome — which is the definition of the thing hiring law forbids, produced by a tool marketed as the cure for it.
And the study delivered a second, more uncomfortable finding for anyone reaching for the obvious fix. The bias did not depend on names being visible. Because these models are trained on the whole sweep of human text, they can infer a candidate’s likely identity from indirect signals — the schools listed, the cities lived in, the dates that hint at age, even the words someone chooses to describe their own work. Strip the name off the top of the CV and the machine still finds the pattern lower down.
When the law has caught up — and when it hasn’t
Regulators have started to notice, but unevenly, and mostly one case at a time rather than through any settled rulebook. The clearest example of enforcement came from the United States. In 2023, the Equal Employment Opportunity Commission settled what it called its first case involving AI-driven hiring discrimination, against iTutorGroup, a company that recruited remote tutors. Its application software had been set to automatically reject female applicants aged 55 or older and male applicants aged 60 or older — screening out more than 200 people on age alone. iTutorGroup agreed to pay $365,000 and to overhaul its practices. Note what that case was and was not: it was blunt, rule-based automation, and it was caught because the discrimination was legible. The subtler, learned bias of a modern model is far harder to prove in the same way, which is part of why systemic enforcement lags so far behind the technology.
Two rules aim more directly at the machinery itself, and their contrast is instructive. New York City’s Local Law 144, in force since 2023, requires any employer using an automated employment decision tool to have it audited for race and gender bias by an independent party every year, to publish a summary of the results, and to tell candidates the tool is being used. It is narrow — it covers a single city, and critics argue the audits can be gamed — but it establishes a principle that matters: if a machine is sorting applicants, someone independent should be checking it for disparate impact, in the open.
Europe went further on paper and then blinked on timing. The EU AI Act explicitly classifies AI used for recruitment and candidate selection as “high-risk” — a category that brings real obligations around risk management, data governance, documentation, human oversight and transparency. Those duties were due to bite from 2 August 2026. Instead, as part of a broader move to ease the compliance timetable, the high-risk obligations covering areas like employment have been postponed to 2 December 2027. In other words, the single most consequential set of protections against biased hiring AI, meant to arrive this very month, has just been pushed back by sixteen months. If that pattern sounds familiar, it is the one we traced in what AI regulation actually protects you from — and what keeps getting delayed: the rule exists, the enforcement recedes.
The steel-man: why anyone automates hiring at all
It would be unfair to pretend the case for these tools is empty. Human hiring is slow, expensive and demonstrably biased in its own right — the very same résumé-with-a-different-name experiments have caught human recruiters discriminating for decades. Faced with thousands of applications for a handful of roles, a company cannot give each a careful human read, and a tool that surfaces plausible candidates faster is genuinely valuable. Done with real care — audited, monitored for disparate impact, used to widen a shortlist rather than to auto-reject — algorithmic screening could in principle be fairer than a tired recruiter on a Friday afternoon, precisely because a model’s bias can be measured and corrected in a way a human’s cannot.
That is the honest best case, and it should be conceded. But it rests on conditions that are mostly absent in practice: independent auditing, transparency to candidates, ongoing monitoring, and a human making the final call with the authority to overrule the machine. Where those are missing — which is nearly everywhere outside a couple of jurisdictions — the tool is not de-biasing hiring. It is automating whatever bias it inherited and lending it the false authority of a number. The problem is not that the technology cannot be fair. It is that fairness is optional, and the default is unaudited.
Why the score is the dangerous part
There is a specific reason algorithmic hiring bias is more corrosive than the human kind, and it is not the bias itself — it is the packaging. When a recruiter passes on you, everyone understands a fallible person made a judgement. When a system returns a 61 out of 100, it arrives dressed as measurement. That veneer of objectivity does two things at once. It makes the decision harder to question, because who argues with a number? And it makes it easier to defer to, because a busy hiring manager handed a ranked list has every incentive to trust the order and move on. The bias does not just survive automation; it gets a lab coat.
This is the same dynamic we keep running into across AI — a system that is confident and quantified but not therefore correct. A résumé score is a probabilistic guess wearing the costume of a fact, and the costume is doing real work: it converts “the model has learned to prefer people like our last hires” into “this candidate scored lower”, which sounds like a property of you rather than a property of the training data.
It also quietly relocates responsibility. A recruiter who rejects you owns that call; a recruiter who accepts a tool’s ranking can tell themselves, and a tribunal, that they merely followed the data. That diffusion is why unaudited scoring is so sticky: it lets a biased outcome happen with nobody in the chain feeling they authored it. The vendor points to the deploying employer, the employer points to the vendor’s model, and the model, of course, points to nothing at all — it has no obligation to explain itself, and in most places no legal duty to let you see the reasons you were filtered out. The comfortable fiction that the number is neutral is precisely what keeps anyone from having to defend it.
What to do about it, whichever side of the CV you’re on
If you are applying for jobs, the practical advice is modest but real. Assume a machine may read your application first, and give it less to misread: use the plain, standard job title as well as any clever one, spell out skills in the words a listing uses, and do not rely on a human to infer what a keyword-matcher will miss. Know your local rights — if you are applying to a New York City employer, a bias audit and a disclosure are owed to you; if you are in the EU, the strongest protections are real but now delayed, so do not assume they are already in force. And if a rejection seems to defy the facts of your experience, it is fair to ask an employer whether an automated tool was used and how it was checked. You may not get a satisfying answer, but the asking is part of how the norm changes.
If you build or buy this software, the checklist is not exotic:
- Audit for disparate impact before deployment and on a schedule after it, ideally by an independent party, and act on what the audit finds.
- Keep a human with genuine authority in the loop — one empowered to overrule the ranking, not just to rubber-stamp it.
- Tell candidates the tool is being used, and give them a route to query or contest a decision.
- Treat “we removed the names” as a starting point, not a defence, and test for the proxies that leak identity anyway.
Above all, resist the score’s false certainty. The uncomfortable truth at the centre of all of this is the one Amazon ran into and had the sense to act on: a hiring model does not learn what a good employee looks like. It learns what your last decisions looked like — and if you would not stand behind the pattern in those decisions, you should not let a machine repeat it, faster, behind a number, on people who never get to see it. When it does cause harm, the question of who is actually liable is still being worked out — which is all the more reason not to wait for the law to decide it for you.
Frequently asked questions
How does an AI hiring tool end up biased in the first place?
Almost always through its training data. These systems learn patterns from historical examples — the CVs a company received and the people it chose to hire or promote. If past hiring favoured one group, whether through overt bias or structural advantage, the model treats that pattern as the target to reproduce. Amazon's scrapped tool is the textbook case: trained on a decade of mostly-male CVs from a male-dominated industry, it learned to downgrade CVs that signalled “woman”, penalising the word “women's” and graduates of some all-women's colleges. The bias was not programmed in; it was learned from us.
Isn't a human recruiter just as biased? Why single out the AI?
Human recruiters are biased too — that is part of why automation was sold as a fix. The difference is scale, speed and opacity. One biased recruiter affects the candidates they personally see; one biased model can filter millions of applications with the same skew, consistently, before any human is involved. And its decisions arrive as a clean score or ranking that looks objective, which makes the bias harder to spot and easier to defer to. Automation does not remove human bias here so much as encode one version of it and apply it at industrial scale.
Does stripping names and photos off applications solve it?
It helps, but it is not a fix, because a model can reconstruct identity from the rest of the CV. Educational institutions, home postcodes, graduation dates, membership of certain organisations and even the vocabulary someone uses to describe their work can all correlate with gender, race or age. Researchers who tested résumé-screening models found bias persisted even when the obvious signals were controlled for. Blind screening addresses the front door while leaving the windows open.
What are my rights if I think an algorithm rejected me unfairly?
It depends heavily on where you are. In New York City, employers using automated hiring tools must commission an annual independent bias audit, publish a summary, and tell candidates the tool is being used. In the EU, recruitment AI is formally “high-risk”, but the obligations that would force documentation, human oversight and transparency have been pushed to December 2027. Almost everywhere else, you fall back on existing anti-discrimination law — which protects you against a discriminatory outcome but rarely gives you a right to see, or contest, the model that produced it.
Sources
- Amazon scraps secret AI recruiting tool that showed bias against women — Reuters
- UW research finds racial and gender bias in AI tools ranking job applicants' names — University of Washington
- Gender, Race, and Intersectional Bias in Resume Screening via Language Model Retrieval — Wilson & Caliskan (AIES 2024)
- iTutorGroup to Pay $365,000 to Settle EEOC Discriminatory Hiring Suit — U.S. EEOC
- Automated Employment Decision Tools (Local Law 144) — NYC (DCWP)
- Annex III: High-Risk AI Systems (and the delay of high-risk obligations to 2 December 2027) — EU Artificial Intelligence Act
