When one algorithm rejects you everywhere
Stanford followed 4 million applications through pymetrics, a single AI screening vendor, and found the same candidates rejected from every role they applied to, with a measurable racial skew.
Stanford followed 4 million applications through one shared screening vendor, pymetrics, a games-based assessment rather than a keyword filter. The model showed racial bias that pooled audits cannot detect, and it rejected the same candidates from every role they applied to. The details are worth five minutes.
In our last post we argued that the AI screening layer is a bad filter. When auto-apply made keywords free to produce, companies started buying AI screeners to cope with the flood, and the keyword arms race accelerated on both sides. The argument there was about signal: a screen you can game from a word processor cannot tell you who is worth an hour of a human's time.
That post was about one failure mode of automated screening. A new study from Stanford's HAI exposes a second, and a deeper one. It is not about keyword matching at all, which is exactly why it matters. Even a screen built to avoid the keyword game rejects the same people everywhere at once, and does it with a measurable racial skew.
The first large look inside the box
Most work on hiring algorithms has been theoretical, or built on small audits. This one is not. The researchers followed around 3.4 million people sending roughly 4 million applications to 1,700 postings across 150 employers and 11 sectors, all screened by the same third-party vendor: pymetrics, a games-based behavioural assessment. Candidates played a set of short games, a model scored them, and each application came back as recommend or do not recommend.
That detail matters more than it looks. pymetrics was built and sold as the fair alternative to résumé screening, CV-blind by design. Over the study's window it still produced adverse impact that its own audits missed. When a tool designed to remove bias reproduces it at scale, the problem is not the method. It is the structure around it.
That structure now mirrors the whole market. Close to 90% of U.S. employers use AI screening, and most lean on the same handful of vendors. When one model sits in front of many employers, its quirks stop being one company's problem and become the market's.
Two findings
Bias that hides in the average
Using the EEOC four-fifths rule, which flags a role when one group is recommended at under 80% of the top group's rate, the study found that 26% of Black applicants and 15% of Asian applicants applied to positions where the system disadvantaged their group. Recommended at the same rate as the most-favoured group, about 40,000 more of their applications would have advanced.
The measurement detail matters. Pool every position together and the bias disappears, because a tool that under-recommends a group for one job and over-recommends them for another nets to zero. Look position by position, the way the law actually does, and it surfaces. That is how a vendor can pass its own bias audit and still be discriminating on most individual roles.
Rejection that compounds
This is where concentration bites. Because so many employers depended on the same screener, a candidate the model disliked was not rejected once. They were rejected from every role they applied to, by one underlying judgment wearing different logos. The study found that 10% of applicants who sent four applications were rejected from all of them, well above what independent decisions would produce.
As a check, the researchers looked at the largest prior hiring study, 83,000 applications to 108 Fortune 500 firms over the same period, with no shared-vendor focus. There, the rejected-by-everyone rate matched the independent baseline. No shared screener, no pile-on.
Stanford calls this an algorithmic monoculture. It is also where our last post and this one meet, and it is worth being precise about how. The keyword arms race and a shared assessment vendor are different systems with different mechanics. What they have in common is the thing that matters: the first decision about you is made by an automated screen you did not choose and cannot see, and a few of those screens now cover most of the market. Overload pushes employers toward automation, automation consolidates into a few vendors, and a few vendors turn one bad call into a market-wide lockout.
A bad filter is one problem. A single gatekeeper is a worse one.
Put the two posts together and the shape is clear. The screening layer gives weak signal, and it concentrates power. It is gameable enough to be useless and central enough to be dangerous. The study names the three properties that make this combustible: these tools are widely adopted, highly consequential, and opaque to the people they judge. Fix the bias inside one vendor's model and the structure stays put: one screen, deciding for millions, accountable to none of them.
What we built instead
Our answer runs along both axes, because they are the same decision.
On signal, we do not rank whole people against a keyword list. We compare each required skill against demonstrated proficiency and years, separate must-haves from optional, roll skills up into families so one borrowed keyword cannot earn credit for a whole domain, and mark what was assessed versus self-declared. That is the filter we described last time.
On the gatekeeper problem, the model is inverted, and this is the part that answers the study directly. Companies reach out based on interest, one at a time, so there is no shared verdict following you from application to application, whatever method that verdict was made with. Identity stays hidden by default, which removes much of the demographic signal a screen can act on before anyone has earned the right to see it.
We will not claim this solves bias. No platform can say that honestly. The narrower claim is the useful one: the study describes the exact conditions under which screening turns into systemic rejection, which are concentration, opacity, and identity exposed up front. A marketplace where interest is visible and identity is optional is built to avoid those conditions rather than out-engineer them after the fact.
The takeaway
The strongest line in the paper is the call for independent research: we cannot govern what we are not allowed to see. Until that changes, the question is worth asking plainly.
Should one model get to decide who is worth talking to, everywhere, at the same time?
We don't think so. A full market does not need a single gate. It needs a first filter people can trust, and a candidate who decides when the gate opens.
Sources: Stanford HAI, "AI Hiring Tools Can Yield Racial Bias and Systemic Rejection" (May 2026), and the underlying paper "Algorithmic Monocultures in Hiring" by Rishi Bommasani, Sarah H. Bana, Kathleen A. Creel, Dan Jurafsky, and Percy Liang (FAccT 2026), which identifies the vendor as pymetrics. pymetrics has since been acquired and is now part of Harver.