What Happened to 'Best Person for the Job'?

The phrase 'best person for the job' has come to imply that structured processes and diverse shortlists are in conflict with finding the most capable candidate.
The assumption is that organisations already select on merit, and that adding process or broadening the pool dilutes it. That assumption is wrong. Merit is a property of a process, not a person. And the processes claiming to select on merit are failing to do so - at three distinct points. The criteria are written to describe whoever last held the role. The screening tools measure credential and network access more reliably than capability. And the performance ratings that selection treats as objective evidence carry documented demographic biases, built on track records generated through unequal access to opportunity.
Definition: Merit and Incumbency
Before any candidate is seen, the criteria themselves encode bias. Gaucher, Friesen and Kay demonstrated that job advertisements in male-dominated fields contain more masculine-coded words - 'competitive,' 'dominant,' 'decisive' - and that this wording reduces how appealing the role feels to women, independent of the actual work. The job description describes the historical incumbent, not the capability required. Lauren Rivera's nine months of participant observation inside an elite professional services firm found evaluators ranking 'culture fit' - shared leisure pursuits, background, self-presentation style - as the most important criterion at interview, often above demonstrable productivity. Sailing. Lacrosse. Classical music. The criteria were real to the evaluators, but unrelated to the work.
The leadership prototype problem compounds this globally. Schein's 'think manager, think male' paradigm - replicated in the US, UK, Germany, China and Japan - established that the traits people ascribe to successful managers match traits they ascribe to men. A competency framework that encodes the historical incumbent of most senior roles. The most precise evidence on this is Uhlmann and Cohen's 'constructed criteria.' In their experiments, evaluators were given male and female candidates for gender-typical roles. They did not see the candidates as having different strengths. Instead they redefined the job requirements to fit the candidate they already preferred - valuing streetwise experience when the man had it, or education when he had that instead. Then they rated themselves more objective for having done so. Pre-committing criteria before seeing any candidate eliminated this entirely.
Measurement: Proxies and Networks
Before an interview, screening typically relies on two proxies that measure privilege more reliably than capability: credentials and networks. The unstructured interview remains the dominant selection method globally - and it is also, the evidence shows, one of the least predictive. Harvard Business School's analysis of 26 million job postings found employers using degree requirements as a proxy for a range of skills, and that degreed and non-degreed workers in the same roles are reported by those employers as roughly equally productive - while degreed workers cost more and turn over faster. The credential requirement inflated beyond what the role requires without improving quality of hire, and disproportionately excluded Black and Hispanic candidates.
Rivera and Tilcsik's field experiment sent 316 law firm offices applications identical in qualifications but manipulating only class signals: an affluent surname, a financial-aid award versus a generic athletic one, sailing and polo versus track and pick-up soccer. The higher-class man received callbacks at more than four times the rate of the other groups. His qualifications were the same. His cultural signals read as belonging. Referral hiring compounds this. Social networks are demographically segregated - people refer people who resemble them - and referral processes narrow the pool to 'best person in the network' rather than 'best person available.' The process feels meritocratic because it relies on trusted information. It is systematically reproducing the composition of whoever built the network.
In Australia, Adamovic and Leibbrandt submitted more than 12,000 applications to real job ads across Sydney, Melbourne and Brisbane. For leadership roles, English-named applicants received callbacks at nearly 57% higher rates than applicants with non-English names on identical resumes. The screening that precedes any interview was already filtering out capable candidates on grounds unrelated to the work.
Evaluation: Ratings and Opportunity
Even when a candidate reaches formal evaluation, the process is measuring something other than pure capability. Castilla and Benard's experiments found that when organisations explicitly describe themselves as meritocratic, managers award larger bonuses to men than to equally-performing women - a bias that disappears when meritocracy is not emphasised. The belief that the system is fair removes the scrutiny the system needs. Performance ratings carry the same problem. Research finds men receive specific, developmental feedback tied to business outcomes, while women receive vague feedback less connected to advancement. Field experiments find that witnessing disrespect reduces bystanders' performance by 22-28% - so the performance being rated is partly a product of the environment's inclusivity, not a pure measure of individual capability.
The 'prove it again' pattern compounds this: 61% of women report needing to demonstrate competence repeatedly for recognition that men receive for a single demonstration. Men are judged on potential; women on demonstrated accomplishment. Attribution research shows success is credited to skill for men and luck for women, while failure runs the opposite way. These effects contaminate the track record that selection later reads as objective evidence of merit. The 'only one in the pool' finding shows how the whole system closes: when a finalist pool contains a single under-represented candidate, the statistical probability of that person being hired approaches zero. A lone different candidate is evaluated against the status-quo template and read as a deviation. With two, the odds change dramatically. This status-quo bias means a process with a single diverse finalist is structurally self-defeating regardless of what the criteria say.
Upstream of all formal evaluation sits a problem selection rarely acknowledges. The track record it relies on is built from access to stretch assignments, high-visibility projects and sponsorship - and that access is distributed unequally. Joan Williams' research found white women are 20% less likely, and women of colour 35% less likely, than white men to receive high-quality assignments. Men receive sponsorship - someone spending political capital to advocate for their advancement - while women receive mentoring: advice without advocacy. Selection reads the resulting track record as evidence of capability. It is partly evidence of who got the chance to build one.
Genuine Meritocracy
The evidence across all three stages points to the same fix: structure that holds the definition of merit constant long enough to apply it consistently.
Fix the criteria in writing before any candidate is seen, weighted in advance and derived from what the role actually requires. Strip degree requirements where the evidence shows degreed and non-degreed workers perform equivalently. Replace 'culture fit' with specific, work-relevant behaviours. Audit job advertisements for coded language. These address the definition problem and prevent criteria from drifting to match whoever the evaluator already prefers.
On measurement: Sackett and colleagues' 2022 re-ranking found structured interviews now top the validity rankings at around .42 - ahead of general cognitive ability - while carrying smaller demographic subgroup gaps than most alternatives. Work sample tests, where candidates demonstrate what the job actually requires, are among the most direct measures of job-relevant capability available. Replace credential screens that cannot be linked to job performance with skills-based assessment of the specific capability the role needs.
Goldin and Rouse's study of US orchestra auditions found adding a screen so evaluators could hear but not see candidates raised the probability of women advancing by around 50%.
On evaluation: ensure at least two under-represented candidates reach the final stage. Separate performance from potential explicitly, and require evidence for both. Audit promotion rates and salary outcomes by demographic after controlling for performance ratings to measure performance-reward bias directly. And audit opportunity distribution upstream - if stretch assignments and sponsorship are not flowing proportionally, the track records you are selecting from are measuring accumulated opportunity as much as innate capability.
The bamboo ceiling research adds a global dimension. Lu, Nisbett and Morris found East Asians are underrepresented in US leadership despite facing less prejudice than other minority groups - the gap is mediated by assertiveness, a culturally variable communication style that Anglo-American selection mistakes for leadership potential. A process that rewards self-promotion as evidence of leadership is measuring acculturation to a communication norm, not merit.
The organisations most committed to the idea that they hire on merit are frequently the ones where the criteria are least defined, the proxies are least examined, and the ratings are least scrutinised. Castilla and Benard called it the paradox of meritocracy. Merit is a property of a process, not a person. The evidence on how to build that process is specific and available. The gap is in whether it is being applied.
The Bottom Line
If you genuinely want the best person for the job, the first question is whether your process is capable of finding them. The evidence says the answer depends entirely on design. Processes with pre-committed criteria, structured assessment, skills-based screening and audited evaluation find talent that unstructured processes routinely miss - as Sackett's validity research, Goldin and Rouse's orchestra study, and Uhlmann and Cohen's constructed-criteria experiments all demonstrate.
The organisations that are actually finding their best people have done the harder work: fixed their criteria before seeing candidates, removed proxies that measure privilege rather than performance, audited where in their process talent is being lost, and stopped reading accumulated opportunity as innate capability. That is what meritocracy requires. What currently travels under that name, in organisations where criteria drift, credentials substitute for capability, and ratings carry unexamined bias, is something considerably more modest.
SOURCES & FURTHER READING
Adamovic M, Leibbrandt A. Field experiment on occupational gender segregation and hiring discrimination. Industrial Relations Journal, 2023
Castilla EJ, Benard S. 'The Paradox of Meritocracy in Organizations.' ASQ 55(4), 2010
Gaucher D, Friesen J, Kay AC. 'Evidence That Gendered Wording in Job Advertisements Exists.' JPSP 101(1), 2011
Goldin C, Rouse C. 'Orchestrating Impartiality.' AER 90(4), 2000
Hsieh C-T et al. 'The Allocation of Talent and U.S. Economic Growth.' Econometrica 87(5), 2019
Johnson SK, Hekman DR, Chan ET. 'If There's Only One Woman in Your Candidate Pool.' HBR, 2016
Lu J, Nisbett RE, Morris MW. 'The Bamboo Ceiling.' PNAS 117(11), 2020
Quillian L et al. 'Meta-analysis of Field Experiments Shows No Change in Racial Discrimination in Hiring.' PNAS 114(41), 2017
Rivera LA. 'Hiring as Cultural Matching.' ASR 77(6), 2012
Rivera LA, Tilcsik A. 'Class Advantage, Commitment Penalty.' ASR 81(6), 2016
Sackett PR et al. 'Revisiting Meta-Analytic Estimates of Validity in Personnel Selection.' JAP 107(9), 2022
Uhlmann EL, Cohen GL. 'Constructed Criteria: Redefining Merit to Justify Discrimination.' Psychological Science 16(6), 2005




Comments