Psychometric testing for hiring gets sold two ways, and both are wrong. One version says it is science, so it settles the question. The other says it is pseudoscience dressed up in percentages.
The useful version sits in between, and it is more actionable than either. These tests measure something real, the measurement is weaker than vendors imply and stronger than sceptics assume, and what you do around the test matters more than which test you buy.
What psychometric testing for hiring measures
The category covers three different things that get lumped together.
Cognitive tests measure reasoning, problem solving and how quickly someone picks up unfamiliar material. Personality assessments describe behavioural tendencies: how a person handles pressure, whether they prefer working alone, how they respond to conflict. Integrity and situational judgment tests sit between the two, asking how someone would act in specific work scenarios.
They predict different things and they are not interchangeable. Treating psychometric testing for hiring as one product is the first mistake most buyers make.
What the research actually says
This is where most articles on the subject go wrong, including the ones citing real papers.
The numbers people quote
The reference point is Schmidt and Hunter’s 1998 meta analysis, which reviewed 85 years of selection research. It reports general mental ability at a validity of .51 and structured interviews also at .51. Unstructured interviews come in at .38. Experience and education, the things a resume actually shows you, sit at .18 and .10.
Note what that means for the popular claim that assessments beat interviews. The paper does not say that. It puts structured interviews at the top of the hierarchy alongside cognitive ability, and it finds the strongest results in combinations. Cognitive ability paired with a structured interview reaches .63. Paired with an integrity test, .65.
Why a percentage is the wrong reading
You will often see a validity coefficient converted into a sentence like “interviews are only 20% predictive”. That conversion is not valid.
These are correlations. A coefficient of .51 does not mean the method is right half the time. It describes how strongly scores and later performance move together across a large population. That is useful for deciding what to include in a process, and useless for predicting any individual hire.
That distinction matters commercially. A method with strong validity still produces plenty of wrong calls on individual candidates, and anyone who tells you otherwise is selling something.
The revision most people skip
The 1998 figures are still quoted everywhere as settled. They were revised. Sackett, Zhang, Berry and Lievens published a reanalysis in 2022 applying more conservative corrections, which brought the cognitive ability estimate down from .51 to roughly .31.
The relative ranking held. Cognitive ability, structured interviews, work samples and integrity tests stayed at the top. Unstructured interviews, years of experience and educational credentials stayed near the bottom. But the headline number is lower than the one on most vendor websites, so it is worth knowing which version you are being quoted.
Where psychometric testing for hiring earns its place
Three situations, specifically.
When the volume makes reading impossible. A structured score lets you rank 300 applicants against the same criteria. No amount of careful reading does that consistently at that scale.
When the role depends on traits a resume cannot show. Front line work often turns on pace tolerance, how someone handles a difficult customer, or whether they need supervision. None of that appears in an employment history.
When you need to compare across reviewers. A score means the same thing whoever is reading it, which is the whole reason screening processes need structure once more than one person is running them.
Where it fails
Four failure modes, and they are all avoidable.
Cultural bias. Some instruments were normed on populations that look nothing like your applicant pool. A test that penalises unfamiliar phrasing is measuring background, not capability.
Coaching and gaming. Candidates can prepare, and practice materials for common assessments are widely available. That compresses the range of scores and makes the top of your distribution less meaningful than it looks.
The static snapshot. A test captures one morning. It does not capture how someone develops over 6 months in your environment.
No benchmark. This is the big one. A score means nothing without a description of what good looks like in that specific role. Teams that buy an assessment before defining the benchmark end up with precise measurements of things that do not matter. It is the same sequencing error that makes resume replacement projects stall.
The combination is the point
If there is one finding worth taking from the research, it is that no single method wins. The strongest results come from pairing methods that fail in different ways.
That is why the framing of assessments versus interviews is a false choice. A structured interview and a psychometric assessment measure overlapping but distinct things, and running both catches candidates that either one alone would misjudge.
What the assessment does change is the order. Running it before interviews rather than after means your interview slots go to people who already resemble your top performers. The interview then confirms and explores rather than screening from scratch.
How Workwolf® uses it
Workwolf® runs the hiring process on your behalf, and psychometric testing for hiring is one stage inside it rather than a product you administer. Packfinder, our assessment, is built and supported by specialists at Self Management Group, and it profiles behaviour against what the role actually demands.
The part that matters is what surrounds it. We build the benchmark with you first, then apply the same profile to every applicant. The assessment result is combined with verified credentials rather than treated as sufficient alone.
We also do not claim the assessment predicts individual outcomes. It improves the odds across a pipeline, which is a real and valuable thing, and it is not the same as knowing who will succeed. Any vendor promising the second is describing something the research does not support.
Where to start
Before evaluating any assessment, write the benchmark for your highest volume role. Which traits actually predict success there, according to the people already doing the job well.
Then ask any vendor two questions. Which validity figures are you quoting and from what year. And what population was the instrument normed on. If either answer is vague, that tells you more than the sales deck will.
If the harder problem is applying any standard consistently across a busy month, that is worth a conversation. Book a call with our team and we will look at where your process currently makes its decisions.

