The Science Behind Predictive Hiring Assessments: What “Validated” Actually Means

Predictive hiring analytics
Search

Try SmoothHiring for free for 14 days

See SmoothHiring in action. Know our features and get insights on how our friendly software helps you with successful hiring. Learn how data and predictive analytics help in hiring the right candidate.

The predictive hiring platform for finding the best employees

Know how predictive hiring helps you find the right people for the right job and increase employee productivity.

Author: SmoothHiring Team

Not finding the best employees?

Schedule a demo to learn how SmoothHiring will help you find the best fit for the job using predictive analytics.

Every assessment vendor, including this one, describes its tests as “validated.” The word gets used constantly and explained rarely. A validated assessment is one with documented statistical evidence that its scores predict or relate to actual job performance, established through methods that professional and legal standards specifically recognize, not a claim based on the test looking professional or having been used for a long time.

This piece explains what that evidence actually looks like: the three recognized types of validity, what a validity coefficient means in practice, why a small business doesn’t need to run its own study, and what questions to ask when a vendor uses the word “validated” without backing it up.

An assessment is validated when there is documented, empirical evidence that scores on it are statistically related to job performance, gathered through one of a small number of recognized methods and, where legally required, formally documented.

That definition comes with real teeth. The U.S. Equal Employment Opportunity Commission’s Uniform Guidelines on Employee Selection Procedures (29 CFR Part 1607) name the acceptable evidence types explicitly, and the Society for Industrial and Organizational Psychology’s Principles for the Validation and Use of Personnel Selection Procedures, an official statement of the American Psychological Association, sets the professional standard for how that evidence should be gathered. “Validated” isn’t a marketing adjective. It’s a claim with a specific, checkable meaning.

The Uniform Guidelines recognize three ways to demonstrate validity, and a rigorous assessment typically leans on at least one of them with real evidence behind it.

This is the most direct form: empirical data showing the test score is statistically correlated with an important measure of job performance, sales numbers, supervisor ratings, error rates, whatever the role’s actual output looks like. It splits into two variants: concurrent validity, which tests current employees and compares their scores to their existing performance, and predictive validity, which tests candidates before hire and checks the correlation once performance data exists. Predictive validity is the stronger evidence of the two, since it mirrors exactly how the test will actually be used.

Content validity asks a narrower question: does the test actually cover the important tasks of the job? It’s established through a formal job analysis that maps test items directly to job duties, which is why a typing test for a data entry role is relatively easy to validate this way, and why a personality inventory almost never can be, personality isn’t a job task in the way typing is.

Construct validity is the most complex of the three: it requires showing that a test measures the underlying trait it claims to measure, conscientiousness, cognitive ability, whatever the construct is, and that the construct itself is genuinely related to job performance. The Uniform Guidelines describe construct validation as an extensive, multi-study undertaking, which is part of why most assessment vendors lean more heavily on criterion-related evidence in practice.

Validity TypeWhat It EstablishesTypical Method
Criterion-relatedScores statistically correlate with actual job performanceCompare test scores to performance data, current or future employees
ContentThe test content reflects real, important job tasksJob analysis maps test items directly to job duties
ConstructThe test measures the underlying trait it claims to measureMultiple studies linking the construct to job-relevant behavior

Validity evidence usually gets reported as a correlation coefficient, written as r, ranging from 0 (no relationship between the score and job performance) to 1 (a perfect relationship, which never happens with human behavior). A widely cited meta-analysis by Schmidt and Hunter, published in Psychological Bulletin, found a validity coefficient of 0.54 for work sample tests and job simulations, the highest of any single selection method they measured, with general cognitive ability tests and structured interviews also scoring well above less formal methods.

Context matters here more than the raw number suggests. In personnel selection research, a coefficient of 0.30 to 0.40 is considered a genuinely useful predictor, and anything above 0.50 is exceptional. That’s a lower bar than fields like physics use for a strong relationship, but human job performance has far more sources of noise, motivation, team dynamics, management quality, than a controlled experiment ever will. A validated hiring assessment isn’t claiming certainty. It’s claiming a real, measurable statistical edge over guessing, which over hundreds of hiring decisions adds up to a large practical difference.

Running an original validity study, especially a criterion-related one, requires a job analysis, a meaningful sample of employees, and performance data to correlate against. That’s realistic for a large employer with hundreds of people in the same role. It isn’t realistic for a company with twelve customer service reps.

This is where validity generalization matters. Decades of meta-analytic research, the same kind of work behind the Schmidt and Hunter coefficients, has shown that validity evidence for a given assessment type transfers across similar jobs and settings, rather than needing to be re-proven at every single company. A cognitive ability test with strong, well-established validity evidence for administrative roles in general doesn’t need a from-scratch study at a 30-person company hiring for the same kind of role. This is also explicitly recognized in the Uniform Guidelines, which allow employers to rely on validity evidence from other sources rather than requiring an original study in every case.

Under the Uniform Guidelines, if a selection procedure produces adverse impact, a meaningfully lower selection rate for a protected group, commonly measured against the four-fifths rule, the employer is required to produce validity evidence for that specific procedure. “We’ve always used this test” or “the vendor said it works” isn’t documentation. Criterion-related, content, or construct validity evidence is.

This is precisely why the word “validated” carries legal weight beyond its scientific meaning. An assessment with real, documented validity evidence gives an employer something to produce if a selection decision is ever challenged. One without it leaves the employer relying on the test being fair, without a way to demonstrate it.

It does not mean the test looks professional. Polished design and a validated score are unrelated. A slick interface says nothing about statistical evidence.

It does not mean it’s been used for years. Longevity isn’t evidence. A test can be widely used and never formally validated for the roles it’s applied to.

It does not mean the vendor says so. A validation claim without a specific coefficient, sample, and method behind it is a marketing statement, not evidence.

It does not mean it passed an internal review. An internal team liking the test’s questions is face validity at best, the weakest and least defensible form of evidence.

It does not mean it’s validated for your specific role. Validity evidence is tied to a job or job family. A test validated for sales roles isn’t automatically validated for engineering ones.

A specific, well-supported validation claim can survive a few direct questions. A marketing claim usually can’t.

Ask for the actual coefficient. A real answer sounds like “0.4 criterion-related validity for this role family.” A real answer is never just “yes, it’s validated.”

Ask which type of validity evidence backs it. Criterion-related, content, or construct, and whether it’s job-specific or generalized from meta-analytic research.

Ask about adverse impact data. A vendor with real validation work will have run, or be able to run, adverse impact analysis on their assessment.

Ask whether the evidence applies to your role. Generalized validity evidence is legitimate, but it should be for a genuinely similar job, not a loosely related one.

It means there’s documented, empirical evidence that scores on the assessment are statistically related to job performance, established through criterion-related, content, or construct validity methods recognized by the EEOC’s Uniform Guidelines and professional standards.

In personnel selection research, 0.30 to 0.40 is considered a genuinely useful predictor and anything above 0.50 is exceptional. Work sample tests, at roughly 0.54 in major meta-analytic research, are near the top of what’s been documented for any single method.

Usually not. Validity generalization, supported by decades of meta-analytic research and recognized in the Uniform Guidelines, allows employers to rely on existing validity evidence for a similar role rather than running an original study from scratch.

Validation evidence becomes a legal requirement specifically when a selection procedure shows adverse impact against a protected group. At that point, the employer must be able to document validity evidence for that procedure under the Uniform Guidelines.

“Validated” is a specific, checkable claim, not a marketing word. It means an assessment has documented statistical evidence, criterion-related, content, or construct, tying its scores to actual job performance, evaluated against standards that both the scientific community and federal law recognize. Understanding what sits behind that word is what separates an assessment worth trusting from one that simply says the right things.

Let us provide you with a detailed tour

Tell us about your problems, and we will present you with the most intriguing choices?

Predictive hiring analytics

Get Started Today

Let us profile your top performers and put together comprehensive WHY data that you can use immediately to hire. Speak with one of our representatives to learn how to save time and money while making dramatically better people decisions:

Get Started Today

Let us profile your top performers and put together comprehensive WHY data that you can use immediately to hire. Speak with one of our representatives to learn how to save time and money while making dramatically better people decisions: