Every assessment vendor, including this one, describes its tests as “validated.” The word gets used constantly and explained rarely. A validated assessment is one with documented statistical evidence that its scores predict or relate to actual job performance, established through methods that professional and legal standards specifically recognize, not a claim based on the test looking professional or having been used for a long time.
This piece explains what that evidence actually looks like: the three recognized types of validity, what a validity coefficient means in practice, why a small business doesn’t need to run its own study, and what questions to ask when a vendor uses the word “validated” without backing it up.
What “Validated” Means, in One Sentence
An assessment is validated when there is documented, empirical evidence that scores on it are statistically related to job performance, gathered through one of a small number of recognized methods and, where legally required, formally documented.
That definition comes with real teeth. The U.S. Equal Employment Opportunity Commission’s Uniform Guidelines on Employee Selection Procedures (29 CFR Part 1607) name the acceptable evidence types explicitly, and the Society for Industrial and Organizational Psychology’s Principles for the Validation and Use of Personnel Selection Procedures, an official statement of the American Psychological Association, sets the professional standard for how that evidence should be gathered. “Validated” isn’t a marketing adjective. It’s a claim with a specific, checkable meaning.
The Three Types of Validity Evidence
The Uniform Guidelines recognize three ways to demonstrate validity, and a rigorous assessment typically leans on at least one of them with real evidence behind it.
Criterion-Related Validity
This is the most direct form: empirical data showing the test score is statistically correlated with an important measure of job performance, sales numbers, supervisor ratings, error rates, whatever the role’s actual output looks like. It splits into two variants: concurrent validity, which tests current employees and compares their scores to their existing performance, and predictive validity, which tests candidates before hire and checks the correlation once performance data exists. Predictive validity is the stronger evidence of the two, since it mirrors exactly how the test will actually be used.
Content Validity
Content validity asks a narrower question: does the test actually cover the important tasks of the job? It’s established through a formal job analysis that maps test items directly to job duties, which is why a typing test for a data entry role is relatively easy to validate this way, and why a personality inventory almost never can be, personality isn’t a job task in the way typing is.
Construct Validity
Construct validity is the most complex of the three: it requires showing that a test measures the underlying trait it claims to measure, conscientiousness, cognitive ability, whatever the construct is, and that the construct itself is genuinely related to job performance. The Uniform Guidelines describe construct validation as an extensive, multi-study undertaking, which is part of why most assessment vendors lean more heavily on criterion-related evidence in practice.
| Validity Type | What It Establishes | Typical Method |
| Criterion-related | Scores statistically correlate with actual job performance | Compare test scores to performance data, current or future employees |
| Content | The test content reflects real, important job tasks | Job analysis maps test items directly to job duties |
| Construct | The test measures the underlying trait it claims to measure | Multiple studies linking the construct to job-relevant behavior |
What a Validity Coefficient Actually Tells You
Validity evidence usually gets reported as a correlation coefficient, written as r, ranging from 0 (no relationship between the score and job performance) to 1 (a perfect relationship, which never happens with human behavior). A widely cited meta-analysis by Schmidt and Hunter, published in Psychological Bulletin, found a validity coefficient of 0.54 for work sample tests and job simulations, the highest of any single selection method they measured, with general cognitive ability tests and structured interviews also scoring well above less formal methods.
Context matters here more than the raw number suggests. In personnel selection research, a coefficient of 0.30 to 0.40 is considered a genuinely useful predictor, and anything above 0.50 is exceptional. That’s a lower bar than fields like physics use for a strong relationship, but human job performance has far more sources of noise, motivation, team dynamics, management quality, than a controlled experiment ever will. A validated hiring assessment isn’t claiming certainty. It’s claiming a real, measurable statistical edge over guessing, which over hundreds of hiring decisions adds up to a large practical difference.
Validity Generalization: Why an SMB Doesn’t Need to Run Its Own Study
Running an original validity study, especially a criterion-related one, requires a job analysis, a meaningful sample of employees, and performance data to correlate against. That’s realistic for a large employer with hundreds of people in the same role. It isn’t realistic for a company with twelve customer service reps.
This is where validity generalization matters. Decades of meta-analytic research, the same kind of work behind the Schmidt and Hunter coefficients, has shown that validity evidence for a given assessment type transfers across similar jobs and settings, rather than needing to be re-proven at every single company. A cognitive ability test with strong, well-established validity evidence for administrative roles in general doesn’t need a from-scratch study at a 30-person company hiring for the same kind of role. This is also explicitly recognized in the Uniform Guidelines, which allow employers to rely on validity evidence from other sources rather than requiring an original study in every case.
Why “Validated” Is a Legal Requirement, Not Just a Best Practice
Under the Uniform Guidelines, if a selection procedure produces adverse impact, a meaningfully lower selection rate for a protected group, commonly measured against the four-fifths rule, the employer is required to produce validity evidence for that specific procedure. “We’ve always used this test” or “the vendor said it works” isn’t documentation. Criterion-related, content, or construct validity evidence is.
This is precisely why the word “validated” carries legal weight beyond its scientific meaning. An assessment with real, documented validity evidence gives an employer something to produce if a selection decision is ever challenged. One without it leaves the employer relying on the test being fair, without a way to demonstrate it.
Five Things “Validated” Does Not Mean
It does not mean the test looks professional. Polished design and a validated score are unrelated. A slick interface says nothing about statistical evidence.
It does not mean it’s been used for years. Longevity isn’t evidence. A test can be widely used and never formally validated for the roles it’s applied to.
It does not mean the vendor says so. A validation claim without a specific coefficient, sample, and method behind it is a marketing statement, not evidence.
It does not mean it passed an internal review. An internal team liking the test’s questions is face validity at best, the weakest and least defensible form of evidence.
It does not mean it’s validated for your specific role. Validity evidence is tied to a job or job family. A test validated for sales roles isn’t automatically validated for engineering ones.
How to Evaluate a Vendor’s Validation Claim
A specific, well-supported validation claim can survive a few direct questions. A marketing claim usually can’t.
Ask for the actual coefficient. A real answer sounds like “0.4 criterion-related validity for this role family.” A real answer is never just “yes, it’s validated.”
Ask which type of validity evidence backs it. Criterion-related, content, or construct, and whether it’s job-specific or generalized from meta-analytic research.
Ask about adverse impact data. A vendor with real validation work will have run, or be able to run, adverse impact analysis on their assessment.
Ask whether the evidence applies to your role. Generalized validity evidence is legitimate, but it should be for a genuinely similar job, not a loosely related one.
Frequently Asked Questions
What does it mean for a hiring assessment to be validated?
What does it mean for a hiring assessment to be validated?
It means there’s documented, empirical evidence that scores on the assessment are statistically related to job performance, established through criterion-related, content, or construct validity methods recognized by the EEOC’s Uniform Guidelines and professional standards.
What is a good validity coefficient for a hiring assessment?
In personnel selection research, 0.30 to 0.40 is considered a genuinely useful predictor and anything above 0.50 is exceptional. Work sample tests, at roughly 0.54 in major meta-analytic research, are near the top of what’s been documented for any single method.
Do small businesses need their own validity study?
Usually not. Validity generalization, supported by decades of meta-analytic research and recognized in the Uniform Guidelines, allows employers to rely on existing validity evidence for a similar role rather than running an original study from scratch.
Is a validated assessment legally required?
Validation evidence becomes a legal requirement specifically when a selection procedure shows adverse impact against a protected group. At that point, the employer must be able to document validity evidence for that procedure under the Uniform Guidelines.
Conclusion
“Validated” is a specific, checkable claim, not a marketing word. It means an assessment has documented statistical evidence, criterion-related, content, or construct, tying its scores to actual job performance, evaluated against standards that both the scientific community and federal law recognize. Understanding what sits behind that word is what separates an assessment worth trusting from one that simply says the right things.





