compose, don't copy
Two Thousand Jobs, No Two Thousand Studies
Defensible hiring criteria for jobs you can't afford to study one at a time — the forty-year-old method that composes validity instead of copying it.
A mid-size company has maybe two thousand distinct jobs. The gold-standard way to know whether a selection tool predicts performance in any one of them is a local validation study: hire people, measure how they do, correlate it with the tool. Run that honestly and it takes months and a sample size most individual jobs will never have. Two thousand jobs, two thousand studies, is not a plan. It's a way of saying you'll never have defensible criteria for anything but the three roles you hire a hundred of.
So companies do one of two things. They validate nothing and hope, or they borrow a meta-analytic average and treat it as local truth — the mistake I've called borrowed validity elsewhere. Both are bad, and the existence of only-two-bad-options is usually a sign the field has a third one you haven't been shown.
It does. It's about forty years old, it has an ungainly name, and it's the quiet engine under any credible attempt to put defensible criteria on jobs you can't afford to study one at a time.
They say you can't validate what you can't study
The objection is real and worth stating at full strength: validity is a property of a tool for a specific job, measured against that job's performance. You can't just assert it travels. That's the whole reason local validation is the gold standard — and the reason "we borrowed the number from a vendor deck" is not.
But look at what actually varies between two jobs. Not the raw title — the components. A job is a bundle of work: the reasoning it demands, the interpersonal load, the physical and procedural pieces. And the accumulated validation literature already tells you, with decades of evidence, how well a given predictor works for each of those components. If you know a job's component profile, and you know how a predictor performs against each component, you can compose the two into an estimate of how the predictor works for that job — without ever running a study on that job specifically.
That's synthetic validity, also called job-component validity, and the key word is synthetic: you synthesize the job's validity from parts whose validity is already established.1
Why it holds up
The reasonable next question is whether the composed estimate is any good — whether "we built it from components" is a euphemism for "we made it up." The evidence says it holds.
When researchers put synthetic coefficients head to head with traditional, locally-derived ones — across eleven job families, twenty-seven components, roughly two thousand people — the synthetic estimates land in the same neighborhood as the real thing.2 The extrapolation isn't a hand-wave; it recovers what a local study would have found. And critically for anyone trying to do this at portfolio scale, it works on standardized job-component data — the kind of public, structured description of work that O*NET provides — with validities recovered in the .35 to .89 range depending on the component.3 That last point is what turns this from a boutique method into a system: if your jobs are described in a consistent component vocabulary, you can extrapolate validity across all of them at once.
There's even a defensibility lineage. Synthetic validity isn't a fringe technique — it's been recognized in the U.S. selection-guidelines tradition for decades under the "J-coefficient."4 (That's context, not legal advice; the guidelines are U.S.-specific and any real compliance question needs a real lawyer.) The point is that "we derived these criteria synthetically" is a sentence you can say to a skeptical procurement reviewer without flinching.
The pattern worth naming
Call it compose, don't copy. The borrowed-validity mistake copies a whole number from one setting to another and prays the settings match. Synthetic validity does the harder, more honest thing: it breaks the job into components, uses established evidence at the component level, and composes an estimate specific to this job. The unit of borrowing shrinks from "the whole job" — which never matches — to "the component," which is exactly what generalizes.
This is why a canonical job reference is worth building rather than buying a title-matcher. If every job in your two thousand carries a consistent, measured component profile — the way JobFrame maps roles into measurable spaces — then defensible criteria stop being a per-job research project and become a property you can derive for the whole catalog. Not two thousand studies. One component vocabulary, and the forty-year-old math that composes it into job-specific answers.
The choice was never "validate everything" or "validate nothing." It was: keep copying numbers you can't defend, or start composing ones you can.
Companion piece: Borrowed Validity — on why a meta-analytic average is a starting estimate, not a local verdict.
Footnotes
-
Scherbaum (2005) overview; Steel, Huffcutt & Kammeyer-Mueller (2006) how-to — validity and criteria can be formally inferred from job components when local studies are infeasible. ↩
-
Johnson & Carter (2010), Personnel Psychology — empirical head-to-head: synthetic coefficients ≈ traditional within-family coefficients, across 11 families / 27 components / N≈1,926. Estimator refinement (meta-analytic WLS) in Steel et al. (2009). ↩
-
Jeanneret & Strong (2003), Personnel Psychology — job-component validity demonstrated on O*NET data specifically, R=.35–.89 from generalized work activities. ↩
-
Trattner (1982) on the J-coefficient's recognition under the Uniform Guidelines; transportability/validity-generalization context in Hoffman & McPhail (1998). U.S.-specific; context, not legal advice. ↩