get2great

the calibration inversion

Structure Is the Tool

The interview is your best selection instrument — and most companies use it wrong on purpose. Structure roughly doubles its accuracy.

The best selection instrument most companies own is the job interview, and most of them are using it wrong on purpose.

Not out of laziness. Out of taste. The unstructured interview — the open conversation, the read on "fit," the gut call at the end — feels like the humane part of hiring, the place where judgment lives and the résumé stops mattering. Structure feels like the opposite: a script, a rubric, a checklist that reduces a person to a set of scored answers. So the field has spent decades treating structure as the thing you add when you don't trust the interviewer, and conversation as the thing you do when you do. The instinct is exactly backward, and the size of the mistake is measurable.

Here is the number that should end the argument. Put structure on an interview — the same questions, in the same order, scored on the same anchored scale — and its validity roughly doubles against the unstructured version.1 Not improves at the margin. Doubles. Two large, independent meta-analyses, one of them pooling 245 validity coefficients across more than eighty thousand people, land in the same place: the conversation you trust is worth about half of the protocol you resist.

They say the interview is where judgment goes

The defense of the open interview is that hiring is about people, and people don't fit rubrics. You have to talk to someone. You have to see how they think on their feet, whether they'd get along with the team, whether there's something the paper doesn't show. Structure, the argument goes, sands all of that off — it optimizes for the candidate who answers the scripted question well, not the one who'd actually do the job.

It's a good-sounding argument, and it survives because nobody keeps score. The unstructured interview is the one selection step in the whole funnel that never gets validated, because it can't be: no two interviews are the same interview, so there's nothing stable to check. The manager who "just knows" after twenty minutes is running an instrument with no calibration, no record, and no way to be wrong out loud. It feels like judgment. What it mostly is, is noise wearing judgment's clothes.

What structure actually does

Structure isn't the enemy of judgment. It's the thing that lets judgment repeat.

Give every candidate the same questions and you can finally compare them to each other instead of to the interviewer's mood. Anchor the rating scale — this answer looks like a 2, this one like a 4, here's what each looks like — and two interviewers start agreeing with each other instead of each running a private scoring system in their head. The reliability jump is not subtle: separate single interviewers agree at about .44, a panel working from anchored scales at about .74.2 That gap — .44 to .74 — is the difference between a coin you're pretending is a compass and an instrument you can actually steer by.

And structure doesn't mean interrogation. The research is specific about what works, and it's more interesting than "read from a card." The strongest single format is the past-behavior question — tell me about a time you had to do the thing this job requires — scored against descriptive anchors; it clears the situational "what would you do" format by a wide margin (validity around .63 against .47).3 The reason is almost obvious once named: what someone actually did under real constraints predicts what they'll do better than what they say they'd do in a hypothetical where they know you're listening. Behavior is the tell. The script is just what keeps you asking every candidate for it.

The deeper reason structured interviews win is that they force you to interview for the constructs that predict the work — the specific skills and dispositions the job actually loads on — instead of the free-floating impression of competence a good talker generates.4 Structure is how the interview stops measuring charisma and starts measuring the job.

The pattern worth naming

Call it the calibration inversion: the part of hiring that feels the most like expert judgment is the part with the least evidence behind it, and the part that feels the most mechanical is the part doing the actual predicting. Every org has this backward somewhere. The senior leader's "I've hired hundreds of people, I know one when I see one" is the tell — it's the exact confidence the record does not support, defended most fiercely by the people with the least reason to.

You don't fix it by trusting people less. You fix it by giving the trust something to stand on: the same questions, past-behavior first, anchors on every scale, a panel that scores before it confers. It is not more bureaucracy. It is the difference between an interview you can defend when a rejected candidate's lawyer asks how you decided, and one where the honest answer is he reminded me of me.

The interview is already your best tool. Structure is what turns it from a feeling into an instrument.

Footnotes

  1. Wiesner & Cronshaw (1988); McDaniel, Whetzel, Schmidt & Maurer (1994), 245 coefficients across N=86,311. Two independent large meta-analyses converge on structure roughly doubling validity over unstructured formats.

  2. Huffcutt, Culbertson & Weyhrauch (2013) on panel-vs-single reliability (.74 vs .44); Taylor & Small (2002) on anchored scales lifting reliability to ~.77–.79.

  3. Taylor & Small (2002): past-behavior questions with descriptive anchors at ρ≈.63 vs situational-with-anchors at ρ≈.47; converging evidence in Pulakos & Schmitt (1995).

  4. Huffcutt, Conway, Roth & Stone (2001), construct taxonomy meta-analysis — structured interviews concentrate on constructs that actually predict performance. Full design playbook: Levashina, Hartwell, Morgeson & Campion (2014), Personnel Psychology.