get2great

the adjective tax

Grading on Vibes

What behaviorally-anchored criteria actually buy you — not validity, but shared meaning, usable feedback, and a decision you can defend.

Ask a manager what a "4 out of 5" means and you will get a pause, and the pause is the whole problem.

The rating scale is the most-used measurement instrument in any company. Every performance review, every interview scorecard, every calibration meeting runs on it. And on most of them, the anchors are adjectives: 5 — outstanding, 4 — exceeds, 3 — meets, 2 — developing, 1 — needs improvement. Which means the scale measures whatever "exceeds" means to the person holding the pen that afternoon. Two evaluators watching the same work will land two numbers apart, not because the work was ambiguous, but because the instrument was. You cannot average your way out of a scale where nobody agrees what the middle is.

The fix has been sitting in the personnel-psychology literature for fifty years, and it's almost aggressively unglamorous. Replace the adjectives with behavior. Instead of "4 — exceeds expectations," the anchor reads: "anticipates the downstream team's needs and adjusts the handoff without being asked." That's a behaviorally-anchored rating scale — BARS — and the point of it is not that it makes the score more scientific. It's that it makes the score mean the same thing to two different people.

They say anchors just dress up the same subjectivity

The skeptical read on BARS is that it's theater: you're still asking a human to judge a human, you've just made them read a paragraph first. And there's a version of that critique the evidence actually supports — so let me concede it before defending the rest, because the honest version of this argument is the persuasive one.

BARS does not reliably beat other rating formats on raw psychometrics. Decades of comparison studies land at rough parity — anchored scales are not meaningfully more valid or more reliable than a well-built graphic scale in a controlled test.1 If someone sells you BARS as a validity upgrade, they're overclaiming, and you should discount the rest of their pitch.

What BARS actually buys you is somewhere else, and it's worth more than the psychometric wash suggests.

What the anchors actually buy

Three things, none of them a validity coefficient.

The first is shared meaning. When the anchor is a behavior, two raters are pointing at the same referent instead of two private definitions of "exceeds." The score becomes a claim about what the person did, checkable against what the anchor says, rather than a claim about how they came across.

The second is feedback that can travel. A "3" tells someone nothing. "You meet the standard but you wait to be asked before adjusting the handoff" tells them exactly what a "4" would look like. The anchors turn a rating into a development map — the same reason BARS has survived fifty years of people predicting its death: managers keep it because it's the rare instrument whose output a human can act on.2

The third, and the one that matters most when the stakes are legal, is defensibility. When a promotion or a termination gets challenged, "the committee felt he was a 2" is not an answer. "He was rated against these specific behavioral criteria, and here is where his documented work fell short of the anchor" is. Behavioral anchors are the difference between a decision you made and a decision you can show your work on.

And they don't require a workshop

The classic objection to BARS is cost: building the anchors means assembling subject-matter experts to brainstorm "critical incidents" for every job, a workshop-heavy exercise most companies do once, for a few roles, and never repeat. That objection is now dated. The anchors can be generated from a corpus of real critical incidents — the accumulated record of what good and bad performance in a role actually looks like — rather than a conference room of managers guessing.3 The scarce input was never the judgment; it was the incidents. Once you have those at scale, the anchors are a derivation, not an offsite.

The pattern worth naming

Call it the adjective tax. Every scale anchored in adjectives instead of behaviors charges you the same hidden fee: rater disagreement you can't diagnose, feedback nobody can use, and decisions you can't defend — paid on every review cycle, invisibly, forever. Behavioral anchors don't make the judgment more scientific. They make it legible — to the second rater, to the person being rated, and to whoever asks you later how you decided.

Stop grading on vibes. Write down what the 4 looks like.

Footnotes

  1. Jacobs, Kafry & Zedeck (1980), Personnel Psychology — BARS at rough psychometric parity with other formats. The honest positioning is usability and defensibility, not superior validity/reliability.

  2. Debnath & Lee (2015) on the 50-year persistence of BARS despite repeated predictions of its obsolescence — utilization and qualitative-criterion advantages (feedback usefulness, defensibility) explain the survival.

  3. Kell, Martin-Raugh, Carney, Inglese, Chen & Feng (2017), ETS Research Report — feasibility of building behavioral anchors from crowdsourced/corpus critical incidents rather than SME-panel workshops.