Your Performance Rating Is Decided Without You [2026]

Signal vs. Noise · Advancement

Calibration Week: Your Rating Was Decided in a Room You've Never Seen

Ninety minutes. Twenty names on a screen. You get about four of those minutes, and you're not there for them.

By Scot Free

Sometime in the next few weeks, at most companies running a calendar-year cycle, your manager will submit a draft rating for you. That draft is not your rating. It's an opening position.

What happens next is called calibration, and almost nobody outside of management and HR has ever had it explained to them.

What actually happens in the room

The mechanics are not secret. They're documented in HR practice literature, and they're more mundane than the mythology around them.

  • Two phases. First, each manager scores their direct reports against the rubric — informed by your self-assessment, goal data, and any 360 input — and produces a draft. Then a panel of managers from the same business unit or function convenes to review those drafts collectively.
  • The sequence. SHRM describes the core of it as managers posting names and proposed ratings where everyone can see them, then discussing each one and adjusting for consistency.
  • The clock. SHRM puts an effective session at about 90 minutes, and practice guidance puts a group at 15 to 25 employees. Work the division: roughly four to six minutes per person, and that includes the ones nobody debates.
  • The wrapper. The full process — HR compiling distribution data beforehand, the session, then documenting what changed — typically spans two to three weeks.

Four to six minutes. One screen. A room of people who manage different teams and have, in most cases, never watched you work.

Why it exists, and the case for it is strong

Before the criticism, the honest part: calibration exists because the alternative is demonstrably worse.

  • Managers rate the same performance differently, by a lot. CEB/Gartner research puts the variance between lenient and strict managers at roughly ±30% without calibration — meaning your rating depends more on who your manager is than on how well you performed.
  • And that's not a rounding error in the model. Research by Kevin Murphy at Colorado State University found that rater tendencies — whether a manager runs lenient or strict — account for more variance in ratings than actual performance differences between employees do.
  • Calibration measurably helps. CEB/Gartner finds managers rating independently are three to five times less accurate than managers who calibrate with peers.

Sit with the Murphy finding for a second, because it reframes the whole thing. The largest single variable in your performance rating, before anyone calibrates anything, is which manager you happen to report to. Calibration is the correction for that. It is a real correction and it works.

It also introduces a different problem.

What the research says goes wrong

A 2025 analysis of tournament-based evaluation systems examined calibration as a proposed remedy for forced ranking and found a specific failure mode:

Calibration sessions in practice devolve into negotiation. Managers report spending hours debating the relative merits of employees they have never met, relying on presentations from their peers.

And the consequence the authors draw from it: advocacy replaces assessment. Managers who are skilled negotiators secure better outcomes for their reports. Errors stop being random and become systematic, because the most persuasive managers consistently win.

That is uncomfortable and it is also obviously true to anyone who has sat in one. When the only information in the room about a person is a slide and a colleague's verbal account, the quality of the account matters.

Two structural pressures make it worse:

  • Distribution targets, where they exist. Some organizations cap the top rating — no more than 15% of employees, say, or a 10/70/20 split. HR practice literature is emphatic that calibration is not forced ranking and that managers shouldn't move people just to produce a curve. The same literature keeps saying it, which tells you how often it happens. Where a cap exists, every top rating argued successfully is one argued away from someone else.
  • Managers find it genuinely hard. Research going back to Schleicher and colleagues in 2009 finds managers perceive forced distribution as more difficult and less fair than traditional systems, with greater stress and role conflict — especially when they believe their team members are uniformly strong.

The four minutes

Here's where this stops being trivia and starts being actionable.

You cannot be in the room. You cannot see the slide. You have no vote and in most companies you will never be told whether your rating was changed. What you can affect is what your manager is able to say during those four minutes, and the research tells you exactly what wins: evidence, delivered by someone who can state it clearly.

A manager defending a rating has two kinds of material available. One works and one doesn't:

What your manager says What the room hears
"She's been great this year. Really strong performer, everyone likes working with her."A manager who likes their employee. Every manager in the room likes their employee. This is noise.
"March through June she led the vendor consolidation — took the close reporting cycle from nine days to four. Treasury asked her to extend the model to their two business units in August."Dates, a number, and a second department's independent action. A peer can push back on an opinion. It's harder to push back on August.

That second version is what the self-review is for. It isn't a reflection exercise — it's the raw material your manager carries into this meeting. Write it so the three strongest items are compressed into sentences a tired person can read aloud under time pressure, and they will get read aloud, because people under time pressure use the words in front of them.

The practice literature backs the mechanism: good sessions require managers to submit written justifications and objective work metrics in advance, and to return to rating definitions and job-related evidence whenever a discussion gets subjective. Evidence is the currency. You can mint some of it.

Three things you can actually do

1. Ask whether your organization calibrates, and at what level. This is not a sensitive question and most managers will answer it. Knowing whether your rating is set by your manager alone, by a panel in your function, or by a panel at the division level tells you how far your evidence has to travel. Ask it in a one-on-one, casually, before the cycle closes.

2. Give your manager three sentences, unprompted. Not a campaign — three lines with a date, a number, and a name in each, in an email they can find during the meeting. Frame it as making their job easier, which is also true. A manager preparing twenty write-ups in a week will remember who handed them usable material.

3. Read your manager's advocacy honestly. This is the uncomfortable one, and it's the one with the most leverage. If the research says persuasive managers win and yours is quiet, conflict-averse, or new to the room, that is a real fact about your situation and it has nothing to do with your work. It is also a legitimate reason to make sure your work is visible to someone else in that room — a skip-level, a peer manager you've delivered for, a function lead who's seen the output. Not politics. Redundancy.

And note where this sits in the season. Next year's seats were argued in a planning room in October. Your rating gets argued in a calibration room in November. Same structure both times: a decision made somewhere you aren't, with a document standing in for you. The only question that ever matters is what's in the document.

The Scot Free Take

I have been in this room. The honest version from the inside is that it isn't a conspiracy and it isn't rigged; it's twenty people in ninety minutes doing their best with thin information, and thin information is exactly the condition under which the loudest well-prepared voice wins.

The finding I'd want people to take away is Murphy's, because it reframes a lot of private grief. If rater tendency explains more variance in ratings than actual performance differences do, then a middling rating is not necessarily a verdict on your work. It may be a verdict on your manager's calibration habits. That's not a license to dismiss feedback, but it is permission to stop taking a number as gospel.

What you do about it is unglamorous. You cannot get into the room. You can make sure the person who goes in is carrying something they can read out loud.

Get the MoneyZoo LAB Report and instant access to The Side Door Playbook. One seat, one receipt, one move, every Monday. Subscribe here.

Sources

  • CEB/Gartner (2023), as cited in HR reference material — ±30% rating variance between lenient and strict managers absent calibration; managers rating independently 3–5x less accurate than those calibrating with peers.
  • Kevin Murphy, Colorado State University — rater tendencies accounting for more variance in ratings than actual performance differences between employees.
  • SHRM, Improving Performance Evaluations Using Calibration — the session sequence (managers post names and proposed ratings, then discuss and adjust); ~90-minute effective session duration.
  • "Tournament-Based Performance Evaluation and Systematic Misallocation: Why Forced Ranking Systems Produce Random Outcomes" (arXiv, 2025) — calibration devolving into negotiation; advocacy replacing assessment; errors becoming systematic rather than random. Cites Schleicher et al. (2009) on manager perceptions of forced distribution.
  • Engagedly, StaffCircle, Deel and PerformSpark practice guidance — two-phase structure, 15–25 employees per group, 60–90 minutes, pre-submitted written justifications and objective metrics, distribution model definitions (10/70/20, 15% caps), and the repeated instruction that calibration is not forced ranking.
  • Cornell ILR School research, as cited in performance management reference material — manager-level variance accounting for a meaningful proportion of rating variance independent of true performance differences.

A note on these sources: the mechanics above come from HR practice literature and vendor guidance, which describes how calibration is supposed to run rather than how it always runs. The two load-bearing research findings — Murphy on rater variance and the 2025 analysis on advocacy — come from academic work, and they point in different directions about whether calibration helps. Both are reported here because both are true.

Previous
Previous

Your Manager Is Your Agent. Give Them Something to Sell [2026]

Next
Next

Your Self-Review Is Evidence. Write It Like Evidence [2026]