Leadify — AI leadership simulation platform

How to Assess Promotion Readiness: An Evidence-Based Framework

A defensible framework for deciding who is ready for the next role — using behavior at the target level, not tenure or performance at the current one.

11 min read Updated Jan 2026
Key takeaways
  • Promotion is a bet that a person can perform the next role, not a reward for the current one. Confusing the two produces the Peter Principle at scale (Peter & Hull, 1969; Benson, Li & Shue, 2019).
  • Readiness must be measured against the target-level competency profile, not extrapolated from current performance.
  • A defensible framework combines observed behavior under target-level pressure, gap analysis, and confidence intervals — not manager narrative alone.
  • Benson, Li & Shue (2019) found the best salespeople become worse-than-average managers after promotion, and firms still promote them — evidence the current-role signal is the wrong one.

Why does promoting your best performer often fail?

Most promotion decisions rest on a fallacy: strong performance in the current role predicts strong performance in the next. It often does not. Benson, Li & Shue (2019) show the best salespeople become worse-than-average managers because the required capabilities are different — the Peter Principle, quantified.

A defensible readiness process treats promotion as a forecast about a different job, evaluated against that different job's competency profile — not as a reward function on the current one.

What should a target-level competency profile contain?

A target-level competency profile is an explicit rubric describing what a strong performer at the next level does that the current level does not. For first-line management this typically covers difficult feedback, prioritization under conflicting demands, and cross-functional influence. Without it, promotion decisions have no defensible measurement target.

For senior leadership, strategic ambiguity, board communication, and organizational design dominate. Every rubric should be documented, versioned, and mapped to observable behaviors — so raters can agree and disagreements can be adjudicated.

How do you measure behavior at a level someone hasn't held yet?

You measure it by simulation. Run scenarios calibrated to the target level — a board challenge for a VP candidate, a layoff conversation for a director candidate — and score behavior against the target-level rubric. Simulation is the only scalable way to observe target-level capability in people who have never held the role.

The core measurement question is not "is this person a strong current-role performer?" but "how do they behave in target-level moments?" Anything that answers only the first question is retrospective; readiness requires a forward-looking signal.

How do you turn readiness evidence into a defensible verdict?

Convert observed evidence into a three-way verdict using the point estimate and the confidence interval. Ready = both the estimate and the lower bound clear the target-level threshold. Ready with caveat = point estimate clears, lower bound does not. Not ready = point estimate below threshold, with a development plan and re-measurement date.

  • Ready: point estimate and lower bound of the confidence interval both clear the target-level threshold.
  • Ready with caveat: point estimate clears the threshold but the lower bound does not; specific coachable gaps identified.
  • Not ready: point estimate below threshold; development plan with re-measurement date.
  • Every verdict is written against evidence, not vibes.
See how Leadify does this

The measurement stack behind this guide.

Read the methodology page for the calibration, scoring, and validity work underneath — or book a demo to see the numbers on your own scenarios.