- Satisfaction surveys ("happy sheets") measure whether learners enjoyed training, not whether capability changed; the correlation with downstream performance is near zero (Alliger et al., 1997).
- Kirkpatrick Level 3 (behavior change) has always been the gold standard; it was expensive to measure at scale until simulation (Kirkpatrick & Kirkpatrick, 2006).
- Pre/post capability measurement produces Cohen's d effect sizes that finance and the CEO recognize as ROI (Cohen, 1988).
- The right L&D metric is not completion rate — it is effect size on measurable business-relevant capabilities.
Why are learner satisfaction surveys not ROI?
Post-training satisfaction surveys measure whether learners enjoyed the session, not whether their behavior changed. Alliger et al. (1997) show the correlation between reaction-level scores and downstream job performance is near zero. Reporting happy-sheet averages as ROI is a category error — it measures the wrong construct.
Every L&D leader knows this. The reason satisfaction surveys persist is that measuring behavior change used to be expensive — until simulation brought the marginal cost down by an order of magnitude.
What is Kirkpatrick Level 3 and why does it matter now?
Kirkpatrick Level 3 measures behavior change on the job — the level the Kirkpatrick model has always identified as most impactful and hardest to measure (Kirkpatrick & Kirkpatrick, 2006). Simulation-based assessment turns Level 3 into a repeatable, scalable measurement using the same calibrated rubric before and after the program.
Most L&D functions report Level 1 (reaction) and Level 2 (learning) because those were the only levels affordably measurable. Level 3 at scale reframes the entire budget conversation from anecdote to evidence.
What effect size counts as a meaningful L&D outcome?
Cohen's d is the standard metric for pre/post capability change: d≈0.2 is small, d≈0.5 is moderate, d≈0.8 is large (Cohen, 1988). Reporting L&D outcomes in effect sizes rather than completion rates lets the CFO compare the return on training against any other capital allocation using a common statistical language.
It also — for the first time — lets L&D defend its budget in a downturn with something other than testimonials.
How do you set up the pre/post measurement loop?
Run a four-step loop: baseline the target capabilities with calibrated simulations before the program, deliver the intervention, re-measure with the same (or parallel-form) scenarios, and report pre/post Cohen's d per competency with confidence intervals at cohort and individual level. The loop is auditable and reproducible.
- Baseline: measure target capabilities with calibrated simulations before the program starts.
- Intervention: run the L&D program.
- Re-measure: run the same scenarios (or parallel-form scenarios calibrated to the same difficulty).
- Report: pre/post effect size per competency, with confidence intervals, at the cohort and individual level.
The measurement stack behind this guide.
Read the methodology page for the calibration, scoring, and validity work underneath — or book a demo to see the numbers on your own scenarios.
