The request usually arrives from finance, phrased as a return on investment question, and it is a fair question that the field has spent seventy years failing to answer cleanly. What follows is what can be measured, what cannot, and where the evaluation effort is best spent.
The four levels, and what they are not
Almost every evaluation framework in use descends from the four levels Donald Kirkpatrick set out in 1959, and the 2017 meta-analysis of leadership training organises its findings around them.
| Level | The question | Usually measured by |
|---|---|---|
| Reactions | Did participants value it? | End-of-programme survey |
| Learning | Do they know or can they do more? | Test, assessment, self-rating |
| Transfer | Are they behaving differently at work? | Ratings from colleagues, observation |
| Results | Did team or organisational outcomes move? | Business metrics |
The framework is frequently presented as a ladder, in which each level causes the next. It is not, and this is the single most consequential misunderstanding in training evaluation.
The meta-analysis reports corrected effects of 0.63 for reactions, 0.73 for learning, 0.82 for transfer and 0.72 for results. Those are four separate estimates of four separate things. Nothing in them establishes that a participant who rated the programme highly is more likely to have learned from it, or that one who learned is more likely to behave differently.
The moderator analyses make the independence visible. Programmes that included feedback produced significantly stronger transfer and showed no significant difference on reactions, learning or results. Programmes based on a needs analysis produced significantly stronger learning and transfer but no significant difference on results. Spaced programmes produced stronger transfer and results with no difference in learning. In each case a design choice moved some levels and not others.
The strength of these effects differs based on various design, delivery, and implementation characteristics.
Which means a programme can score well at one level and fail at the next, and the end-of-course questionnaire cannot tell you which happened.
The measurement everyone does is the weakest one
Level one is measured almost universally because it is nearly free: the participants are in the room, they have five minutes, and the form is already printed.
It is also the level with the loosest connection to anything an organisation cares about. A high score establishes that a programme was well run and enjoyable, which is worth knowing and is not the reason it was funded.
The more uncomfortable version of this point comes from the attendance finding in the meta-analysis. Voluntary programmes produced significantly stronger transfer than mandatory ones, while mandatory programmes produced significantly stronger results. Voluntary attendance selects people who already wanted to change, which raises both satisfaction scores and the odds of behaviour change, and neither of those is caused by the programme. An evaluation that compares a voluntary cohort's scores to a benchmark is partly measuring who signed up.
What a defensible measurement looks like
Three things separate an evaluation that survives scrutiny from one that does not, and none of them is expensive.
A baseline taken before the programme. Not a retrospective rating of how things were, which is reconstructed after the fact and contaminated by the experience of the programme. The same instrument, administered to the same raters, before anything happens.
A comparison group. The single biggest improvement available to most evaluation designs, and usually free, because leadership programmes are almost always rolled out in waves. The people scheduled for the next cohort are a ready-made comparison, matched on the criteria that got them selected. Measuring both groups at the same two points is the difference between observing that scores went up and knowing whether the programme moved them.
Behaviour rated by someone other than the participant. Self-ratings of one's own leadership behaviour after a leadership programme are the least reliable measure in the set. Ratings from direct reports are the ones that carry weight in a room where the number is being contested.
Why the results level rarely works
Level four is what finance actually wants and it is the hardest to deliver honestly.
The meta-analysis is itself constrained here. Several moderators could not be tested against results because too few primary studies measured them, and the comparison between face-to-face and virtual delivery could not be assessed at the results level at all for the same reason. If the published research struggles to get enough observations at this level, an internal evaluation of one cohort will not do better.
The specific difficulties are structural. Team and unit outcomes move for reasons unconnected to the manager, on timescales longer than the programme, and attributing a share of them to a two-day workshop requires assumptions that will not survive a sceptical reading. Where a results-level claim is made and defended, it is usually because the organisation ran waves and compared them, which is the comparison group point again.
An honest alternative is to state the chain explicitly: the programme changed these behaviours, measured this way, and these behaviours are associated with these outcomes in the research literature. That is a weaker claim than a calculated return and it has the advantage of being true.
Measure the design, not just the outcome
There is a cheaper move available, and for most organisations it is the higher-value one.
The meta-analysis identifies which design features predict transfer, across 335 samples. A formal needs analysis. Information, demonstration and practice combined rather than any one alone. Sessions spaced across weeks rather than compressed into a block. Feedback during the programme. These are the features that separated programmes that changed behaviour from those that did not.
Auditing whether a programme has those features costs an afternoon and requires no instrument at all. It will not tell you what the programme achieved. It will tell you, before it runs, whether it belongs to the category of programmes that achieve things, and that is available at the point where it can still be changed.
Given that the money spent per trained person has fallen about a quarter in real terms since 2005, and that practice and spacing are the first things cut when hours shrink, this audit frequently finds the answer before any evaluation is run. A programme with no needs analysis, delivered in one block, with no practice and no feedback, has predictable results and does not need measuring to establish them.
One design feature is worth checking against the evidence rather than assumed: 360-degree feedback showed no significant advantage over single-source feedback on learning, transfer or results, which makes it an expensive component to justify in an evaluation plan.
Common questions
How do you measure the ROI of leadership development? A defensible financial return is rarely achievable, because organisational outcomes move for many reasons over longer timescales than the programme. What is achievable is a measured change in behaviour against a comparison group, with the link to outcomes stated explicitly rather than calculated.
What are the four levels of training evaluation? Reactions, learning, transfer and results, set out by Donald Kirkpatrick in 1959. They measure whether participants valued the programme, whether they learned, whether their behaviour changed, and whether organisational outcomes moved.
Do the four levels predict each other? No. They are four separate measures. In the meta-analytic evidence, design features that improved one level frequently made no difference to the others.
Why are end-of-course surveys a weak measure? They measure reactions, which is the level with the loosest connection to behaviour change or organisational outcomes. They are used because they are nearly free, not because they are informative.
What is the cheapest way to improve an evaluation? Use the next cohort as a comparison group. Leadership programmes are almost always run in waves, so a matched comparison group already exists and costs nothing to measure.
Should participants rate their own behaviour change? Not on their own. Self-ratings of leadership behaviour after a leadership programme are the least reliable measure available. Ratings from direct reports carry more weight.
Can we predict whether a programme will work before running it? Partly. Across 335 samples, programmes with a needs analysis, combined information, demonstration and practice, spaced sessions and in-programme feedback produced significantly stronger behaviour change. Auditing a design against that list costs almost nothing.
Sources
- Lacerenza, Reyes, Marlow, Joseph & Salas (2017), Leadership training design, delivery, and implementation: A meta-analysis, Journal of Applied Psychology 102(12)
- Full text of the leadership training meta-analysis, Doerr Institute, Rice University
- Eurostat, Cost of CVT courses per participant (trng_cvt_19s)



