Coaching has an unusual evidence problem. It is not that nobody has studied it. It is that the three serious attempts to pool the studies each drew a different boundary around what counts as coaching, and got a different answer as a result.
Read in sequence, the three tell a coherent story. Read individually, whichever one supports the point being made, they do not.
The three reviews
Theeboom, Beersma and van Vianen, 2014. Eighteen studies, published in the Journal of Positive Psychology. Effects by outcome ranged from 0.43 for coping to 0.74 for goal-directed self-regulation, with performance and skills at 0.60 and work attitudes at 0.54.
Jones, Woods and Guillaume, 2016. Seventeen studies, in the Journal of Occupational and Organizational Psychology. This review deliberately restricted itself to workplace coaching delivered by internal or external coaches, and excluded manager-to-subordinate and peer coaching. Overall effect 0.36, with skill-based outcomes at 0.28, affective outcomes at 0.51 and individual-level results at 1.24.
Cannon-Bowers and colleagues, 2023. Eleven studies, ten after removing an outlier, in Frontiers in Psychology. Overall effect 0.430 with a confidence interval from 0.301 to 0.558. Skill outcomes 0.72, affective outcomes 0.41.
| Theeboom 2014 | Jones 2016 | Cannon-Bowers 2023 | |
|---|---|---|---|
| Studies pooled | 18 | 17 | 11 |
| Scope | Coaching in an organisational context | Workplace coaching only, internal or external coaches | Workplace coaching, tight methodological criteria |
| Headline effect | 0.43 to 0.74 by outcome | 0.36 overall | 0.430 overall |
The pattern is in the inclusion criteria
The 2014 review has the largest numbers and the loosest boundary. Of its eighteen studies, six were general life coaching and one was health coaching. That is not a flaw the authors hid; it is a consequence of asking whether coaching works rather than whether workplace coaching works, and the later reviews name it explicitly as the reason they drew the line differently.
Once the line moves to workplace coaching by a designated coach, the pooled effect drops to around 0.36. Once the methodological bar rises again, it settles at 0.43 with a tight confidence interval on a much smaller set of studies.
This is the ordinary shape of a maturing evidence base and it should be reassuring rather than alarming. The effect does not vanish. It lands, across three independent teams with three different rulebooks, somewhere between 0.35 and 0.45 for workplace coaching in general. By convention that is a small to moderate effect: real, worth buying, and considerably less than the language used to sell it.
Our analyses indicated that coaching had positive effects on organizational outcomes overall, and on specific forms of outcome criteria.
Three things all three agree on
Where the reviews converge is more useful to a buyer than where they differ, because agreement across different samples and different methods is the strongest signal this literature produces.
The number of sessions does not predict the outcome. Theeboom found session count unrelated to effectiveness. Jones found no moderation by number of sessions or by the length of the engagement. Cannon-Bowers found number of sessions not a significant predictor, and hours of coaching not a significant predictor either. Three teams, three datasets, the same null.
That finding is commercially inconvenient and is treated accordingly. Packages are still sold and priced by session count, and what that implies for how coaching should be bought is the most practical conclusion in this whole literature.
The format does not predict the outcome. Jones found no moderation by format, comparing face-to-face coaching with blended face-to-face and e-coaching. Cannon-Bowers compared face-to-face at 0.48 with virtual at 0.35 and found no significant difference. Remote coaching has never had to be defended on evidence; it simply performed the same.
The affective effects are larger than the skill effects, except when they are not. Jones found affective outcomes at 0.51 against skill-based at 0.28. Cannon-Bowers found the opposite, skills at 0.72 against affective at 0.41. This is the one place the reviews disagree outright, and with eleven and seventeen studies respectively neither is in a position to settle it.
The finding that should worry providers
The Jones review tested two moderators that the coaching industry would rather it had not.
Coaching delivered by internal coaches produced stronger effects than coaching by external coaches. And engagements that used multi-source feedback produced smaller positive effects than those that did not.
Both cut against how corporate coaching is normally bought: external, senior, expensive, and wrapped around a 360-degree instrument. The sample is small and neither finding has been replicated, so they are a reason to ask questions rather than to restructure a programme. But the direction is the opposite of the sales argument, and that is worth knowing before the next renewal. The internal and external comparison is examined in what the evidence says about internal coaches, and the feedback finding lines up with the evidence on 360-degree feedback from the leadership training literature.
How thin the base really is
The three reviews between them draw on fewer than fifty distinct primary studies, with overlap between them. The most rigorous of the three could find eleven.
For comparison, the equivalent meta-analysis of leadership training pooled 335 independent samples. Coaching is bought at scale by large organisations and has been for thirty years, and the research base is an order of magnitude smaller than for the adjacent intervention.
The consequences are concrete. Nobody can say with confidence which coachees benefit most, which coach qualifications predict results, or how long an effect lasts after the engagement ends, because there are not enough studies to test those questions. The 2023 review names the recurring gap directly: the primary studies frequently lack detail on the coaches, the participants and what actually happened in the sessions.
So the fair summary is that workplace coaching has a positive effect of moderate size, that the effect is better established than most workplace interventions and much less well established than the marketing implies, and that the two variables buyers spend the most time negotiating, session count and delivery format, are the two the evidence says do not matter.
Common questions
Does workplace coaching work? Yes, with a small to moderate effect. Three meta-analyses since 2014 report pooled effects between roughly 0.36 and 0.6 depending on which studies they include, all positive and all statistically significant.
Why do the meta-analyses report different numbers? Because they define coaching differently. The 2014 review included general life coaching and health coaching alongside workplace coaching; the 2016 review restricted itself to workplace coaching by a designated coach; the 2023 review applied tighter methodological criteria and found only eleven eligible studies.
How many coaching sessions are needed to see an effect? The evidence does not support any particular number. All three meta-analyses tested session count or duration as a moderator and none found a significant relationship with outcomes.
Is remote coaching less effective than face-to-face? No measurable difference has been found. One review found no moderation by format between face-to-face and blended or electronic delivery; another compared 0.48 for face-to-face against 0.35 for virtual and found the difference not significant.
Are external coaches better than internal ones? The one meta-analysis that tested it found the opposite: internal coaches produced stronger effects. The finding rests on a small number of studies and has not been replicated, so it is best treated as a question to ask rather than a conclusion.
How does coaching compare to leadership training on the evidence? Leadership training has a much larger evidence base, 335 independent samples against fewer than fifty studies, and larger reported effects, around 0.7 to 0.8 against 0.36 to 0.43. The two are not measuring identical outcomes, so the comparison is indicative rather than direct.
What is a 0.4 effect size in practical terms? By the usual convention 0.2 is small, 0.5 moderate and 0.8 large. An effect of 0.4 means the average coached person ends up better off than about 66 percent of an uncoached comparison group on the outcome measured.
Sources
- Theeboom, Beersma & van Vianen (2014), Does coaching work? A meta-analysis on the effects of coaching on individual level outcomes in an organizational context, Journal of Positive Psychology 9(1), 1-18
- Jones, Woods & Guillaume (2016), The effectiveness of workplace coaching: A meta-analysis of learning and performance outcomes from coaching, Journal of Occupational and Organizational Psychology
- Cannon-Bowers, Bowers, Carlson, Doherty, Evans & Hall (2023), Workplace coaching: a meta-analysis and recommendations for advancing the science of coaching, Frontiers in Psychology 14:1204166



