Does AI coaching work as well as a human coach?

The study behind the parity claim compared two trials run two years apart. A direct randomised comparison found effects only for human coaching.

Two people in silhouette facing each other across a table in a brightly lit room.
The question the widely cited trial could not answer is what happens when the two are compared directly. Photo: Heng Chiu on Unsplash

Every vendor selling AI coaching to an enterprise buyer eventually reaches the same slide. It says that a randomised controlled trial found no significant difference in goal attainment between an AI coach and a human coach, and it cites a 2022 paper in PLOS ONE.

The paper exists, the finding is real, and the slide is misleading. The reason is in the study design, which almost nobody who cites it describes.

What the 2022 study actually compared

Terblanche, Molyn, de Haan and Nilsson ran two longitudinal randomised trials, each about ten months long, each with the same measurement instruments.

The first ran from October 2017 to July 2018. It had a human coaching group of 105 and a control group of 105. The second ran from November 2019 to August 2020, with an AI chatbot coach called Vici, 134 participants, and its own control group of 134.

Two trials, two years apart, two different cohorts of people. Nobody was ever randomised between a human coach and a chatbot.

The result reported is that goal attainment improved substantially in both experimental groups and that the two never significantly differed from each other. The effect sizes were 0.265 for the human coaching group and 0.269 for the AI group, against 0.16 and 0.11 in the two control groups.

That is a legitimate finding and the authors describe their design accurately. But comparing effect sizes across separate trials run in different years with different participants is not the same as a head-to-head test, and the difference matters: any variation between the two cohorts, any change in circumstances between 2018 and 2020, and the entire onset of the pandemic during the second trial all sit inside the comparison.

The paper is also explicit about two further limitations. The participants in both studies were undergraduate students. And all the outcomes were self-scored.

The direct comparison, and what it found

In 2026, three researchers, two of whom were on the 2022 paper, published the test the field had been missing.

de Haan, Terblanche and Nowack randomised 114 coachees between accredited human coaches, an automated AI coach and a waitlist control. Both coaching conditions were an hour a week. The measures were validated instruments for goal progress, motivation, resilience and wellbeing.

The study demonstrated substantial effectiveness in diverse coaching outcomes but only for human coaching, with effect sizes in the mid-to-high range.

Only for human coaching. When the comparison was made properly, within one trial, with randomisation between the two conditions, the parity disappeared.

The paper also reports that outcomes were strongly predicted by the coachee's own initial self-efficacy and hope, and that dropout was barely affected by which condition people were assigned to. The second point is worth holding on to: the chatbot did not fail because people refused to use it.

Two findings, one direction of travel

Set side by side, the picture is consistent rather than contradictory.

2022, PLOS ONE2026, Human Resource Development International
DesignTwo separate trials, two years apartOne trial, randomised between conditions
ParticipantsUndergraduate studentsCoachees, 114 in total
ComparisonEffect sizes compared across studiesHuman coaching against AI coaching directly
OutcomesGoal attainment, self-scoredGoal progress, motivation, resilience, wellbeing
ResultNo significant difference between conditionsSubstantial effects for human coaching only

The weaker design found parity. The stronger design did not. That is the ordinary pattern when a promising effect is tested more rigorously, and it is the same pattern visible across the coaching literature generally: as the inclusion criteria tighten, the measured effect shrinks.

The disclosure nobody makes

One fact about this literature belongs in any honest summary of it.

The AI coach used in the 2022 study, Vici, was created by Nicky Terblanche, who is the first author of that paper and an author on the 2026 one. He is Associate Professor of Leadership Coaching at Stellenbosch Business School and has published extensively on AI in coaching; the chatbot came out of that research.

This is disclosed in the academic record and is not a scandal. Researchers building the thing they study is normal in applied fields, and the fact that the same group went on to run the harder test, and published a result that undercut the earlier one, is the system working as intended.

It is worth knowing anyway, because the 2022 result is cited by companies with a commercial interest and no such connection to disclose. A buyer told that independent research proves AI coaching matches human coaching is being told something inaccurate twice over: the research was not independent of the tool, and it did not compare the two directly.

What the evidence supports today

Three things can be said with reasonable confidence.

AI coaching produces measurable gains against doing nothing. Both experimental groups in the 2022 study substantially outperformed their controls. A chatbot that prompts someone to set a goal, review it weekly and record progress is not inert.

It has not been shown to match a trained human coach in a fair test. The one direct randomised comparison found effects for human coaching and not for the chatbot.

The evidence base is thin on both sides. The 2022 study used students. The 2026 study used 114 people. The most recent meta-analysis of workplace coaching found only eleven studies rigorous enough to include, for human coaching, after four decades of practice. Anyone claiming certainty in either direction is ahead of the data.

For a procurement decision the practical reading is that AI coaching is best evaluated as a scale intervention rather than a substitute: something that reaches people a coaching budget would never have reached, judged against the alternative of reaching them with nothing. Framed that way the evidence supports it. Framed as an equivalent to a human coach at a fraction of the price, it does not, and the study usually cited to support that framing does not say what it is said to say.

The compliance position for tools of this kind changed in 2025 and again in 2026, and is covered in what the AI Act requires of coaching and HR tools.

Common questions

Is AI coaching as effective as human coaching? The only direct randomised comparison, published in 2026, found substantial effects for human coaching and not for the AI chatbot. The earlier study often cited for parity compared two separate trials rather than randomising people between the two conditions.

What did the 2022 PLOS ONE study find? That goal attainment improved significantly in both a human coaching group and an AI chatbot group, with effect sizes of 0.265 and 0.269, and that the two experimental groups never significantly differed. The two groups were in separate trials run two years apart.

Why does the study design matter so much here? Because comparing results across two trials held in different years with different people leaves every difference between those cohorts and those periods inside the comparison. Randomising people between the conditions within one trial removes that problem, and when it was done the result changed.

Who were the participants in these studies? In the 2022 study, undergraduate students scoring their own outcomes, which the authors name as a limitation. In the 2026 study, 114 coachees randomised between accredited human coaches, an AI coach and a waitlist control.

Does AI coaching do better than nothing? On the available evidence, yes. Both experimental groups in the 2022 study substantially outperformed their control groups.

Is there a conflict of interest in this research? The AI coach used in the 2022 study was built by its first author. That is disclosed in the academic record. The same author was also on the 2026 paper, which reported the less favourable result.

How large is the evidence base overall? Small. The most recent meta-analysis of workplace coaching found eleven studies methodologically strong enough to include. For AI coaching specifically there are a handful of trials, the largest of which had 268 participants across two separate studies.

Sources

  1. Terblanche, Molyn, de Haan & Nilsson (2022), Comparing artificial intelligence and human coaching goal attainment efficacy, PLoS ONE 17(6): e0270255
  2. de Haan, Terblanche & Nowack (2026), A randomised controlled comparison of the effectiveness of human and AI chatbot coaching with goal attainment, wellbeing and self-efficacy, Human Resource Development International 29(4), 754-783
  3. Record for the 2026 study at Vrije Universiteit Amsterdam, including the abstract
  4. Cannon-Bowers et al. (2023), Workplace coaching: a meta-analysis and recommendations for advancing the science of coaching, Frontiers in Psychology 14:1204166

More from the Journal