In more than a third of tested cases, feedback made performance worse

Across 607 effects, feedback raised performance by 0.41 on average and lowered it in over 38 percent of cases. Praise and threats to self-esteem weakened it.

A hand holding a red pen above a blank spiral notebook on a wooden table.
In the 1996 meta-analysis of feedback, more than 38 percent of the measured effects on performance were negative. Photo: Kelly Sikkema on Unsplash

Feedback lowered performance in more than 38 percent of the 607 effects that Avraham Kluger and Angelo DeNisi pooled in 1996. On average it helped, by 0.41 of a standard deviation. Both numbers come from the same distribution, and the first is the one that feedback guides leave out.

A 2020 meta-analysis describes it as among the most comprehensive reviews of feedback. Its abstract says "over a third". We read the full text, where the exact figure, the checks behind it and the moderators that explain it are set out. The finding holds, and it is more specific than the one-line version.

607 effects, 131 studies, and 38 percent below zero

The authors started from roughly 3,000 papers and technical reports. Only 131 papers, 5 percent, met their criteria: an unconfounded feedback treatment, a control group that received no feedback, and a measure of performance. About 37 percent of the papers they considered had manipulated feedback without a control group at all.

From those 131 papers they extracted 607 effect sizes, based on 12,652 participants and 23,663 observations. The paper in Psychological Bulletin reports a weighted mean effect of 0.41 and states that over 38 percent of the effects were negative.

MeasureValue
Papers consideredabout 3,000
Papers meeting the criteria131 (5%)
Effect sizes607
Participants12,652
Weighted mean effect0.41
Share of effects below zeroover 38%
Share below zero without one lab's 91 effects33%

An effect of 0.41 is moderate. By the usual convention it means the average person who received feedback ended up ahead of about 66 percent of the comparison group. A distribution with that mean and more than a third of its mass below zero describes a tool that helps under some conditions and harms under others.

Removing the most negative lab still leaves a third below zero

The authors tested the obvious objection, that a few odd studies drag the result about. Ninety-one of the effects came from one researcher's laboratory experiments, which used deliberately extreme negative feedback. Those effects averaged minus 0.39. The rest averaged 0.47, and 33 percent of them were still negative.

Sampling error cannot explain the spread either. The weighted variance of the effects was 0.97, against an expected sampling error variance of 0.09. Most of the variation is real, which is why the authors went looking for what drives it.

Whether feedback was positive or negative made no difference

The intuitive explanation is that negative feedback hurts and positive feedback helps. The data do not support it. The abstract states that the result cannot be explained by feedback sign, and in the moderator analysis the effect of sign became nonsignificant once the extreme laboratory studies were set aside.

What did predict the result was where the feedback pointed attention. The authors propose three levels: the details of the task, motivation for the task, and the self. Their abstract summarises the pattern.

The results suggest that FI effectiveness decreases as attention moves up the hierarchy closer to the self and away from the task.

The moderators that survived every exclusion show the shape. Praise weakened the effect of feedback. So did feedback designed to discourage, and feedback that threatened self-esteem. Feedback delivered verbally by another person did less than feedback from a computer, which was marginally stronger. Feedback that supplied the correct solution, or showed the rate at which performance was changing, made the effect stronger.

Praise is the uncomfortable result. It is positive, it is pleasant to give, and on the authors' account it pulls attention away from the task and toward the person. Telling someone they did well tells them about themselves. Telling them the third paragraph lost the argument tells them about the work.

The frequency result is the one the authors doubted

Feedback given more often did show stronger effects after all the exclusions. That looks like support for regular feedback, and the authors warn against reading it that way. They write that the distribution of the frequency variable was poor, that the result in the extreme quartiles ran in the opposite direction, and that the frequency effect may be an artifact.

Goal setting did better. Feedback combined with a goal was more effective, at a significance level the authors treat with caution, and the combination helped most when the feedback message on its own was hard to interpret. A figure such as "you produced 200 units" means little until there is a target to set it against.

The nature of the task also mattered, and here the paper is candid. Feedback helped memory tasks more and physical tasks less, and, marginally, simple tasks more than complex ones. In the authors' words the task characteristics that moderate feedback are still poorly understood.

At work, managers' 360 ratings improved by 0.15

Most of the 1996 studies were experiments, many of them in laboratories. The closest workplace test is multi-source feedback for managers. James Smither, Manuel London and Richard Reilly pooled 24 longitudinal studies in Personnel Psychology in 2005. The abstract calls the improvement in ratings over time generally small, and the full text gives the numbers.

Rater sourceEffect sizesRateesCorrected mean effect
Direct reports217,7050.15
Peers75,3310.05
Supervisors105,3580.15
Self113,684minus 0.04

For every rater source except supervisors, the 95 percent confidence interval included zero. Whether a facilitator helped recipients interpret their report was not significantly related to the size of the improvement.

A worked example makes 0.15 concrete. Take a firm of 300 people that runs a 360 for its 30 managers, and assume the managers start, on average, at the middle of the distribution of direct report ratings. An improvement of 0.15 moves the average manager from the 50th percentile to about the 56th. That is a real change and a modest one, for a process that takes weeks of rater time. Whether the multi-source version beats single-source feedback is a separate question, and the largest leadership training meta-analysis found no advantage on any outcome it could test.

The 2005 authors also review evidence on who improves, and it matches the 1996 theory. They conclude improvement is most likely when the feedback indicated a need to change, the recipient reacted positively, believed change was feasible, and set goals. Feedback that lands on the person without reaching a goal has little to work with.

Twenty-four years later, education found the same spread

The same question has been asked of classrooms at a larger scale. A 2020 meta-analysis in Frontiers in Psychology pooled 435 studies and more than 61,000 learners and found a medium effect of 0.48. Its abstract reports significant heterogeneity and concludes that feedback cannot be understood as a single consistent form of treatment. Effects were larger on cognitive and motor skills than on motivational and behavioural outcomes.

Three tests of the sandwich, three different answers

The feedback sandwich is praise, then critique, then praise. The 2013 study below describes it as commonly recommended despite scant evidence of its efficacy. We found three papers that test it, and they do not agree.

In two studies of written peer feedback among third-year medical students, with 20 and 350 participants, Parkes, Abercrombie and McCarty found that students believed sandwiches improved their later performance when there was no evidence that they did.

In a 2020 experiment, Prochazka, Ovcari and Durinik randomly assigned 91 university students who had solved maths problems to corrective computer feedback, the same correction sandwiched between two general positive statements, or no feedback. The sandwich group spent more time preparing and solved more problems in the second round. The authors call their result partial evidence and ask for replications.

The third comes closest to a workplace. Bottini and Gillis trained participants to run a simple behavioural assessment and compared sandwich feedback with constructive then positive feedback, against a control of feedback given during the session. Feedback during the session produced the highest fidelity in the first role play, and by the third and final role play there were no significant differences between the conditions.

None of the three papers is large, and none measured performance on a real job. On the 1996 evidence the sandwich has an internal problem: its bread is praise, the moderator that weakened feedback, and its filling is the only part that points at the task.

Which tasks feedback harms is still unknown

The evidence supports a narrower rule than the one in circulation. Feedback about the task, with the correct solution and a goal to measure it against, tends to help. Feedback about the person, flattering or threatening, tends to help less and sometimes hurts. Leadership programmes that include feedback show their gain in changed behaviour at work, and not on the other outcomes measured.

What the 1996 authors could not explain, and what the workplace studies we found do not test, is which tasks turn feedback negative. They named task characteristics as the moderator still poorly understood. Thirty years on, we found no controlled workplace study that would tell a manager in advance on which side of that 38 percent a given piece of feedback will fall.

Common questions

Does feedback improve performance? On average, yes. The largest meta-analysis, of 607 effects from 131 studies published in 1996, found a weighted mean improvement of 0.41 of a standard deviation. The same study found that more than 38 percent of the effects were negative.

How often does feedback make performance worse? In over 38 percent of the 607 effects in the 1996 meta-analysis. After removing 91 effects from one laboratory that used extreme negative feedback, 33 percent of the remaining effects were still negative.

Is negative feedback worse than positive feedback? Not on this evidence. The 1996 meta-analysis found that the sign of the feedback did not explain its effect. What mattered was whether the feedback directed attention to the task or to the person.

Does praise improve performance? In the 1996 meta-analysis praise weakened the effect of feedback on performance. The authors explain it as praise drawing attention toward the self and away from the task.

Does the feedback sandwich work? The evidence is thin and split. A 2013 study of medical students found no effect on performance even though students believed it helped. A 2020 experiment with 91 students found the sandwich group solved more problems afterwards. A 2021 training study found no significant difference from other sequences by the final role play.

How much does 360-degree feedback improve managers? A 2005 meta-analysis of 24 longitudinal studies found small improvements in ratings over time: a corrected effect of 0.15 from direct reports and supervisors, 0.05 from peers and minus 0.04 in self-ratings.

Does giving feedback more often help? The 1996 meta-analysis found stronger effects for more frequent feedback, but the authors warned that the result may be an artifact of how the variable was distributed. It is the weakest of their findings.

What makes feedback more effective? In the 1996 data, feedback that gave the correct solution, showed the rate of change in performance, or came with a goal performed better. Feedback that praised, discouraged or threatened self-esteem performed worse.

Sources

  1. Kluger & DeNisi (1996), The effects of feedback interventions on performance: A historical review, a meta-analysis, and a preliminary feedback intervention theory, Psychological Bulletin 119(2), 254-284
  2. Smither, London & Reilly (2005), Does performance improve following multisource feedback? A theoretical model, meta-analysis, and review of empirical findings, Personnel Psychology 58(1), 33-66
  3. Wisniewski, Zierer & Hattie (2020), The power of feedback revisited: A meta-analysis of educational feedback research, Frontiers in Psychology 10:3087
  4. Parkes, Abercrombie & McCarty (2013), Feedback sandwiches affect perceptions but not performance, Advances in Health Sciences Education 18(3), 397-407
  5. Prochazka, Ovcari & Durinik (2020), Sandwich feedback: The empirical evidence of its effectiveness, Learning and Motivation 71, 101649
  6. Bottini & Gillis (2021), A comparison of the feedback sandwich, constructive-positive feedback, and within session feedback for training preference assessment implementation, Journal of Organizational Behavior Management 41(1), 83-93

More from the Journal