A supervisor has two sentences ready for a review on Thursday. The first one is about the work. The reconciliation in column three does not tie to the general ledger, and here is the query that would have caught it. The second one is about the employee. You have been careless with numbers lately.
The supervisor says both and believes both. The second sentence is the one the review form asks for, under a heading called development areas.
The supervisor and the review are invented. Putting the scene on nobody in particular means I am not describing anyone I work with, and we are looking at the shape of an ordinary review.
Most people believe two things about feedback. That giving it helps, and that when it goes wrong the delivery went wrong.
Here is what the evidence supports. In the largest pooling of the research, more than a third of the measured effects ran the wrong way, leaving the people who got feedback behind the people who got none. And where the negative motivational effects have been sorted by feedback type, they sat in feedback that carried no information about the work.
A review takes an hour and that hour is the only part of this anybody counts. What it can cost is work that gets worse afterward, delivered by a supervisor who did everything the training said, and nobody traces a quarter's output back to one sentence in a Thursday meeting.
What was measured
Avraham Kluger and Angelo DeNisi pooled the feedback literature in 1996. They kept 131 papers, five percent of what their search returned, and those papers produced 607 separate effects drawn from 12,652 people and 23,663 observations, because many of the studies measured the same person more than once.
Each effect is reported as a single number, where zero means the people who got feedback and the people who did not came out the same.
The average effect was positive. Weighted by sample size it came to 0.41, which is a moderate improvement.
Over 38 percent of the 607 effects were negative, meaning the people who got feedback performed worse than the people who did not.
The spread across those 607 effects is wider than chance would produce. The variance across the effects was 0.97. The variance expected from sampling error alone was 0.09. That gap cuts both ways. Sampling error cannot account for the negative tail, and an average of 0.41 is a weak summary of a set scattered that far.
One research program, Mikulincer's laboratory experiments, supplied 91 of the 607 effects and averaged minus 0.39, against an average of 0.47 across the remaining effects. Drop Mikulincer's studies entirely and 33 percent of what is left is still negative.
The explanation the authors offered
Kluger and DeNisi proposed that feedback works on attention, and that attention has three levels.
At the bottom are task-learning processes, which is attention on the details of the job in front of the person. Above that are task-motivation processes, which is attention on the job as a whole. At the top are meta-task processes, which is attention on the self.
Their summary sentence: "FI effectiveness decreases as attention moves up the hierarchy closer to the self and away from the task."
That is their theory, and they are careful about its standing. They say in the paper that they can test the effects of a feedback intervention and cannot test the processes underneath, because the original studies were not designed to test the theory and contain no measure of where anyone's attention went. The word preliminary is in the title.
Their other cautions are worth carrying. They did not correct for unreliability in the measures, because only two of the 131 studies reported enough to allow it. And they estimate the literature they searched is a fraction of what exists, putting the number of dissertations meeting their search criteria above three thousand.
I could not read their moderator tables at a resolution I trust, so nothing here rests on any single moderator in that paper, including praise.
A later pooling, in schools
Benedikt Wisniewski, Klaus Zierer and John Hattie pooled the research again in 2020, working from 435 studies, 994 effects and more than 61,000 people. Their average came out at 0.55 across everything, and at 0.48 with an interval of 0.44 to 0.51 after they excluded 35 extreme values. The 1996 average of 0.41 has no such exclusion, so the two are not a clean comparison. In either version the average is positive.
The useful part of the 2020 work is that it sorts the effects by what the feedback contained.
Feedback that was nothing but reward or punishment came in at 0.24, with an interval running from 0.06 to 0.43. Feedback that corrected an error came in at 0.46, with an interval of 0.39 to 0.55. The category the researchers call high-information feedback came in at 0.99.
On motivational outcomes, 21 percent of the effects were negative, and 86 percent of the interventions producing those negative effects were, in the authors' word, uninformative, being rewards or punishments.
Those three categories were not three arms of one experiment. They are groupings assembled after the fact from different studies, so the 0.99 is the largest pooled average among the groupings rather than a treatment anybody administered.
What the top category actually contains
The 0.99 category is not purely about the task, and a reader who takes the phrase at face value will get this wrong.
The authors define it as feedback that carries everything corrective feedback carries and "additionally contained information on self-regulation from monitoring attention, emotions, or motivation during the learning process." So the best-scoring grouping in that pooling does direct some of the recipient's attention inward.
What it does not do is name a quality the recipient has. Hattie and Timperley drew that line in 2007, splitting feedback four ways: about the task, about how the task is being processed, about the person's self-regulation, and about the self as a person. The first three describe what is happening. The fourth describes who somebody is.
"On the third pass through the schedule you skipped the tie-out query" falls in the first three. "You have been careless with numbers lately" falls in the fourth.
What the workplace evidence says, and where it cuts against me
Both poolings above are laboratory and classroom work. The workplace literature is separate, it exists, and it complicates this post rather than confirming it.
Jetmir Zyberaj published a meta-analysis of supervisory feedback in 2026, covering 24 studies, 75 effect sizes and 595,950 observations across 26 characteristics of supervisory feedback. The pooled association with how employees process feedback is 0.36, with an interval of 0.30 to 0.41. The most studied characteristic, appearing in 12 of the 24 studies, is the credibility of the supervisor giving the feedback, and it is among the strongest.
That is a result about the source rather than the sentence, and it is the first thing that should give a reader pause about everything above.
The second is a field experiment. Shankar Naskar and Prathiba Natesan Batley ran six waves of performance feedback on 331 managers at a manufacturing firm in India over six years, randomly assigning each manager in a fully crossed design that varied who gave the feedback and what it contained. Their finding on the comparison: "the feedback source exerts more impact than the feedback content over time." Developmental feedback did nothing in the short run and helped in the long run. Low performers improved more than high performers, and by the end of six years the scores converged, which the authors read as a ceiling on what feedback interventions can do.
A critical review of the whole workplace literature by Heine, Stouten and Liden, covering 173 studies, appeared in the Journal of Organizational Behavior in 2026 and describes the field as fragmented, with contested definitions of the word feedback itself. Frederik Anseel and Elad Sherf reviewed twenty-five years of it in 2025 and wrote that the science of feedback at work "is not yet a story of coherent and cumulative progress."
So the sentence I was tempted to write, that delivery is the wrong variable and content is the right one, does not survive the workplace evidence. Who is talking matters, and in the one long-run randomized test it mattered more.
Putting specific information about the work into the sentence has support in both literatures. Leaving out a judgment about the person has support in a theory its own authors could not test.
Two things this does not say
The obvious reading is that this is an argument for being kinder, and it is not.
Hattie and Timperley put the level version plainly: "Feedback at the self or personal level (usually praise), on the other hand, is rarely effective." Praise is feedback about the person, and the grouping that scored lowest in 2020 was rewards and punishments together.
It is also not an argument against telling somebody their work is wrong. Corrective feedback came in at nearly twice the value of reward and punishment, though the intervals around the two overlap.
What to do on Thursday
Take each sentence you have ready and ask what it points at.
A sentence pointing at the work names the piece of work, says what is wrong with it, and says what would have caught the mistake. A sentence about how the work went names something the person did on this job. A sentence pointing at the person names a quality the person has.
Two different claims sit behind that ordering and they are not equally supported, so here is which is which.
Putting information about the work into the sentence is the measured part. Uninformative feedback scored lowest and produced most of the motivational harm in a pooling drawn entirely from schools and universities, and supervisory feedback quality shows up in the workplace meta-analysis as well.
Leaving out the quality is the theory part. It comes from Kluger and DeNisi's hierarchy, and their hierarchy is the piece of the 1996 paper they say they could not test. "You have been careless with numbers lately" is not a reward and not a punishment, so it is not what the 2020 harm finding measured.
On the form, a box labeled development areas invites a sentence about the employee and does not require one. Reconciliation process, review of supporting schedules, and escalation of unresolved variances all fit in that box, and all three are things somebody can go do on Friday. Put the work in the box.