Most criticism of AI is not actually about quality. This seems to be contradicted by the legion of social media posts about “AI-slop,” but such a pejorative is increasingly becoming a way to dismiss people and their creative capacity, rather than any legitimate identification. Just this week, I saw a photographer friend of mine defend their work because people were calling it AI, when really it was about the skills of camera angles, lighting, and the sheer enormity of the ignorance of how, yes, the Badlands of South Dakota do in fact have greenery.
It is about authorship. That distinction matters more than the AI debate usually admits, not least because people generally have a constantly moving and typically based-in-ignorance rubric in judging what AI (often undefined as well) is and what is a program that people have been using for years for creative output like Photoshop and various writing grammar helpers.
Nowhere is this distinction more visible than in a small, uncomfortable body of research on therapists’ rating psychological advice.
The Finding that Blindsided
In 2025, a research team in Sweden ran a blinded study that should disturb anyone confident about any unique quality inherent to therapeutic language. Ludwig Franke Föyen and colleagues took real advice-column questions from a Swedish national newspaper, along with the expert responses that had actually been published, and generated matched AI responses to the same questions. They then recruited licensed psychologists, psychotherapists, and physicians and asked them to rate the responses on quality and empathy, without being told which was which.
The clinicians could not reliably tell the difference. Worse, for anyone invested in the supposed inherent quality of human clinical warmth, the AI-generated responses were rated significantly higher on emotional and motivational empathy than the expert-written ones, and statistically comparable on scientific quality and cognitive empathy. This was not a fluke result. A pooled review by Alastair Howcroft and colleagues, synthesizing fifteen studies comparing chatbot and healthcare-professional empathy, found AI rated as significantly more empathic in thirteen of them — an advantage roughly equivalent to two points on a ten-point scale. A widely cited 2023 study of the r/AskDocs subreddit found licensed clinicians preferring chatbot answers over physician answers 78.6% of the time.
But wait, there’s more. The researchers didn’t just ask clinicians to rate blind responses. They also manipulated perceived authorship, telling some raters a response came from an expert when it hadn’t, or vice versa. The effect was large. Believing a response had been written by a human expert raised its ratings substantially — enough to swing the study’s core comparison. This was true regardless of who had actually written it. The label did the work that the words themselves hadn’t earned.
Let’s be clear: expert clinicians, trained to detect empathic attunement (the basis for rapport-building) for a living, could be induced to rate a paragraph as more caring by being told a human wrote it, independent of anything in the paragraph itself. Authorship became the signal to override expertise.
Naming What This Is
Now, this is not a story about AI being secretly wonderful at therapy. It’s a story about what the word “human” is doing when clinicians, and so many others, reach for it as a quality determinant rather than a description. That distinction may be getting lost, so let’s try being as exact as possible: quality is a property of the artifact, highlighting its coherence, precision, and responsiveness to what was actually said. Authorship is a fact about origin. When origin starts functioning as a substitute for quality even after the artifact has been blind-tested and found equal or better, something other than quality assessment is happening. Let’s call a spade a spade: what’s happening is a tax levied on the artifact for not having the right kind of author, paid in the currency of a rating scale, with quality being left behind. Put in the vernacular of the social critic, this is virtue-signaling at its finest, where declaring oneself in alignment with a socially-constrained ethical prerogative is more important than providing a qualitatively better judgment.
Despite my rhetoric, I don’t honestly think this bias is without some legitimacy. A great deal of what makes human clinical relationships valuable, like being known over time, is real and is not measured by a single vignette rating. The Swedish study, like most other current similar studies, tested isolated written responses, not a sustained relationship. That’s a limitation, and I’m not dismissing it. But notice what the bias is not: it is not a claim about relationship, continuity, or embodiment. It’s a preference for the response believed to be human, holding the actual content (including its helpfulness and possibly even accuracy) of the response as less important. That’s a narrower and stranger thing than “human relationship matters,” and it’s worth being honest about rather than ignoring it by throwing it under the bus of a broader self-serving ethical claim.
The Applause Line
I watched this same mechanism play out live a few weeks ago. I was sitting through a presentation where the presenter opened by announcing — with visible pride — that they hadn’t used AI anywhere in building their slides. The room applauded as if the person had just declared the cure for cancer. With cheers, before a single slide of content had been evaluated on its own terms.
The deck that followed was bad. Not bad in an interesting or debatable way, but bad in the textbook sense that guidelines exist to prevent. Slides were dense with paragraphs of unreadable body text. The backgrounds were busy templates that fought with the text sitting on top of them. Images were dropped in because they were decorative, not because they clarified anything, and several actively pulled attention away from the point being made. The examples cited were often years out of date, and the reasoning connecting slide to slide had gaps that should have been caught in a rough draft. By any standard, a presentation-design course would apply — text load, visual hierarchy, signal-to-noise, currency of the evidence — it failed.
You know what would have been helpful? The tool that would have addressed many of the problems and offered directions for addressing them? AI. But the room had already applauded. Not as a response to the content, it couldn’t have been, since the content hadn’t been seen yet. It was a response to the performative statement: I made this without AI. That announcement did the same work in that room that “written by a human expert” did in the Karolinska raters’ heads — it pre-loaded the evaluation before the evaluation had anything to evaluate. The only difference is that the therapists in the study were at least judging real content and getting the rating wrong. The conference audience applauded, and therefore set up an affirmative judgment about the content, before the content existed for them at all.
I don’t think anyone in that room actually believed the deck was good. It’s not as if they could actually read any of it on the stage anyway. I think most of them didn’t consciously register that it was bad, because the applause had already closed the question the way all virtue-signaling closes critical inquiry: by converting “did not use AI” into a virtue that stands on its own, independent of whether the resulting artifact does its job. That’s the authorship tax again, just moved out of the lab and into a conference room, and it indicates the same issue the clinicians inadvertently displayed: that preloading social solidarity will undermine objectivity, even if it means being poorer because of it. There was just a slide deck, doing its job badly, applauded for the one thing that had nothing to do with whether it worked.
The Critique that Supports Mediocrity
Here is where I think the AI-critical conversations go wrong. The standard move is to treat any evidence that AI output is rated well as evidence that the rating instrument is broken, that raters were fooled, that empathy “isn’t really” what’s being measured, or that the study can’t capture what matters. Sometimes those are legitimate methodological points. Often, they function as an escape hatch: a way of preserving the conclusion that human work (again, lacking in nuance since a definitive quality of being human is the use of tools, which means human work should always include an asterisk) is intrinsically better, without doing the harder work of showing it on the actual dimension being tested.
This is where criticism of AI devolves into a defense of mediocrity. Not because AI is categorically superior — as saying so would break a fundamental therapeutic tool for noting cognitive errors, namely black/white thinking — but because the reasoning pattern stops being about output and starts being about protecting the person from disconfirming evidence. If a clinician’s response is judged less empathic than a language model’s on a task both were asked to do, the honest response is to sit with that, not to shoot the messenger and declare what feels better to be the winner. A profession that only accepts evaluations it passes isn’t holding a quality standard; it’s simply declaring itself above criticism. While admittedly, I think this is often occurring throughout the field of psychology, particularly therapeutic psychology, it doesn’t make the practice any less ethically vacuous.
I want to be careful here, because often people will take this criticism and slide into “well then AI should just replace therapists,” which is not only not what I’m saying, but isn’t supported by the data anyway. In-the-moment empathy ratings and therapeutic outcomes are not the same measurement, and a chatbot’s higher score on a single written exchange says nothing about whether it can competently engage in confrontation, tolerate a client’s anger at it, or hold a treatment plan together across months of relapse and repair. Psychology Today has an article that adds further caution: AI empathy may be more template-driven, where human responses, even when rated as “less empathic” in aggregate, carry more linguistic diversity. This points to how a single-point measurement may capture one rating, but long-term relationship-building would capture a different one. These are legit limits on what the empathy studies lead to, though they’re also a cautionary story of what incremental improvements of AI could accomplish. None of the limits, though, is an approval for the authorship tax. They’re arguments for better measurement, not arguments for rating the same words higher once you’re told a human wrote them.
What Mediocrity Actually Costs
I work with many who have been burned by poor therapy, whose personal and relational lives have been led astray by incredibly poor interventions under the guise of expertise, and often given self-righteous weight by socio-political activist ideologies that pre-judge the client into categories for easy dismissal. Much of my writing is also critical of a therapeutic establishment that has become the secular equivalent of a priesthood, and has quite quickly become the go-to space for a society that has found itself unable to deal with adversity, form nuanced friendships, or accept that reality isn’t theirs to shape in any way they want. I have also sat with clients who received technically correct but empathically hollow care from human clinicians, and I’ve watched colleagues produce boilerplate reassurance dressed in a caring tone because caring tone was all they were trained to do. The authorship bias isn’t hypothetical for me. It’s the same mechanism that lets a mediocre human intervention pass because it came from a human, while a genuinely well-constructed piece of psychoeducation gets dismissed because it is believed to have come from a machine.
If empathy ratings can be moved by a label, then “I can tell the difference” is not a reliable clinical skill, nor is it a reliable tool for judgment by anyone else. It is, instead, a belief that answers a question that isn’t even being asked, by faking an answer to one that is. The professionally responsible response to research findings that indicate a potential problem in the field is not defensiveness, but empathically critical reflection.
Where this Leaves Us
None of this is a simple case for AI-delivered therapy, certainly in any expansive form currently. Empathy ratings on isolated vignettes settle nothing about relationship, safety, or long-term care. The head researcher himself, in commenting on his own findings, has been explicit that AI shouldn’t replace human relationships or professional care, partly because of exactly the safety failures that make consumer chatbots dangerous in a genuine crisis. What it is a case for is separating two questions that keep getting collapsed into one: is this good, and who made it? The Swedish clinicians in this study were good at their jobs and still couldn’t keep those two questions apart when the label was switched on them. That’s not an indictment of therapists. It’s a demonstration that authorship bias operates below the level of professional training, and that a critique of AI built on that bias, instead of on the actual, defensible limits of what these studies measure, isn’t defending humanistic care. It’s defending the label. And the mediocrity that results doesn’t help those who need quality care.
Sources:
Additional context drawn from the fabulous writing of Charlotte Blease and Psychology Today’s “The Shadow Artificial Third.”
Follow me on Threads or Instagram for more psychology, humor, and photography.
Schedule a Consultation
Reach out to schedule a mental health session regarding issues related to identity and decision-making, emotional regulation, difficulties in relationships, religious trauma, developing meaning/purpose, and the struggles of communication.
How You Can Support the Newsletter
This post was free to read for all. If you like what I’m doing with Humanity’s Values and want to support my work, there are several ways to do so.
Like and Restack: Click the buttons at the top or bottom of the page to boost the post’s visibility on Substack.
Share: Send the post to friends or share it on social media.
Dialogue: Expand the space for dialogue by engaging with others on the topics discussed here and help build the community.
Upgrade to Paid: A paid subscription gets you:
Full access to all new posts and the archive
Full access to all online classes
Full access to weekly presentations on various topics of psychology
Promotional discount on coaching/consultation services
The ability to post comments and engage in chat with the growing Humanity’s Values community
If you could do any of the above, I’d be very grateful. Readers like you help keep this newsletter going and growing.
Thanks!
David




