A moral Turing test: How belief and source shape detection of and agreement with LLM judgments
People can spot AI-written moral justifications only somewhat better than chance, and they trust content less once they believe it is AI-generated, regardless of who actually wrote it.
Shows that deploying LLM moral judgments in sensitive systems faces a trust penalty tied to perceived source, separate from the judgment's actual quality.
Basile Garcia · Crystal Qian · Stefano Palminteri — meet the researchers →
Two readings, equal authority
How to choose: The paper’s words is verbatim — use it when you need to quote, or to judge how they write. Plain language is a paraphrase written for comprehension — use it when you want the idea fast. Neither is a summary of the other; they are two doors into the same room.
“As large language models (LLMs) are increasingly integrated into decision-making systems (e.g., autonomous vehicles and medical devices), understanding how humans perceive and evaluate AI-generated judgments is crucial. To investigate this, we conducted a series of experiments in which participants evaluated justifications for moral and non-moral choices, generated either by humans or LLMs. Participants attempted to identify the source of each justification (either human or LLM) and indicated their agreement with its content. We found that while detection accuracy was consistently above chance, it remained below 75%. In terms of agreement, there was no overall preference for human-generated responses, even though machine-generated justifications were favored in particularly challenging moral scenarios. Notably, we observed a systematic anti-AI bias: participants were less likely to agree with judgments they believed were AI-generated, regardless of the true source. Linguistic cues, such as response length, typos, first-person pronouns, and cost-benefit language markers (e.g., “lives,” “save”), influenced both detection and agreement. Participants tended to disagree with cost-benefit calculations, possibly due to an expectation that AI would favor such reasoning. These findings highlight the influence of motivated belief and ingroup/outgroup bias in shaping human evaluation of AI-generated content, particularly in morally sensitive contexts.”
Google DeepMind researchers had people judge whether moral and non-moral justifications were written by a human or an AI, and rate how much they agreed with each one. People could tell the difference somewhat, but not reliably, and they trusted a justification less once they believed it came from AI, even when it had not. Wording features like length, typos, and cost-benefit phrasing swayed both judgments.
What this paper defines
Every definition below is the paper’s own sentence, with its locator. The plain gloss is a reading aid and is marked as one.
anti-AI bias
“we observed a systematic anti-AI bias: participants were less likely to agree with judgments they believed were AI-generated, regardless of the true source.”Abstract
In plain terms: People trusted a judgment less simply because they believed it came from an AI, even in cases where it actually came from a human.
What they actually did
Each step is a synthesis. Open any step to see the paper’s own sentence it was derived from, with its locator — so nothing here floats free of the source.
- Assembled a set of justifications for moral and non-moral choices, some written by humans and some generated by LLMs, to use as experimental stimuli.
Trace this step to the paper
“we conducted a series of experiments in which participants evaluated justifications for moral and non-moral choices, generated either by humans or LLMs.”Abstract
- Had participants guess whether each justification was written by a human or an LLM, and separately rate how much they agreed with its content.
Trace this step to the paper
“Participants attempted to identify the source of each justification (either human or LLM) and indicated their agreement with its content.”Abstract
- Examined whether surface linguistic features of the justifications, such as length, typos, first-person pronouns, and cost-benefit wording, related to participants' source-detection and agreement judgments.
Trace this step to the paper
“Linguistic cues, such as response length, typos, first-person pronouns, and cost-benefit language markers (e.g., “lives,” “save”), influenced both detection and agreement.”Abstract
Exactly what was run, and how
What they reported — and what they left out
The page does not name which LLM(s) generated the justifications, nor does it report any model settings such as temperature, reasoning effort, or deployment mode; models_studied is empty because none are specified in this text.
The numbers they report
Participants could distinguish AI-written from human-written justifications better than chance, but well short of reliably.
above chance, below 75%
See it in the paper
“We found that while detection accuracy was consistently above chance, it remained below 75%.”Abstract
Overall, participants did not systematically prefer human-generated justifications over machine-generated ones.
See it in the paper
“there was no overall preference for human-generated responses”Abstract
In particularly difficult moral scenarios, participants actually favored the machine-generated justifications.
See it in the paper
“machine-generated justifications were favored in particularly challenging moral scenarios”Abstract
Participants were less likely to agree with a judgment once they believed it was AI-generated, whether or not it actually was.
See it in the paper
“participants were less likely to agree with judgments they believed were AI-generated, regardless of the true source.”Abstract
Justification features such as length, typos, first-person language, and cost-benefit wording affected both whether people correctly detected the source and whether they agreed with it.
See it in the paper
“Linguistic cues, such as response length, typos, first-person pronouns, and cost-benefit language markers (e.g., “lives,” “save”), influenced both detection and agreement.”Abstract
Participants tended to push back against cost-benefit style reasoning, apparently because they expected AI to reason that way.
See it in the paper
“Participants tended to disagree with cost-benefit calculations, possibly due to an expectation that AI would favor such reasoning.”Abstract
What they assert, beside what they showed
Left is the claim in the paper’s own words. Right is the data offered for it. Where the two do not fully meet, a gold band names the distance.
Humans can tell AI-generated moral justifications apart from human-written ones, but only somewhat reliably.
“We found that while detection accuracy was consistently above chance, it remained below 75%.”
“We found that while detection accuracy was consistently above chance, it remained below 75%.”
AbstractContent believed to be AI-generated is held to a harsher standard than the same content would receive if attributed to a human.
“we observed a systematic anti-AI bias: participants were less likely to agree with judgments they believed were AI-generated, regardless of the true source.”
“participants were less likely to agree with judgments they believed were AI-generated, regardless of the true source.”
AbstractPeople are suspicious of cost-benefit moral reasoning partly because they associate that style of reasoning with how they expect AI to reason.
“Participants tended to disagree with cost-benefit calculations, possibly due to an expectation that AI would favor such reasoning.”
“Participants tended to disagree with cost-benefit calculations, possibly due to an expectation that AI would favor such reasoning.”
AbstractHow they frame it, and what they want next
Their framing
The authors frame their findings as evidence that trust in AI-generated moral judgments is shaped by motivated reasoning and ingroup/outgroup dynamics rather than by the actual quality of the content, especially in morally sensitive contexts.
Register: The abstract states most findings directly ('We found...', 'Notably, we observed...') but hedges its one causal interpretation with 'possibly due to'.
Where they hedge
“possibly due to an expectation that AI would favor such reasoning”Abstract
What they say it means
- Because people apply a harsher standard to content they believe is AI-generated, deploying LLM judgments in morally sensitive systems may face resistance independent of the judgments' actual quality.
the paper’s words
“These findings highlight the influence of motivated belief and ingroup/outgroup bias in shaping human evaluation of AI-generated content, particularly in morally sensitive contexts.”Abstract
Moves worth stealing
The title repurposes the well-known 'Turing test' concept to instantly signal the paper's question, translating a familiar AI-history frame into a new moral-judgment setting.
“A moral Turing test: How belief and source shape detection of and agreement with LLM judgments”
The abstract opens with real-world stakes (autonomous vehicles, medical devices) before narrowing to the lab experiment, lending outsized relevance to a fairly abstract detection-and-agreement study.
“As large language models (LLMs) are increasingly integrated into decision-making systems (e.g., autonomous vehicles and medical devices), understanding how humans perceive and evaluate AI-generated judgments is crucial.”
Where else this leads
Same people
- Real-Time Group Dynamics with LLM Facilitation: Evidence from a Charity Allocation Task Google DeepMind
shares Crystal Qian
Published alongside it
The nearest publications in time, across all three labs.
- Ten advances in mathematics and theoretical computer science OpenAI
2026-08-01 - Learning more about Claude's mathematical capabilities Anthropic
2026-08-10 - How enabling two settings tripled our scores on the ARC-AGI-3 benchmark OpenAI
2026-07-29 - Patterns and problems in emerging multiagent systems Anthropic
2026-08-13
What this page was built from
Working only from Google DeepMind's publication listing page (title, date, abstract, author list, and the venue name PLOS One) plus site navigation chrome; the full PLOS One paper's methods, results, and figures are not present in this source text.