Google DeepMindP152026-08-05abstract onlymoral judgmenthuman-ai trustai detectionanti-ai biasllm evaluation

A moral Turing test: How belief and source shape detection of and agreement with LLM judgments

People can spot AI-written moral justifications only somewhat better than chance, and they trust content less once they believe it is AI-generated, regardless of who actually wrote it.

Shows that deploying LLM moral judgments in sensitive systems faces a trust penalty tied to perceived source, separate from the judgment's actual quality.

Basile Garcia · Crystal Qian · Stefano Palminteri — meet the researchers →

How much of this do you want?
Orient me keeps four things: the abstract, the method, the claim↔evidence panel, and where it leads next. Everything adds constructs, the model table, every reported statistic, the discussion framing and the style moves. Switching hides nothing permanently and never changes what a section says — it only changes how many are on screen.
Abstract

Two readings, equal authority

How to choose: The paper’s words is verbatim — use it when you need to quote, or to judge how they write. Plain language is a paraphrase written for comprehension — use it when you want the idea fast. Neither is a summary of the other; they are two doors into the same room.

“As large language models (LLMs) are increasingly integrated into decision-making systems (e.g., autonomous vehicles and medical devices), understanding how humans perceive and evaluate AI-generated judgments is crucial. To investigate this, we conducted a series of experiments in which participants evaluated justifications for moral and non-moral choices, generated either by humans or LLMs. Participants attempted to identify the source of each justification (either human or LLM) and indicated their agreement with its content. We found that while detection accuracy was consistently above chance, it remained below 75%. In terms of agreement, there was no overall preference for human-generated responses, even though machine-generated justifications were favored in particularly challenging moral scenarios. Notably, we observed a systematic anti-AI bias: participants were less likely to agree with judgments they believed were AI-generated, regardless of the true source. Linguistic cues, such as response length, typos, first-person pronouns, and cost-benefit language markers (e.g., “lives,” “save”), influenced both detection and agreement. Participants tended to disagree with cost-benefit calculations, possibly due to an expectation that AI would favor such reasoning. These findings highlight the influence of motivated belief and ingroup/outgroup bias in shaping human evaluation of AI-generated content, particularly in morally sensitive contexts.”

Constructs

What this paper defines

Every definition below is the paper’s own sentence, with its locator. The plain gloss is a reading aid and is marked as one.

anti-AI bias

“we observed a systematic anti-AI bias: participants were less likely to agree with judgments they believed were AI-generated, regardless of the true source.”Abstract

In plain terms: People trusted a judgment less simply because they believed it came from an AI, even in cases where it actually came from a human.

Method

What they actually did

Each step is a synthesis. Open any step to see the paper’s own sentence it was derived from, with its locator — so nothing here floats free of the source.

The three-stage design described in the abstract: stimuli generation, then a source-detection judgment, then an agreement rating on the same justification.
Click any box to open it.
  1. Assembled a set of justifications for moral and non-moral choices, some written by humans and some generated by LLMs, to use as experimental stimuli.
    Trace this step to the paper
    “we conducted a series of experiments in which participants evaluated justifications for moral and non-moral choices, generated either by humans or LLMs.”Abstract
  2. Had participants guess whether each justification was written by a human or an LLM, and separately rate how much they agreed with its content.
    Trace this step to the paper
    “Participants attempted to identify the source of each justification (either human or LLM) and indicated their agreement with its content.”Abstract
  3. Examined whether surface linguistic features of the justifications, such as length, typos, first-person pronouns, and cost-benefit wording, related to participants' source-detection and agreement judgments.
    Trace this step to the paper
    “Linguistic cues, such as response length, typos, first-person pronouns, and cost-benefit language markers (e.g., “lives,” “save”), influenced both detection and agreement.”Abstract
The models under study

Exactly what was run, and how

What they reported — and what they left out

The page does not name which LLM(s) generated the justifications, nor does it report any model settings such as temperature, reasoning effort, or deployment mode; models_studied is empty because none are specified in this text.

Results

The numbers they report

Participants could distinguish AI-written from human-written justifications better than chance, but well short of reliably.

above chance, below 75%

See it in the paper
“We found that while detection accuracy was consistently above chance, it remained below 75%.”Abstract

Overall, participants did not systematically prefer human-generated justifications over machine-generated ones.

See it in the paper
“there was no overall preference for human-generated responses”Abstract

In particularly difficult moral scenarios, participants actually favored the machine-generated justifications.

See it in the paper
“machine-generated justifications were favored in particularly challenging moral scenarios”Abstract

Participants were less likely to agree with a judgment once they believed it was AI-generated, whether or not it actually was.

See it in the paper
“participants were less likely to agree with judgments they believed were AI-generated, regardless of the true source.”Abstract

Justification features such as length, typos, first-person language, and cost-benefit wording affected both whether people correctly detected the source and whether they agreed with it.

See it in the paper
“Linguistic cues, such as response length, typos, first-person pronouns, and cost-benefit language markers (e.g., “lives,” “save”), influenced both detection and agreement.”Abstract

Participants tended to push back against cost-benefit style reasoning, apparently because they expected AI to reason that way.

See it in the paper
“Participants tended to disagree with cost-benefit calculations, possibly due to an expectation that AI would favor such reasoning.”Abstract
Claim ↔ evidence

What they assert, beside what they showed

Left is the claim in the paper’s own words. Right is the data offered for it. Where the two do not fully meet, a gold band names the distance.

The claim

Humans can tell AI-generated moral justifications apart from human-written ones, but only somewhat reliably.

“We found that while detection accuracy was consistently above chance, it remained below 75%.”

The evidence

“We found that while detection accuracy was consistently above chance, it remained below 75%.”

Abstract
Mind the gap: The abstract states only that accuracy fell between chance and 75%, without giving the exact figure, so the claim of 'detection' is bounded rather than strong.
The claim

Content believed to be AI-generated is held to a harsher standard than the same content would receive if attributed to a human.

“we observed a systematic anti-AI bias: participants were less likely to agree with judgments they believed were AI-generated, regardless of the true source.”

The evidence

“participants were less likely to agree with judgments they believed were AI-generated, regardless of the true source.”

Abstract
Mind the gap: Claim and evidence are the same reported finding restated; the abstract offers no separate effect-size statistic for the bias.
The claim

People are suspicious of cost-benefit moral reasoning partly because they associate that style of reasoning with how they expect AI to reason.

“Participants tended to disagree with cost-benefit calculations, possibly due to an expectation that AI would favor such reasoning.”

The evidence

“Participants tended to disagree with cost-benefit calculations, possibly due to an expectation that AI would favor such reasoning.”

Abstract
Mind the gap: The authors explicitly hedge this causal mechanism with 'possibly due to' rather than demonstrating it directly.
Discussion & after

How they frame it, and what they want next

Their framing

The authors frame their findings as evidence that trust in AI-generated moral judgments is shaped by motivated reasoning and ingroup/outgroup dynamics rather than by the actual quality of the content, especially in morally sensitive contexts.

Register: The abstract states most findings directly ('We found...', 'Notably, we observed...') but hedges its one causal interpretation with 'possibly due to'.

Where they hedge

“possibly due to an expectation that AI would favor such reasoning”Abstract

What they say it means

  • Because people apply a harsher standard to content they believe is AI-generated, deploying LLM judgments in morally sensitive systems may face resistance independent of the judgments' actual quality.
    the paper’s words
    “These findings highlight the influence of motivated belief and ingroup/outgroup bias in shaping human evaluation of AI-generated content, particularly in morally sensitive contexts.”Abstract
For your own writing

Moves worth stealing

The title repurposes the well-known 'Turing test' concept to instantly signal the paper's question, translating a familiar AI-history frame into a new moral-judgment setting.

“A moral Turing test: How belief and source shape detection of and agreement with LLM judgments”

The abstract opens with real-world stakes (autonomous vehicles, medical devices) before narrowing to the lab experiment, lending outsized relevance to a fairly abstract detection-and-agreement study.

“As large language models (LLMs) are increasingly integrated into decision-making systems (e.g., autonomous vehicles and medical devices), understanding how humans perceive and evaluate AI-generated judgments is crucial.”
Connected

Where else this leads

What this page was built from

Working only from Google DeepMind's publication listing page (title, date, abstract, author list, and the venue name PLOS One) plus site navigation chrome; the full PLOS One paper's methods, results, and figures are not present in this source text.