Real-Time Group Dynamics with LLM Facilitation: Evidence from a Charity Allocation Task
Across two studies (N=879), LLM facilitation of group deliberation did not improve consensus, but it measurably steered charity-allocation outcomes and created a false sense of inclusion.
It shows people can trust and prefer an AI facilitator that is quietly steering real-money group decisions without making the group's participation any more equitable.
Aaron Parisi · Nithum Thain · Alden Hallak · Vivian Tsai · Crystal Qian — meet the researchers →
Two readings, equal authority
How to choose: The paper’s words is verbatim — use it when you need to quote, or to judge how they write. Plain language is a paraphrase written for comprehension — use it when you want the idea fast. Neither is a summary of the other; they are two doors into the same room.
“As large language models (LLMs) evolve from single-user assistants to active participants in civic and workplace deliberation, evaluating their effects on collective decision making becomes a governance challenge. We present two empirical studies (N=879) of real-time, text-based group deliberation in an incentive-compatible charity allocation task with real financial stakes ($7,200 USD). Groups of three allocate a donation budget under varying LLM facilitation conditions: Study 1 (N=204) compares three frontier models; Study 2 (N=675) compares facilitator strategies against a no-facilitation baseline. Across both studies, LLM facilitation did not significantly improve group consensus in either study, yet participants consistently preferred facilitated discussion. We additionally identify two governance-relevant risks. First, algorithmic steering : facilitators shifted select charity-level allocations by up to 5.5 percentage points—directly affecting the final charitable payout—even when aggregate agreement metrics remained unchanged. Second, an illusion of inclusion : participants cited inclusivity as their primary reason for preferring LLM facilitators, yet neither survey nor transcript-based measures of participation equity improved. Notably, participants reported greater trust in the process under the same conditions where facilitators exerted directional influence on outcomes. Together, these findings show that in AI-mediated group deliberation, perceived procedural improvement can coexist with measurable steering and unchanged participation inequality, motivating evaluation practices that treat collective outcomes, interaction dynamics, and participant perceptions as distinct governance targets.”
Researchers ran two real-money experiments with 879 people split into three-person groups deciding how to allocate charity donations, sometimes with an AI facilitator helping the discussion. Having an AI facilitator did not make groups reach agreement any more often, but people liked having one anyway. The AI's involvement nudged the actual charity payouts by several percentage points even though standard agreement scores looked unchanged, and people felt the process was more inclusive and trustworthy even though real measures of who participated didn't actually improve.
What this paper defines
Every definition below is the paper’s own sentence, with its locator. The plain gloss is a reading aid and is marked as one.
algorithmic steering
“algorithmic steering : facilitators shifted select charity-level allocations by up to 5.5 percentage points—directly affecting the final charitable payout—even when aggregate agreement metrics remained unchanged”Abstract
In plain terms: When an AI facilitator's involvement measurably shifts what a group actually decides, even though standard agreement scores show no change.
illusion of inclusion
“an illusion of inclusion : participants cited inclusivity as their primary reason for preferring LLM facilitators, yet neither survey nor transcript-based measures of participation equity improved”Abstract
In plain terms: When people feel a discussion was more inclusive because an AI facilitated it, even though objective measures of who actually participated didn't improve.
What they actually did
Each step is a synthesis. Open any step to see the paper’s own sentence it was derived from, with its locator — so nothing here floats free of the source.
- Ran two empirical studies with a combined 879 participants on a real-time, text-based group deliberation task with real financial stakes.
Trace this step to the paper
“We present two empirical studies (N=879) of real-time, text-based group deliberation in an incentive-compatible charity allocation task with real financial stakes ($7,200 USD).”Abstract
- Had groups of three participants allocate a real donation budget together under varying LLM facilitation conditions.
Trace this step to the paper
“Groups of three allocate a donation budget under varying LLM facilitation conditions”Abstract
- In Study 1, with 204 participants, compared three frontier models as facilitators.
Trace this step to the paper
“Study 1 (N=204) compares three frontier models”Abstract
- In Study 2, with 675 participants, compared facilitator strategies against a no-facilitation baseline.
Trace this step to the paper
“Study 2 (N=675) compares facilitator strategies against a no-facilitation baseline”Abstract
- Measured participation equity with both survey-based and transcript-based methods, alongside consensus and allocation outcomes.
Trace this step to the paper
“neither survey nor transcript-based measures of participation equity improved”Abstract
Exactly what was run, and how
What they reported — and what they left out
The abstract reports that Study 1 compared three frontier models as facilitators but does not name them, and gives no temperature, effort level, or deployment details for any model used.
The numbers they report
The two studies together enrolled 879 participants.
N=879
See it in the paper
“We present two empirical studies (N=879)”Abstract
Study 1 compared three frontier models with 204 participants.
N=204
See it in the paper
“Study 1 (N=204) compares three frontier models”Abstract
Study 2 compared facilitator strategies against a no-facilitation baseline with 675 participants.
N=675
See it in the paper
“Study 2 (N=675) compares facilitator strategies against a no-facilitation baseline”Abstract
The charity allocation task carried real financial stakes.
$7,200 USD
See it in the paper
“real financial stakes ($7,200 USD)”Abstract
LLM facilitation did not produce a statistically significant improvement in group consensus in either study.
See it in the paper
“LLM facilitation did not significantly improve group consensus in either study”Abstract
Facilitators shifted specific charity-level allocation percentages by a meaningful margin even as aggregate agreement metrics stayed flat.
up to 5.5 percentage points
See it in the paper
“facilitators shifted select charity-level allocations by up to 5.5 percentage points”Abstract
What they assert, beside what they showed
Left is the claim in the paper’s own words. Right is the data offered for it. Where the two do not fully meet, a gold band names the distance.
Perceived procedural improvement in AI-mediated deliberation can coexist with real steering of outcomes and unchanged participation inequality.
“Together, these findings show that in AI-mediated group deliberation, perceived procedural improvement can coexist with measurable steering and unchanged participation inequality, motivating evaluation practices that treat collective outcomes, interaction dynamics, and participant perceptions as distinct governance targets.”
“facilitators shifted select charity-level allocations by up to 5.5 percentage points—directly affecting the final charitable payout—even when aggregate agreement metrics remained unchanged”
AbstractParticipants experience an illusion of inclusion, valuing LLM facilitation for its perceived inclusivity despite no real gain in participation equity.
“Second, an illusion of inclusion : participants cited inclusivity as their primary reason for preferring LLM facilitators, yet neither survey nor transcript-based measures of participation equity improved.”
“neither survey nor transcript-based measures of participation equity improved”
AbstractParticipant trust in the process rose specifically under the same conditions where facilitators were exerting directional influence on outcomes.
“Notably, participants reported greater trust in the process under the same conditions where facilitators exerted directional influence on outcomes.”
“facilitators shifted select charity-level allocations by up to 5.5 percentage points—directly affecting the final charitable payout”
AbstractParticipants preferred facilitated discussion even though facilitation did not significantly improve consensus.
“LLM facilitation did not significantly improve group consensus in either study, yet participants consistently preferred facilitated discussion.”
“yet participants consistently preferred facilitated discussion”
AbstractHow they frame it, and what they want next
Their framing
The authors frame their two named risks as a caution for AI governance: an LLM facilitator can feel more inclusive and trustworthy to participants while simultaneously exerting real, unmeasured influence over group outcomes, so evaluation should separate outcomes, interaction dynamics, and perceptions rather than treating a good 'feel' as proof of a good process.
Register: The abstract states its null result on consensus plainly rather than downplaying it, but frames the steering and inclusion findings assertively as named, numbered risks ('First, ... Second, ...') rather than as tentative observations.
Where they hedge
“LLM facilitation did not significantly improve group consensus in either study”Abstract
What they say it means
- Standard aggregate agreement metrics can miss real, consequential steering by an AI facilitator, so governance evaluations need finer-grained, outcome-level measures.
the paper’s words
“even when aggregate agreement metrics remained unchanged”Abstract
- Perceived inclusivity and trust are not reliable proxies for actual participation equity when an LLM facilitates group deliberation.
the paper’s words
“participants cited inclusivity as their primary reason for preferring LLM facilitators, yet neither survey nor transcript-based measures of participation equity improved”Abstract
What they call for next
- Evaluation practices for AI-mediated deliberation should treat collective outcomes, interaction dynamics, and participant perceptions as distinct governance targets rather than one combined measure.
the paper’s words
“motivating evaluation practices that treat collective outcomes, interaction dynamics, and participant perceptions as distinct governance targets”Abstract
Moves worth stealing
Numbers its two key risk findings explicitly ('First, ... Second, ...') and gives each a short named construct, making the contribution easy to cite and remember.
“First, algorithmic steering : facilitators shifted select charity-level allocations by up to 5.5 percentage points”
Opens with a governance framing sentence before any methods or results, establishing why the reader should care before describing the study.
“evaluating their effects on collective decision making becomes a governance challenge”
States the null result on the headline metric (consensus) up front and unhedged, rather than burying a non-significant finding.
“LLM facilitation did not significantly improve group consensus in either study, yet participants consistently preferred facilitated discussion.”
Where else this leads
Same people
- A moral Turing test: How belief and source shape detection of and agreement with LLM judgments Google DeepMind
shares Crystal Qian
Published alongside it
The nearest publications in time, across all three labs.
- Bridging the Scale Gap: Augmenting Human Red-Teaming to Uncover Latent Risks in T2I Models Google DeepMind
2026-06-26 - Going PLACES: Participatory Localized Red Teaming for Text-to-Image Safety in the Global South Google DeepMind
2026-06-25 - Introducing GeneBench-Pro OpenAI
2026-06-30 - Towards Structural Understanding of LLM Overthinking Google DeepMind
2026-07-02
What this page was built from
This record is built only from the DeepMind publication landing page's abstract, author list, and venue; the full paper (FAccT 2026) is behind ACM's paywall and was not available for extraction, and the page's separate 'Download' link target was not captured in the extracted text.