Google DeepMindP312026-06-26abstract onlyred-teamingtext-to-imageai-safetybiasadversarial-prompts

Bridging the Scale Gap: Augmenting Human Red-Teaming to Uncover Latent Risks in T2I Models

Seed2Harvest expands human-authored adversarial prompts using sociolinguistic attack strategies, achieving ~20x more demographic and geographic coverage in T2I red-teaming without more human effort.

It offers T2I safety teams a way to multiply red-teaming coverage across cultures and demographics without diluting the human judgment that catches subtle harms.

Jessica Quaye · Alicia Parrish · Charvi Rastogi · Minsuk Kahng · Oana Inel · Lora Aroyo · Vijay Janapa Reddi — meet the researchers →

How much of this do you want?
Orient me keeps four things: the abstract, the method, the claim↔evidence panel, and where it leads next. Everything adds constructs, the model table, every reported statistic, the discussion framing and the style moves. Switching hides nothing permanently and never changes what a section says — it only changes how many are on screen.
Abstract

Two readings, equal authority

How to choose: The paper’s words is verbatim — use it when you need to quote, or to judge how they write. Plain language is a paraphrase written for comprehension — use it when you want the idea fast. Neither is a summary of the other; they are two doors into the same room.

“Human red-teaming is essential for identifying culturally-specific and context-dependent harms in Text-to-Image (T2I) models, yet it faces a fundamental "scale gap": human insight is resource-intensive, while automated approaches lack the sociolinguistic nuance to detect subtle failures. This leaves models vulnerable to "implicitly adversarial" prompts -- inputs that appear benign but trigger unsafe or biased generations, disproportionately affecting users from underrepresented communities. We introduce Seed2Harvest, a hybrid framework that bridges this gap by operationalizing human expertise rather than replacing it: human-authored adversarial prompts serve as "seeds" systematically expanded using sociolinguistic attack strategies distilled through reflexive thematic analysis of 3,748 human-crafted adversarial prompts. These human-derived strategies provide the structured guidance directing prompt expansion, distinguishing our approach from zero-shot synthetic generation. Our approach achieves what neither paradigm accomplishes alone: balanced threat discovery across harm categories, without proportional increases in human auditor effort. This pattern holds across three evaluation datasets (Adversarial Nibbler, I2P, and CoPro), with expanded datasets preserving attack effectiveness comparable to human baselines while increasing geographic and demographic coverage by a factor of ~20x on average. Our work demonstrates that an effective path to comprehensive T2I safety evaluation is not replacing human auditors with automation, but systematically amplifying what makes them irreplaceable.”

Constructs

What this paper defines

Every definition below is the paper’s own sentence, with its locator. The plain gloss is a reading aid and is marked as one.

scale gap

“it faces a fundamental "scale gap": human insight is resource-intensive, while automated approaches lack the sociolinguistic nuance to detect subtle failures”Abstract

In plain terms: The tradeoff between human red-teamers, who catch nuance but are too few, and automated tools, which scale but miss cultural subtlety.

implicitly adversarial prompts

“"implicitly adversarial" prompts -- inputs that appear benign but trigger unsafe or biased generations, disproportionately affecting users from underrepresented communities”Abstract

In plain terms: Prompts that look harmless but still make a model produce unsafe or biased images, hurting underrepresented groups the most.

Seed2Harvest

“a hybrid framework that bridges this gap by operationalizing human expertise rather than replacing it: human-authored adversarial prompts serve as "seeds" systematically expanded using sociolinguistic attack strategies distilled through reflexive thematic analysis of 3,748 human-crafted adversarial prompts”Abstract

In plain terms: A framework that expands human-written risky prompts into many more using attack patterns learned from analyzing thousands of other human-written risky prompts.

Method

What they actually did

Each step is a synthesis. Open any step to see the paper’s own sentence it was derived from, with its locator — so nothing here floats free of the source.

A pipeline from human-seed prompts through thematic analysis of attack strategies to strategy-guided expansion and benchmark evaluation.
Click any box to open it.
  1. Distilled sociolinguistic attack strategies from a reflexive thematic analysis of thousands of human-crafted adversarial prompts.
    Trace this step to the paper
    “distilled through reflexive thematic analysis of 3,748 human-crafted adversarial prompts”Abstract
  2. Used human-authored adversarial prompts as seed prompts, then systematically expanded them via the distilled attack strategies.
    Trace this step to the paper
    “human-authored adversarial prompts serve as "seeds" systematically expanded using sociolinguistic attack strategies”Abstract
  3. Distinguished this seeded-expansion approach from purely zero-shot synthetic prompt generation, arguing the human-derived strategies are what direct the expansion.
    Trace this step to the paper
    “These human-derived strategies provide the structured guidance directing prompt expansion, distinguishing our approach from zero-shot synthetic generation.”Abstract
  4. Evaluated the expanded prompt sets on three benchmark datasets, comparing attack effectiveness and demographic/geographic coverage against human baselines.
    Trace this step to the paper
    “This pattern holds across three evaluation datasets (Adversarial Nibbler, I2P, and CoPro), with expanded datasets preserving attack effectiveness comparable to human baselines while increasing geographic and demographic coverage by a factor of ~20x on average.”Abstract
  5. Concluded that comprehensive T2I safety evaluation is better served by amplifying human auditors than by replacing them with automation.
    Trace this step to the paper
    “an effective path to comprehensive T2I safety evaluation is not replacing human auditors with automation, but systematically amplifying what makes them irreplaceable”Abstract
The models under study

Exactly what was run, and how

What they reported — and what they left out

The abstract does not name any specific T2I model, temperature, effort level, or deployment setting; it names only the evaluation datasets used (Adversarial Nibbler, I2P, CoPro), not the generative models evaluated against them.

Results

The numbers they report

The attack strategies were distilled from a large corpus of human-crafted adversarial prompts.

3,748 human-crafted adversarial prompts

See it in the paper
“reflexive thematic analysis of 3,748 human-crafted adversarial prompts”Abstract

Expanded prompt sets substantially increased demographic and geographic coverage relative to the human-only baseline.

~20x average increase in geographic and demographic coverage

See it in the paper
“increasing geographic and demographic coverage by a factor of ~20x on average”Abstract

Expanded datasets kept attack effectiveness on par with the original human-authored prompts.

See it in the paper
“expanded datasets preserving attack effectiveness comparable to human baselines”Abstract
Claim ↔ evidence

What they assert, beside what they showed

Left is the claim in the paper’s own words. Right is the data offered for it. Where the two do not fully meet, a gold band names the distance.

The claim

The framework achieves balanced threat discovery across harm categories without a proportional increase in human auditor effort.

“Our approach achieves what neither paradigm accomplishes alone: balanced threat discovery across harm categories, without proportional increases in human auditor effort.”

The evidence

“This pattern holds across three evaluation datasets (Adversarial Nibbler, I2P, and CoPro), with expanded datasets preserving attack effectiveness comparable to human baselines while increasing geographic and demographic coverage by a factor of ~20x on average.”

Abstract
Mind the gap: No labor-hours or cost figures for human auditor effort are reported in either condition, so 'without proportional increases in effort' is asserted rather than shown with a matching statistic.
The claim

Amplifying human expertise, rather than replacing it with automation, is the effective path to comprehensive T2I safety evaluation.

“Our work demonstrates that an effective path to comprehensive T2I safety evaluation is not replacing human auditors with automation, but systematically amplifying what makes them irreplaceable.”

The evidence

“human-authored adversarial prompts serve as "seeds" systematically expanded using sociolinguistic attack strategies distilled through reflexive thematic analysis of 3,748 human-crafted adversarial prompts”

Abstract
The claim

Structured, human-derived attack strategies produce a meaningfully different result than zero-shot synthetic prompt generation.

“These human-derived strategies provide the structured guidance directing prompt expansion, distinguishing our approach from zero-shot synthetic generation.”

The evidence

“expanded datasets preserving attack effectiveness comparable to human baselines while increasing geographic and demographic coverage by a factor of ~20x on average”

Abstract
Mind the gap: The abstract asserts a distinction from zero-shot synthetic generation but reports no head-to-head comparison against an actual zero-shot baseline's own coverage or effectiveness numbers.
Discussion & after

How they frame it, and what they want next

Their framing

The authors frame their contribution as resolving a false choice between two paradigms that each fail alone -- thorough-but-unscalable human red-teaming and scalable-but-nuance-blind automation -- by using automation to amplify human insight rather than substitute for it.

Register: The abstract is written in confident, unhedged declarative sentences ('We introduce', 'Our approach achieves', 'Our work demonstrates'), with its only explicit scope qualifier being the three named evaluation datasets.

Where they hedge

“This pattern holds across three evaluation datasets (Adversarial Nibbler, I2P, and CoPro)”Abstract

What they say it means

  • Scaling red-team coverage across languages, cultures, and demographics may not require abandoning human judgment, only amplifying it more efficiently.
    the paper’s words
    “an effective path to comprehensive T2I safety evaluation is not replacing human auditors with automation, but systematically amplifying what makes them irreplaceable”Abstract
For your own writing

Moves worth stealing

Names the core problem with a memorable metaphor ('scale gap') and defines it in the same breath, giving readers a handle for the thesis before any results appear.

“it faces a fundamental "scale gap": human insight is resource-intensive, while automated approaches lack the sociolinguistic nuance to detect subtle failures”

States the contribution as resolving a false binary between two paradigms, positioning the method as synthesis rather than competition.

“Our approach achieves what neither paradigm accomplishes alone: balanced threat discovery across harm categories, without proportional increases in human auditor effort.”

Closes on a quotable, values-forward thesis sentence rather than a results summary, framing the technical contribution as a statement about the value of human expertise.

“systematically amplifying what makes them irreplaceable”
Connected

Where else this leads

Same people

Published alongside it

The nearest publications in time, across all three labs.

What this page was built from

This record is built only from the DeepMind publication landing page's abstract, author list, and venue; the full paper (FAccT 2026) is behind ACM's paywall and was not available for extraction.