Bridging the Scale Gap: Augmenting Human Red-Teaming to Uncover Latent Risks in T2I Models
Seed2Harvest expands human-authored adversarial prompts using sociolinguistic attack strategies, achieving ~20x more demographic and geographic coverage in T2I red-teaming without more human effort.
It offers T2I safety teams a way to multiply red-teaming coverage across cultures and demographics without diluting the human judgment that catches subtle harms.
Jessica Quaye · Alicia Parrish · Charvi Rastogi · Minsuk Kahng · Oana Inel · Lora Aroyo · Vijay Janapa Reddi — meet the researchers →
Two readings, equal authority
How to choose: The paper’s words is verbatim — use it when you need to quote, or to judge how they write. Plain language is a paraphrase written for comprehension — use it when you want the idea fast. Neither is a summary of the other; they are two doors into the same room.
“Human red-teaming is essential for identifying culturally-specific and context-dependent harms in Text-to-Image (T2I) models, yet it faces a fundamental "scale gap": human insight is resource-intensive, while automated approaches lack the sociolinguistic nuance to detect subtle failures. This leaves models vulnerable to "implicitly adversarial" prompts -- inputs that appear benign but trigger unsafe or biased generations, disproportionately affecting users from underrepresented communities. We introduce Seed2Harvest, a hybrid framework that bridges this gap by operationalizing human expertise rather than replacing it: human-authored adversarial prompts serve as "seeds" systematically expanded using sociolinguistic attack strategies distilled through reflexive thematic analysis of 3,748 human-crafted adversarial prompts. These human-derived strategies provide the structured guidance directing prompt expansion, distinguishing our approach from zero-shot synthetic generation. Our approach achieves what neither paradigm accomplishes alone: balanced threat discovery across harm categories, without proportional increases in human auditor effort. This pattern holds across three evaluation datasets (Adversarial Nibbler, I2P, and CoPro), with expanded datasets preserving attack effectiveness comparable to human baselines while increasing geographic and demographic coverage by a factor of ~20x on average. Our work demonstrates that an effective path to comprehensive T2I safety evaluation is not replacing human auditors with automation, but systematically amplifying what makes them irreplaceable.”
Human reviewers catch subtle, culturally-specific harms in AI image generators that automated tools miss, but there are never enough human reviewers to cover everything. The authors built Seed2Harvest, which takes prompts humans already flagged as risky and multiplies them using attack patterns learned from studying thousands of those human-written prompts, instead of having an AI invent new risky prompts from scratch. Across three benchmark datasets, the expanded prompts stayed as effective as the human originals while covering roughly 20 times more geographic and demographic ground.
What this paper defines
Every definition below is the paper’s own sentence, with its locator. The plain gloss is a reading aid and is marked as one.
scale gap
“it faces a fundamental "scale gap": human insight is resource-intensive, while automated approaches lack the sociolinguistic nuance to detect subtle failures”Abstract
In plain terms: The tradeoff between human red-teamers, who catch nuance but are too few, and automated tools, which scale but miss cultural subtlety.
implicitly adversarial prompts
“"implicitly adversarial" prompts -- inputs that appear benign but trigger unsafe or biased generations, disproportionately affecting users from underrepresented communities”Abstract
In plain terms: Prompts that look harmless but still make a model produce unsafe or biased images, hurting underrepresented groups the most.
Seed2Harvest
“a hybrid framework that bridges this gap by operationalizing human expertise rather than replacing it: human-authored adversarial prompts serve as "seeds" systematically expanded using sociolinguistic attack strategies distilled through reflexive thematic analysis of 3,748 human-crafted adversarial prompts”Abstract
In plain terms: A framework that expands human-written risky prompts into many more using attack patterns learned from analyzing thousands of other human-written risky prompts.
What they actually did
Each step is a synthesis. Open any step to see the paper’s own sentence it was derived from, with its locator — so nothing here floats free of the source.
- Distilled sociolinguistic attack strategies from a reflexive thematic analysis of thousands of human-crafted adversarial prompts.
Trace this step to the paper
“distilled through reflexive thematic analysis of 3,748 human-crafted adversarial prompts”Abstract
- Used human-authored adversarial prompts as seed prompts, then systematically expanded them via the distilled attack strategies.
Trace this step to the paper
“human-authored adversarial prompts serve as "seeds" systematically expanded using sociolinguistic attack strategies”Abstract
- Distinguished this seeded-expansion approach from purely zero-shot synthetic prompt generation, arguing the human-derived strategies are what direct the expansion.
Trace this step to the paper
“These human-derived strategies provide the structured guidance directing prompt expansion, distinguishing our approach from zero-shot synthetic generation.”Abstract
- Evaluated the expanded prompt sets on three benchmark datasets, comparing attack effectiveness and demographic/geographic coverage against human baselines.
Trace this step to the paper
“This pattern holds across three evaluation datasets (Adversarial Nibbler, I2P, and CoPro), with expanded datasets preserving attack effectiveness comparable to human baselines while increasing geographic and demographic coverage by a factor of ~20x on average.”Abstract
- Concluded that comprehensive T2I safety evaluation is better served by amplifying human auditors than by replacing them with automation.
Trace this step to the paper
“an effective path to comprehensive T2I safety evaluation is not replacing human auditors with automation, but systematically amplifying what makes them irreplaceable”Abstract
Exactly what was run, and how
What they reported — and what they left out
The abstract does not name any specific T2I model, temperature, effort level, or deployment setting; it names only the evaluation datasets used (Adversarial Nibbler, I2P, CoPro), not the generative models evaluated against them.
The numbers they report
The attack strategies were distilled from a large corpus of human-crafted adversarial prompts.
3,748 human-crafted adversarial prompts
See it in the paper
“reflexive thematic analysis of 3,748 human-crafted adversarial prompts”Abstract
Expanded prompt sets substantially increased demographic and geographic coverage relative to the human-only baseline.
~20x average increase in geographic and demographic coverage
See it in the paper
“increasing geographic and demographic coverage by a factor of ~20x on average”Abstract
Expanded datasets kept attack effectiveness on par with the original human-authored prompts.
See it in the paper
“expanded datasets preserving attack effectiveness comparable to human baselines”Abstract
What they assert, beside what they showed
Left is the claim in the paper’s own words. Right is the data offered for it. Where the two do not fully meet, a gold band names the distance.
The framework achieves balanced threat discovery across harm categories without a proportional increase in human auditor effort.
“Our approach achieves what neither paradigm accomplishes alone: balanced threat discovery across harm categories, without proportional increases in human auditor effort.”
“This pattern holds across three evaluation datasets (Adversarial Nibbler, I2P, and CoPro), with expanded datasets preserving attack effectiveness comparable to human baselines while increasing geographic and demographic coverage by a factor of ~20x on average.”
AbstractAmplifying human expertise, rather than replacing it with automation, is the effective path to comprehensive T2I safety evaluation.
“Our work demonstrates that an effective path to comprehensive T2I safety evaluation is not replacing human auditors with automation, but systematically amplifying what makes them irreplaceable.”
“human-authored adversarial prompts serve as "seeds" systematically expanded using sociolinguistic attack strategies distilled through reflexive thematic analysis of 3,748 human-crafted adversarial prompts”
AbstractStructured, human-derived attack strategies produce a meaningfully different result than zero-shot synthetic prompt generation.
“These human-derived strategies provide the structured guidance directing prompt expansion, distinguishing our approach from zero-shot synthetic generation.”
“expanded datasets preserving attack effectiveness comparable to human baselines while increasing geographic and demographic coverage by a factor of ~20x on average”
AbstractHow they frame it, and what they want next
Their framing
The authors frame their contribution as resolving a false choice between two paradigms that each fail alone -- thorough-but-unscalable human red-teaming and scalable-but-nuance-blind automation -- by using automation to amplify human insight rather than substitute for it.
Register: The abstract is written in confident, unhedged declarative sentences ('We introduce', 'Our approach achieves', 'Our work demonstrates'), with its only explicit scope qualifier being the three named evaluation datasets.
Where they hedge
“This pattern holds across three evaluation datasets (Adversarial Nibbler, I2P, and CoPro)”Abstract
What they say it means
- Scaling red-team coverage across languages, cultures, and demographics may not require abandoning human judgment, only amplifying it more efficiently.
the paper’s words
“an effective path to comprehensive T2I safety evaluation is not replacing human auditors with automation, but systematically amplifying what makes them irreplaceable”Abstract
Moves worth stealing
Names the core problem with a memorable metaphor ('scale gap') and defines it in the same breath, giving readers a handle for the thesis before any results appear.
“it faces a fundamental "scale gap": human insight is resource-intensive, while automated approaches lack the sociolinguistic nuance to detect subtle failures”
States the contribution as resolving a false binary between two paradigms, positioning the method as synthesis rather than competition.
“Our approach achieves what neither paradigm accomplishes alone: balanced threat discovery across harm categories, without proportional increases in human auditor effort.”
Closes on a quotable, values-forward thesis sentence rather than a results summary, framing the technical contribution as a statement about the value of human expertise.
“systematically amplifying what makes them irreplaceable”
Where else this leads
Same people
- Going PLACES: Participatory Localized Red Teaming for Text-to-Image Safety in the Global South Google DeepMind
shares Charvi Rastogi, Minsuk Kahng, Alicia Parrish, Jessica Quaye, Vijay Janapa Reddi, Lora Aroyo
Same territory
- Going PLACES: Participatory Localized Red Teaming for Text-to-Image Safety in the Global South Google DeepMind
red-teaming ai-safety - Fine-Tuned Lie Detectors Failed to Generalize Anthropic
ai-safety - Diffuse AI Control on Fuzzy Tasks Anthropic
red-teaming - Solipsistic superintelligence is unlikely to be cooperative Google DeepMind
evaluation - Gram: Assessing sabotage propensities via automated alignment auditing Google DeepMind
evaluation
Published alongside it
The nearest publications in time, across all three labs.
- Real-Time Group Dynamics with LLM Facilitation: Evidence from a Charity Allocation Task Google DeepMind
2026-06-26 - Going PLACES: Participatory Localized Red Teaming for Text-to-Image Safety in the Global South Google DeepMind
2026-06-25 - Introducing GeneBench-Pro OpenAI
2026-06-30 - Towards Structural Understanding of LLM Overthinking Google DeepMind
2026-07-02
What this page was built from
This record is built only from the DeepMind publication landing page's abstract, author list, and venue; the full paper (FAccT 2026) is behind ACM's paywall and was not available for extraction.