AnthropicP102026-08-15lab post onlyai safetyllm-as-judgebenchmarksresearch evaluationhuman-ai agreement

TASTE: Can AI Models Judge AI Safety Research Proposals?

Anthropic's TASTE benchmark tests whether AI models can judge AI safety research proposals as well as expert humans; the best model still trails human researchers.

It shows current frontier models can't yet reliably judge research quality, a capability automating AI safety research would require.

Hasan Baig · Hailey Joren · Joe Benton — meet the researchers →

How much of this do you want?
Orient me keeps four things: the abstract, the method, the claim↔evidence panel, and where it leads next. Everything adds constructs, the model table, every reported statistic, the discussion framing and the style moves. Switching hides nothing permanently and never changes what a section says — it only changes how many are on screen.
Abstract

Two readings, equal authority

How to choose: The paper’s words is verbatim — use it when you need to quote, or to judge how they write. Plain language is a paraphrase written for comprehension — use it when you want the idea fast. Neither is a summary of the other; they are two doors into the same room.

This source carries no verbatim abstract.

Constructs

What this paper defines

Every definition below is the paper’s own sentence, with its locator. The plain gloss is a reading aid and is marked as one.

TASTE (The AI Safety Taste Evaluation)

“we present TASTE (The AI Safety Taste Evaluation), a benchmark of experienced AI safety researchers’ preferences over empirical safety research proposals”Building a Research Judgment Benchmark (TASTE)

In plain terms: TASTE is a benchmark of paired AI safety research proposals scored by whether models agree with expert human preferences.

Pair-discussion protocol

“For each prompt, four researchers gave their feedback individually, discussed disagreements in pairs, then revised their feedback.”Pair-discussion protocol (figure caption)

In plain terms: Four researchers rate proposals alone, discuss disagreements in pairs, then update their scores.

Strong-confidence filtering

“filtering researchers’ labels for self-reported "strong" confidence”tl;dr

In plain terms: Only keeping human preference labels where the rater reported high confidence in their own judgment.

Standard setup (model evaluation)

“a standard setup where the model sees two proposals in context and gives a probability to each proposal winning which we binarize to give a preference”Evaluating Models’ Research Judgment

In plain terms: The model reads both proposals together and estimates the probability each one wins, which is converted into a single preferred choice.

Single-proposal scoring setup

“a tougher single-proposal scoring setup where the model sees each proposal individually in context, scores it, and is assessed on the implied preferences from its scores against the preferences in TASTE”Evaluating Models’ Research Judgment

In plain terms: The model scores each proposal on its own, without seeing its pair, and its scores are compared afterward to infer a preference.

Anchor preferences

“we compare the preferences from that condition (anchor preferences) with the preferences of the two researchers in the opposing discussion pair”figure caption after 'Strong-confidence, post-discussion preferences...'

In plain terms: The preference labels from one condition, used as the reference point checked for agreement against a different pair of researchers.

Method

What they actually did

Each step is a synthesis. Open any step to see the paper’s own sentence it was derived from, with its locator — so nothing here floats free of the source.

The TASTE construction pipeline, from human-written proposals and reverse-engineered prompts through model-generated proposals, individual and post-discussion human scoring, confidence/gap filtering, to the final 92-pair benchmark used to evaluate models.
Click any box to open it.
  1. The team started from 93 human-written research proposals from Anthropic's Fellows Program and used Claude Opus 4.6 to reverse-engineer prompts that could have motivated each one, supplementing with self-written prompts.
    Trace this step to the paper
    “we started with a collection of 93 human-written research proposals presented in Anthropic’s Fellows Program. We used Claude Opus 4.6 to reverse-engineer prompts that could motivate each human proposal, and supplemented this set with prompts we wrote ourselves.”Building a Research Judgment Benchmark (TASTE)
  2. Using those prompts, they generated AI-written research proposals with a prompt scaffold that varied paper summaries in context for diversity.
    Trace this step to the paper
    “We generated research proposals using a prompt scaffold that varied paper summaries in context for diversity.”Building a Research Judgment Benchmark (TASTE)
  3. AI safety researchers individually scored each proposal on three axes (overall, high-level, approach), ranked the proposals with ties allowed, and reported per-prompt confidence.
    Trace this step to the paper
    “Researchers scored each proposal 1–5 on three axes (overall, high-level, approach), ranked the three proposals with ties allowed, and reported a per-prompt confidence.”TASTE construction pipeline (figure caption)
  4. For each prompt, four researchers gave individual feedback, then discussed disagreements in pairs, then revised their scores.
    Trace this step to the paper
    “For each prompt, four researchers gave their feedback individually, discussed disagreements in pairs, then revised their feedback.”Pair-discussion protocol (figure caption)
  5. The team filtered the resulting preferences down to those with self-reported 'strong' confidence, taken after the discussion stage.
    Trace this step to the paper
    “filtering for self-reported “strong” confidence, raises estimated human agreement by 15 percentage points, from 53% pre-discussion to 68% for strong-confidence, post-discussion preferences”Building a Research Judgment Benchmark (TASTE)
  6. They additionally filtered for pairs where the two proposals' overall scores differed by at least two points on the five-point scale.
    Trace this step to the paper
    “To produce TASTE, we take strong-confidence, post-discussion preferences over proposals which may come from the same or different prompts, and filter for pairs where scores differ by at least two points.”Building a Research Judgment Benchmark (TASTE)
  7. They capped how many times any single proposal could appear across pairs at ten, to keep data points independent, yielding the final 92-pair benchmark.
    Trace this step to the paper
    “We further cap the number of times proposals can appear at ten times, for independence between data points, giving 92 pairs with 77% estimated human agreement.”Building a Research Judgment Benchmark (TASTE)
  8. To estimate the human-agreement baseline, they randomly sampled one researcher's overall score per proposal from the opposing discussion pair and took the higher-scored proposal as preferred.
    Trace this step to the paper
    “To estimate human agreement, we randomly select a researcher’s “overall score” per proposal from the opposing discussion pair, and take the proposal with the higher sampled score as preferred.”Building a Research Judgment Benchmark (TASTE)
  9. Models were evaluated in a standard paired setup, seeing both proposals together and outputting a win probability for each, binarized into a preference.
    Trace this step to the paper
    “a standard setup where the model sees two proposals in context and gives a probability to each proposal winning which we binarize to give a preference”Evaluating Models’ Research Judgment
  10. Models were also evaluated in a tougher setup where they score each proposal individually without seeing its pair, with preferences inferred afterward from those scores.
    Trace this step to the paper
    “a tougher single-proposal scoring setup where the model sees each proposal individually in context, scores it, and is assessed on the implied preferences from its scores against the preferences in TASTE”Evaluating Models’ Research Judgment
The models under study

Exactly what was run, and how

ModelDeveloperTempEffort / reasoningDeploymentOther settings
Claude Opus 4.6not reportednot reportedunstated
Fable 5not reportednot reportedunstated
Opus 5not reportednot reportedunstated
GPT-5.6-Solnot reportednot reportedunstated
Source for Claude Opus 4.6 settings
“we use a prompt scaffold with Claude Opus 4.6 to generate AI safety research proposals”Building a Research Judgment Benchmark (TASTE)
Source for Fable 5 settings
“the best model performs worse than our human researchers — Fable 5 achieves 60% whereas we estimate researcher performance at 77%”Evaluating Models’ Research Judgment
Source for Opus 5 settings
“Opus 5 and GPT-5.6-Sol perform near chance on TASTE despite being at the frontier on general agentic benchmarks.”Model performance on TASTE by release date and model provider (figure caption)
Source for GPT-5.6-Sol settings
“Opus 5 and GPT-5.6-Sol perform near chance on TASTE despite being at the frontier on general agentic benchmarks.”Model performance on TASTE by release date and model provider (figure caption)

What they reported — and what they left out

The paper names several evaluated models (Fable 5, Opus 5, GPT-5.6-Sol) and the model used to generate proposals (Claude Opus 4.6), reports only accuracy figures for them, and states no temperature, reasoning-effort level, deployment mode, or explicit developer/lab attribution for any model.

Results

The numbers they report

TASTE contains 92 pairwise comparisons with an estimated human agreement rate of 77%.

92 pairs; 77% estimated human agreement

See it in the paper
“Our benchmark contains 92 pairwise comparisons, and we estimate human researcher agreement with our benchmark’s labels at 77%.”Building a Research Judgment Benchmark (TASTE)

The discussion stage combined with strong-confidence filtering raised estimated human agreement by 15 percentage points.

53% pre-discussion -> 68% post-discussion, strong-confidence (+15 percentage points)

See it in the paper
“We find that this discussion stage, in combination with filtering for self-reported “strong” confidence, raises estimated human agreement by 15 percentage points, from 53% pre-discussion to 68% for strong-confidence, post-discussion preferences.”Building a Research Judgment Benchmark (TASTE)

Filtering for a larger score gap between paired proposals further increased estimated agreement, for both same-prompt and different-prompt pairs.

min score-gap of >=2 points (5-point scale) improves agreement; 95% CIs via bootstrap resampling over prompts

See it in the paper
“Filtering for a gap in “overall score” of at least two points, on a five-point scale, improves estimated agreement for pairs of proposals drawn from the same prompt and pairs drawn from different prompts (right). Error bars show 95% confidence intervals from bootstrap resampling over prompts.”figure caption, Building a Research Judgment Benchmark (TASTE)

A stricter inter-rater agreement measure, using only pairs where the same researcher scored both proposals, gives a higher agreement rate but covers far fewer pairs.

83% agreement; validates 50 of 92 pairs

See it in the paper
“A more typical measure of inter-rater agreement – comparing only the pairs where another researcher scored both proposals – gives 83% agreement but only validates 50 of the 92 pairs.”Building a Research Judgment Benchmark (TASTE)

The best-performing model tested, Fable 5, scored well below the estimated human researcher agreement rate in the standard evaluation setup.

Fable 5: 60% vs. estimated human researcher performance: 77%

See it in the paper
“the best model performs worse than our human researchers — Fable 5 achieves 60% whereas we estimate researcher performance at 77%”Evaluating Models’ Research Judgment

Fable 5 performed better specifically on pairs of proposals drawn from different motivating prompts than on the full benchmark.

Fable 5: 69% accuracy on 74 cross-prompt pairs

See it in the paper
“Fable 5 achieves 69% accuracy on 74 pairs of proposals drawn from different prompts”Evaluating Models’ Research Judgment

Almost all models tested performed close to chance, including two named frontier models, with wide per-model confidence intervals given the limited number of pairs.

Fable 5: 60%; most models within 2 SD of chance; per-model CIs ~ ±10 percentage points

See it in the paper
“Fable 5 achieves 60%, and almost all models perform within 2 standard deviations of chance. Opus 5 and GPT-5.6-Sol perform near chance on TASTE despite being at the frontier on general agentic benchmarks. Per-model confidence intervals span roughly ±10 percentage points, as we have a limited number of preference pairs making it difficult to draw conclusions about relative model performance.”Model performance on TASTE by release date and model provider (figure caption)

The proposal pool used to build the benchmark started from a fixed set of human-written proposals.

93 human-written research proposals

See it in the paper
“we started with a collection of 93 human-written research proposals presented in Anthropic’s Fellows Program”Building a Research Judgment Benchmark (TASTE)

Proposal reuse across pairs was capped to preserve independence between data points in the final benchmark.

cap of 10 appearances per proposal

See it in the paper
“We further cap the number of times proposals can appear at ten times, for independence between data points, giving 92 pairs with 77% estimated human agreement.”Building a Research Judgment Benchmark (TASTE)
Claim ↔ evidence

What they assert, beside what they showed

Left is the claim in the paper’s own words. Right is the data offered for it. Where the two do not fully meet, a gold band names the distance.

The claim

Current AI models judge AI safety research proposals worse than experienced human researchers do.

“We find models perform worse than human researchers on TASTE (Fable 5, 60%).”

The evidence

“the best model performs worse than our human researchers — Fable 5 achieves 60% whereas we estimate researcher performance at 77%”

Evaluating Models’ Research Judgment
The claim

A pair-discussion stage and filtering for self-reported strong confidence were both important for building a high-agreement human preference benchmark.

“Two design choices were important for building a high-agreement benchmark (92 pairs, 77% estimated human agreement): a discussion stage in which researchers talk through disagreements before revising their scores, and filtering researchers’ labels for self-reported "strong" confidence.”

The evidence

“this discussion stage, in combination with filtering for self-reported “strong” confidence, raises estimated human agreement by 15 percentage points, from 53% pre-discussion to 68% for strong-confidence, post-discussion preferences”

Building a Research Judgment Benchmark (TASTE)
The claim

Future models could close the current performance gap with human researchers on judging research proposals.

“While there is currently a noticeable gap versus human performance, we think future models could close it”

The evidence

“Fable 5 achieves 69% accuracy on 74 pairs of proposals drawn from different prompts, and we find some evidence that models over-focus on how well proposals answer the motivating question when shown on pairs from the same prompt.”

Evaluating Models’ Research Judgment
Mind the gap: The cited evidence is one model's subset performance plus a preliminary observation about same-prompt bias, not trend or scaling data about future models — the optimism about future models closing the gap is not itself directly supported by these numbers.
The claim

Most current frontier models perform near chance at judging AI safety research proposals, even models at the frontier of general agentic capability.

“Opus 5 and GPT-5.6-Sol perform near chance on TASTE despite being at the frontier on general agentic benchmarks.”

The evidence

“Fable 5 achieves 60%, and almost all models perform within 2 standard deviations of chance.”

Model performance on TASTE by release date and model provider (figure caption)
Mind the gap: The same figure caption notes per-model confidence intervals of roughly ±10 percentage points given the limited number of preference pairs, 'making it difficult to draw conclusions about relative model performance' -- a limitation on how strongly the near-chance claim can be read for any individual named model.
The claim

The random-sampling method used to estimate human agreement is an adequate proxy for how well human researchers agree on TASTE.

“This allows us to evaluate pairs of proposals where no single researcher rated both of them.”

The evidence

“This does not perfectly capture human performance, and instead approximates "I assign a different person to score each proposal; the ‘preferred’ proposal is the one that receives the higher score".”

Building a Research Judgment Benchmark (TASTE)
Mind the gap: The authors themselves state this method 'does not perfectly capture human performance'; a stricter same-rater comparison gives a different figure (83%) but covers only 50 of the 92 pairs, so the two agreement estimates are not directly interchangeable.
Discussion & after

How they frame it, and what they want next

Their framing

The authors frame the work as motivated by a forward-looking, conditional need: if automated AI research and development outpaces the ability to mitigate misalignment and misuse, models will need reliable judgment on hard-to-verify safety research tasks. They present the current human-versus-model gap as real but not fixed, and close by pointing to both benchmark scale-up and human-labeling methodology as next steps.

Register: The authors write cautiously, repeatedly flagging statistical limitations (small pair counts, wide confidence intervals) and explicitly labeling their own human-agreement estimate as an approximation rather than a precise measurement.

Where they hedge

“This does not perfectly capture human performance”Building a Research Judgment Benchmark (TASTE)
“Per-model confidence intervals span roughly ±10 percentage points, as we have a limited number of preference pairs making it difficult to draw conclusions about relative model performance.”Model performance on TASTE by release date and model provider (figure caption)
“While there is currently a noticeable gap versus human performance, we think future models could close it”Evaluating Models’ Research Judgment

What they say it means

  • If automating AI safety research becomes necessary, models will first need reliable, human-level judgment on hard-to-verify tasks such as evaluating research proposals.
    the paper’s words
    “If we want to automate AI safety research — which might become necessary if automated AI research and development outpaces our ability to mitigate the risk of misalignment and misuse — we need reliable measurements of models’ cap abilities on the hard-to-verify parts of safety research.”Background
  • Future efforts to collect human preference labels for fuzzy, hard-to-verify tasks can improve label quality by adding a pair-discussion stage and filtering for self-reported confidence.
    the paper’s words
    “Our findings suggest that future data-collection efforts can improve human label quality by including a pair-discussion stage and by filtering labels for self-reported confidence.”Conclusion

What they call for next

  • The authors call for more diverse and larger-scale evaluations of models' AI safety research capabilities.
    the paper’s words
    “More diverse and larger-scale evaluations of models’ AI safety research capabilities will be needed going forward.”Conclusion
  • The authors invite AI safety researchers to request access to the TASTE dataset via a form.
    the paper’s words
    “We are sharing TASTE with AI safety researchers. For access, please fill out this form.”closing line

Limitations they state

“This does not perfectly capture human performance, and instead approximates "I assign a different person to score each proposal; the ‘preferred’ proposal is the one that receives the higher score".”Building a Research Judgment Benchmark (TASTE)
“A more typical measure of inter-rater agreement – comparing only the pairs where another researcher scored both proposals – gives 83% agreement but only validates 50 of the 92 pairs.”Building a Research Judgment Benchmark (TASTE)
“Per-model confidence intervals span roughly ±10 percentage points, as we have a limited number of preference pairs making it difficult to draw conclusions about relative model performance.”Model performance on TASTE by release date and model provider (figure caption)
“However, humans often disagree, making it unclear what the ground truth should be.”Background
For your own writing

Moves worth stealing

Leads with a one-paragraph tl;dr that states the headline finding and method up front, before any background section, so a skimming reader gets the punchline immediately.

“We built TASTE (The AI Safety Taste Evaluation) — a benchmark measuring how well models can judge pairs of AI safety research proposals, scored by agreement with the preferences of experienced human researchers.”

Frames the motivating risk scenario as conditional and hedged rather than as an assertion that automation is already needed.

“which might become necessary if automated AI research and development outpaces our ability to mitigate the risk of misalignment and misuse”

Reports a competing, stricter version of its own headline statistic (83% on 50 pairs) right alongside the main figure (77% on 92 pairs), rather than presenting only the number that favors the benchmark's design.

“A more typical measure of inter-rater agreement – comparing only the pairs where another researcher scored both proposals – gives 83% agreement but only validates 50 of the 92 pairs.”

Repeatedly defers methodological detail to the underlying paper, keeping the blog post itself compressed to headline findings and figures.

“See the paper for more details.”

What this page was built from

This is a blog-post rendering of the paper (manifest text_grade: partial), carrying a tl;dr in place of a formal abstract and inline figure captions rather than a captured PDF; the post's own links to the full paper PDF and the dataset-access form are not resolved to explicit URLs in this plain-text extract.