AnthropicP092026-08-15lab post onlyai safetybenchmarksconceptual reasoningdecision theorymodel evaluation

Introducing the Conceptual Reasoning Index

Researchers built three benchmarks measuring AI models' conceptual reasoning (argument judging, logical consistency, decision theory) and combined them into a Conceptual Reasoning Index; top models remain far below the ceiling.

It gives a concrete, updating way to track whether models' ability to do hard-to-verify AI-safety reasoning is keeping pace with their general capability growth.

Emery Cooper · Caspar Oesterheld · Chi Nguyen · Alex Kastner · Joe Benton · Ethan Perez — meet the researchers →

How much of this do you want?
Orient me keeps four things: the abstract, the method, the claim↔evidence panel, and where it leads next. Everything adds constructs, the model table, every reported statistic, the discussion framing and the style moves. Switching hides nothing permanently and never changes what a section says — it only changes how many are on screen.
Abstract

Two readings, equal authority

How to choose: The paper’s words is verbatim — use it when you need to quote, or to judge how they write. Plain language is a paraphrase written for comprehension — use it when you want the idea fast. Neither is a summary of the other; they are two doors into the same room.

This source carries no verbatim abstract.

Constructs

What this paper defines

Every definition below is the paper’s own sentence, with its locator. The plain gloss is a reading aid and is marked as one.

Conceptual reasoning

“Given these properties, efforts to reduce risk from advanced AI may particularly benefit from an improved ability to reason about questions where empirical evidence is limited, there is no (practically) verifiable answer, and one therefore has to rely heavily on argumentation. We refer to this as conceptual reasoning.”Background

In plain terms: Reasoning well about important questions that can't be checked against data or math, so you mostly have to argue your way to an answer.

Conceptual Reasoning Index (CRI)

“We aggregate the benchmarks into the Conceptual Reasoning Index (CRI), available at conceptualreasoning.ai , where you can also find more details on our methodology.”tl;dr

In plain terms: A single combined score built from the three benchmarks below, meant to summarize a model's overall conceptual-reasoning ability.

LMCA (Language Model Conceptual Argumentation)

“LMCA (Language Model Conceptual Argumentation) is a dataset of curated and expert-rated conceptual arguments on a diverse range of topics, including decision theory, philosophy, and risks from advanced AI.”Our benchmarks > LMCA

In plain terms: A dataset of arguments for and against philosophical/AI-risk positions, rated by experts, used to test how well models judge argument quality.

ACCoRD (Assessment of Consistency in Conceptual Reasoning Domains)

“ACCoRD (Assessment of Consistency in Conceptual Reasoning Domains) measures the extent to which models' reported beliefs and preferences on conceptual issues are logically consistent.”Our benchmarks > ACCoRD

In plain terms: Tests whether a model's own stated probabilities and preferences contradict each other, as a proxy for trustworthy reasoning.

DTBench capabilities (Decision Theory Benchmark)

“DTBench capabilities (Decision Theory Benchmark) is a dataset of 407 handcrafted multiple-choice questions designed to measure models' ability to reason about decision-theoretic situations that involve faithful predictions of a model's own behavior or interactions with (near) copies.”Our benchmarks > DTBench

In plain terms: Multiple-choice questions testing whether a model can reason correctly about decision theory, including scenarios involving copies of itself.

Method

What they actually did

Each step is a synthesis. Open any step to see the paper’s own sentence it was derived from, with its locator — so nothing here floats free of the source.

Three separate benchmarks (LMCA, ACCoRD, DTBench capabilities) are combined into a single weighted score, the CRI.
Click any box to open it.
  1. Built three separate benchmarks, each targeting a different facet of conceptual reasoning.
    Trace this step to the paper
    “Improving this capability requires being able to measure it, so we built three benchmarks: LMCA , ACCoRD , and DTBench capabilities .”Background
  2. Compiled the LMCA dataset from position texts and arguments against them, rated by expert conceptual researchers.
    Trace this step to the paper
    “The dataset contains 560 position texts with 1,461 arguments against these position texts. Nearly all 2 arguments were rated by conceptual researcher Emery Cooper, and some were independently rated by at least one other researcher, for a total of 2,140 ratings.”Our benchmarks > LMCA
  3. Scored models on LMCA by comparing the model's ratings of arguments to the human expert ratings.
    Trace this step to the paper
    “We measure how good models are at judging arguments against position texts by comparing their ratings to ours.”Our benchmarks > LMCA
  4. Validated the human rating process itself using a subset independently rated by several people and then discussed at length.
    Trace this step to the paper
    “This includes a validation set of roughly 50 arguments, each rated independently by 4–6 people and then discussed for 7–8 hours total.”Our benchmarks > LMCA
  5. Tested models' argument-generation ability by having one model produce a new argument and a second model rate it using the same rubric and few-shot examples.
    Trace this step to the paper
    “We then give model B the rubric and few-shot prompt it with the three existing arguments and their ratings, asking it to rate model A's new argument.”Our benchmarks > LMCA
  6. Generated a large pool of ACCoRD consistency constraints using models, filtered them through an automated checker, then manually approved a subset.
    Trace this step to the paper
    “The ACCoRD dataset contains close to 14,000 model-generated consistency constraints, which are distributed across 18 constraint types and have gone through an automated checker pipeline. Of these, 567 were further checked and approved by us.”Our benchmarks > ACCoRD
  7. Tested logical consistency by asking models for numeric probabilities or preference orderings and checking whether the answers satisfy logical constraints.
    Trace this step to the paper
    “All consistency constraints in the dataset ask models for either numeric probability estimates or preference orderings.”Our benchmarks > ACCoRD
  8. Built DTBench capabilities from handcrafted multiple-choice decision-theory questions, authored mainly by one domain expert and validated by another.
    Trace this step to the paper
    “All questions were independently validated by Emery Cooper, another domain expert.”Our benchmarks > DTBench
  9. Ran all evaluated models at their maximum token limits and effort levels, substituting a fallback model for one model's refusals.
    Trace this step to the paper
    “All models were run at their maximum token limits and effort levels. To compute Fable 5's score, we used Opus 5 as a fallback in cases where Fable 5 refused to answer a question.”Results
  10. Combined the three benchmarks into the single CRI score using fixed weights.
    Trace this step to the paper
    “The CRI is currently a weighted average of LMCA (60%), ACCoRD (20%), and DTBench capabilities (20%).”Results
The models under study

Exactly what was run, and how

ModelDeveloperTempEffort / reasoningDeploymentOther settings
Opus 5not reportedmaximumunstatedRun at maximum token limits and effort levels, per the paper's general statement about all evaluated models; described as the highest-scoring model overall.
Claude Fable 5not reportedmaximumunstatedOpus 5 was substituted as a fallback score whenever Fable 5 refused to answer a question.
Muse Spark 1.2not reportedmaximumunstatedIncluded as a top-performing model on external benchmarks, though not the top scorer on the CRI.
Gemini 3.6 Flashnot reportedmaximumunstatedIncluded as a top-performing model on external benchmarks, though not the top scorer on the CRI.
GPT-4not reportednot reportedunstatedIts ACCoRD score is based on incomplete data because it refused to fully answer a portion of the items.
Source for Opus 5 settings
“All models were run at their maximum token limits and effort levels.”Results
Source for Claude Fable 5 settings
“To compute Fable 5's score, we used Opus 5 as a fallback in cases where Fable 5 refused to answer a question.”Results
Source for Muse Spark 1.2 settings
“We also include scores for Claude Fable 5, Muse Spark 1.2, and Gemini 3.6 Flash, which are their respective companies’ top-performing models on many external benchmarks, though not on the CRI.”Results
Source for Gemini 3.6 Flash settings
“We also include scores for Claude Fable 5, Muse Spark 1.2, and Gemini 3.6 Flash, which are their respective companies’ top-performing models on many external benchmarks, though not on the CRI.”Results
Source for GPT-4 settings
“The ACCoRD score of GPT-4 is based on incomplete data since the model refused to fully answer 18% of the benchmark's items.”Results

What they reported — and what they left out

The paper names the models compared (Opus 5, Claude Fable 5, Muse Spark 1.2, Gemini 3.6 Flash, GPT-4) and states all were run at maximum token limits and effort levels, but it never explicitly states which company developed each named model, nor does it report temperature, sampling settings, or exact deployment method (API vs. web UI).

Results

The numbers they report

LMCA is built from several hundred position texts with over a thousand rated arguments against them.

560 position texts; 1,461 arguments

See it in the paper
“The dataset contains 560 position texts with 1,461 arguments against these position texts.”Our benchmarks > LMCA

LMCA's arguments received over two thousand total expert ratings.

2,140 ratings

See it in the paper
“Nearly all 2 arguments were rated by conceptual researcher Emery Cooper, and some were independently rated by at least one other researcher, for a total of 2,140 ratings.”Our benchmarks > LMCA

A small validation subset of LMCA was rated by multiple people and discussed extensively to check reliability.

~50 arguments; 4–6 raters each; 7–8 hours of discussion

See it in the paper
“This includes a validation set of roughly 50 arguments, each rated independently by 4–6 people and then discussed for 7–8 hours total.”Our benchmarks > LMCA

ACCoRD's raw pool of model-generated consistency constraints spans many constraint types.

~14,000 constraints across 18 constraint types

See it in the paper
“The ACCoRD dataset contains close to 14,000 model-generated consistency constraints, which are distributed across 18 constraint types and have gone through an automated checker pipeline.”Our benchmarks > ACCoRD

Only a manually-approved subset of ACCoRD's constraints is used in the CRI score.

567 constraints

See it in the paper
“Of these, 567 were further checked and approved by us. We include only those 567 constraints in our aggregate conceptual reasoning performance metric, the CRI.”Our benchmarks > ACCoRD

DTBench capabilities consists of several hundred handcrafted questions.

407 questions

See it in the paper
“DTBench capabilities (Decision Theory Benchmark) is a dataset of 407 handcrafted multiple-choice questions designed to measure models' ability to reason about decision-theoretic situations that involve faithful predictions of a model's own behavior or interactions with (near) copies.”Our benchmarks > DTBench

An additional set of DTBench questions about models' decision-theoretic attitudes exists but is excluded from the CRI.

130 additional questions (excluded)

See it in the paper
“The full DTBench suite includes an additional 130 questions that measure models' decision-theoretic attitudes. We do not include these in the CRI.”Our benchmarks > DTBench

The CRI weights the three benchmarks unevenly, favoring LMCA.

LMCA 60%, ACCoRD 20%, DTBench capabilities 20%

See it in the paper
“The CRI is currently a weighted average of LMCA (60%), ACCoRD (20%), and DTBench capabilities (20%).”Results

The authors estimate that even a maximally good model would not reach a perfect score on LMCA or DTBench capabilities, due to rating noise and the LMCA scoring scheme.

estimated per-benchmark ceilings: LMCA ~85, DTBench capabilities ~100

See it in the paper
“Because human ratings are noisy, we expect that a model giving maximally good LMCA ratings would score roughly 85 rather than 100, which we estimate based on expert inter-rater agreement. Meanwhile, we expect that giving the correct answer to every DTBench capabilities question would yield a score of 100 or extremely close to 100.”Results

Combining the per-benchmark ceiling estimates, the authors estimate an overall ceiling for the CRI itself.

~91 (estimated CRI ceiling)

See it in the paper
“Overall, this leads us to estimate ceiling performance on the CRI to be around 91.”Results

The best-scoring model on the CRI is still well below the estimated ceiling.

73.6 (95% CI: ± 2.1), model: Opus 5

See it in the paper
“The highest-scoring model, Opus 5, is still well below this ceiling, with a score of 73.6 (95% CI: ± 2.1).”Results

CRI scores across models have risen steadily over roughly two years with no sign of leveling off.

See it in the paper
“Scores have been increasing roughly linearly since late 2024, with no signs of flattening.”Results

DTBench capabilities performance is nearly saturated for the best model.

98% (Fable 5)

See it in the paper
“DTBench capabilities scores are already close to the ceiling, with Fable 5 getting 98% of questions right.”Results

The authors project LMCA will begin saturating in roughly a year.

~1 year (projected)

See it in the paper
“Extrapolating from scores to date, we loosely estimate that LMCA will start saturating about a year from now.”Results

GPT-4's ACCoRD score is based on incomplete data because it refused to answer a notable share of items.

18% refusal rate on ACCoRD items

See it in the paper
“The ACCoRD score of GPT-4 is based on incomplete data since the model refused to fully answer 18% of the benchmark's items.”Results
Claim ↔ evidence

What they assert, beside what they showed

Left is the claim in the paper’s own words. Right is the data offered for it. Where the two do not fully meet, a gold band names the distance.

The claim

Improving models' safety-relevant reasoning is urgent because such work may become increasingly AI-driven.

“We think improving models' ability to do work that mitigates risk from advanced AI is important and urgent.”

The evidence

“Once models can perform work that reduces AI risk at the level of human experts, AI(-assisted) output in the area might dwarf unassisted human output. This suggests that a major determinant of whether we address AI risks in time is how early we can automate or uplift this work, relative to high-risk capabilities.”

Background
Mind the gap: This is a strategic/values claim about urgency, supported by a hypothetical scenario about future capability trajectories rather than by any empirical result from the paper's own benchmarks.
The claim

Having one model rate another model's newly generated argument, using the rubric and few-shot examples, produces trustworthy ratings.

“This methodology produces fairly accurate ratings from model B.”

The evidence

“Ratings follow a detailed rubric. On arguments rated by at least two people, inter-rater agreement is high compared to agreement between humans and models.”

Our benchmarks > LMCA
Mind the gap: No specific accuracy statistic is given for model B's ratings of model-generated arguments; the cited inter-rater-agreement statement describes human-vs-human and human-vs-model agreement generally, not this specific rate-a-new-argument methodology.
The claim

Human raters agree with each other more than models agree with humans, which the authors treat as validating LMCA's human ratings as ground truth.

“On arguments rated by at least two people, inter-rater agreement is high compared to agreement between humans and models.”

The evidence

“This includes a validation set of roughly 50 arguments, each rated independently by 4–6 people and then discussed for 7–8 hours total.”

Our benchmarks > LMCA
Mind the gap: No numeric inter-rater-agreement statistic (e.g., a percentage or correlation) is given in the text — 'high' is asserted, and the validation-set description explains how the set was built, not what the measured agreement level actually was.
The claim

The current best model's CRI score still falls well short of what's achievable.

“The highest-scoring model, Opus 5, is still well below this ceiling, with a score of 73.6 (95% CI: ± 2.1).”

The evidence

“Overall, this leads us to estimate ceiling performance on the CRI to be around 91.”

Results
Mind the gap: The 'well below ceiling' framing rests on the authors' own estimated ceiling (~91), which is itself derived from assumptions about human-rating noise and question answerability rather than from a directly measured maximum score.
The claim

Decision-theoretic capability, as measured by DTBench, is nearly solved by the best current model.

“DTBench capabilities scores are already close to the ceiling, with Fable 5 getting 98% of questions right.”

The evidence

“DTBench capabilities scores are already close to the ceiling, with Fable 5 getting 98% of questions right.”

Results
The claim

GPT-4's ACCoRD score should be read with caution because it is based on incomplete responses.

“The ACCoRD score of GPT-4 is based on incomplete data since the model refused to fully answer 18% of the benchmark's items.”

The evidence

“* Partial data for GPT-4: refused to fully answer 18% of ACCoRD items.”

Results
Discussion & after

How they frame it, and what they want next

Their framing

The authors frame conceptual reasoning as an important but under-measured capability for AI-safety work, positioning the CRI as a first step toward tracking it rather than a finished verdict. They repeatedly stress that the index is a living project: benchmarks may be added, retired, or reweighted as models improve.

Register: The authors write cautiously and hedge estimates explicitly ('loosely estimate,' 'very uncertain,' 'roughly'), while stating institutional facts (dataset sizes, weights, top scores) plainly and without hedging.

Where they hedge

“We're very uncertain about when ACCoRD will saturate.”Results
“Extrapolating from scores to date, we loosely estimate that LMCA will start saturating about a year from now.”Results
“Because human ratings are noisy, we expect that a model giving maximally good LMCA ratings would score roughly 85 rather than 100, which we estimate based on expert inter-rater agreement.”Results
“Currently, only models' performance at judging arguments goes into the CRI, but we hope to add a measurement of models' argumentation ability in the future.”Our benchmarks > LMCA

What they say it means

  • How early models' safety-relevant reasoning skills are improved, relative to risky general capabilities, may determine whether AI risks are addressed in time.
    the paper’s words
    “This suggests that a major determinant of whether we address AI risks in time is how early we can automate or uplift this work, relative to high-risk capabilities.”Background
  • Targeted skill-building, such as training on AI governance and cooperation-failure avoidance, could be a deliberate lever for improving safety-relevant reasoning.
    the paper’s words
    “One way to influence this might be to selectively improve models' relevant skills, such as reasoning about how to govern and align AI and how to avoid catastrophic cooperation failures involving AI.”Background

What they call for next

  • Invites researchers to request access to the primary LMCA dataset.
    the paper’s words
    “You can request access to our primary conceptual dataset, LMCA, through this form .”tl;dr
  • Directs readers to the CRI website for ongoing, updated scores and methodology details.
    the paper’s words
    “For more information and live scores on the CRI, please visit conceptualreasoning.ai .”Conclusion

Limitations they state

“Currently, only models' performance at judging arguments goes into the CRI, but we hope to add a measurement of models' argumentation ability in the future.”Our benchmarks > LMCA
“We're very uncertain about when ACCoRD will saturate.”Results
“The ACCoRD score of GPT-4 is based on incomplete data since the model refused to fully answer 18% of the benchmark's items.”Results
“A very small number of arguments (16 out of 1,461) were not rated by Emery Cooper.”Notes
For your own writing

Moves worth stealing

Opens with a plain-language tl;dr framing the motivating problem before any technical detail, so a busy reader gets the point immediately.

“A core hope for managing AI risks is that AIs will help us understand our situation, plan for what lies ahead, and develop risk mitigations.”

States quantitative uncertainty (confidence intervals, estimated ceilings) alongside headline numbers instead of presenting a bare score.

“The highest-scoring model, Opus 5, is still well below this ceiling, with a score of 73.6 (95% CI: ± 2.1).”

Explicitly separates what the benchmark currently measures from what it does not yet measure, signaling an evolving research program rather than a finished claim.

“Currently, only models' performance at judging arguments goes into the CRI, but we hope to add a measurement of models' argumentation ability in the future.”

Uses numbered footnotes to carry data caveats (partial ratings, refusals, related work) without interrupting the main narrative flow.

“A very small number of arguments (16 out of 1,461) were not rated by Emery Cooper. All of them were rated by Caspar Oesterheld, who is also a conceptual researcher.”

What this page was built from

Working from a saved plain-text copy of this Anthropic Alignment Science Blog post (co-authored with Redwood Research researchers), about 10,300 characters, graded 'partial' in the corpus manifest; the underlying CRI chart/figure image is not present in the text, so figure-only data beyond what the surrounding prose describes is unavailable, and the date used here (2026-08-12) is the date printed in the post itself, which differs from the manifest's listed pub_date (2026-08-15).