Introducing the Conceptual Reasoning Index
Researchers built three benchmarks measuring AI models' conceptual reasoning (argument judging, logical consistency, decision theory) and combined them into a Conceptual Reasoning Index; top models remain far below the ceiling.
It gives a concrete, updating way to track whether models' ability to do hard-to-verify AI-safety reasoning is keeping pace with their general capability growth.
Emery Cooper · Caspar Oesterheld · Chi Nguyen · Alex Kastner · Joe Benton · Ethan Perez — meet the researchers →
Two readings, equal authority
How to choose: The paper’s words is verbatim — use it when you need to quote, or to judge how they write. Plain language is a paraphrase written for comprehension — use it when you want the idea fast. Neither is a summary of the other; they are two doors into the same room.
This source carries no verbatim abstract.
AI models may need to reason about high-stakes questions, like AI safety strategy, that have no clean empirical feedback loop; the authors call this 'conceptual reasoning.' Working with Anthropic, Redwood Research built three benchmarks testing argument evaluation, logical consistency, and decision theory, then combined them into one score, the Conceptual Reasoning Index (CRI). Current top models are improving steadily on the CRI but remain well below the authors' estimated ceiling.
What this paper defines
Every definition below is the paper’s own sentence, with its locator. The plain gloss is a reading aid and is marked as one.
Conceptual reasoning
“Given these properties, efforts to reduce risk from advanced AI may particularly benefit from an improved ability to reason about questions where empirical evidence is limited, there is no (practically) verifiable answer, and one therefore has to rely heavily on argumentation. We refer to this as conceptual reasoning.”Background
In plain terms: Reasoning well about important questions that can't be checked against data or math, so you mostly have to argue your way to an answer.
Conceptual Reasoning Index (CRI)
“We aggregate the benchmarks into the Conceptual Reasoning Index (CRI), available at conceptualreasoning.ai , where you can also find more details on our methodology.”tl;dr
In plain terms: A single combined score built from the three benchmarks below, meant to summarize a model's overall conceptual-reasoning ability.
LMCA (Language Model Conceptual Argumentation)
“LMCA (Language Model Conceptual Argumentation) is a dataset of curated and expert-rated conceptual arguments on a diverse range of topics, including decision theory, philosophy, and risks from advanced AI.”Our benchmarks > LMCA
In plain terms: A dataset of arguments for and against philosophical/AI-risk positions, rated by experts, used to test how well models judge argument quality.
ACCoRD (Assessment of Consistency in Conceptual Reasoning Domains)
“ACCoRD (Assessment of Consistency in Conceptual Reasoning Domains) measures the extent to which models' reported beliefs and preferences on conceptual issues are logically consistent.”Our benchmarks > ACCoRD
In plain terms: Tests whether a model's own stated probabilities and preferences contradict each other, as a proxy for trustworthy reasoning.
DTBench capabilities (Decision Theory Benchmark)
“DTBench capabilities (Decision Theory Benchmark) is a dataset of 407 handcrafted multiple-choice questions designed to measure models' ability to reason about decision-theoretic situations that involve faithful predictions of a model's own behavior or interactions with (near) copies.”Our benchmarks > DTBench
In plain terms: Multiple-choice questions testing whether a model can reason correctly about decision theory, including scenarios involving copies of itself.
What they actually did
Each step is a synthesis. Open any step to see the paper’s own sentence it was derived from, with its locator — so nothing here floats free of the source.
- Built three separate benchmarks, each targeting a different facet of conceptual reasoning.
Trace this step to the paper
“Improving this capability requires being able to measure it, so we built three benchmarks: LMCA , ACCoRD , and DTBench capabilities .”Background
- Compiled the LMCA dataset from position texts and arguments against them, rated by expert conceptual researchers.
Trace this step to the paper
“The dataset contains 560 position texts with 1,461 arguments against these position texts. Nearly all 2 arguments were rated by conceptual researcher Emery Cooper, and some were independently rated by at least one other researcher, for a total of 2,140 ratings.”Our benchmarks > LMCA
- Scored models on LMCA by comparing the model's ratings of arguments to the human expert ratings.
Trace this step to the paper
“We measure how good models are at judging arguments against position texts by comparing their ratings to ours.”Our benchmarks > LMCA
- Validated the human rating process itself using a subset independently rated by several people and then discussed at length.
Trace this step to the paper
“This includes a validation set of roughly 50 arguments, each rated independently by 4–6 people and then discussed for 7–8 hours total.”Our benchmarks > LMCA
- Tested models' argument-generation ability by having one model produce a new argument and a second model rate it using the same rubric and few-shot examples.
Trace this step to the paper
“We then give model B the rubric and few-shot prompt it with the three existing arguments and their ratings, asking it to rate model A's new argument.”Our benchmarks > LMCA
- Generated a large pool of ACCoRD consistency constraints using models, filtered them through an automated checker, then manually approved a subset.
Trace this step to the paper
“The ACCoRD dataset contains close to 14,000 model-generated consistency constraints, which are distributed across 18 constraint types and have gone through an automated checker pipeline. Of these, 567 were further checked and approved by us.”Our benchmarks > ACCoRD
- Tested logical consistency by asking models for numeric probabilities or preference orderings and checking whether the answers satisfy logical constraints.
Trace this step to the paper
“All consistency constraints in the dataset ask models for either numeric probability estimates or preference orderings.”Our benchmarks > ACCoRD
- Built DTBench capabilities from handcrafted multiple-choice decision-theory questions, authored mainly by one domain expert and validated by another.
Trace this step to the paper
“All questions were independently validated by Emery Cooper, another domain expert.”Our benchmarks > DTBench
- Ran all evaluated models at their maximum token limits and effort levels, substituting a fallback model for one model's refusals.
Trace this step to the paper
“All models were run at their maximum token limits and effort levels. To compute Fable 5's score, we used Opus 5 as a fallback in cases where Fable 5 refused to answer a question.”Results
- Combined the three benchmarks into the single CRI score using fixed weights.
Trace this step to the paper
“The CRI is currently a weighted average of LMCA (60%), ACCoRD (20%), and DTBench capabilities (20%).”Results
Exactly what was run, and how
| Model | Developer | Temp | Effort / reasoning | Deployment | Other settings |
|---|---|---|---|---|---|
| Opus 5 | — | not reported | maximum | unstated | Run at maximum token limits and effort levels, per the paper's general statement about all evaluated models; described as the highest-scoring model overall. |
| Claude Fable 5 | — | not reported | maximum | unstated | Opus 5 was substituted as a fallback score whenever Fable 5 refused to answer a question. |
| Muse Spark 1.2 | — | not reported | maximum | unstated | Included as a top-performing model on external benchmarks, though not the top scorer on the CRI. |
| Gemini 3.6 Flash | — | not reported | maximum | unstated | Included as a top-performing model on external benchmarks, though not the top scorer on the CRI. |
| GPT-4 | — | not reported | not reported | unstated | Its ACCoRD score is based on incomplete data because it refused to fully answer a portion of the items. |
Source for Opus 5 settings
“All models were run at their maximum token limits and effort levels.”Results
Source for Claude Fable 5 settings
“To compute Fable 5's score, we used Opus 5 as a fallback in cases where Fable 5 refused to answer a question.”Results
Source for Muse Spark 1.2 settings
“We also include scores for Claude Fable 5, Muse Spark 1.2, and Gemini 3.6 Flash, which are their respective companies’ top-performing models on many external benchmarks, though not on the CRI.”Results
Source for Gemini 3.6 Flash settings
“We also include scores for Claude Fable 5, Muse Spark 1.2, and Gemini 3.6 Flash, which are their respective companies’ top-performing models on many external benchmarks, though not on the CRI.”Results
Source for GPT-4 settings
“The ACCoRD score of GPT-4 is based on incomplete data since the model refused to fully answer 18% of the benchmark's items.”Results
What they reported — and what they left out
The paper names the models compared (Opus 5, Claude Fable 5, Muse Spark 1.2, Gemini 3.6 Flash, GPT-4) and states all were run at maximum token limits and effort levels, but it never explicitly states which company developed each named model, nor does it report temperature, sampling settings, or exact deployment method (API vs. web UI).
The numbers they report
LMCA is built from several hundred position texts with over a thousand rated arguments against them.
560 position texts; 1,461 arguments
See it in the paper
“The dataset contains 560 position texts with 1,461 arguments against these position texts.”Our benchmarks > LMCA
LMCA's arguments received over two thousand total expert ratings.
2,140 ratings
See it in the paper
“Nearly all 2 arguments were rated by conceptual researcher Emery Cooper, and some were independently rated by at least one other researcher, for a total of 2,140 ratings.”Our benchmarks > LMCA
A small validation subset of LMCA was rated by multiple people and discussed extensively to check reliability.
~50 arguments; 4–6 raters each; 7–8 hours of discussion
See it in the paper
“This includes a validation set of roughly 50 arguments, each rated independently by 4–6 people and then discussed for 7–8 hours total.”Our benchmarks > LMCA
ACCoRD's raw pool of model-generated consistency constraints spans many constraint types.
~14,000 constraints across 18 constraint types
See it in the paper
“The ACCoRD dataset contains close to 14,000 model-generated consistency constraints, which are distributed across 18 constraint types and have gone through an automated checker pipeline.”Our benchmarks > ACCoRD
Only a manually-approved subset of ACCoRD's constraints is used in the CRI score.
567 constraints
See it in the paper
“Of these, 567 were further checked and approved by us. We include only those 567 constraints in our aggregate conceptual reasoning performance metric, the CRI.”Our benchmarks > ACCoRD
DTBench capabilities consists of several hundred handcrafted questions.
407 questions
See it in the paper
“DTBench capabilities (Decision Theory Benchmark) is a dataset of 407 handcrafted multiple-choice questions designed to measure models' ability to reason about decision-theoretic situations that involve faithful predictions of a model's own behavior or interactions with (near) copies.”Our benchmarks > DTBench
An additional set of DTBench questions about models' decision-theoretic attitudes exists but is excluded from the CRI.
130 additional questions (excluded)
See it in the paper
“The full DTBench suite includes an additional 130 questions that measure models' decision-theoretic attitudes. We do not include these in the CRI.”Our benchmarks > DTBench
The CRI weights the three benchmarks unevenly, favoring LMCA.
LMCA 60%, ACCoRD 20%, DTBench capabilities 20%
See it in the paper
“The CRI is currently a weighted average of LMCA (60%), ACCoRD (20%), and DTBench capabilities (20%).”Results
The authors estimate that even a maximally good model would not reach a perfect score on LMCA or DTBench capabilities, due to rating noise and the LMCA scoring scheme.
estimated per-benchmark ceilings: LMCA ~85, DTBench capabilities ~100
See it in the paper
“Because human ratings are noisy, we expect that a model giving maximally good LMCA ratings would score roughly 85 rather than 100, which we estimate based on expert inter-rater agreement. Meanwhile, we expect that giving the correct answer to every DTBench capabilities question would yield a score of 100 or extremely close to 100.”Results
Combining the per-benchmark ceiling estimates, the authors estimate an overall ceiling for the CRI itself.
~91 (estimated CRI ceiling)
See it in the paper
“Overall, this leads us to estimate ceiling performance on the CRI to be around 91.”Results
The best-scoring model on the CRI is still well below the estimated ceiling.
73.6 (95% CI: ± 2.1), model: Opus 5
See it in the paper
“The highest-scoring model, Opus 5, is still well below this ceiling, with a score of 73.6 (95% CI: ± 2.1).”Results
CRI scores across models have risen steadily over roughly two years with no sign of leveling off.
See it in the paper
“Scores have been increasing roughly linearly since late 2024, with no signs of flattening.”Results
DTBench capabilities performance is nearly saturated for the best model.
98% (Fable 5)
See it in the paper
“DTBench capabilities scores are already close to the ceiling, with Fable 5 getting 98% of questions right.”Results
The authors project LMCA will begin saturating in roughly a year.
~1 year (projected)
See it in the paper
“Extrapolating from scores to date, we loosely estimate that LMCA will start saturating about a year from now.”Results
GPT-4's ACCoRD score is based on incomplete data because it refused to answer a notable share of items.
18% refusal rate on ACCoRD items
See it in the paper
“The ACCoRD score of GPT-4 is based on incomplete data since the model refused to fully answer 18% of the benchmark's items.”Results
What they assert, beside what they showed
Left is the claim in the paper’s own words. Right is the data offered for it. Where the two do not fully meet, a gold band names the distance.
Improving models' safety-relevant reasoning is urgent because such work may become increasingly AI-driven.
“We think improving models' ability to do work that mitigates risk from advanced AI is important and urgent.”
“Once models can perform work that reduces AI risk at the level of human experts, AI(-assisted) output in the area might dwarf unassisted human output. This suggests that a major determinant of whether we address AI risks in time is how early we can automate or uplift this work, relative to high-risk capabilities.”
BackgroundHaving one model rate another model's newly generated argument, using the rubric and few-shot examples, produces trustworthy ratings.
“This methodology produces fairly accurate ratings from model B.”
“Ratings follow a detailed rubric. On arguments rated by at least two people, inter-rater agreement is high compared to agreement between humans and models.”
Our benchmarks > LMCAHuman raters agree with each other more than models agree with humans, which the authors treat as validating LMCA's human ratings as ground truth.
“On arguments rated by at least two people, inter-rater agreement is high compared to agreement between humans and models.”
“This includes a validation set of roughly 50 arguments, each rated independently by 4–6 people and then discussed for 7–8 hours total.”
Our benchmarks > LMCAThe current best model's CRI score still falls well short of what's achievable.
“The highest-scoring model, Opus 5, is still well below this ceiling, with a score of 73.6 (95% CI: ± 2.1).”
“Overall, this leads us to estimate ceiling performance on the CRI to be around 91.”
ResultsDecision-theoretic capability, as measured by DTBench, is nearly solved by the best current model.
“DTBench capabilities scores are already close to the ceiling, with Fable 5 getting 98% of questions right.”
“DTBench capabilities scores are already close to the ceiling, with Fable 5 getting 98% of questions right.”
ResultsGPT-4's ACCoRD score should be read with caution because it is based on incomplete responses.
“The ACCoRD score of GPT-4 is based on incomplete data since the model refused to fully answer 18% of the benchmark's items.”
“* Partial data for GPT-4: refused to fully answer 18% of ACCoRD items.”
ResultsHow they frame it, and what they want next
Their framing
The authors frame conceptual reasoning as an important but under-measured capability for AI-safety work, positioning the CRI as a first step toward tracking it rather than a finished verdict. They repeatedly stress that the index is a living project: benchmarks may be added, retired, or reweighted as models improve.
Register: The authors write cautiously and hedge estimates explicitly ('loosely estimate,' 'very uncertain,' 'roughly'), while stating institutional facts (dataset sizes, weights, top scores) plainly and without hedging.
Where they hedge
“We're very uncertain about when ACCoRD will saturate.”Results
“Extrapolating from scores to date, we loosely estimate that LMCA will start saturating about a year from now.”Results
“Because human ratings are noisy, we expect that a model giving maximally good LMCA ratings would score roughly 85 rather than 100, which we estimate based on expert inter-rater agreement.”Results
“Currently, only models' performance at judging arguments goes into the CRI, but we hope to add a measurement of models' argumentation ability in the future.”Our benchmarks > LMCA
What they say it means
- How early models' safety-relevant reasoning skills are improved, relative to risky general capabilities, may determine whether AI risks are addressed in time.
the paper’s words
“This suggests that a major determinant of whether we address AI risks in time is how early we can automate or uplift this work, relative to high-risk capabilities.”Background
- Targeted skill-building, such as training on AI governance and cooperation-failure avoidance, could be a deliberate lever for improving safety-relevant reasoning.
the paper’s words
“One way to influence this might be to selectively improve models' relevant skills, such as reasoning about how to govern and align AI and how to avoid catastrophic cooperation failures involving AI.”Background
What they call for next
- Invites researchers to request access to the primary LMCA dataset.
the paper’s words
“You can request access to our primary conceptual dataset, LMCA, through this form .”tl;dr
- Directs readers to the CRI website for ongoing, updated scores and methodology details.
the paper’s words
“For more information and live scores on the CRI, please visit conceptualreasoning.ai .”Conclusion
Limitations they state
“Currently, only models' performance at judging arguments goes into the CRI, but we hope to add a measurement of models' argumentation ability in the future.”Our benchmarks > LMCA
“We're very uncertain about when ACCoRD will saturate.”Results
“The ACCoRD score of GPT-4 is based on incomplete data since the model refused to fully answer 18% of the benchmark's items.”Results
“A very small number of arguments (16 out of 1,461) were not rated by Emery Cooper.”Notes
Moves worth stealing
Opens with a plain-language tl;dr framing the motivating problem before any technical detail, so a busy reader gets the point immediately.
“A core hope for managing AI risks is that AIs will help us understand our situation, plan for what lies ahead, and develop risk mitigations.”
States quantitative uncertainty (confidence intervals, estimated ceilings) alongside headline numbers instead of presenting a bare score.
“The highest-scoring model, Opus 5, is still well below this ceiling, with a score of 73.6 (95% CI: ± 2.1).”
Explicitly separates what the benchmark currently measures from what it does not yet measure, signaling an evolving research program rather than a finished claim.
“Currently, only models' performance at judging arguments goes into the CRI, but we hope to add a measurement of models' argumentation ability in the future.”
Uses numbered footnotes to carry data caveats (partial ratings, refusals, related work) without interrupting the main narrative flow.
“A very small number of arguments (16 out of 1,461) were not rated by Emery Cooper. All of them were rated by Caspar Oesterheld, who is also a conceptual researcher.”
Where else this leads
Same people
- TASTE: Can AI Models Judge AI Safety Research Proposals? Anthropic
shares Joe Benton - Diffuse AI Control on Fuzzy Tasks Anthropic
shares Joe Benton - SLEIGHT-Bench: Finding Blind Spots in AI Monitors Anthropic
shares Joe Benton
Same territory
- TASTE: Can AI Models Judge AI Safety Research Proposals? Anthropic
ai safety benchmarks - Training a Misaligned Reward Seeker Anthropic
ai safety - Would This Change Your Answer? Evaluating Explanations of LLM Behavior in the Wild with Counterfactual Experiments Anthropic
ai safety - Patterns and problems in emerging multiagent systems Anthropic
ai safety - Visual prompt engineering for video models Google DeepMind
benchmarks - Project Pilot: Can AI control a drone? Anthropic
benchmarks
Published alongside it
The nearest publications in time, across all three labs.
- Automated Researchers Can Mitigate Well-Characterized Alignment Failures Anthropic
2026-08-15 - Characterizing interference weights in a tiny language model Anthropic
2026-08-15 - Fine-Tuned Lie Detectors Failed to Generalize Anthropic
2026-08-15 - TASTE: Can AI Models Judge AI Safety Research Proposals? Anthropic
2026-08-15
What this page was built from
Working from a saved plain-text copy of this Anthropic Alignment Science Blog post (co-authored with Redwood Research researchers), about 10,300 characters, graded 'partial' in the corpus manifest; the underlying CRI chart/figure image is not present in the text, so figure-only data beyond what the surrounding prose describes is unavailable, and the date used here (2026-08-12) is the date printed in the post itself, which differs from the manifest's listed pub_date (2026-08-15).