Introducing LifeSciBench
Two titles, both real. The heading above is how the lab announced this work. The document actually behind it is titled “LifeSciBench: Evaluating Language Models on Realistic, Expert-Level Tasks in the Life Sciences” — everything below is read from that document.
LifeSciBench tests 5 frontier models on 750 expert-written life-science research tasks; the best model passes only 36.1% of them.
It shows expert-level biology research work remains a hard, unsaturated benchmark even for the strongest current models.
Amelia Liu · Andrew Ho · Anne Marie Droste · David Martin · Edmund Wong · Edward Zhou · Isabelle Zhou · Joshua Park · Joy Jiao · Katie-Rose Skelly · Kenny Kim · Kevin Rao · Masatoshi Uehara · Max Marion · … — meet the researchers →
Two readings, equal authority
How to choose: The paper’s words is verbatim — use it when you need to quote, or to judge how they write. Plain language is a paraphrase written for comprehension — use it when you want the idea fast. Neither is a summary of the other; they are two doors into the same room.
“We introduce LifeSciBench, a benchmark of 750 expert-authored tasks designed to evaluate whether language models can handle realistic life science research work. At present, the vast majority of biological benchmarks fail to capture the complexity of research-level work; questions are typically narrowly scoped and purely knowledge-based, while real-world work is often ambiguous and requires multiple judgment calls. Additionally, almost all existing benchmarks are scoped to at best a small collection of scientific domains. There is no existing benchmark in the life sciences with both the requisite breadth and depth required to convincingly measure proficiency in real-world professional settings. LifeSciBench addresses this gap by spanning seven scientific workflows and seven life science domains, with each task paired with an expert-written rubric. Across five frontier and domain-specialized models, GPT-Rosalind performs best with a problem-weighted normalized score of 0.576 and a 36.1% task pass rate, but the benchmark remains far from saturated: no model passes 171 tasks (22.8%), and 261 tasks (34.8%) have a best-model pass rate below 20%. LifeSciBench serves as a high-resolution evaluation of practical scientific reasoning and operational decision-making in biology.”
The authors built a 750-task benchmark of realistic, expert-written life-science research problems spanning seven workflows and seven biological domains, each graded against a detailed expert rubric. Testing five frontier models, the best one (GPT-Rosalind) only solved about a third of the tasks, showing the benchmark still has a lot of headroom left.
What this paper defines
Every definition below is the paper’s own sentence, with its locator. The plain gloss is a reading aid and is marked as one.
Normalized Rubric Score
“For each response, we divide the awarded rubric points by the total possible points for that task.”Section 5.2, Metrics — Normalized Rubric Score
In plain terms: The fraction of all possible rubric points a model's answer actually earned, averaged across tasks.
Task Pass Rate
“We define task pass rate as the fraction of tasks for which a model response meets or exceeds the task-specific pass threshold of 70%.”Section 5.2, Metrics — Task Pass Rate
In plain terms: The share of tasks where a model scored at least 70% of the rubric points, counted as a full success.
Workflow taxonomy
“We began by defining a taxonomy of problem types in life sciences by surveying practicing scientists about the workflows they use most often in applied research settings, then grouping their responses into seven central categories.”Section 3.1, Benchmark Organization & Coverage
In plain terms: Seven categories of everyday scientific work (evidence handling, analysis, design, etc.) used to organize and stratify all the tasks.
Best-model task pass rate (benchmark headroom)
“We quantified the remaining benchmark headroom by computing the highest task pass rate achieved by any evaluated model on each task.”Section 6.3.1, Benchmark Headroom
In plain terms: For each task, the pass rate of whichever model did best on it — used to measure how much room the benchmark has left before it counts as solved.
Rubric criteria
“For each task, rubric criteria describe attributes of a response that should be rewarded or penalized.”Section 3.3, Task Formulation — Rubrics
In plain terms: The individual, point-valued checklist items experts write for each task to score specific facts, reasoning steps, or outputs in a response.
What they actually did
Each step is a synthesis. Open any step to see the paper’s own sentence it was derived from, with its locator — so nothing here floats free of the source.
- Built a taxonomy of seven scientific workflow categories by surveying practicing life scientists about their most common work.
Trace this step to the paper
“We began by defining a taxonomy of problem types in life sciences by surveying practicing scientists about the workflows they use most often in applied research settings, then grouping their responses into seven central categories.”Section 3.1, Benchmark Organization & Coverage
- Recruited 173 PhD-level expert scientists, each with at least two years of industry experience, to author benchmark tasks.
Trace this step to the paper
“Experts were required to have completed a Ph.D. in a relevant discipline, such as biochemistry, molecular biology, neuroscience, immunology, pharmacology, medicinal chemistry, computational biology, or a related field, and have at least two years of experience as practicing scientists in the biotechnology or pharmaceutical industries.”Section 3.2, Expert Writer Cohort
- Formulated each task as a natural-language prompt plus any supporting artifacts plus a task-specific grading rubric.
Trace this step to the paper
“Each LifeSciBench task consists of a prompt, any supporting artifacts needed to answer the question, and a task-specific grading rubric.”Section 3.3, Task Formulation
- Put every task through an uncapped multi-round expert review process before acceptance.
Trace this step to the paper
“Tasks could undergo as many revision cycles as needed before acceptance, with no fixed cap on the number of rounds; accepted tasks averaged six self-directed automated review cycles and completed at least two rounds of expert reviews.”Section 3.4, Review Process
- Required review decisions to be anchored in a verifiable correct answer or at least 90% expert agreement before a task was accepted.
Trace this step to the paper
“Reviews were anchored either in a verifiable correct answer or strong expert consensus, requiring at least 90% agreement among domain experts.”Section 3.4, Review Process
- Ran a separate independent validation study with reviewers distinct from the task writers to rate task realism and quality.
Trace this step to the paper
“To validate the usefulness and scientific quality of LifeSciBench, we conducted an independent expert review of the benchmark tasks. This validation was separate from the task construction and review process described above.”Section 4, Benchmark Validation
- Evaluated five frontier and domain-specialized models single-turn, giving each the full prompt and artifacts and allowing unrestricted internet browsing.
Trace this step to the paper
“Models are instructed to answer the task directly and to include reasoning, calculations, caveats, or assumptions when useful for the final answer. Unrestricted Internet browsing is permitted.”Section 5.1, Evaluation Protocol
- Graded every model response against its task's expert-written rubric, then aggregated points into the paper's headline metrics.
Trace this step to the paper
“Model outputs are graded against the expert-written rubric associated with each task.”Section 5.1, Evaluation Protocol
Exactly what was run, and how
| Model | Developer | Temp | Effort / reasoning | Deployment | Other settings |
|---|---|---|---|---|---|
| GPT-5.4 | OpenAI | not reported | not reported | unstated | Single-turn evaluation with unrestricted internet browsing permitted; no per-model temperature, reasoning-effort, or sampling settings reported. |
| GPT-5.5 | OpenAI | not reported | not reported | unstated | Single-turn evaluation with unrestricted internet browsing permitted; no per-model temperature, reasoning-effort, or sampling settings reported. |
| GPT-Rosalind | OpenAI | not reported | not reported | unstated | Described as a domain-specialized model; single-turn evaluation with unrestricted internet browsing permitted; no temperature, reasoning-effort, or sampling settings reported. |
| Gemini 3.1 Pro | Google DeepMind | not reported | not reported | unstated | Single-turn evaluation with unrestricted internet browsing permitted; no per-model temperature, reasoning-effort, or sampling settings reported. |
| Grok 4.3 | xAI | not reported | not reported | unstated | Single-turn evaluation with unrestricted internet browsing permitted; no per-model temperature, reasoning-effort, or sampling settings reported. |
Source for GPT-5.4 settings
“We evaluate a set of frontier general-purpose and domain-specialized language models on LifeSciBench, namely GPT-5.4, GPT-5.5, GPT-Rosalind, Gemini 3.1 Pro, and Grok 4.3.”Section 5, Experimental Setup & Grader Validation
Source for GPT-5.5 settings
“We evaluate a set of frontier general-purpose and domain-specialized language models on LifeSciBench, namely GPT-5.4, GPT-5.5, GPT-Rosalind, Gemini 3.1 Pro, and Grok 4.3.”Section 5, Experimental Setup & Grader Validation
Source for GPT-Rosalind settings
“We evaluate a set of frontier general-purpose and domain-specialized language models on LifeSciBench, namely GPT-5.4, GPT-5.5, GPT-Rosalind, Gemini 3.1 Pro, and Grok 4.3.”Section 5, Experimental Setup & Grader Validation
Source for Gemini 3.1 Pro settings
“We evaluate a set of frontier general-purpose and domain-specialized language models on LifeSciBench, namely GPT-5.4, GPT-5.5, GPT-Rosalind, Gemini 3.1 Pro, and Grok 4.3.”Section 5, Experimental Setup & Grader Validation
Source for Grok 4.3 settings
“We evaluate a set of frontier general-purpose and domain-specialized language models on LifeSciBench, namely GPT-5.4, GPT-5.5, GPT-Rosalind, Gemini 3.1 Pro, and Grok 4.3.”Section 5, Experimental Setup & Grader Validation
What they reported — and what they left out
The paper names the five evaluated models and their developers by implication, but reports no temperature, reasoning-effort, sampling, or deployment-interface settings for any of them, only that evaluation was single-turn with unrestricted internet browsing permitted.
The numbers they report
GPT-Rosalind had the highest aggregate score and pass rate of the five models tested.
GPT-Rosalind 0.576 / 36.1%; GPT-5.5 0.519 / 25.7%; Gemini 3.1 Pro 0.515 / 23.6%; GPT-5.4 0.479 / 20.7%; Grok 4.3 0.399 / 13.0%
See it in the paper
“GPT-Rosalind was the strongest overall system, with a problem-weighted mean score of 0.576 and a task pass rate of 36.1%, compared with 0.519 / 25.7% for GPT-5.5, 0.515 / 23.6% for Gemini 3.1 Pro, 0.479 / 20.7% for GPT-5.4, and 0.399 / 13.0% for Grok 4.3.”Section 6.1, Results
GPT-Rosalind scored highest on just over half of all tasks individually.
386 of 750 tasks
See it in the paper
“GPT-Rosalind had the highest per-task mean score on 386 of 750 tasks.”Section 6.1, Results
No model solved nearly a quarter of tasks, and about a third of tasks had a best-model pass rate under 20%.
171/750 tasks (22.8%) zero-pass; 261/750 tasks (34.8%) best-model pass rate below 20%
See it in the paper
“no model passes 171 tasks (22.8%), and 261 tasks (34.8%) have a best-model pass rate below 20%.”Abstract
More than half of all tasks had a best-model pass rate below 50%.
422/750 tasks (56.3%) below 50%; 261/750 tasks (34.8%) below 20%
See it in the paper
“Using these bins, 422 tasks (56.3%) had a best-model pass rate below 50%, including 261 tasks (34.8%) with a best-model pass rate below 20%.”Section 6.3.1, Benchmark Headroom
GPT-Rosalind's pass rate dropped sharply on tasks that required attached artifacts.
44.5% (text-only) vs 28.6% (artifact-heavy)
See it in the paper
“GPT-Rosalind achieved a 44.5% pass rate on text-only tasks but dropped to 28.6% on tasks requiring attached artifacts.”Section 6.3.2, Artifact Heavy & Operationally Constrained Tasks
GPT-5.5 showed the same artifact-related performance drop as GPT-Rosalind.
29.5% (text-only) vs 22.2% (artifact-heavy)
See it in the paper
“GPT-5.5 showed the same pattern, dropping from 29.5% on text-only tasks to 22.2% on attached-artifact … tasks.”Section 6.3.2, Artifact Heavy & Operationally Constrained Tasks
Tasks requiring exact sequence or structure outputs were the hardest problem type for every model.
46.9% (GPT-Rosalind) to 18.0% (Grok) criterion success
See it in the paper
“such questions were among the lowest-scoring problem types for every model, with sequence/structure criterion success ranging from 46.9% for GPT-Rosalind to 18.0% for Grok.”Section 6.3.3, Exact and Construct Level Outputs
GPT-Rosalind barely improved over GPT-5.5 on tasks requiring generated/constructed outputs, unlike its large gains on judgment tasks.
+0.001 on generate/construct items
See it in the paper
“GPT-Rosalind’s improvement over GPT-5.5 was large for scientific judgment categories but minimal for generate/construct items (+0.001)”Section 6.3.3, Exact and Construct Level Outputs
GPT-Rosalind's highest-scoring workflow was Translation.
0.712 mean score
See it in the paper
“GPT-Rosalind reached a mean score of 0.712 on Translation, where tasks require models to connect preclinical or biological evidence to clinical relevance, safety, trial design, or other translational implications.”Section 6.2.1, Workflow and Rubric Level Strengths
GPT-Rosalind also scored highly on Scientific Communication tasks.
0.718 mean score
See it in the paper
“Scientific Communication was also high-scoring, with GPT-Rosalind reaching a mean score of 0.718 on tasks requiring models to explain or summarize scientific findings for a specified audience or decision context.”Section 6.2.1, Workflow and Rubric Level Strengths
GPT-Rosalind's biggest rubric-level gains over GPT-5.5 were in mechanism explanation, experiment design, and critique/validation.
+0.086 (mechanisms), +0.079 (experiment design), +0.078 (critique/validation)
See it in the paper
“GPT-Rosalind’s largest gains over GPT-5.5 were in criteria related to explaining mechanisms (+0.086), designing experiments (+0.079), and critique or validation (+0.078).”Section 6.2.1, Workflow and Rubric Level Strengths
Gemini 3.1 Pro won outright on a substantial number of individual tasks despite lower aggregate scores.
214 tasks
See it in the paper
“Gemini uniquely led on 214 tasks, indicating complementary strengths across task types.”Section 6.2.2, Model Specific Profiles
GPT-Rosalind often made substantial partial progress on tasks it still failed overall.
109 tasks
See it in the paper
“For GPT-Rosalind, 109 tasks had pass rates below 20% while still receiving at least 50% rubric score.”Section 6.3.4, Partial Progress Without Full Task Success
The independent validation panel was itself highly credentialed.
453 reviewers; 97% held a PhD or equivalent; average 12 years field experience and 14 peer-reviewed publications; 88% received an award or fellowship
See it in the paper
“453 expert reviewers participated, 97% held a Ph.D. or equivalent doctorate, reviewers had an average of 12 years of field experience and 14 peer-reviewed publications, and 88% reported receiving at least one award or fellowship.”Section 4, Benchmark Validation
Independent reviewers broadly agreed the tasks were strong, useful life-science evaluation items.
79.1% strongly agreed / 96.6% agreed overall
See it in the paper
“for overall usefulness, 79.1% strongly agreed and 96.6% agreed overall that the task was a strong life science evaluation item.”Section 4, Benchmark Validation
What they assert, beside what they showed
Left is the claim in the paper’s own words. Right is the data offered for it. Where the two do not fully meet, a gold band names the distance.
No prior life-science benchmark combines enough breadth and depth to convincingly measure real-world professional proficiency, which is the gap LifeSciBench fills.
“There is no existing benchmark in the life sciences with both the requisite breadth and depth required to convincingly measure proficiency in real-world professional settings.”
“LifeSciBench contains 750 tasks spanning seven core life-science workflows, seven biological domains, and multiple stages of the research process.”
Section 3.5, Dataset CompositionThe benchmark is far from saturated, leaving substantial room to measure future model progress.
“the benchmark remains far from saturated”
“422 tasks (56.3%) had a best-model pass rate below 50%, including 261 tasks (34.8%) with a best-model pass rate below 20%.”
Section 6.3.1, Benchmark HeadroomGPT-Rosalind is the strongest overall system on LifeSciBench.
“GPT-Rosalind was the strongest overall system”
“with a problem-weighted mean score of 0.576 and a task pass rate of 36.1%, compared with 0.519 / 25.7% for GPT-5.5, 0.515 / 23.6% for Gemini 3.1 Pro, 0.479 / 20.7% for GPT-5.4, and 0.399 / 13.0% for Grok 4.3.”
Section 6.1, ResultsCurrent frontier models can already be useful for scientific synthesis and expert-facing interpretation work.
“current frontier systems can be useful on scientific synthesis and expert-facing interpretation”
“Scientific Communication was also high-scoring, with GPT-Rosalind reaching a mean score of 0.718 on tasks requiring models to explain or summarize scientific findings for a specified audience or decision context.”
Section 6.2.1, Workflow and Rubric Level StrengthsLifeSciBench tasks are recognizable to practicing experts as realistic, scientifically grounded assessments.
“these results suggest that LifeSciBench tasks are not only technically gradeable, but also recognizable to practicing experts as realistic, scientifically grounded assessments of the reasoning and judgment required in applied life science research”
“For real-world relevance, 86.8% of reviewers strongly agreed and 98.3% agreed overall that tasks reflected realistic life science work.”
Section 4, Benchmark ValidationProducing specific, exact scientific outputs (like genomic sequences or constructs) is a fruitful direction for further model development because current models lag there.
“life science workflows generally do require specific, actionable outputs such as genomic sequences or constructs intended for downstream use, suggesting that this is a fruitful direction for further model development”
“GPT-Rosalind’s improvement over GPT-5.5 was large for scientific judgment categories but minimal for generate/construct items (+0.001)”
Section 6.3.3, Exact and Construct Level OutputsHow they frame it, and what they want next
Their framing
The authors present LifeSciBench as closing a gap between narrow, fact-recall biology benchmarks and the messy, judgment-heavy reality of applied life-science research. They frame current frontier models as showing real but partial progress: strong on scientific synthesis and communication-style tasks, but still unreliable on artifact-heavy analysis, exact technical outputs, and turning partial reasoning into a complete, actionable decision. Throughout, they emphasize that the benchmark remains far from saturated, positioning it as a forward-looking yardstick for future model progress rather than a solved test.
Register: The authors write with cautious, hedge-qualified confidence: bold quantitative claims about state-of-the-art performance are consistently paired with explicit caveats about small sample categories, single-turn design limits, and the gap between benchmark score and real-world research impact.
Where they hedge
“Scientific Communication is a small category, so its estimate should be interpreted cautiously, but the pattern is consistent with the broader rubric-level results.”Section 6.2.1, Workflow and Rubric Level Strengths
“These results should be interpreted with some caution”Section 6.3.3, Exact and Construct Level Outputs
“LifeSciBench is intended to measure model performance on realistic, self-contained life-science tasks, but it does not directly measure the impact of AI systems in live research environments.”Section 7, Limitations & Future Work
“due to the tremendous complexity and breadth of the life sciences, it is not feasible for a single benchmark to cover every single type of problem that scientists might practically encounter”Section 7, Limitations & Future Work
“performance on LifeSciBench should be interpreted as evidence of task-level capability under realistic scientific constraints rather than as a direct estimate of downstream research impact”Section 7, Limitations & Future Work
What they say it means
- Model choice should depend on the specific scientific task rather than aggregate benchmark rank, since models show complementary rather than uniformly ranked strengths.
the paper’s words
“A model that performs slightly worse overall may still be better suited for particular workflows, domains, or output formats.”Section 6.2.2, Model Specific Profiles
- Improving how models handle attached artifacts (files, images, sequences) is a priority for real lab usefulness.
the paper’s words
“artifact use remains a major bottleneck”Section 6.3.2, Artifact Heavy & Operationally Constrained Tasks
- Benchmark scores should eventually be linked to real research-productivity outcomes in deployed settings, not treated as a proxy on their own.
the paper’s words
“Such studies could establish a tighter relationship between the progression of model capabilities and whether AI systems actually improve scientific productivity.”Section 7, Limitations & Future Work
What they call for next
- Expand the benchmark's coverage to more specialized scientific workflows and domains.
the paper’s words
“Future work should expand coverage to a larger range of specialized workflows and scientific domains.”Section 7, Limitations & Future Work
- Correlate benchmark performance with results from real deployment studies in live research settings.
the paper’s words
“Ideally, benchmark performance could even be correlated with the results of deployment studies in … live research settings.”Section 7, Limitations & Future Work
Limitations they state
“LifeSciBench is intended to measure model performance on realistic, self-contained life-science tasks, but it does not directly measure the impact of AI systems in live research environments.”Section 7, Limitations & Future Work
“The evaluation is also conducted in a single-turn setting: each model receives a task and any associated artifacts then produces one final response.”Section 7, Limitations & Future Work
“it does not capture the full dynamics of deployed research programs where outcomes depend on cross-field collaboration, generation of novel data, and financial or operational constraints”Section 7, Limitations & Future Work
“Public release of tasks, rubrics, artifacts, or evaluation materials may be limited by licensing, privacy, proprietary information, or biological safety considerations.”Appendix A.5, Data Availability and Safety Disclosure
Moves worth stealing
Leads the abstract with a concrete headline number (best model's score and pass rate) immediately followed by the remaining-headroom statistic, framing the paper as 'progress plus gap' rather than just 'progress.'
“GPT-Rosalind performs best with a problem-weighted normalized score of 0.576 and a 36.1% task pass rate, but the benchmark remains far from saturated”
Uses named Related Work subsections to position the new benchmark against two clearly delineated prior lines of work before claiming the specific gap it fills.
“LifeSciBench addresses the remaining gap: a combination of expert-level scientific reasoning and data analysis across a broad swathe of applied life science research.”
Publishes full worked example tasks with complete rubrics in an appendix so a reader can audit exactly what 'expert-level' means instead of taking the label on faith.
“The model receives Visium FFPE data from a cervical cancer slide and must perform unsupervised clustering, cell-type annotation, antigen expression analysis, and evidence-based therapy recommendation.”
States institutional and disclosure caveats explicitly in a dedicated appendix (AI use, contributor compensation, institutional conflict of interest) rather than burying them in a footnote.
“LifeSciBench was developed by OpenAI, and the evaluated systems include OpenAI models. Results should be interpreted with this institutional context in mind.”
Where else this leads
Same people
- Scientific computing in the age of agentic AI OpenAI
shares Suyash Shringarpure, Andrew Ho
Same territory
- How Claude is accelerating protein design and analytical chemistry Anthropic
agentic-ai life-sciences - Towards Structural Understanding of LLM Overthinking Google DeepMind
llm-evaluation - Introducing GeneBench-Pro OpenAI
benchmark - SLEIGHT-Bench: Finding Blind Spots in AI Monitors Anthropic
benchmark
Published alongside it
The nearest publications in time, across all three labs.
- A near-autonomous AI chemist improves a challenging reaction in medicinal chemistry OpenAI
2026-06-17 - Diffuse AI Control on Fuzzy Tasks Anthropic
2026-06-15 - Artificial Minds, Human Disagreement: The Politics of AI Consciousness Google DeepMind
2026-06-15 - From AGI to ASI Google DeepMind
2026-06-12
What this page was built from
Extraction is drawn from the full paper text (manifest text_grade: full), including all numbered sections, tables, figures, and appendices A through E plus references.