How Claude is accelerating protein design and analytical chemistry
Claude autonomously designed working protein binders for 14 of 15 targets and matched a lab's own NMR/LC-MS chemistry analysis in under 25 minutes.
Shows AI models may already compress weeks-to-months of specialist drug-discovery and lab-analysis work into hours, with dual-use safety implications attached.
Two readings, equal authority
How to choose: The paper’s words is verbatim — use it when you need to quote, or to judge how they write. Plain language is a paraphrase written for comprehension — use it when you want the idea fast. Neither is a summary of the other; they are two doors into the same room.
“In this post, we share two results that show how Claude can help life scientists increase the pace of their research. In the first, we tested Claude’s ability to design protein binders from scratch, a key task representative of the early parts of the drug design process and one that has historically taken a specialist weeks or months per target. Claude (Mythos Preview and Opus 4.8) designed protein binders against 15 targets, and succeeded against 14 of them. Between 22% and 35% of its individual designs bound successfully, depending on the setup, compared to the 10-15% that is typical in protein design campaigns today. Some of its strongest designs bound several times more tightly than the best previously published result. In the second example, we evaluated whether Claude can accelerate chemical analysis. Claude Opus 5, a generally available model, was given NMR and LC-MS data (the data that allows chemists to assess the identity and purity of the compounds they work with). Provided with only a contract lab’s raw files and a two-sentence prompt, Claude returned finished results in 23 and 19 minutes, matching the lab’s own analysis on hydrogen counts and purity (96.4% versus 96.33%). These examples demonstrate how Claude can reduce the time and computational expertise currently required to make progress on complex scientific tasks. The pace of AI-enabled discoveries has quickened over the past few months. The bulk of these discoveries have been in areas where verification is relatively fast. In mathematics, for example, agents have begun to work their way through unsolved problems: Erdős problems that have stood for decades are falling at a rate of several a month, and we recently shared how Claude improved on a longstanding lower bound on the Riemann zeta function .”
Anthropic tested Claude on two lab-science tasks: designing new proteins that bind specific therapeutic targets, and interpreting raw chemistry instrument data. Claude designed working binders for 14 of 15 protein targets at a higher success rate than typical industry campaigns, and separately processed a chemistry sample's NMR and mass-spectrometry data in under 25 minutes, matching a professional lab's own results.
What this paper defines
Every definition below is the paper’s own sentence, with its locator. The plain gloss is a reading aid and is marked as one.
minibinder
“A minibinder is a small protein designed to latch tightly onto a target protein.”Claude designs proteins
In plain terms: A small, purpose-built protein engineered to stick tightly to one specific target protein.
de novo design
“Designing a new binder (known as de novo design) has historically taken protein engineers months of computation, optimization, and screening per target.”Claude designs proteins
In plain terms: Designing a brand-new binder protein from scratch rather than modifying an existing one.
affinity
“Affinity is a measure of how strongly a protein binds to its target; high-affinity binders are generally needed to achieve a therapeutic effect because they make the drug effective at lower doses, reducing the risk of side effects and the cost to manufacture them.”Claude designs proteins
In plain terms: How tightly a designed protein grips its target; tighter grip generally means a more effective, lower-dose drug.
high-affinity (operational threshold)
“We consider binders to be high-affinity if they have at most single-digit nanomolar equilibrium dissociation constants (KD < 10 nM).”Footnotes
In plain terms: The paper's cutoff for calling a binder 'high-affinity': a dissociation constant below 10 nanomolar.
co-folding models
“co-folding models (models that predict the structure of a protein, together with whatever it binds, in a single pass)”The campaign
In plain terms: AI models that predict, in one step, the shape of a protein together with whatever it is bound to.
hit rate
“Mythos Preview and Opus 4.8 achieve overall hit rates—how many of the designs are, in fact, binders—of 26.7% and 22.6%, respectively, when designing against all targets simultaneously in a 48-hour session.”Claude designs proteins
In plain terms: The fraction of designed candidates that actually turn out to bind the target in lab testing.
NMR spectrum
“An NMR spectrum is a series of peaks, each corresponding to a hydrogen atom, or a group of equivalent hydrogens, somewhere in the molecule.”Claude runs the analytical chemistry workflow
In plain terms: A chart of peaks where each peak represents one or more hydrogen atoms in the molecule being studied.
LC-MS
“liquid chromatography–mass spectrometry (LC-MS), which first separates the sample into its individual components as they flow through a column, then records how much of each is present based on its ultraviolet absorbance, before measuring the molecular mass of each one.”Claude runs the analytical chemistry workflow
In plain terms: A lab technique that separates a sample into its components and then measures how much of each is present and how heavy each one is.
What they actually did
Each step is a synthesis. Open any step to see the paper’s own sentence it was derived from, with its locator — so nothing here floats free of the source.
- Selected 16 protein targets, including two novel ones absent from published competitions, to test genuine de novo design rather than recall of memorized solutions.
Trace this step to the paper
“We also chose two novel targets, 15-PGDH and GDF-8 , from Adaptyv Bio’s most recent competitions to ensure Claude was able to design against targets without drawing upon pre-recorded successes in its training data or from online search”The campaign
- Ran two design modes: one session designing against all targets at once, and separate parallel sessions each addressing a single target.
Trace this step to the paper
“The first was a multi-target mode, where Claude designed against all targets simultaneously in a single Claude Science session. The second was a single-target mode, in which each session addressed one target and sessions for all targets ran in parallel.”The campaign
- Equipped Claude with a detailed design prompt, internet and literature access, connectors to Drive/Slack/Gmail/BioRxiv, GPU access for specialist models, and no token or sub-agent budget cap within the time window.
Trace this step to the paper
“An extensive protein design prompt … that was also included in the agent context; Access to the internet and a corpus of resources, such as papers, on protein design; Connectors for Google Drive, Slack, Gmail, and BioRxiv; Access to GPUs for running specialized protein design and folding models; No limits on token and sub-agent budget within the allotted time”The campaign
- Let Claude run the entire campaign autonomously, providing no further scientific, technical, or operational guidance after launch.
Trace this step to the paper
“After giving Claude the prompt, we left the model to execute autonomously. We provided no additional scientific, technical, or operational guidance after we initiated the campaigns.”The campaign
- Claude chose binding sites, generated candidate structures and sequences by orchestrating specialist structure-, sequence-, and co-folding models, ran cycles of in silico optimization, and screened for viable, diverse candidates.
Trace this step to the paper
“It chose where on each protein target to design against; generated candidate structures and sequences by orchestrating several structure design, sequence design, and co-folding models … ran the designs through multiple cycles of in silico optimization; and computationally screened for novel, diverse candidates that would express, stay soluble, and bind.”The campaign
- Generated 30 candidate binder designs per target and sent them to external labs Adaptyv Bio and Twist Bioscience for wet-lab production and testing.
Trace this step to the paper
“For each of the 15 targets, we asked Claude to design 30 protein binders. … Claude’s designs were then sent to Adaptyv Bio and Twist Bioscience to validate.”The campaign
- Gave Claude Opus 5 only a contract lab's raw NMR and LC-MS instrument files plus a short plain-language prompt, with no vendor software or human operator.
Trace this step to the paper
“Supplied with only a contract lab’s raw files for a routine quality-control sample and a short plain-language prompt, … with no vendor software and no operator, Claude, working within Claude Science, returned processed NMR and LC-MS results in 23 and 19 minutes, respectively, working in parallel.”Claude runs the analytical chemistry workflow
- For NMR, Claude processed the raw signal into a spectrum, flagged ambiguous peaks, proposed a confirmatory heavy-water exchange test, and used its results to correct an error in its own first-pass reading.
Trace this step to the paper
“Given the raw file from the heavy-water run, Claude quantified what had changed in the data, caught and corrected an overstatement in its first reading (its first pass reported that all four flagged peaks had disappeared, but its own self-check showed that only two had), and arrived at the same conclusion as the lab’s operator.”Claude runs the analytical chemistry workflow
- For LC-MS, Claude reverse-engineered an undocumented vendor binary format, verified its decoding by reproducing the instrument's own recorded scan totals, then produced the full standard set of chemist-facing analysis outputs.
Trace this step to the paper
“Claude worked out how the data was encoded, then confirmed it had read the file correctly by reproducing the instrument's own recorded totals for all 2,664 scans before analyzing anything. It then delivered all the outputs a chemist would expect: the separation trace, mass and UV spectra, a purity table, the compound’s molecular mass”Claude runs the analytical chemistry workflow
- Compared Claude's NMR and LC-MS results directly against the contract lab's own manual analysis of the identical sample.
Trace this step to the paper
“Its results matched the lab’s own processing—hydrogen counts per peak were within 0.08 ¹H of the lab’s, and its purity was measured at 96.4% versus the 96.33% of the lab.”Claude runs the analytical chemistry workflow
Exactly what was run, and how
| Model | Developer | Temp | Effort / reasoning | Deployment | Other settings |
|---|---|---|---|---|---|
| Mythos Preview | Anthropic | not reported | not reported | unstated | Run within Claude Science; multi-target mode: 48 hours wall time, up to 12,500 NVIDIA H100 hours; single-target mode: 24 hours wall time, up to 2,500 H100 hours per target; no token/sub-agent budget limits; fast mode enabled. |
| Opus 4.8 | Anthropic | not reported | not reported | unstated | Run within Claude Science in multi-target mode (48h, up to 12,500 H100 hours) plus single-target mode against three targets. |
| Claude Opus 5 | Anthropic | not reported | not reported | unstated | Described as 'a generally available model'; run within Claude Science; given only raw instrument files and the two prompts quoted here, with no vendor software and no human operator. |
Source for Mythos Preview settings
“We ran Opus 4.8 and Mythos Preview in multi-target mode with 48 hours of wall time and up to 12,500 NVIDIA H100 hours of compute for running specialized protein design and folding models. We also ran Mythos Preview in single-target mode with 24 hours of wall time and up to 2,500 NVIDIA H100 hours of compute for each target.”The campaign
Source for Opus 4.8 settings
“Opus 4.8 was run in single-target mode against three targets: TNFα, latent GDF-8, and mature GDF-8.”Footnotes
Source for Claude Opus 5 settings
“The NMR prompt, in full: “i have a raw 1H FID: process it: FT, phase, baseline-correct. show me the spectrum. then pick peaks and integrate: give me a table with δ (ppm), multiplicity, J (Hz), and integral.””Footnotes
What they reported — and what they left out
The post gives the exact task prompts (footnoted verbatim) and rough compute/time budgets for the protein-design models, but reports no temperature, sampling, or reasoning-effort settings for any model, and never specifies what kind of model Mythos Preview is beyond naming it.
The numbers they report
Claude designed successful binders for nearly all targets attempted.
14 of 15 targets succeeded
See it in the paper
“Claude (Mythos Preview and Opus 4.8) designed protein binders against 15 targets, and succeeded against 14 of them.”Summary
Claude's per-design success rate exceeded typical industry rates.
22%–35% (Claude) vs 10–15% (typical)
See it in the paper
“Between 22% and 35% of its individual designs bound successfully, depending on the setup, compared to the 10-15% that is typical in protein design campaigns today.”Summary
In multi-target mode, both models cleared roughly a quarter hit rate in a single 48-hour session.
Mythos Preview 26.7%, Opus 4.8 22.6% (multi-target, 48h)
See it in the paper
“Mythos Preview and Opus 4.8 achieve overall hit rates—how many of the designs are, in fact, binders—of 26.7% and 22.6%, respectively, when designing against all targets simultaneously in a 48-hour session.”Claude designs proteins
Focusing on one target at a time improved Mythos Preview's hit rate further.
35.1% (single-target mode)
See it in the paper
“Mythos Preview achieves an overall hit rate of 35.1% when designing against each target separately using multiple 24-hour sessions.”Claude designs proteins
Performance varied hugely by target.
target-level hit rates ranged 0%–90%
See it in the paper
“Mythos Preview and Opus 4.8 achieve hit rates over 20% across a set of 15 targets, with target-level hit rates ranging from as high as 90% to as low as 0%.”Claude designs proteins
The campaign produced a large volume of confirmed binders relative to the total designs tried.
354 binders / 1,320 designs / 14 of 15 targets
See it in the paper
“we produced 354 binders against 14 of 15 targets using a total of 1,320 designs.”Claude's performance on the targets
The output is compared to the size of existing public de novo binder collections.
~770 binders / 5,700 designs / 40 targets (public corpora)
See it in the paper
“the two largest collections, proteinbase.com and the collection curated by Overath et al. , consist of approximately 770 binders out of 5,700 designs against 40 targets.”Claude's performance on the targets
Against RBX1, Claude far outperformed competition participants.
40% hit rate (Claude) vs 3.7% (competition participants)
See it in the paper
“Mythos Preview in single-target mode achieved a 40% hit rate, compared to a 3.7% hit rate among participants.”Claude's designs are competitive with entries in Adaptyv Bio's protein design competition
Claude's best RBX1 design beat the competition's winning entry among hundreds of submissions.
top design outperformed winner among 245 entered designs
See it in the paper
“Its top-ranked design was a high-affinity binder that outperformed the winning design, which was among 245 designs entered.”Claude's designs are competitive with entries in Adaptyv Bio's protein design competition
Claude produced multiple confirmed binders with a harder-to-design secondary structure.
15 confirmed β-sheet binders across 6 targets, ≥20% β-strand
See it in the paper
“Claude designed 15 confirmed binders across six targets that contain at least 20% β-strand, demonstrating its ability to reason about protein structure.”Claude designs fold-diverse binders with β-sheets
Claude had only weak success against two especially hard targets.
BBF-14: 3 binders, sub-µM–µM affinity; MBP: 0 of 90 designs confirmed bound
See it in the paper
“Claude still managed to produce three independent BBF-14 binders—one from each design arm, and each built on a different backbone—with modest (sub-micromolar to micromolar) affinities.”Claude struggled against some targets
Claude finished both chemistry analyses in well under half an hour, run in parallel.
23 minutes (NMR), 19 minutes (LC-MS)
See it in the paper
“Claude, working within Claude Science, returned processed NMR and LC-MS results in 23 and 19 minutes, respectively, working in parallel.”Claude runs the analytical chemistry workflow
Claude's chemistry results closely matched the lab's own manual analysis.
H-count within 0.08 ¹H; purity 96.4% (Claude) vs 96.33% (lab)
See it in the paper
“hydrogen counts per peak were within 0.08 ¹H of the lab’s, and its purity was measured at 96.4% versus the 96.33% of the lab.”Claude runs the analytical chemistry workflow
Claude verified it had correctly decoded an undocumented file format before analyzing it.
reproduced totals for all 2,664 scans
See it in the paper
“confirmed it had read the file correctly by reproducing the instrument's own recorded totals for all 2,664 scans”Claude runs the analytical chemistry workflow
Claude identified the compound's retention time, purity, and mass from the LC-MS run.
4.34 min retention; 96.4% UV signal; 504 daltons
See it in the paper
“a single component at 4.34 minutes carrying 96.4% of the UV signal, with a molecular mass of 504 daltons.”Claude runs the analytical chemistry workflow
What they assert, beside what they showed
Left is the claim in the paper’s own words. Right is the data offered for it. Where the two do not fully meet, a gold band names the distance.
Claude can design protein binders as well as, or better than, leading human experts.
“showing that Claude can design protein binders against a variety of targets as well as (or even better than) leading human experts.”
“Mythos Preview in single-target mode achieved a 40% hit rate, compared to a 3.7% hit rate among participants.”
Claude's designs are competitive with entries in Adaptyv Bio's protein design competitionClaude's overall hit rate beats what is typical in protein design campaigns today.
“Between 22% and 35% of its individual designs bound successfully, depending on the setup, compared to the 10-15% that is typical in protein design campaigns today.”
“Mythos Preview and Opus 4.8 achieve overall hit rates—how many of the designs are, in fact, binders—of 26.7% and 22.6%, respectively, when designing against all targets simultaneously in a 48-hour session.”
Claude designs proteinsThe protein design campaign was carried out with minimal human involvement.
“This campaign was carried out with minimal human involvement 3 beyond the information we provided Claude in our initial prompt”
“Our only involvement was granting access approvals (such as network access requests) and monitoring the infrastructure to ensure the sessions were running.”
The campaignClaude Opus 5's chemistry analysis matched the contract lab's own results.
“Its results matched the lab’s own processing—hydrogen counts per peak were within 0.08 ¹H of the lab’s, and its purity was measured at 96.4% versus the 96.33% of the lab.”
“confirmed it had read the file correctly by reproducing the instrument's own recorded totals for all 2,664 scans”
Claude runs the analytical chemistry workflowClaude's chemistry run showed a degree of scientific judgment.
“Claude’s run also showed a degree of scientific judgment, for instance in proposing the very same follow-up experiment the contract lab had independently run.”
“It then proposed the standard check: add heavy water to the NMR sample, which swaps those hydrogens out so their peaks shrink or vanish. (Independently, the lab had run this same check three days after the first measurement.)”
Claude runs the analytical chemistry workflowSome of Claude's protein designs bound several times more tightly than the best previously published result.
“Some of its strongest designs bound several times more tightly than the best previously published result.”
“These include high-affinity binders 1 against at least six targets, and binders matching or exceeding the best reported affinity against at least four targets.”
Claude designs proteinsHow they frame it, and what they want next
Their framing
Anthropic frames both results as concrete early evidence that Claude can compress weeks-to-months of specialist scientific labor into hours, positioning this as one step toward an end-to-end AI-run drug development pipeline. They temper the protein design results by flagging targets where Claude struggled and by noting minibinders are not themselves a therapeutic product, while also naming the dual-use safety risk of increasingly autonomous biological research directly.
Register: The authors write with confident, results-forward language ('successfully orchestrates,' 'state-of-the-art affinity') but consistently pair strong claims with explicit uncertainty admissions, disclosed failure cases, and stated caveats rather than unqualified assertions.
Where they hedge
“We’re not sure why Opus 4.8 was successful on this target and Mythos Preview was not.”Claude designs species cross-reactive binders against TNFα, a challenging, therapeutically relevant target
“To better understand how well Claude performed across these design campaigns, we intend to follow our experiments with more extensive characterization to confirm our hit rates and affinity measurements.”Claude struggled against some targets
“alongside its own list of caveats about the trustworthiness of the results (it noted, for example, that this class of instrument gives the mass only to the nearest whole unit).”Claude runs the analytical chemistry workflow
“Protein minibinders are not a standard therapeutic modality for drugs and even for the common drug modalities, such as monoclonal antibodies and small molecules, designing a high-affinity binder is just the first step in the process of generating a drug-like molecule.”Conclusion
What they say it means
- Increasingly autonomous AI research capability could substantially speed development of new therapies and scientific discoveries.
the paper’s words
“The uplift provided by the increasingly autonomous research capabilities of AI models will undoubtedly speed the development of human therapies and fundamental scientific discoveries.”Agentic biological discovery is dual-use
- The same autonomous biological research capability is dual-use and could enable dangerous research absent safeguards.
the paper’s words
“such capabilities are also dual-use: without robust safety measures, they could enable bad actors to perform dangerous research, such as the development of bioweapons.”Agentic biological discovery is dual-use
- Even generally available (non-frontier) models can already take on routine, time-intensive laboratory analysis work.
the paper’s words
“demonstrating how general-access models can support the routine and time-intensive aspects of research.”Introduction
What they call for next
- Invites scientists to try the analytical chemistry workflow themselves in Claude Science.
the paper’s words
“To try this yourself in Claude Science , give Claude a raw NMR or LC-MS file and ask the model to confirm the compound’s identity and purity.”Claude runs the analytical chemistry workflow
- Shares the prompts and generated data from the protein campaign for outside scrutiny.
the paper’s words
“we are sharing the prompts we used for these campaigns, as well as all in vitro and in silico data we generated.”Claude struggled against some targets
Limitations they state
“Claude still managed to produce three independent BBF-14 binders—one from each design arm, and each built on a different backbone—with modest (sub-micromolar to micromolar) affinities.”Claude struggled against some targets
“Against MBP, however, none of the 90 designs was confirmed to have bound to the target, although one demonstrated a weak, reproducible binding signal.”Claude struggled against some targets
“We’re not sure why Opus 4.8 was successful on this target and Mythos Preview was not.”Claude designs species cross-reactive binders against TNFα, a challenging, therapeutically relevant target
“the experimental data for one target, GDF-8 (Mature), were inconclusive due to target aggregation and non-specific stickiness.”Footnotes
“life science research tasks are currently blocked in our most capable model”Introduction
Moves worth stealing
Opens with a dense, numbers-forward Summary paragraph stating both headline results and their key statistics before any narrative framing.
“Claude (Mythos Preview and Opus 4.8) designed protein binders against 15 targets, and succeeded against 14 of them.”
Leans on independent third-party validators rather than in-house verification alone to certify the core result.
“Our external evaluators, Adaptyv Bio and Twist Bioscience , independently produced and tested Claude’s designs in the lab, finding that of the 15 targets we designed against, Claude successfully designed binders against 14 of them.”
Devotes an explicit, plainly labeled section to failure cases rather than only showcasing wins.
“Claude struggled against some targets”
Names the safety/dual-use risk of its own capability directly, alongside the access-control response, rather than omitting it.
“such capabilities are also dual-use: without robust safety measures, they could enable bad actors to perform dangerous research, such as the development of bioweapons.”
Backs qualitative claims with precise, checkable figures down to specific decimal places rather than rounded approximations.
“hydrogen counts per peak were within 0.08 ¹H of the lab’s, and its purity was measured at 96.4% versus the 96.33% of the lab.”
Discloses the exact prompt text used to elicit a result, in a footnote, rather than only describing the prompt's intent.
“The NMR prompt, in full: “i have a raw 1H FID: process it: FT, phase, baseline-correct. show me the spectrum. then pick peaks and integrate: give me a table with δ (ppm), multiplicity, J (Hz), and integral.””
Where else this leads
Same territory
- Introducing LifeSciBench OpenAI
life-sciences agentic-ai
Published alongside it
The nearest publications in time, across all three labs.
- Automated Researchers Can Mitigate Well-Characterized Alignment Failures Anthropic
2026-08-15 - Characterizing interference weights in a tiny language model Anthropic
2026-08-15 - Fine-Tuned Lie Detectors Failed to Generalize Anthropic
2026-08-15 - Introducing the Conceptual Reasoning Index Anthropic
2026-08-15
What this page was built from
This is Anthropic's research blog post itself, captured in full including footnotes and figure captions; the separately linked protein-design and chemical-analysis technical reports it points to are not part of this source text.