22 publications · 2026-05-15 – 2026-09-04

Anthropic

Most papers in the window, and the shortest window. Interpretability, alignment science, and a hard turn into using Claude as a research instrument.

Annotated bibliography

Every paper, one sentence each

Newest first. Each line is a TLDR written from the paper itself — enough to decide whether to open it. Click any row for the full read.

22 papers
At a glance

What this lab is working on

ai safety 7mechanistic interpretability 3benchmarks 3interpretability 3mathematics 2ai-alignment 2generalization 2reinforcement learning 2alignment 2agentic ai 2dual-use risk 2agentic misalignment 2ai alignment 2formal-verification 1autoformalization 1lean 1multi-agent-systems 1prove2me 1protein-design 1drug-discovery 1agentic-ai 1analytical-chemistry 1