The three labs, read closely.
Every empirical research publication Anthropic, Google DeepMind and OpenAI put out in the last five months — taken apart into what they defined, what they ran, what they measured, what they claimed, and where the claim and the evidence do not quite meet.
Four ways in
The Abstract Library
Every abstract in one place. Filter by lab, search the text, and pin up to four side by side to compare how three labs open an argument.
compare & pinAn annotated bibliography per lab
Each lab hub lists its papers with a one-sentence TLDR, so you can scan fifty studies without opening one.
scan firstHow to use this site
What each section means, what the colour coding tells you, and how to read the claim-versus-evidence panels.
the mapHow these 50 were chosen
The enumeration, the exclusion rule, the five contaminations that were caught and repaired, and the honest gaps.
the receiptsWho published, and how much
Anthropic
Most papers in the window, and the shortest window. Interpretability, alignment science, and a hard turn into using Claude as a research instrument.
Google DeepMind
The most conventionally academic of the three: peer-reviewed venues, long author lists, and a strong sociotechnical and evaluation streak.
OpenAI
Fewest paper-genre publications of the three. Heavy on science acceleration, agentic evaluation, and results announced as company posts.
Anthropic wins on density, not just count
Google DeepMind’s window opens earliest (23 April) and OpenAI’s next (29 April). Anthropic’s earliest paper in the set lands in May — the shortest window of the three — and it still contributes the most papers. Its publication rate over these five months is the highest of the Big Three.
Caveat kept in view: 16 of Anthropic’s 22 carry month-only dates, because its Circuits and Alignment surfaces date by month. Those are sorted at the 15th as a midpoint, so the imprecision does not systematically age them up or down.
Every paper on one line each
Each dot is a publication; hover for its title, click to open it. Dot size shows how much text was retrievable — big is the full paper, small is abstract-only. Where dots stack, they are fanned so none is hidden.
Who tells you what they ran → — a cross-lab count of which labs disclose temperature, reasoning effort and deployment, and which stay silent.
What you are actually holding
Not every publication yields the same depth of text. Each paper carries a badge saying which tier it is, so no claim on this site rests on more than it should.
Full text
The complete paper, followed through to arXiv, ACL or the lab’s own PDF.
Lab post only
The lab’s own write-up, with no separate full paper published behind it.
Abstract only
Full text sits behind a paywall or has no external link. Read these as abstracts.
Deep reads in progress
49 of 50 papers have their full structured read built. The rest have their corpus entry and a working link to the paper; their pages fill in as the reads land.