Who tells you what they ran?
A paper that never states its temperature is making a choice about reproducibility. This page counts those choices, per lab, across every paper read so far. The silences are the finding.
What the papers say about the models they studied
Denominator is every model mentioned across that lab’s read papers — so a paper studying twelve models counts twelve times. A low bar means the lab describes what it ran without describing how.
| Lab | Papers read | Models named | Temperature stated | Effort / reasoning stated | Deployment stated |
|---|---|---|---|---|---|
| Anthropic | 22 | 116 | 0%0/116 | 8%9/116 | 28%32/116 |
| Google DeepMind | 18 | 27 | 15%4/27 | 7%2/27 | 59%16/27 |
| OpenAI | 9 | 23 | 0%0/23 | 13%3/23 | 22%5/23 |
How each lab builds a paper
| Lab | Has a real abstract | States limitations | Claim–evidence pairs found | …of which show a gap |
|---|---|---|---|---|
| Anthropic | 32%7/22 | 100%22/22 | 123 | 71%87/123 |
| Google DeepMind | 100%18/18 | 78%14/18 | 92 | 80%74/92 |
| OpenAI | 67%6/9 | 78%7/9 | 47 | 70%33/47 |
How to read these numbers honestly
“Has a real abstract” is a genre fact, not a quality one. Anthropic publishes much of its research as lab posts that open with a tl;dr instead of a labelled abstract. That is a different publishing form, not a missing section.
The gap column is the softest number here. A gap was recorded when the claim reached past the evidence offered for it. That is a judgment made per pair, and it is the one column on this site that is not simply counting what the paper said.