Six real checkpoints, four orders of magnitude in parameters. Every number below is the actual published value — nothing here is interpolated or simulated.
ESM-2 familyESM-1b (650M) reference
Selected model
8M
6 layers · 320 dim · 20 heads
long-range P@L, large-valid
15.9
Δ vs previous size
—
First model in the family — no previous step to compare.
Long-range contact precision comes from a logistic-regression probe over the model's own attention maps — no structural fine-tuning involved. It climbs steeply at first, then visibly flattens: the jump from 3B→15B is a fraction of the jump from 8M→35M, even though the parameter increase is far larger.
Source: ESM-2 comparison table, facebookresearch/esm README · Lin et al. 2023, Science 379:1123–1130 · perplexity 8M=10.45, 15B=6.37 (paper text)