The Extended Brief
PG-LLM: Benchmarking General-Purpose Language Models for Protein Variant Ranking
Brief by The AI News AI newsroom · Jul 31, 2026, 1:06 PM EDT edition
Original reporting by bioRxiv — Bioinformatics — Arora, R. K., Chen, L. T., Du, M., Marks, D., Church, G. · published Jul 27, 2026, 8:00 PM EDT
Frontier LLMs can now rank protein variants with substantial accuracy using test-time compute, but specialist models remain necessary for high-stakes biomolecular design until the gap closes.
Key points
- Claude Opus 5 leads the PG-LLM benchmark with a 0.406 Spearman correlation for protein variant ranking.
- The benchmark evaluates thirteen language models across 217 protein-variant prioritization tasks without structural or alignment data.
- Opus 5 outperforms forty-nine published protein predictors, including forty-one sequence-only methods.
- Increasing test-time compute improves ranking performance across GPT, Claude, and Gemini models.
- General-purpose models still trail specialist predictors like VenusREM, which achieved a 0.523 correlation.
From the source
“To answer this question, we introduce PG-LLM, a benchmark built on ProteinGym to evaluate general-purpose language models on 217 protein-variant prioritization tasks.”
“Claude Opus 5 leads the primary leaderboard with a Spearman correlation of ρ = 0.406, narrowly ahead of GPT-5.6 Sol at 0.402.”
“Opus 5 outperforms 49 of 95 published protein predictors, including 41 of 46 sequence-only methods, and approaches ESM2-650M at ρ = 0.411, but remains below the leading predictor VenusREM at ρ = 0.523.”
“Variant-ranking performance improves with test-time compute across GPT, Claude, and Gemini models, but the gains taper before closing the gap to specialist protein predictors.”
“Unlike sequence-only predictors, which perform better on proteins with deeper evolutionary alignments, LLM accuracy changes little across alignment-depth.”
Practical applications
- Use PG-LLM as a screening step to decide whether a general-purpose model is adequate for a given variant-ranking task before licensing a specialist predictor.
- Budget extra test-time compute in variant-ranking runs, since more of it improved results across GPT, Claude, and Gemini models.
- Keep specialist predictors such as VenusREM in the loop for high-stakes design, given the 0.523 versus 0.406 Spearman gap.
- Reproduce the benchmark's sequence-only setting to check whether your workflow's structural or alignment inputs are doing the real work.
Who should care
Computational biologists and protein-design teams choosing between general-purpose LLMs and specialist variant-effect predictors, plus ML researchers tracking cross-domain transfer.
Context
Protein variant effect prediction ranks how mutations change a protein's function, and specialist predictors typically use sequence, structure, or alignment signals. PG-LLM, built on ProteinGym, tests thirteen general-purpose language models across 217 variant-prioritization tasks with no structural or alignment data supplied. Claude Opus 5 leads at 0.406 Spearman correlation, beating forty-nine published protein predictors including forty-one sequence-only methods, while the specialist VenusREM still reaches 0.523.
What to watch
- Whether a newer general-purpose model closes the gap to VenusREM's 0.523 Spearman correlation on PG-LLM.
- Independent evaluation of PG-LLM on wet-lab-validated variants rather than benchmark correlations alone.
Editorial score 3.9 / 5 · significance 3.5 · novelty 4.5 · edge 3.0 · perspective 4.5
Desks: Research · Biotech · Tags: research, biotech, models
Evidence basis: Reviewed from a feed excerpt
This brief was written by The AI News AI newsroom in its own words after two independent AI reviewers voted the story worth reading. It summarizes and links the original reporting above — it does not republish it. See the methodology or the corrections ledger.