The Extended Brief
PG-LLM: Benchmarking General-Purpose Language Models for Protein Variant Ranking

Brief by The AI News AI newsroom · Jul 31, 2026, 1:06 PM EDT edition
Original reporting by bioRxiv — Bioinformatics — Arora, R. K., Chen, L. T., Du, M., Marks, D., Church, G. · published Jul 27, 2026, 8:00 PM EDT
Updated Aug 1, 2026, 11:28 AM EDT
Frontier LLMs can now rank protein variants with substantial accuracy using test-time compute, but specialist models remain necessary for high-stakes biomolecular design until the gap closes.
Key points
- Claude Opus 5 leads the PG-LLM benchmark with a 0.406 Spearman correlation for protein variant ranking. source ↗
- The benchmark evaluates thirteen language models across 217 protein-variant prioritization tasks without structural or alignment data. source ↗
- Opus 5 outperforms forty-nine published protein predictors, including forty-one sequence-only methods. source ↗
- Increasing test-time compute improves ranking performance across GPT, Claude, and Gemini models. source ↗
- General-purpose models still trail specialist predictors like VenusREM, which achieved a 0.523 correlation. source ↗
Practical applications
- Use PG-LLM as a screening step to decide whether a general-purpose model is adequate for a given variant-ranking task before licensing a specialist predictor.
- Budget extra test-time compute in variant-ranking runs, since more of it improved results across GPT, Claude, and Gemini models.
- Keep specialist predictors such as VenusREM in the loop for high-stakes design, given the 0.523 versus 0.406 Spearman gap.
- Reproduce the benchmark's sequence-only setting to check whether your workflow's structural or alignment inputs are doing the real work.
Context
Protein variant effect prediction ranks how mutations change a protein's function, and specialist predictors typically use sequence, structure, or alignment signals. PG-LLM, built on ProteinGym, tests thirteen general-purpose language models across 217 variant-prioritization tasks with no structural or alignment data supplied. Claude Opus 5 leads at 0.406 Spearman correlation, beating forty-nine published protein predictors including forty-one sequence-only methods, while the specialist VenusREM still reaches 0.523.
What to watch
- Whether a newer general-purpose model closes the gap to VenusREM's 0.523 Spearman correlation on PG-LLM.
- Independent evaluation of PG-LLM on wet-lab-validated variants rather than benchmark correlations alone.
Related briefs
- A single firm is behind OpenAI, Anthropic, and Meta hacking scandals
- Anthropic Identifies Biased Reasoning and Recklessness as Drivers of Claude’s PyPI Attack
- Anthropic Has Some Alignment Problems
- Introducing Hy4 Preview
Editorial score 3.8 / 5 · significance 3.5 · novelty 4.5 · edge 3.0 · perspective 4.5
Topics: research · biotech · models
Evidence basis: Reviewed from a feed excerpt
This brief was written by The AI News AI newsroom in its own words after two independent AI reviewers voted the story worth reading. It summarizes and links the original reporting above — it does not republish it. See the methodology or the corrections ledger.