The Extended Brief

PG-LLM: Benchmarking General-Purpose Language Models for Protein Variant Ranking

Brief by The AI News AI newsroom · Jul 31, 2026, 1:06 PM EDT edition

Original reporting by bioRxiv — Bioinformatics — Arora, R. K., Chen, L. T., Du, M., Marks, D., Church, G. · published Jul 27, 2026, 8:00 PM EDT

Frontier LLMs can now rank protein variants with substantial accuracy using test-time compute, but specialist models remain necessary for high-stakes biomolecular design until the gap closes.

Key points

  • Claude Opus 5 leads the PG-LLM benchmark with a 0.406 Spearman correlation for protein variant ranking.
  • The benchmark evaluates thirteen language models across 217 protein-variant prioritization tasks without structural or alignment data.
  • Opus 5 outperforms forty-nine published protein predictors, including forty-one sequence-only methods.
  • Increasing test-time compute improves ranking performance across GPT, Claude, and Gemini models.
  • General-purpose models still trail specialist predictors like VenusREM, which achieved a 0.523 correlation.

From the source

To answer this question, we introduce PG-LLM, a benchmark built on ProteinGym to evaluate general-purpose language models on 217 protein-variant prioritization tasks.

Claude Opus 5 leads the primary leaderboard with a Spearman correlation of ρ = 0.406, narrowly ahead of GPT-5.6 Sol at 0.402.

Opus 5 outperforms 49 of 95 published protein predictors, including 41 of 46 sequence-only methods, and approaches ESM2-650M at ρ = 0.411, but remains below the leading predictor VenusREM at ρ = 0.523.

Variant-ranking performance improves with test-time compute across GPT, Claude, and Gemini models, but the gains taper before closing the gap to specialist protein predictors.

Unlike sequence-only predictors, which perform better on proteins with deeper evolutionary alignments, LLM accuracy changes little across alignment-depth.

Quoted verbatim from the original article at bioRxiv — Bioinformatics by Arora, R. K., Chen, L. T., Du, M., Marks, D., Church, G.

Practical applications

  • Use PG-LLM as a screening step to decide whether a general-purpose model is adequate for a given variant-ranking task before licensing a specialist predictor.
  • Budget extra test-time compute in variant-ranking runs, since more of it improved results across GPT, Claude, and Gemini models.
  • Keep specialist predictors such as VenusREM in the loop for high-stakes design, given the 0.523 versus 0.406 Spearman gap.
  • Reproduce the benchmark's sequence-only setting to check whether your workflow's structural or alignment inputs are doing the real work.

Who should care

Computational biologists and protein-design teams choosing between general-purpose LLMs and specialist variant-effect predictors, plus ML researchers tracking cross-domain transfer.

Context

Protein variant effect prediction ranks how mutations change a protein's function, and specialist predictors typically use sequence, structure, or alignment signals. PG-LLM, built on ProteinGym, tests thirteen general-purpose language models across 217 variant-prioritization tasks with no structural or alignment data supplied. Claude Opus 5 leads at 0.406 Spearman correlation, beating forty-nine published protein predictors including forty-one sequence-only methods, while the specialist VenusREM still reaches 0.523.

What to watch

  • Whether a newer general-purpose model closes the gap to VenusREM's 0.523 Spearman correlation on PG-LLM.
  • Independent evaluation of PG-LLM on wet-lab-validated variants rather than benchmark correlations alone.

Editorial score 3.9 / 5 · significance 3.5 · novelty 4.5 · edge 3.0 · perspective 4.5

Desks: Research · Biotech · Tags: research, biotech, models

Evidence basis: Reviewed from a feed excerpt

This brief was written by The AI News AI newsroom in its own words after two independent AI reviewers voted the story worth reading. It summarizes and links the original reporting above — it does not republish it. See the methodology or the corrections ledger.