Abbindgym Reveals Benchmark-Context-Dependent Protein Language Model Selection for Antibody-Antigen Binding
Protein language models (PLMs) are increasingly used to prioritize antibody variants, but whether benchmark performance reliably guides PLM choice for antibody engineering remains unclear. Here we present AbBindGym, a dual-paradigm benchmark comparing 42 pretrained PLMs across 10 antibody--antige
Protein language models (PLMs) are increasingly used to prioritize antibody variants, but whether benchmark performance reliably guides PLM choice for antibody engineering remains unclear. Here we present AbBindGym, a dual-paradigm benchmark comparing 42 pretrained PLMs across 10 antibody–antigen datasets, including a single-laboratory BCR-seq/ELISA dataset (AbELA) for assessing benchmark transferability. Models were evaluated by zero-shot mutation scoring and supervised frozen-backbone regression on assay-specific binding readouts. Rankings were strongly context-dependent: structure-aware PLMs were enriched among leading zero-shot models, whereas high-capacity general-purpose PLMs occupied many leading positions under supervised probing. Current antibody-specific PLMs did not consistently outperform strong general-purpose models. Rankings also shifted with assay readout, split structure and dataset scale/composition, and public-benchmark rankings did not reliably transfer to AbELA. These findings indicate that generic leaderboards are insufficient for antibody-engineering model selection and support deployment-matched evaluation as a practical standard.
📤 Share this page
Found this useful? Share it with your network.
Files are hosted on the source repository. Click download to access the full dataset.