Case Study
AIntibody Benchmark: Testing AI Antibody Discovery Against Affinity and Developability

Inquiry
AIntibody Benchmark: Testing AI Antibody Discovery Against Affinity and Developability - CD ComputaBio

AIntibody Benchmark: Testing AI Antibody Discovery Against Affinity and Developability

CD ComputaBio provides software-based computational services to support research and development. We do not offer free software packages.

Overview

This article reviews the 2026 Nature Biotechnology study "A blinded, prospective benchmark of in silico antibody discovery anchored to experimental affinity and developability." The AIntibody challenge was designed to test a question that retrospective benchmarks cannot answer reliably: when computational teams receive realistic discovery data but do not know the experimental outcomes, can their models deliver antibodies that combine strong binding with drug-relevant biophysical behavior?

Twenty-nine organizations submitted 511 designed or ranked antibody sequences across three tasks. All submitted constructs were produced as full-length IgGs and evaluated under shared experimental protocols. The benchmark found real but task-dependent progress. Several groups produced developable antibodies below 100 pM, and one affinity-maturation pipeline generated a 95 pM design representing an approximately 2,000-fold improvement over the parental antibody. Yet most methods did not generalize across tasks, cluster-level ranking was usually worse than random clone picking, and out-of-library design produced a broad mixture of exceptional binders, nonbinders and candidates with developability liabilities.

511AI-designed or computationally selected antibody submissions
29participating organizations spanning academia, biotech, pharma and big tech
3distinct tasks covering maturation, ranking and out-of-library CDR design
5 assaysused to score hydrophobicity, self-interaction, polyreactivity and thermal behavior

Why a Prospective Antibody Benchmark Was Needed

Computational antibody design is commonly evaluated retrospectively, often on public structures or affinity data already available to method developers. Such evaluations can be useful for method development, but they cannot fully separate transferable capability from target-specific tuning, data leakage or favorable case selection. Antibody engineering also differs from structure prediction because there is no single geometric answer. A therapeutic candidate must balance affinity and kinetics with expression, stability, aggregation, hydrophobicity, polyreactivity and manufacturability.

AIntibody therefore adopted a prospective and blinded design inspired by CASP. Teams worked from sequencing outputs and limited experimental information, submitted sequences before the organizers disclosed test results, and had the constructs evaluated by independent laboratories. SARS-CoV-2 receptor-binding domain (RBD) served as the antigen. The authors emphasize that this was a favorable, data-rich target with unusually extensive structural and sequence knowledge, so the results should be interpreted as an upper bound for this target class rather than evidence of general performance across arbitrary antigens.

Three Tasks Probe Different Parts of the Discovery Pipeline

1

Affinity maturation

Use phase 1 CDR-library sequencing outputs to redesign five CDRs while keeping HCDR3 and the framework fixed.

2

Cluster ranking

Select the highest-affinity developable clone in each of three HCDR3 clusters containing 2,524, 3,554 and 410 sequences.

3

Out-of-library design

Generate new CDR combinations absent from 32,059 unique selection sequences while leaving antibody frameworks unchanged.

The three AIntibody benchmark tasks: in silico affinity maturation, HCDR3-cluster ranking, and out-of-library CDR design
Figure 1. The three benchmark tasks differ in their design freedom and available experimental context.

Affinity and Developability Were Measured Together

The benchmark did not reward affinity in isolation. Submitted IgGs were first measured by high-throughput surface plasmon resonance (SPR), followed by rank-order single-point kinetic exclusion assay (KinExA) and standard KinExA for selected high-affinity candidates. Rank ordering was consistent across platforms: Spearman correlations were 0.94 between SPR and single-point KinExA, 0.97 between single-point and standard KinExA, and 0.92 between standard KinExA and SPR. Absolute values differed, with KinExA generally reporting stronger affinities, but the concordance supported robust ranking.

Each antibody also entered a five-assay developability panel: affinity-capture self-interaction nanoparticle spectroscopy (AC-SINS), baculovirus particle polyreactivity, hydrophobic interaction chromatography retention, melting temperature and aggregation temperature. Each assay received a pass, questionable or fail score of 0, 1 or 2. A summed score of 3 or less was considered developable. This composite system created a practical multiparameter screen, although the authors later identified an important weakness: an extreme failure in one assay could be offset by good performance elsewhere.

Challenge 1: AI Could Mature a Defined Antibody

Challenge 1 was the clearest success. Twenty-five organizations submitted 165 sequences. Of these, 118 bound RBD and 47 were nonbinders; 22 reached single-digit nanomolar affinity, five were subnanomolar and one was below 100 pM. Overall, 86.1% passed the developability threshold and 63.0% were both binders and developable.

Aureka's winning antibody reached 95 pM by KinExA, approximately 2,000 times stronger than the parental antibody. Its confidence interval overlapped that of the best experimentally affinity-matured antibody at 113 pM, so the two measurements were statistically indistinguishable. The sequence was also relatively distant from the parent and experimental set, indicating that the model was not merely returning the most common observed clone.

At the same time, a simple non-ML consensus baseline from ProBioGen ranked third by KinExA at 540 pM and achieved a perfect developability score of zero. That result is strategically important: with deep selection data, positional enrichment and consensus analysis can remain highly competitive. AI value should therefore be judged against strong statistical baselines, not only against the unmodified parent.

Challenge 2: Clone Ranking Was Usually Worse Than Random

Challenge 2 asked models to find better clones inside three HCDR3 clusters. Binding rates were high because the source library was already enriched for functional antibodies, but identifying improvement over the most abundant cluster control proved difficult. Only 10.3%, 13.8% and 9.8% of submissions in clusters 27F, 28F and 47F, respectively, improved on the control. By comparison, 39% of randomly picked clones from these clusters had higher affinity than the control.

The best predictions were individually impressive: 9.2 pM for 27F, 105-106 pM for 28F and 50 pM for 47F, corresponding to 4.0-fold, 1.7-fold and 7.6-fold improvements. However, only the WashU approach exceeded the random baseline overall, and it did not win across every cluster. Developability also depended strongly on HCDR3 context. All 47F submissions passed the composite threshold, whereas 41.4% of 28F submissions were disqualified for poor developability. The results show that a high binder hit rate is not the same as reliable affinity ranking and that cluster composition can dominate apparent model performance.

Challenge 3: Exceptional Designs Coexisted with High Failure Rates

For out-of-library design, teams received 32,059 unique selection sequences and affinity measurements for 142 antibodies, then proposed new CDR sequences or combinations. Among 168 tested submissions, 57 were subnanomolar and 24 were below 100 pM. The highest-affinity design measured 2.9 pM, and another developable design measured 8.69 pM, statistically comparable with the best experimental antibody at 9.2 pM.

The distribution was nevertheless broad. Approximately 30.4% of submissions were nonbinders and 16% were nondevelopable binders; together, 46.4% failed one of these basic requirements. Only 53.6% were both binders and developable. The nominal 2.9 pM winner also failed severely in hydrophobic interaction chromatography. Under the published composite score it could still win, but the authors concluded that such a liability would probably disqualify it from therapeutic development. Future benchmarks will need explicit go/no-go gates rather than allowing very high affinity to compensate for a critical biophysical red flag.

No Single Method Class Generalized Across Tasks

Submissions included statistical or heuristic approaches, classical machine learning, protein-language-model and attention-based systems, and structure-aware pipelines using tools such as AlphaFold, ProteinMPNN, AntiFold, RFdiffusion and pairformer architectures. The winning class changed with the task. A combined language-model and structure-aware pipeline won affinity maturation; classical ML, custom transformers and combined sequence-structure methods shared cluster-level wins; and a protein-language-model pipeline won out-of-library design.

Cross-task rankings were inconsistent even when the same organization or algorithm participated in multiple challenges. This argues against a universal "one model for everything" interpretation. The useful question for an antibody program is narrower: which method is appropriate for the available data, the permitted sequence changes and the decision being made? Data-rich local maturation, cluster ranking and novel-sequence generation are related but distinct computational problems.

Comparison of computational antibody method classes and cross-challenge performance in the AIntibody study
Figure 2. Method classes and cross-challenge performance show that no approach dominated every task.

Practical Lessons for Antibody Discovery Programs

DecisionEvidence from the benchmarkPractical implication
Choose the modeling regimeWinning method classes differed across maturation, ranking and out-of-library design.Match the algorithm to the data regime and allowed design space instead of selecting by headline performance.
Set a competitive baselineA consensus sequence ranked third in challenge 1, while random picking outperformed most cluster-ranking models.Compare AI against consensus, abundance-based and random baselines before attributing value to model complexity.
Keep wet-lab gatesStrong binders could be nondevelopable, and the nominal 2.9 pM winner had a severe HIC liability.Affinity, specificity and developability assays must remain integrated into design-build-test cycles.
Test transferabilitySuccesses rarely transferred across all tasks or HCDR3 clusters.Validate on new antigens, libraries and data depths before treating a result as platform-level performance.

Limitations and Outlook

The study was a first benchmark rather than a final verdict on computational antibody design. It tested one heavily characterized antigen, used deeper sequencing than many practical campaigns, and supplied affinity data in challenges 2 and 3 that would normally appear late in discovery. The organizers also participated in reporting the results, and blinding depended on consortium integrity rather than informatic safeguards. Several prominent computational groups did not participate, limiting representativeness.

Even within these favorable conditions, consistency and generalization remained the central gaps. Future rounds should broaden antigens, separate organizers from evaluators, secure blinding technically, test data-sparse settings and enforce hard developability gates. The most constructive conclusion is not that AI antibody design succeeded or failed as a category. Rather, it can produce valuable candidates in defined, biologically grounded regimes, but reliable deployment still requires task-specific validation, competitive baselines and uniform experimental measurement.

Overview of What CD ComputaBio Can Provide

Computational antibody projects benefit from linking sequence design and structural analysis to explicit experimental decisions. The following capabilities relate to the technical questions raised by this study; they do not imply that any computational workflow can replace prospective affinity and developability testing.

Research NeedRelated CD ComputaBio SupportConnection to This Article
Generate or refine antibody sequencesAntibody Drug Design ServiceSupports sequence-level design strategies aligned to defined antigen, format and optimization objectives.
Build and compare Fv structuresAntibody Modeling ServicesProvides structural models for CDR conformation review and downstream interaction analysis.
Evaluate antigen recognitionAntibody-Antigen Docking ServiceExamines plausible binding poses and interface contacts for prioritized candidates.
Optimize an established leadAntibody Drug Optimization ServiceConnects affinity, stability and sequence-liability considerations during lead refinement.
Assess conformational stabilityAntibody Molecular Dynamics SimulationExplores flexibility, interface persistence and structural behavior beyond a single static model.
Plan a de novo programAntibody De Novo Design ServiceFrames candidate generation around target epitope, sequence constraints and validation-ready ranking.

Reference

  1. Erasmus, M. F. et al. A blinded, prospective benchmark of in silico antibody discovery anchored to experimental affinity and developability. Nature Biotechnology (2026). https://doi.org/10.1038/s41587-026-03238-6.

* For Research Use Only.

Related Services

Online Inquiry

Submit your project details below, and our team will respond within 24 hours.

x
Need help getting the data you need?

Talk to our technical team about your project!

I Want To Talk