AIntibody Benchmark: Testing AI Antibody Discovery Against Affinity and Developability
CD ComputaBio provides software-based computational services to support research and development. We do not offer free software packages.
Overview
This article reviews the 2026 Nature Biotechnology study "A blinded, prospective benchmark of in silico antibody discovery anchored to experimental affinity and developability." The AIntibody challenge was designed to test a question that retrospective benchmarks cannot answer reliably: when computational teams receive realistic discovery data but do not know the experimental outcomes, can their models deliver antibodies that combine strong binding with drug-relevant biophysical behavior?
Twenty-nine organizations submitted 511 designed or ranked antibody sequences across three tasks. All submitted constructs were produced as full-length IgGs and evaluated under shared experimental protocols. The benchmark found real but task-dependent progress. Several groups produced developable antibodies below 100 pM, and one affinity-maturation pipeline generated a 95 pM design representing an approximately 2,000-fold improvement over the parental antibody. Yet most methods did not generalize across tasks, cluster-level ranking was usually worse than random clone picking, and out-of-library design produced a broad mixture of exceptional binders, nonbinders and candidates with developability liabilities.
Why a Prospective Antibody Benchmark Was Needed
Computational antibody design is commonly evaluated retrospectively, often on public structures or affinity data already available to method developers. Such evaluations can be useful for method development, but they cannot fully separate transferable capability from target-specific tuning, data leakage or favorable case selection. Antibody engineering also differs from structure prediction because there is no single geometric answer. A therapeutic candidate must balance affinity and kinetics with expression, stability, aggregation, hydrophobicity, polyreactivity and manufacturability.
AIntibody therefore adopted a prospective and blinded design inspired by CASP. Teams worked from sequencing outputs and limited experimental information, submitted sequences before the organizers disclosed test results, and had the constructs evaluated by independent laboratories. SARS-CoV-2 receptor-binding domain (RBD) served as the antigen. The authors emphasize that this was a favorable, data-rich target with unusually extensive structural and sequence knowledge, so the results should be interpreted as an upper bound for this target class rather than evidence of general performance across arbitrary antigens.
Three Tasks Probe Different Parts of the Discovery Pipeline
Affinity maturation
Use phase 1 CDR-library sequencing outputs to redesign five CDRs while keeping HCDR3 and the framework fixed.
Cluster ranking
Select the highest-affinity developable clone in each of three HCDR3 clusters containing 2,524, 3,554 and 410 sequences.
Out-of-library design
Generate new CDR combinations absent from 32,059 unique selection sequences while leaving antibody frameworks unchanged.
Affinity and Developability Were Measured Together
The benchmark did not reward affinity in isolation. Submitted IgGs were first measured by high-throughput surface plasmon resonance (SPR), followed by rank-order single-point kinetic exclusion assay (KinExA) and standard KinExA for selected high-affinity candidates. Rank ordering was consistent across platforms: Spearman correlations were 0.94 between SPR and single-point KinExA, 0.97 between single-point and standard KinExA, and 0.92 between standard KinExA and SPR. Absolute values differed, with KinExA generally reporting stronger affinities, but the concordance supported robust ranking.
Each antibody also entered a five-assay developability panel: affinity-capture self-interaction nanoparticle spectroscopy (AC-SINS), baculovirus particle polyreactivity, hydrophobic interaction chromatography retention, melting temperature and aggregation temperature. Each assay received a pass, questionable or fail score of 0, 1 or 2. A summed score of 3 or less was considered developable. This composite system created a practical multiparameter screen, although the authors later identified an important weakness: an extreme failure in one assay could be offset by good performance elsewhere.
Challenge 1: AI Could Mature a Defined Antibody
Challenge 1 was the clearest success. Twenty-five organizations submitted 165 sequences. Of these, 118 bound RBD and 47 were nonbinders; 22 reached single-digit nanomolar affinity, five were subnanomolar and one was below 100 pM. Overall, 86.1% passed the developability threshold and 63.0% were both binders and developable.
Aureka's winning antibody reached 95 pM by KinExA, approximately 2,000 times stronger than the parental antibody. Its confidence interval overlapped that of the best experimentally affinity-matured antibody at 113 pM, so the two measurements were statistically indistinguishable. The sequence was also relatively distant from the parent and experimental set, indicating that the model was not merely returning the most common observed clone.
At the same time, a simple non-ML consensus baseline from ProBioGen ranked third by KinExA at 540 pM and achieved a perfect developability score of zero. That result is strategically important: with deep selection data, positional enrichment and consensus analysis can remain highly competitive. AI value should therefore be judged against strong statistical baselines, not only against the unmodified parent.
Challenge 2: Clone Ranking Was Usually Worse Than Random
Challenge 2 asked models to find better clones inside three HCDR3 clusters. Binding rates were high because the source library was already enriched for functional antibodies, but identifying improvement over the most abundant cluster control proved difficult. Only 10.3%, 13.8% and 9.8% of submissions in clusters 27F, 28F and 47F, respectively, improved on the control. By comparison, 39% of randomly picked clones from these clusters had higher affinity than the control.
The best predictions were individually impressive: 9.2 pM for 27F, 105-106 pM for 28F and 50 pM for 47F, corresponding to 4.0-fold, 1.7-fold and 7.6-fold improvements. However, only the WashU approach exceeded the random baseline overall, and it did not win across every cluster. Developability also depended strongly on HCDR3 context. All 47F submissions passed the composite threshold, whereas 41.4% of 28F submissions were disqualified for poor developability. The results show that a high binder hit rate is not the same as reliable affinity ranking and that cluster composition can dominate apparent model performance.
Challenge 3: Exceptional Designs Coexisted with High Failure Rates
For out-of-library design, teams received 32,059 unique selection sequences and affinity measurements for 142 antibodies, then proposed new CDR sequences or combinations. Among 168 tested submissions, 57 were subnanomolar and 24 were below 100 pM. The highest-affinity design measured 2.9 pM, and another developable design measured 8.69 pM, statistically comparable with the best experimental antibody at 9.2 pM.
The distribution was nevertheless broad. Approximately 30.4% of submissions were nonbinders and 16% were nondevelopable binders; together, 46.4% failed one of these basic requirements. Only 53.6% were both binders and developable. The nominal 2.9 pM winner also failed severely in hydrophobic interaction chromatography. Under the published composite score it could still win, but the authors concluded that such a liability would probably disqualify it from therapeutic development. Future benchmarks will need explicit go/no-go gates rather than allowing very high affinity to compensate for a critical biophysical red flag.
No Single Method Class Generalized Across Tasks
Submissions included statistical or heuristic approaches, classical machine learning, protein-language-model and attention-based systems, and structure-aware pipelines using tools such as AlphaFold, ProteinMPNN, AntiFold, RFdiffusion and pairformer architectures. The winning class changed with the task. A combined language-model and structure-aware pipeline won affinity maturation; classical ML, custom transformers and combined sequence-structure methods shared cluster-level wins; and a protein-language-model pipeline won out-of-library design.
Cross-task rankings were inconsistent even when the same organization or algorithm participated in multiple challenges. This argues against a universal "one model for everything" interpretation. The useful question for an antibody program is narrower: which method is appropriate for the available data, the permitted sequence changes and the decision being made? Data-rich local maturation, cluster ranking and novel-sequence generation are related but distinct computational problems.
Practical Lessons for Antibody Discovery Programs
| Decision | Evidence from the benchmark | Practical implication |
|---|---|---|
| Choose the modeling regime | Winning method classes differed across maturation, ranking and out-of-library design. | Match the algorithm to the data regime and allowed design space instead of selecting by headline performance. |
| Set a competitive baseline | A consensus sequence ranked third in challenge 1, while random picking outperformed most cluster-ranking models. | Compare AI against consensus, abundance-based and random baselines before attributing value to model complexity. |
| Keep wet-lab gates | Strong binders could be nondevelopable, and the nominal 2.9 pM winner had a severe HIC liability. | Affinity, specificity and developability assays must remain integrated into design-build-test cycles. |
| Test transferability | Successes rarely transferred across all tasks or HCDR3 clusters. | Validate on new antigens, libraries and data depths before treating a result as platform-level performance. |
Limitations and Outlook
The study was a first benchmark rather than a final verdict on computational antibody design. It tested one heavily characterized antigen, used deeper sequencing than many practical campaigns, and supplied affinity data in challenges 2 and 3 that would normally appear late in discovery. The organizers also participated in reporting the results, and blinding depended on consortium integrity rather than informatic safeguards. Several prominent computational groups did not participate, limiting representativeness.
Even within these favorable conditions, consistency and generalization remained the central gaps. Future rounds should broaden antigens, separate organizers from evaluators, secure blinding technically, test data-sparse settings and enforce hard developability gates. The most constructive conclusion is not that AI antibody design succeeded or failed as a category. Rather, it can produce valuable candidates in defined, biologically grounded regimes, but reliable deployment still requires task-specific validation, competitive baselines and uniform experimental measurement.
Overview of What CD ComputaBio Can Provide
Computational antibody projects benefit from linking sequence design and structural analysis to explicit experimental decisions. The following capabilities relate to the technical questions raised by this study; they do not imply that any computational workflow can replace prospective affinity and developability testing.
| Research Need | Related CD ComputaBio Support | Connection to This Article |
|---|---|---|
| Generate or refine antibody sequences | Antibody Drug Design Service | Supports sequence-level design strategies aligned to defined antigen, format and optimization objectives. |
| Build and compare Fv structures | Antibody Modeling Services | Provides structural models for CDR conformation review and downstream interaction analysis. |
| Evaluate antigen recognition | Antibody-Antigen Docking Service | Examines plausible binding poses and interface contacts for prioritized candidates. |
| Optimize an established lead | Antibody Drug Optimization Service | Connects affinity, stability and sequence-liability considerations during lead refinement. |
| Assess conformational stability | Antibody Molecular Dynamics Simulation | Explores flexibility, interface persistence and structural behavior beyond a single static model. |
| Plan a de novo program | Antibody De Novo Design Service | Frames candidate generation around target epitope, sequence constraints and validation-ready ranking. |
Reference
- Erasmus, M. F. et al. A blinded, prospective benchmark of in silico antibody discovery anchored to experimental affinity and developability. Nature Biotechnology (2026). https://doi.org/10.1038/s41587-026-03238-6.
* For Research Use Only.