Generative AI for Antibody CDR Design: Protein Language Models for CLDN18.2 Antibody Optimization
CD ComputaBio provides software-based computational services to support research and development. We do not offer free software packages.
Overview
This article examines the 2026 PLOS Computational Biology study by Qu and colleagues, which introduces cdrGPT for scaffold-constrained generation and prioritization of antibody heavy-chain complementarity-determining region 3 (CDRH3) sequences. The design case centers on zolbetuximab, a clinically validated antibody against claudin 18.2 (CLDN18.2). Rather than redesigning a complete antibody, the authors keep the zolbetuximab-derived framework and light chain fixed while varying CDRH3, a loop that often makes important antigen contacts.
The workflow combines GPT-2-style autoregressive generation, pretraining on 637,887 curated Observed Antibody Space sequences, rejection-sampling fine-tuning, and computational filters for predicted affinity, charge, hydrophobicity, and MHC class II presentation risk. From 50,000 FT2-generated sequences, 313 met all zolbetuximab-guided thresholds. Seven candidates were then examined structurally with AlphaFold 3 (AF3), and two underwent 100 ns molecular-dynamics (MD) simulations. All reported candidate performance remains computational: the study did not experimentally measure expression, binding, specificity, developability, or biological activity.
Why CLDN18.2 Is a Relevant Test Case
CLDN18.2 is normally restricted largely to differentiated gastric mucosa but becomes exposed and aberrantly expressed in gastric and other selected solid tumors. The clinical activity and regulatory approval of zolbetuximab establish that the antigen can support therapeutic intervention. For computational antibody engineering, however, a validated target does not remove the multi-objective problem: sequence changes intended to improve predicted binding may also alter solution behavior, clearance risk, immunogenicity-related signals, or specificity.
The authors therefore use zolbetuximab as both a structural scaffold and a numerical reference. Generated CDRH3 loops are grafted into its heavy-chain variable region, while the native light chain is retained. This narrows the search space and makes comparisons more controlled, but it also means the work is best understood as CDRH3-focused optimization around one parental architecture, not unrestricted de novo antibody discovery.
From Antibody Repertoires to a CDRH3 Generator
The source repertoire was downloaded from OAS and standardized with IMGT numbering. The authors removed sequences containing ambiguous residues and retained CDRH3 lengths of 8-14 amino acids. Their final pretraining set contained 637,887 non-redundant sequences. The model uses eight decoder blocks, 256-dimensional attention outputs, and feed-forward layers that expand to 1,024 dimensions before projecting back to 256. Causal masking trains it to predict the next amino-acid token from the preceding sequence context.
Model checkpoint selection was important. At epoch 6, generated lengths were concentrated mainly between 10 and 14 residues and were reported not to differ significantly from natural OAS sequences on the evaluated metrics (P > 0.05). Epochs 8 and 10 produced abnormal sequences longer than 100 residues, which the authors interpreted as overfitting. Subsequent candidate generation therefore used epoch-6 weights. On four NVIDIA A100 GPUs for training and one A100 for inference, generation plus computational evaluation of 20,000 candidates took about four hours.
Multi-Objective Fine-Tuning and Candidate Selection
cdrGPT does not optimize an experimentally measured endpoint. Instead, the authors assemble each generated CDRH3 in the zolbetuximab variable-region context and score several computational proxies. AlphaBind supplies a predicted binding score, with lower values treated as stronger predicted affinity. FvNetCharge and the charge-symmetry parameter FvCSP are used as viscosity-associated descriptors, while the summed hydrophobicity index HISum is used as a clearance-associated descriptor. NetMHCIIpan 4.0 evaluates 19 overlapping 15-mers against 34 HLA class II alleles; a higher minimum percentage rank is interpreted as lower predicted presentation risk.
Two rounds of rejection-sampling fine-tuning moved the generated distributions toward the specified profile. The final FT2 screen required a predicted AlphaBind score better than the zolbetuximab reference, FvNetCharge below 18.1, FvCSP above 22.6, HISum between 0 and 4, and MHC II minPR above 1.29. Applying all criteria to 50,000 sequences left 313 unique candidates. These thresholds create a transparent funnel, but they are model-derived filters rather than experimentally calibrated acceptance specifications.
| Stage | Input | Computational decision | Output |
|---|---|---|---|
| Pretraining | 637,887 curated OAS CDRH3 sequences | Learn natural-sequence statistics by next-token prediction | Epoch-6 prior model |
| Generation | Zolbetuximab scaffold context | Sample new 8-14-residue CDRH3 sequences | 50,000 FT2 sequences |
| Multi-property screen | Full variable-region constructs | Apply predicted affinity, charge, hydrophobicity, and MHC II thresholds | 313 candidates |
| Representation analysis | ESM2 embeddings | PCA, clustering, and LASSO-based ranking | Seven cluster-0 representatives |
| Structural review | Grafted antibody candidates and CLDN18.2 | AF3 modeling; selected 100 ns MD simulations | Computational structural hypotheses |
Novelty, Generalization, and Metric-Specific Trade-Offs
The pretrained model produced uniqueness of 1.000, novelty of 0.999, and an average pairwise Levenshtein distance of 9.54. Fine-tuned models retained high uniqueness and novelty but showed lower diversity values of 4.741 and 5.459, consistent with task-specific optimization constraining sequence space. A HER2 exercise grafted 5,000 generated CDRH3 sequences onto trastuzumab and found predicted-property distributions aligned with the chosen trastuzumab benchmarks. Additional tests used adalimumab, bevacizumab, and panitumumab, but the preferred properties differed by antibody context.
A shared zolbetuximab benchmark against AbLang and AntiBERTy further prevents a simple “best model” conclusion. The authors report that cdrGPT had the highest HISum success rate and a slightly better FvNetCharge success rate, while its MHC II minPR performance was weaker. These are computational, metric-dependent comparisons. They support selecting or combining models according to project objectives rather than assuming that one generator dominates every developability dimension.
Clustering and Structural Assessment of Prioritized Designs
ESM2 embeddings separated the 313 candidates into three clusters: 95 sequences in cluster 0, 126 in cluster 1, and 92 in cluster 2. Zolbetuximab co-localized with cluster 0. Within-cluster cosine similarities exceeded 0.99, although such embedding similarity does not establish shared binding function. LASSO analysis ranked contributions across features and samples; seven cluster-0 sequences among the top 20 were selected for structural analysis.
AF3 models of the seven candidates were superposed on the zolbetuximab-derived template. The paper reports overall RMSDs within its predefined 2.0 Å threshold and a CDRH3-loop RMSD of 1.331 Å. For the selected seq2 and seq6 complexes, AF3 placed the designed loops near the CLDN18.2 extracellular region and suggested local contacts. These residues are predicted interface neighbors, not experimentally mapped paratope or epitope determinants. Moreover, repeated AF3 comparisons against CLDN18.1 and CLDN18.2 did not provide strong evidence of consistent isoform discrimination.
What the Simulations Add - and What They Do Not
The authors ran 100 ns MD simulations for seq2 and seq6 in complex with CLDN18.2. Neither model showed sustained dissociation or progressive unfolding during the simulated interval. The reported average chain A-chain B interface distances were approximately 6.56 nm for seq2 and 5.47 nm for seq6, with stable temperature, pressure, and potential-energy profiles used to support equilibration.
These simulations provide a dynamic consistency check for modeled complexes and can reveal obvious instability that a single structure misses. They do not validate the original pose, quantify binding affinity, demonstrate CLDN18.2-over-CLDN18.1 selectivity, or predict manufacturability. Their strongest interpretation is compatibility of the selected grafts with the parental framework and modeled complex over the simulated timescale.
R&D Implications, Limitations, and Next Experiments
For antibody teams, the study illustrates a useful funnel: restrict generation to a tractable region, evaluate candidates in full variable-domain context, apply multiple computational filters, inspect representation-space diversity, and reserve expensive structure or simulation work for a smaller set. It also shows why predicted affinity alone is insufficient. Charge, hydrophobicity, immunogenicity-related screening, scaffold integrity, and cross-reactivity hypotheses should be tracked independently so that a gain in one proxy does not hide deterioration in another.
The decisive limitation is the absence of wet-lab validation. No generated candidate was expressed or assayed for CLDN18.2 binding, kinetics, functional activity, aggregation, viscosity, stability, pharmacokinetics, or immune response. AlphaBind values are predictions rather than measured dissociation constants; MHC II binding is not equivalent to clinical immunogenicity; OAS pretraining is not a formal humanness assessment; and AF3/MD models cannot confirm specificity. The current model is also described as suitable for human IgG-like antibodies rather than VHH antibody scaffolds.
A prospective program would therefore synthesize a sequence-diverse subset, confirm expression and monomeric state, measure binding kinetics, test CLDN18.1 cross-reactivity, and add functional assays appropriate to the intended mechanism. Developability measurements should include aggregation and thermal stability, viscosity at relevant concentrations, nonspecific binding, and early pharmacokinetic risk. Experimental results can then be used to recalibrate thresholds or retrain the ranking layer instead of treating the initial computational scores as final evidence.
Overview of What CD ComputaBio Can Provide
The paper maps naturally to several computational research needs. The services below describe related support without implying that CD ComputaBio reproduced cdrGPT, validated its candidates, or can guarantee the reported outcomes.
| Research Need | Related Support | Connection |
|---|---|---|
| Generate candidates under antigen and scaffold constraints | Antibody De Novo Design Service | Supports computational sequence and structure exploration around defined design objectives. |
| Review framework and CDR structures | Antibody Modeling Services | Provides structural models for comparing loop conformations and scaffold integrity. |
| Evaluate candidate complex poses | Antibody-Antigen Docking Service | Explores plausible binding orientations and interface geometries for prioritized pairs. |
| Map predicted contact residues | Antibody-Antigen Interaction Modeling | Compares residue-level contacts relevant to paratope and epitope hypotheses. |
| Prioritize multi-property variants | Antibody Drug Optimization Service | Integrates computational criteria for iterative lead prioritization. |
| Probe modeled complex dynamics | Antibody Molecular Dynamics Simulation | Examines flexibility and interface persistence beyond a static prediction. |
References
- Qu T, Yuan L, Cui W, et al. CLDN18.2 antibody design with protein language models: A deep learning optimization framework. PLOS Computational Biology. 2026;22(8):e1014499. https://doi.org/10.1371/journal.pcbi.1014499.
- Olsen TH, Boyles F, Deane CM. Observed Antibody Space: A diverse database of cleaned, annotated, and translated unpaired and paired antibody sequences. Protein Science. 2022;31:141-146.
- Jumper J, Evans R, Pritzel A, et al. Highly accurate protein structure prediction with AlphaFold. Nature. 2021;596:583-589.
- Reynisson B, Alvarez B, Paul S, Peters B, Nielsen M. NetMHCIIpan-4.1 and NetMHCIIpan-4.0: improved predictions of MHC antigen presentation by concurrent motif deconvolution and integration of MS MHC eluted ligand data. Nucleic Acids Research. 2020;48:W449-W454.
Article and figure license. The primary study is Open Access under the Creative Commons Attribution 4.0 International License, which permits commercial reuse and adaptation with attribution. Figures 1 and 2 on this page are cropped from Figures 1 and 11 of Qu et al.; no scientific content within the figure panels was altered.
* For Research Use Only.