Generative AI for Antibody CDR Design: Protein Language Models for CLDN18.2 Antibody Optimization

CD ComputaBio provides software-based computational services to support research and development. We do not offer free software packages.

Overview

This article examines the 2026 PLOS Computational Biology study by Qu and colleagues, which introduces cdrGPT for scaffold-constrained generation and prioritization of antibody heavy-chain complementarity-determining region 3 (CDRH3) sequences. The design case centers on zolbetuximab, a clinically validated antibody against claudin 18.2 (CLDN18.2). Rather than redesigning a complete antibody, the authors keep the zolbetuximab-derived framework and light chain fixed while varying CDRH3, a loop that often makes important antigen contacts.

The workflow combines GPT-2-style autoregressive generation, pretraining on 637,887 curated Observed Antibody Space sequences, rejection-sampling fine-tuning, and computational filters for predicted affinity, charge, hydrophobicity, and MHC class II presentation risk. From 50,000 FT2-generated sequences, 313 met all zolbetuximab-guided thresholds. Seven candidates were then examined structurally with AlphaFold 3 (AF3), and two underwent 100 ns molecular-dynamics (MD) simulations. All reported candidate performance remains computational: the study did not experimentally measure expression, binding, specificity, developability, or biological activity.

637,887unique OAS CDRH3 sequences used for pretraining after curation
50,000FT2 sequences screened in the zolbetuximab framework
313unique candidates meeting every computational selection criterion
7cluster-0 representatives advanced to AF3 structural assessment

Why CLDN18.2 Is a Relevant Test Case

CLDN18.2 is normally restricted largely to differentiated gastric mucosa but becomes exposed and aberrantly expressed in gastric and other selected solid tumors. The clinical activity and regulatory approval of zolbetuximab establish that the antigen can support therapeutic intervention. For computational antibody engineering, however, a validated target does not remove the multi-objective problem: sequence changes intended to improve predicted binding may also alter solution behavior, clearance risk, immunogenicity-related signals, or specificity.

The authors therefore use zolbetuximab as both a structural scaffold and a numerical reference. Generated CDRH3 loops are grafted into its heavy-chain variable region, while the native light chain is retained. This narrows the search space and makes comparisons more controlled, but it also means the work is best understood as CDRH3-focused optimization around one parental architecture, not unrestricted de novo antibody discovery.

From Antibody Repertoires to a CDRH3 Generator

The source repertoire was downloaded from OAS and standardized with IMGT numbering. The authors removed sequences containing ambiguous residues and retained CDRH3 lengths of 8-14 amino acids. Their final pretraining set contained 637,887 non-redundant sequences. The model uses eight decoder blocks, 256-dimensional attention outputs, and feed-forward layers that expand to 1,024 dimensions before projecting back to 256. Causal masking trains it to predict the next amino-acid token from the preceding sequence context.

Model checkpoint selection was important. At epoch 6, generated lengths were concentrated mainly between 10 and 14 residues and were reported not to differ significantly from natural OAS sequences on the evaluated metrics (P > 0.05). Epochs 8 and 10 produced abnormal sequences longer than 100 residues, which the authors interpreted as overfitting. Subsequent candidate generation therefore used epoch-6 weights. On four NVIDIA A100 GPUs for training and one A100 for inference, generation plus computational evaluation of 20,000 candidates took about four hours.

cdrGPT workflow from curated CDRH3 sequences through language-model generation, multi-objective fine-tuning, clustering, and structural evaluation
Figure 1. cdrGPT connects OAS-based CDRH3 generation with predicted-property filtering and structural prioritization. Source: Qu et al. (2026), Figure 1, CC BY 4.0.

Multi-Objective Fine-Tuning and Candidate Selection

cdrGPT does not optimize an experimentally measured endpoint. Instead, the authors assemble each generated CDRH3 in the zolbetuximab variable-region context and score several computational proxies. AlphaBind supplies a predicted binding score, with lower values treated as stronger predicted affinity. FvNetCharge and the charge-symmetry parameter FvCSP are used as viscosity-associated descriptors, while the summed hydrophobicity index HISum is used as a clearance-associated descriptor. NetMHCIIpan 4.0 evaluates 19 overlapping 15-mers against 34 HLA class II alleles; a higher minimum percentage rank is interpreted as lower predicted presentation risk.

Two rounds of rejection-sampling fine-tuning moved the generated distributions toward the specified profile. The final FT2 screen required a predicted AlphaBind score better than the zolbetuximab reference, FvNetCharge below 18.1, FvCSP above 22.6, HISum between 0 and 4, and MHC II minPR above 1.29. Applying all criteria to 50,000 sequences left 313 unique candidates. These thresholds create a transparent funnel, but they are model-derived filters rather than experimentally calibrated acceptance specifications.

StageInputComputational decisionOutput
Pretraining637,887 curated OAS CDRH3 sequencesLearn natural-sequence statistics by next-token predictionEpoch-6 prior model
GenerationZolbetuximab scaffold contextSample new 8-14-residue CDRH3 sequences50,000 FT2 sequences
Multi-property screenFull variable-region constructsApply predicted affinity, charge, hydrophobicity, and MHC II thresholds313 candidates
Representation analysisESM2 embeddingsPCA, clustering, and LASSO-based rankingSeven cluster-0 representatives
Structural reviewGrafted antibody candidates and CLDN18.2AF3 modeling; selected 100 ns MD simulationsComputational structural hypotheses

Novelty, Generalization, and Metric-Specific Trade-Offs

The pretrained model produced uniqueness of 1.000, novelty of 0.999, and an average pairwise Levenshtein distance of 9.54. Fine-tuned models retained high uniqueness and novelty but showed lower diversity values of 4.741 and 5.459, consistent with task-specific optimization constraining sequence space. A HER2 exercise grafted 5,000 generated CDRH3 sequences onto trastuzumab and found predicted-property distributions aligned with the chosen trastuzumab benchmarks. Additional tests used adalimumab, bevacizumab, and panitumumab, but the preferred properties differed by antibody context.

A shared zolbetuximab benchmark against AbLang and AntiBERTy further prevents a simple “best model” conclusion. The authors report that cdrGPT had the highest HISum success rate and a slightly better FvNetCharge success rate, while its MHC II minPR performance was weaker. These are computational, metric-dependent comparisons. They support selecting or combining models according to project objectives rather than assuming that one generator dominates every developability dimension.

Clustering and Structural Assessment of Prioritized Designs

ESM2 embeddings separated the 313 candidates into three clusters: 95 sequences in cluster 0, 126 in cluster 1, and 92 in cluster 2. Zolbetuximab co-localized with cluster 0. Within-cluster cosine similarities exceeded 0.99, although such embedding similarity does not establish shared binding function. LASSO analysis ranked contributions across features and samples; seven cluster-0 sequences among the top 20 were selected for structural analysis.

AF3 models of the seven candidates were superposed on the zolbetuximab-derived template. The paper reports overall RMSDs within its predefined 2.0 Å threshold and a CDRH3-loop RMSD of 1.331 Å. For the selected seq2 and seq6 complexes, AF3 placed the designed loops near the CLDN18.2 extracellular region and suggested local contacts. These residues are predicted interface neighbors, not experimentally mapped paratope or epitope determinants. Moreover, repeated AF3 comparisons against CLDN18.1 and CLDN18.2 did not provide strong evidence of consistent isoform discrimination.

Predicted structures of selected CDRH3 designs, CLDN18.2 complexes, interface close-ups, and interaction diagrams
Figure 2. AF3-based structural assessment proposes candidate binding geometries but does not experimentally establish affinity or isoform selectivity. Source: Qu et al. (2026), Figure 11, CC BY 4.0.

What the Simulations Add - and What They Do Not

The authors ran 100 ns MD simulations for seq2 and seq6 in complex with CLDN18.2. Neither model showed sustained dissociation or progressive unfolding during the simulated interval. The reported average chain A-chain B interface distances were approximately 6.56 nm for seq2 and 5.47 nm for seq6, with stable temperature, pressure, and potential-energy profiles used to support equilibration.

These simulations provide a dynamic consistency check for modeled complexes and can reveal obvious instability that a single structure misses. They do not validate the original pose, quantify binding affinity, demonstrate CLDN18.2-over-CLDN18.1 selectivity, or predict manufacturability. Their strongest interpretation is compatibility of the selected grafts with the parental framework and modeled complex over the simulated timescale.

R&D Implications, Limitations, and Next Experiments

For antibody teams, the study illustrates a useful funnel: restrict generation to a tractable region, evaluate candidates in full variable-domain context, apply multiple computational filters, inspect representation-space diversity, and reserve expensive structure or simulation work for a smaller set. It also shows why predicted affinity alone is insufficient. Charge, hydrophobicity, immunogenicity-related screening, scaffold integrity, and cross-reactivity hypotheses should be tracked independently so that a gain in one proxy does not hide deterioration in another.

The decisive limitation is the absence of wet-lab validation. No generated candidate was expressed or assayed for CLDN18.2 binding, kinetics, functional activity, aggregation, viscosity, stability, pharmacokinetics, or immune response. AlphaBind values are predictions rather than measured dissociation constants; MHC II binding is not equivalent to clinical immunogenicity; OAS pretraining is not a formal humanness assessment; and AF3/MD models cannot confirm specificity. The current model is also described as suitable for human IgG-like antibodies rather than VHH antibody scaffolds.

A prospective program would therefore synthesize a sequence-diverse subset, confirm expression and monomeric state, measure binding kinetics, test CLDN18.1 cross-reactivity, and add functional assays appropriate to the intended mechanism. Developability measurements should include aggregation and thermal stability, viscosity at relevant concentrations, nonspecific binding, and early pharmacokinetic risk. Experimental results can then be used to recalibrate thresholds or retrain the ranking layer instead of treating the initial computational scores as final evidence.

Overview of What CD ComputaBio Can Provide

The paper maps naturally to several computational research needs. The services below describe related support without implying that CD ComputaBio reproduced cdrGPT, validated its candidates, or can guarantee the reported outcomes.

Research NeedRelated SupportConnection
Generate candidates under antigen and scaffold constraintsAntibody De Novo Design ServiceSupports computational sequence and structure exploration around defined design objectives.
Review framework and CDR structuresAntibody Modeling ServicesProvides structural models for comparing loop conformations and scaffold integrity.
Evaluate candidate complex posesAntibody-Antigen Docking ServiceExplores plausible binding orientations and interface geometries for prioritized pairs.
Map predicted contact residuesAntibody-Antigen Interaction ModelingCompares residue-level contacts relevant to paratope and epitope hypotheses.
Prioritize multi-property variantsAntibody Drug Optimization ServiceIntegrates computational criteria for iterative lead prioritization.
Probe modeled complex dynamicsAntibody Molecular Dynamics SimulationExamines flexibility and interface persistence beyond a static prediction.

References

  1. Qu T, Yuan L, Cui W, et al. CLDN18.2 antibody design with protein language models: A deep learning optimization framework. PLOS Computational Biology. 2026;22(8):e1014499. https://doi.org/10.1371/journal.pcbi.1014499.
  2. Olsen TH, Boyles F, Deane CM. Observed Antibody Space: A diverse database of cleaned, annotated, and translated unpaired and paired antibody sequences. Protein Science. 2022;31:141-146.
  3. Jumper J, Evans R, Pritzel A, et al. Highly accurate protein structure prediction with AlphaFold. Nature. 2021;596:583-589.
  4. Reynisson B, Alvarez B, Paul S, Peters B, Nielsen M. NetMHCIIpan-4.1 and NetMHCIIpan-4.0: improved predictions of MHC antigen presentation by concurrent motif deconvolution and integration of MS MHC eluted ligand data. Nucleic Acids Research. 2020;48:W449-W454.

Article and figure license. The primary study is Open Access under the Creative Commons Attribution 4.0 International License, which permits commercial reuse and adaptation with attribution. Figures 1 and 2 on this page are cropped from Figures 1 and 11 of Qu et al.; no scientific content within the figure panels was altered.

* For Research Use Only.

Related Services

Online Inquiry

Submit your project details below, and our team will respond within 24 hours.

x
Need help getting the data you need?

Talk to our technical team about your project!

I Want To Talk