Overview
Selecting the right therapeutic target is one of the highest-leverage decisions in drug discovery. A sophisticated molecule cannot rescue a programme built around weak disease biology, inadequate safety margins or an inaccessible target. Artificial intelligence does not remove this biological uncertainty, but it can make the evidence-gathering and prioritization process more systematic, scalable and transparent.
A 2026 review by Pun and colleagues describes how machine learning, knowledge graphs, foundation models, structural prediction and automated experimentation are being integrated across target identification and assessment. The central message is not that a single algorithm can “discover a target.” Instead, AI is most useful when it connects heterogeneous evidence—human genetics, multi-omics, perturbation data, imaging, literature, clinical records and molecular structure—into testable therapeutic hypotheses.
Why Target Identification Remains a Bottleneck
The human genome contains roughly 20,000 protein-coding genes. Prior analyses estimate that several thousand may be druggable, while the number of targets through which approved drugs act is much smaller. This gap represents opportunity, but it also illustrates why target discovery is difficult: association with disease is not equivalent to causal relevance, and causal relevance is not equivalent to therapeutic tractability.
Protein-coding genes
The broad search space from which disease mechanisms and candidate targets may emerge.
Potentially druggable genes
An estimate cited in the review, not a list of targets that are already validated or clinically accessible.
Biological evidence
Genetics, causal inference, disease models and perturbation studies must support a coherent therapeutic hypothesis.
Translational evidence
Modality fit, safety, biomarkers, patient stratification and experimental feasibility determine whether the hypothesis can become a programme.
Target selection therefore requires several questions to be answered together: Does the target causally influence disease? Which direction of modulation is required? Is the target accessible to an appropriate modality? Can efficacy be separated from mechanism-based toxicity? Is there a biomarker strategy? Does the programme have a defensible point of differentiation?

The Evidence Foundation for AI-Driven Target Discovery
Human genetics and causal biology
Genome-wide association studies, expression or protein quantitative trait loci and Mendelian randomization can connect genetic variation to molecular traits and disease outcomes. These methods are valuable because they move the analysis from correlation toward causal inference, although linkage, pleiotropy and population representation must be handled carefully. Genetic support can increase confidence in a mechanism, but it does not by itself define the best modality, dose or patient population.
Multi-omics and perturbation data
Genomics, transcriptomics, proteomics, metabolomics and epigenomics describe disease at different molecular layers. Machine-learning models can integrate these layers to identify convergent pathways, disease subtypes and candidate regulators. CRISPR screens, RNA interference, cellular perturbation atlases and dependency maps add functional information that can distinguish a passenger signal from a vulnerability worth testing.
Cellular imaging and phenotypic data
High-content imaging captures organelle state, morphology, localization and stress responses at scale. Deep-learning embeddings can turn these images into quantitative phenotypes, enabling target nomination from compound perturbations or genetically defined cellular models. Imaging is especially useful when the disease phenotype cannot be summarized by one marker.
Knowledge graphs, literature and clinical information
Biological knowledge graphs connect genes, proteins, pathways, diseases, drugs and phenotypes. Natural-language processing can extract evidence from publications, patents and trial records, while graph models can prioritize relationships that are not obvious in a single dataset. However, these systems inherit publication bias, annotation errors and the tendency to overemphasize well-studied genes.
What AI Models Contribute
| Model family | Typical target-discovery use | Important limitation |
|---|---|---|
| Tree-based and statistical learning | Feature ranking, target classification, causal evidence scoring and interpretable baseline models. | Performance depends on feature engineering and may miss complex nonlinear biology. |
| Graph neural networks | Learning from protein-interaction networks, pathways and heterogeneous biomedical knowledge graphs. | Graph incompleteness and degree bias can make familiar biology appear more important. |
| Transformers and foundation models | Literature mining, sequence representation, cellular image embeddings and multimodal evidence synthesis. | Outputs require provenance, calibration and domain-specific validation. |
| Generative models | Generating hypotheses, structures or molecules after a target and design objective have been defined. | Generated candidates still require physical, biological and experimental validation. |
| Explainable AI | Identifying which evidence, features or pathways drive a target score. | Post-hoc explanations do not automatically establish causality. |
The most credible implementations pair prediction with traceable evidence. A useful target-ranking system should show why a target was prioritized, how uncertainty was estimated, which datasets contributed and what experiment would most efficiently challenge the hypothesis.

From Target Nomination to Target Assessment
Druggability and modality fit
Structural prediction can help assess whether a protein contains a ligandable pocket or an interface suitable for a biologic. AlphaFold-family models and related tools have expanded structural coverage, but a predicted structure is not automatically a screening-ready receptor. Conformational state, cofactors, oligomerization, membrane context and uncertainty around flexible regions all influence practical druggability.
Modality expands the concept of tractability. A target that is difficult for a conventional small molecule may be reachable through an antibody, peptide, oligonucleotide, targeted degradation strategy, cell therapy or gene-based intervention. The preferred modality must be matched to localization, desired pharmacology, tissue access and safety.
Safety and off-target risk
Target-associated safety risk can arise from normal physiology, tissue expression or pathway redundancy. Compound-associated risk adds off-target pharmacology and exposure-dependent liabilities. Computational approaches can compare binding pockets, predict secondary pharmacology and integrate expression or genetic loss-of-function evidence, but these predictions should guide—not replace—experimental safety assessment.
Novelty, confidence and differentiation
A mature target may offer abundant validation but intense competition. A novel target may provide differentiation while carrying greater biological and translational uncertainty. AI can organize the evidence landscape, yet the final decision remains a portfolio trade-off involving disease need, competitive timing, intellectual property and the feasibility of generating decisive data.
Clinical-Stage Examples: What They Demonstrate
The review discusses programmes in which AI contributed to identifying or supporting a target hypothesis. These examples are informative, but they should not be interpreted as proof that AI-originated targets have a uniformly higher clinical success rate.
| Target | AI contribution | Translational lesson |
|---|---|---|
| TNIK | Multimodal disease data and network-based models prioritized TNIK for idiopathic pulmonary fibrosis. | Target ranking can be connected to generative chemistry and rapid experimental iteration, but clinical efficacy still requires prospective testing. |
| APLNR/APJ | Longitudinal human ageing data supported apelin signalling as a metabolic and functional ageing axis. | Human data can strengthen a therapeutic hypothesis, while unexpected safety findings can still reshape a programme. |
| PIKfyve | Patient-derived multi-omics, protein networks and phenotypic data reinforced the target hypothesis in ALS. | Compelling computational and preclinical evidence may not translate into an adequate clinical risk–benefit profile. |
| DRD2 | A Bayesian model helped deconvolute a target for a compound originally found through phenotypic screening. | AI can support reverse target identification as well as target-first discovery. |
Challenges and Future Directions
- Data quality: missing metadata, inconsistent identifiers, batch effects and non-reproducible findings can propagate through a model.
- Representation bias: rare diseases and underrepresented populations may have insufficient data, while highly studied genes dominate literature-derived graphs.
- Interpretability: a target score must be linked to evidence and uncertainty if it is to guide expensive experiments.
- Benchmarking: retrospective accuracy is not enough; prospective and time-split validation better reflects real discovery performance.
- Experimental closure: predictions become more valuable when integrated with perturbation assays, automated laboratories and iterative model updating.

How CD ComputaBio Can Support Target Discovery Programs
AI-enabled target discovery is most useful when the computational work is designed around a concrete decision: which target to validate, which evidence gap to close, which modality to pursue or which experiment to run next. CD ComputaBio supports research-stage programmes through integrated computational workflows.
| Research need | Related support | Connection to the workflow |
|---|---|---|
| Rank disease-relevant targets | Target Identification and Validation Service | Integrates literature, databases and biological evidence to define a testable target shortlist. |
| Interpret molecular datasets | Bioinformatics Services | Supports processing and interpretation of omics and other biological data used in target nomination. |
| Enable structure-based assessment | Protein Structure Modeling Service | Provides structural models for feasibility assessment and downstream interaction analysis. |
| Evaluate ligandability | Binding Pocket and Druggability Modeling | Examines candidate pockets and structural features relevant to small-molecule targeting. |
| Assess secondary pharmacology | Drug Off-Target Effect Prediction Service | Prioritizes potential off-target interactions for focused experimental follow-up. |
| Connect a target to a specific programme | TNIK Targeting Services | Supports target-focused computational strategy development for a clinically relevant example discussed in the review. |
Contact Us
If you are building a target discovery programme from genetics, multi-omics, literature or structural data, CD ComputaBio can help convert heterogeneous evidence into a prioritized and experimentally actionable workflow. Contact our scientific team to discuss your disease area, available data and decision criteria.
References
- Pun FW, Podolskiy D, Izumchenko E, et al. Target identification and assessment in the era of AI. Nature Reviews Drug Discovery. 2026;25:534–552. https://doi.org/10.1038/s41573-026-01412-8
- Finan C, Gaulton A, Kruger FA, et al. The druggable genome and support for target identification and validation in drug development. Science Translational Medicine. 2017;9:eaag1166. https://doi.org/10.1126/scitranslmed.aag1166
- Hingorani AD, Kuan V, Finan C, et al. Improving the odds of drug development success through human genomics: modelling study. Scientific Reports. 2019;9:18911. https://doi.org/10.1038/s41598-019-54849-w
- Hempel K, Mergenthaler P, Riese D, et al. Improving target assessment in biomedical research: the GOT-IT recommendations. Nature Reviews Drug Discovery. 2021;20:64–81. https://doi.org/10.1038/s41573-020-0087-3
For Research Use Only. This page summarizes published research and describes computational research services. It does not constitute medical advice, a clinical claim or a guarantee of experimental or therapeutic success.