Technical Bulletin · Prediction Reliability

How Reliable Are In Silico ADMET Predictions?

A practical framework for deciding how much weight to place on favorable or unfavorable predictions—and when confidence, applicability and decision consequence require experimental evidence.

Evaluate Your Prediction Strategy
Opening answer

Reliability is endpoint-, domain- and decision-specific

An in silico ADMET prediction is most useful when its endpoint is precisely defined, the compound lies within the model's applicability domain, the training measurements are relevant and well curated, and the model has been tested on genuinely independent chemistry. A favorable prediction can lower the priority of an experimental test in early, reversible ranking decisions; an unfavorable prediction can flag a liability or help select compounds for confirmation. Neither result is a laboratory measurement. Confidence should therefore be judged separately from the direction of the prediction. A model can be confident and wrong for a novel scaffold, or uncertain while still correctly flagging a risky region of chemical space. The practical question is not whether "ADMET AI" is reliable in general. It is whether a particular model output is reliable enough for a particular compound, endpoint and decision. As decision consequence rises—from virtual triage to lead nomination—the required evidence should progress from model checks and consensus to endpoint-specific experiments, exposure-aware interpretation and, where relevant, in vivo or clinical evidence.

Model output

Read the output type before reading the result

Classification

Category or threshold

Examples include inhibitor/non-inhibitor, high/low permeability or alert/no alert. Review the decision threshold, class balance, sensitivity, specificity, precision and calibration. A probability near 0.5 carries a different meaning from one near a well-calibrated extreme.

Regression

Estimated continuous value

Examples include logS, intrinsic clearance, fraction unbound or IC50. Interpret the prediction with units, transformation, experimental range, prediction error and interval. Error in log units can translate into a substantial fold difference.

Score or alert

Relative priority

Similarity scores, structural alerts and ensemble rankings often support prioritization rather than direct biological interpretation. Their meaning depends on how the score was derived and whether it maps to an experimentally defined endpoint.

Seven reliability factors

What determines whether a prediction deserves weight?

Applicability domain

Similarity, leverage, local density or distance-based criteria indicate whether the query is represented by the model's training experience.

Chemical-space coverage

Global dataset size matters less than coverage near the relevant scaffold, ionization pattern, lipophilicity range, stereochemistry and modality.

Training-data quality

Correct structures, consistent units, assay metadata, replicate handling, censored values and duplicate resolution shape the learnable signal.

Endpoint performance

Performance must be reviewed for the specific endpoint, metric, validation split, class distribution and operating threshold used in the decision.

Assay heterogeneity

Different protocols, biological systems, concentrations and readouts can create endpoint labels that are nominally similar but experimentally distinct.

Uncertainty calibration

Confidence scores become decision-useful when observed error rates track predicted confidence and uncertainty increases for unfamiliar chemistry.

Endpoint-specific interpretation

The same model quality does not transfer across ADMET endpoints

Endpoint biology and assay design set different ceilings on predictability. Physicochemical properties measured under defined conditions may support relatively direct structure–property learning. Cellular, organ-level and toxicity endpoints integrate multiple mechanisms and frequently contain greater biological and protocol variability. The validation question must therefore match the endpoint and intended use.

Endpoint familyTypical outputMajor sources of variabilityUseful validation viewPractical follow-up
Solubility / lipophilicity / ionizationContinuous value or category at a stated pHSolid form, pH, kinetic versus thermodynamic protocol, salts and aggregationExternal error distribution by property range and chemical seriesMeasure under formulation- and route-relevant conditions
Permeability / absorptionApparent permeability, absorption class or efflux riskCell line, transporter expression, recovery, pH gradient and solubilitySeries-aware validation with threshold performance and recovery checksCell-based permeability/efflux assay with concentration context
Metabolic stability / clearanceHalf-life, intrinsic clearance or stability classSpecies, matrix, protein concentration, nonspecific binding and extrahepatic pathwaysSeparate validation by assay system, species and clearance rangeMicrosomal or hepatocyte stability; IVIVE when exposure translation is needed
CYP and transporter liabilitiesSubstrate/inhibitor probability, IC50/Ki estimate or alertIsoform, probe substrate, incubation design, time dependence, metabolites and exposurePer-isoform metrics, calibrated thresholds and external chemistryMechanistic in vitro testing followed by exposure-based interpretation
Toxicity endpointsAlert, probability, potency estimate or multi-label profileCell type, dose, duration, species, mechanism, class imbalance and label aggregationEndpoint-specific sensitivity, precision, calibration and scaffold/time splitOrthogonal assays selected around the predicted mechanism and expected exposure
PK parametersClearance, volume, bioavailability, half-life or exposure estimateCoupled ADME processes, species, dose, formulation and routeProspective or external compounds with parameter-wise and profile-level errorCombine measured inputs, IVIVE, animal PK or PBPK according to the question
Decision component

Prediction × Confidence × Decision Criticality Matrix

First determine whether the output is favorable or unfavorable. Then evaluate confidence from applicability domain, local support, model agreement, calibration and endpoint-specific validation. Finally, scale the action to the consequence and reversibility of the decision.

Prediction & confidence
Low-criticality decision
Reversible ranking
Medium-criticality decision
Design or series selection
High-criticality decision
Lead or candidate gate
Favorable · High confidence
Use for prioritizationAdvance within the ranked set while retaining routine quality controls.
Confirm the key endpointUse one fit-for-purpose assay if the property drives design.
Build convergent evidenceAdd repeat or orthogonal measurement and exposure context.
Favorable · Low confidence
Keep, but flagDo not discard the compound; reduce ranking weight and retain diversity.
Measure before optimizationResolve domain or data uncertainty before using the result as a design objective.
Treat as an evidence gapObtain experimental data before the gate decision.
Unfavorable · High confidence
Deprioritize selectivelyApply the alert while checking series diversity and threshold sensitivity.
Confirm mechanism and magnitudeRun concentration-response or endpoint-specific testing.
Quantify risk and marginUse orthogonal evidence, exposure comparison and mitigation options.
Unfavorable · Low confidence
Avoid hard exclusionUse as a sampling flag rather than a binary filter.
Resolve the conflictCheck structure, domain, model disagreement and assay definition, then test.
Escalate immediatelyGenerate decision-grade evidence for the liability and document uncertainty.

Key distinction: the direction of a prediction and confidence in that prediction are separate variables. Decision criticality determines how much independent evidence is required.

Chemical space

Out-of-domain compounds require a different interpretation pathway

Signals that a compound may be out of domain

  • Low similarity or high distance from training compounds
  • A new scaffold, uncommon functional group or extreme property range
  • Unrepresented charge state, stereochemical class or molecular size
  • Covalent, metal-containing, macrocyclic, peptide-like or degrader chemistry in a conventional small-molecule model
  • Strong disagreement among independently trained models

Recommended response

  • Verify structure standardization and the molecular form submitted
  • Inspect nearest neighbors and their assay provenance
  • Report the prediction together with domain status and uncertainty
  • Select representative compounds for measurement
  • Update local models only after consistent, curated experimental data become available

Chemical-space coverage should be assessed locally. A large training set can still provide little support around a particular scaffold. Random train/test splits often place close analogues on both sides of the split and can therefore overstate performance expected on future or structurally novel chemistry. Scaffold, cluster and time-based splits usually provide a more demanding view of generalization.

Data and assay quality

Models learn the experimental record—including its inconsistencies

Structure layer

Duplicates, salts, mixtures, tautomer handling, stereochemistry and incorrect structures can create contradictory examples.

Measurement layer

Units, transformations, limits of quantification, censored values, replicate variation and protocol differences affect the target.

Dataset layer

Class imbalance, narrow chemistry, series dominance, missing metadata and selective publication influence apparent performance.

Validation layer

Data leakage, analogue leakage, metric selection and threshold tuning can make a model appear stronger than its prospective use.

Assay heterogeneity is particularly important when public data are aggregated. Two records labeled "solubility," "CYP3A4 inhibition" or "hepatotoxicity" may differ in protocol, biological system, concentration, time point and endpoint definition. Curation should preserve provenance and separate incompatible measurements instead of forcing them into one target. For classification, examine class prevalence and the cost of false negatives versus false positives. For regression, inspect absolute and fold error across the measured range rather than relying on correlation alone.

Project checklist

What to request with every prediction package

  1. Endpoint definition. Biological meaning, units, transformation, assay system and threshold.
  2. Model output type. Class, probability, continuous estimate, interval, ranking score or structural alert.
  3. Training-set scope. Number of compounds, chemical diversity, endpoint distribution and assay provenance.
  4. Validation design. External, time, scaffold or cluster split and any hyperparameter-selection separation.
  5. Endpoint metrics. Metrics matched to class balance, range and intended operating threshold.
  6. Applicability domain. Domain method, nearest-neighbor support and compound-level status.
  7. Uncertainty and calibration. Prediction interval, ensemble spread or calibrated probability with observed error behavior.
  8. Conflict handling. Rules for model disagreement, outliers, structural alerts and data-quality concerns.
  9. Experimental follow-up. Assay, concentration range, controls and decision trigger for confirmation.
  10. Traceability. Model version, structure-processing workflow, input form and date of analysis.
References

Scientific and Regulatory Sources

  1. Organisation for Economic Co-operation and Development. Guidance Document on the Validation of (Quantitative) Structure–Activity Relationship [(Q)SAR] Models. OECD Series on Testing and Assessment No. 69, OECD Publishing, 2014. doi:10.1787/9789264085442-en.
  2. International Council for Harmonisation. ICH M7(R2): Assessment and Control of DNA Reactive (Mutagenic) Impurities in Pharmaceuticals to Limit Potential Carcinogenic Risk. 2023.
  3. Wu, Z.; Ramsundar, B.; Feinberg, E. N.; et al. MoleculeNet: a benchmark for molecular machine learning. Chemical Science 2018, 9, 513–530. doi:10.1039/C7SC02664A.
  4. Sheridan, R. P. Time-split cross-validation as a method for estimating the goodness of prospective prediction. Journal of Chemical Information and Modeling 2013, 53, 783–790. doi:10.1021/ci400084k.
  5. Wallach, I.; Heifets, A. Most ligand-based classification benchmarks reward memorization rather than generalization. Journal of Chemical Information and Modeling 2018, 58, 916–932. doi:10.1021/acs.jcim.7b00403.
  6. Fourches, D.; Muratov, E.; Tropsha, A. Trust, but verify: on the importance of chemical structure curation in cheminformatics and QSAR modeling research. Journal of Chemical Information and Modeling 2010, 50, 1189–1204. doi:10.1021/ci100176x.
  7. Kramer, C.; Kalliokoski, T.; Gedeck, P.; Vulpetti, A. The experimental uncertainty of heterogeneous public Ki data. Journal of Medicinal Chemistry 2012, 55, 5165–5173. doi:10.1021/jm300131x.
  8. Yang, K.; Swanson, K.; Jin, W.; et al. Analyzing learned molecular representations for property prediction. Journal of Chemical Information and Modeling 2019, 59, 3370–3388. doi:10.1021/acs.jcim.9b00237.

Online Inquiry

Submit your project details below, and our team will respond within 24 hours.

x
Need help getting the data you need?

Talk to our technical team about your project!

I Want To Talk