Case Study
Single-Cell Biomarker Discovery for Patient Stratification

Inquiry
Single-Cell Biomarker Discovery for Patient Stratification - CD ComputaBio
The scale mismatch
Single-cell resolution expands biological detail—but the clinical decision still belongs to the patient.

Thousands of cells can reveal a rare disease-associated state, yet they do not create thousands of independent clinical observations. A biomarker program succeeds only when cellular evidence can be encoded once per patient, tested without information leakage and reproduced in a cohort that did not participate in discovery.

CellsReveal heterogeneity, rare states and within-sample biology.
SamplesDefine technical and anatomical replicates.
PatientsProvide independent units for clinical association.
CohortsTest portability across settings and populations.
Design rule: feature discovery may occur at cellular resolution, but model training, data splitting, confidence intervals and validation must respect patient identity.
Core component

Roadmap: from cell-level discovery to patient-level validation

The route is deliberately staged. Each stop changes the unit of analysis, reduces flexibility and raises the evidence standard. A candidate should not move forward merely because it is statistically significant in the discovery atlas.

Define use

Clinical stratification question

Specify the intended population, specimen, decision, endpoint, timing and comparator. Distinguish diagnosis, prognosis and treatment-response prediction before selecting features.

Gate: clinically meaningful contrast
Discover cells

Cell state and composition signals

Identify reproducible cell populations, state programs, differential abundance, pathway activity or communication features using donor-aware comparisons.

Gate: effect across patients
Encode patients

Patient-level feature construction

Convert cellular evidence into one feature vector per patient: pseudobulk expression, cell-state proportions, module scores, diversity measures or hybrid signatures.

Gate: defined calculation
Reduce and lock

Parsimonious signature selection

Remove unstable and redundant features, establish missing-data rules, select an assay-compatible panel and freeze preprocessing, coefficients and thresholds.

Gate: specification locked
Challenge internally

Patient-separated validation

Use nested cross-validation or a held-out set with all cells and visits from each patient confined to one partition. Quantify uncertainty and calibration.

Gate: no leakage
Transfer externally

Independent cohort and assay validation

Apply the locked signature once in an external population, then test whether it survives platform, specimen and workflow changes needed for practical deployment.

Gate: reproducible utility
Patient encoding

What can become a stratification feature?

A marker gene is only one option. Single-cell data support features that preserve cell identity, abundance and biological state while remaining calculable for every patient.

Cell abundance
Cell-level discoveryDisease-associated expansion, depletion or continuous neighborhood shift.Patient-level encodingComposition-aware proportion, log-ratio or model-derived abundance.
Cell-state expression
Cell-level discoveryGenes or programs altered within a defined lineage or state.Patient-level encodingCell-type pseudobulk expression or a compact module score.
Functional activity
Cell-level discoveryPathway, regulon, exhaustion, activation or metabolic state.Patient-level encodingRobust aggregate score within the relevant population.
Heterogeneity
Cell-level discoveryCoexisting states, clonal diversity or transitional continua.Patient-level encodingEntropy, dispersion, state mixture or resistant-cell burden.
Multicellular context
Cell-level discoveryCo-occurring cell states or mechanism-linked niches.Patient-level encodingRatio, interaction module or composite microenvironment score.
Prefer biology-preserving compression: choose the simplest patient-level representation that retains the relevant cell context and can be recalculated in validation data. Complex embeddings may be useful, but their reference, normalization and transfer behavior must be explicitly frozen.
Lock before testing

Discovery flexibility must end before validation begins

Pre-lock decisions

Specification document

  1. Eligible patients, samples and cells
  2. QC and annotation procedure
  3. Feature formulas and transformations
  4. Missing-feature and low-cell-count rules
  5. Model coefficients and cut point
  6. Primary endpoint and performance metrics
  7. Subgroup and sensitivity analyses
Why it matters

A validation set is not a second discovery set

If genes, cell states, thresholds or preprocessing are retuned after viewing external outcomes, the analysis remains exploratory. Report that honestly and reserve a new cohort for confirmation.

Locking also exposes practical gaps early: a feature may require a cell population absent from many biopsies, depend on an unstable annotation boundary or be impossible to reproduce on the intended assay platform.

Evidence architecture

Different datasets answer different validation questions

Evidence stagePrimary purposePermitted flexibilityRequired separationDecision
Discovery cohortFind cell states, candidate features and plausible mechanisms.Broad exploration with multiplicity control and transparent provenance.Patient-aware estimation; avoid cell-level pseudoreplication.Nominate
Internal resamplingEstimate optimism and select model complexity.Feature selection may occur only inside each training fold.All cells, samples and visits from one patient stay in one fold.Tune
Internal holdoutEvaluate the final pipeline in the discovery setting.No outcome-driven changes after evaluation begins.Untouched patients; preferably temporally or institutionally separated.Check
External cohortTest transportability to a different population or workflow.Only prespecified technical mapping; no coefficient or threshold refit.Independent enrollment, outcomes and data production.Validate
Assay transferShow the signature can be measured in the intended specimen and platform.Analytical bridging must be predefined and documented.Independent runs, operators, lots and relevant specimen variability.Translate
Cell-level splitCells from one patient appear in both training and test sets.
Global feature selectionOutcome-associated genes are selected before cross-validation.
Threshold huntingThe best external cut point is reported as validation.
Batch–outcome couplingResponse groups were processed in different runs or sites.
Incomplete-case biasPatients lacking a rare cell population are silently removed.
Stratification performance

AUROC alone is not enough

A useful classifier must distinguish groups, assign credible risks and add information beyond standard clinical variables. Performance should be reported with confidence intervals and with the intended prevalence in mind.

Discrimination

AUROC, AUPRC, sensitivity and specificity at a locked threshold.

Calibration

Agreement between predicted probability and observed outcome, including calibration slope and intercept.

Robustness

Performance across centers, batches, demographic groups, sample quality and clinically relevant subgroups.

Incremental value

Improvement over baseline clinical factors, established biomarkers or a simpler model.

For patient stratification, a smaller reproducible signature with a clear measurement route may be more valuable than a high-dimensional model that cannot be locked, interpreted or transferred.
Open-access evidence

Cellular signatures can refine patient groups beyond bulk subtypes

Khaliq and colleagues used single-cell analysis to characterize colorectal cancer cell states, then evaluated cell-specific signatures in two independent bulk transcriptomic cohorts. Their analysis linked cancer-associated fibroblast and C1Q-positive tumor-associated macrophage enrichment to disease-free survival and further separated patients within established consensus molecular subtypes.

This example illustrates an important translation route: use single-cell data to define a biologically specific cellular signature, encode it in larger patient cohorts, adjust for relevant clinical factors and test whether it adds stratification beyond an existing classification. It does not make single-cell discovery automatically clinical; rather, it shows how independent cohort evidence can challenge and refine the proposed biomarker.

Survival analysis of cell-specific signatures in two independent colorectal cancer cohorts
Figure 7, Khaliq et al., 2022. Cell-type signature associations and disease-free survival analyses in two independent bulk datasets. Source: Genome Biology 23, 113. Reproduced under CC BY 4.0.
Decision-ready package

What the project should deliver

Every candidate should remain traceable from its cellular origin to its patient-level calculation and validation evidence.

01
Cellular discovery atlasRelevant states, abundance changes, expression programs and donor-level evidence.
02
Patient feature dictionaryExact formulas, reference populations, transformations and missing-data rules.
03
Locked signature specificationFeatures, coefficients, preprocessing, cut point and software version.
04
Validation reportPatient-separated resampling, external performance, calibration and subgroup stability.
05
Mechanistic interpretationCell-state origin, pathways, alternative explanations and orthogonal evidence.
06
Assay translation planTarget specimen, platform, analytes, controls and bridging study requirements.

Minimum project inputs

Provide the clinical use case, patient and sample identifiers, outcome definitions, collection time points, covariates, batch metadata, analysis-ready counts or object, and any intended validation cohort or assay platform. Early alignment prevents a biologically interesting signature from becoming impossible to test.

Start a Biomarker Discovery Discussion
References

Selected scientific references

  1. Khaliq AM, et al. Refining colorectal cancer classification and clinical stratification through a single-cell atlas. Genome Biol. 2022;23:113. doi:10.1186/s13059-022-02677-z.
  2. Pinhasi A, Yizhak K. Uncovering gene and cellular signatures of immune checkpoint response via machine learning and single-cell RNA-seq. npj Precis Oncol. 2025;9:95. doi:10.1038/s41698-025-00883-z.
  3. Gambardella G, et al. A single-cell analysis of breast cancer cell lines to study tumour heterogeneity and drug response. Nat Commun. 2022;13:1714. doi:10.1038/s41467-022-29358-6.
  4. Crowell HL, et al. muscat detects subpopulation-specific state transitions from multi-sample multi-condition single-cell transcriptomics data. Nat Commun. 2020;11:6077. doi:10.1038/s41467-020-19894-4.
  5. Büttner M, et al. scCODA is a Bayesian model for compositional single-cell data analysis. Nat Commun. 2021;12:6876. doi:10.1038/s41467-021-27150-6.
  6. Dann E, et al. Differential abundance testing on single-cell data using k-nearest neighbor graphs. Nat Biotechnol. 2022;40:245–253. doi:10.1038/s41587-021-01033-z.
  7. Sade-Feldman M, et al. Defining T cell states associated with response to checkpoint immunotherapy in melanoma. Cell. 2018;175:998–1013.e20. doi:10.1016/j.cell.2018.10.038.
  8. Reshef YA, et al. Co-varying neighborhood analysis identifies cell populations associated with phenotypes of interest from single-cell transcriptomics. Nat Biotechnol. 2022;40:355–363. doi:10.1038/s41587-021-01066-4.

Online Inquiry

Submit your project details below, and our team will respond within 24 hours.

x
Need help getting the data you need?

Talk to our technical team about your project!

I Want To Talk