Case Study
Single-Cell RNA-Seq Data Analysis Service

Inquiry
Single-Cell RNA-Seq Data Analysis Service - CD ComputaBio
Bioinformatics Services

Single-Cell RNA-Seq Data Analysis Service

From raw barcodes to replicate-aware cellular programs—an analysis framework built around your biological question.

Start Your Project

Overview

Resolve cellular heterogeneity without losing the experimental unit

Single-cell RNA sequencing can reveal cell types, transient states and condition-associated programs that are obscured in bulk tissue. Reliable interpretation, however, depends on separating biological heterogeneity from empty droplets, damaged cells, ambient RNA, doublets, library depth, donor differences and processing batches. CD ComputaBio develops a question-led workflow from raw reads or count matrices through cell-level quality control, annotation and replicate-aware inference.

We analyze single cells or nuclei from human, mouse, rat, non-human primate and other organisms with suitable references. Supported study types include fresh or frozen tissue, cell lines, organoids, tumors, blood and immune samples, perturbation screens and longitudinal or multi-condition cohorts. Results are reported as computational cell identities, states and associations. A cluster is not a new cell type merely because it is separated on a UMAP, and a predicted ligand–receptor pair is not proof of physical signaling. High-value findings receive independent-cohort, spatial, protein-level or functional validation recommendations. Projects can also connect to broader Bioinformatics Services.

Study unitDonor, animal, organoid, culture or other independent biological replicate.
Assay unitLibrary, lane, capture batch, multiplex pool and sample barcode.
Cell unitBarcode, cell or nucleus, QC metrics and assigned identity.
ComparisonCondition, tissue, time, genotype, dose and prespecified contrast.
DecisionPopulation change, cellular program, candidate marker or validation target.

Data Inputs

Raw data, processed matrices and metadata

Accepted platforms and formats

We accept FASTQ files, BAM/CRAM where compatible, filtered and raw feature–barcode matrices, H5/H5AD, Loom, Seurat objects, SingleCellExperiment objects and tabular count matrices. Common platforms include 10x Genomics Chromium 3′ or 5′ gene expression, droplet methods such as Drop-seq, plate-based Smart-seq/Smart-seq2 and single-nucleus RNA-seq. Feature Barcode or CITE-seq layers, hashtag oligos and V(D)J data can be incorporated when supplied and within project scope.

Raw processing is matched to the assay chemistry and reference. For 10x data, barcode processing, alignment and UMI counting may use a verified Cell Ranger workflow. Full-length plate data require a different alignment and quantification strategy. The reference genome, transcript annotation, chemistry, read layout, strandedness and intronic-read policy are recorded.

Required metadata

Sample and donor IDs, biological group, batch, tissue, dissociation or nuclei method, library chemistry, lanes, expected cells, sequencing run, sex, genotype, treatment, dose, time, pairing, exclusions and known technical events.

Design review

We inspect whether treatment is confounded with batch, whether donors contribute to all relevant contrasts, and whether replication supports the intended inference. More cells from one donor do not replace independent donors.

Optional evidence

Flow cytometry, histology, sorted-cell references, public atlases, bulk RNA-seq, protein markers, genotype data, CRISPR guide assignments or spatial datasets can strengthen annotation and validation.

Privacy and transfer

Human data should be de-identified before transfer. Access, retention and permitted secondary use are defined for the project; certifications or regulatory status are not assumed.

Six-stage single-cell RNA-seq analysis workflow
Six-stage workflow from barcode-aware quantification and cell-level QC to annotated populations and replicate-aware biological interpretation.

Six-Stage Workflow

A vertical decision path from reads to biology

1. Data and reference audit

Confirm platform, chemistry, read structure, sample sheet, reference build, expected cell recovery and biological contrasts.

2. Quantification and cell calling

Process barcodes and UMIs, align or map reads, generate raw and filtered matrices, and review per-library sequencing and saturation metrics.

3. Cell-level quality control

Evaluate detected genes, counts, mitochondrial or ribosomal fractions, complexity and sample-specific outliers; model empty droplets, ambient RNA and doublets where justified.

4. Normalization and integration

Normalize expression, select features, reduce dimensions and integrate batches only when required. Preserve condition-linked biology and document sensitivity to alternative settings.

5. Annotation and state analysis

Combine canonical markers, differential signatures, reference mapping and tissue context. Resolve broad lineages before fine states and record ambiguous or mixed populations.

6. Replicate-aware inference

Test abundance or expression changes using donors or samples as the experimental unit, then perform pathway, trajectory or communication analyses aligned with the study design.

Quality Decisions

QC thresholds are learned from each sample, not copied blindly

RiskEvidence reviewedDecision ruleResidual limitation
Low-quality cells or nucleiCounts, genes, complexity, mitochondrial fraction, intronic content and sample distributions.Use multivariate, sample-aware outlier detection plus biological review; retain thresholds and removal counts.Stressed or metabolically active populations can resemble poor-quality cells.
Empty droplets and ambient RNABarcode rank, empty-droplet profiles and contamination signatures.Evaluate cell calling and contamination correction when ambient transcripts materially affect interpretation.Correction cannot reconstruct RNA never captured from a cell.
Doublets and multipletsExpected loading rate, simulated-neighbor scores, unusually high complexity and incompatible markers.Combine algorithmic score and cluster context; do not remove plausible transitional states solely for co-expression.Homotypic doublets are harder to identify.
Batch effectsLibrary, donor, preparation day, chemistry, sequencing and cluster composition.Integrate for visualization or shared-state alignment only when justified; perform inference on appropriate uncorrected/count representations.No algorithm can separate treatment from a perfectly confounded batch.
Rare populationsCell count per sample, marker coherence, QC, reproducibility and reference support.Require presence across replicates or independent evidence before strong claims.A few cells from one library provide weak frequency and state estimates.

Analysis Modules

Modular depth for different biological questions

Cell Type and State Annotation

Hierarchical annotation integrates marker panels, cluster signatures, reference atlases and tissue context. We distinguish lineage identity from activation, cycling, stress, interferon response and other cross-lineage states. Automated labels are reviewed rather than accepted as ground truth, and uncertainty is retained for poorly resolved populations.

Composition Analysis

Compare cell-type or state proportions across samples using models appropriate for compositional data and unequal recovery. Changes in captured proportions are associations and may reflect dissociation or sampling as well as biology.

Differential State Analysis

Within each cell type, aggregate counts by biological replicate for pseudobulk testing or use justified mixed models. Report effect sizes, confidence intervals, adjusted P values and the number of contributing samples and cells.

Trajectory and Velocity

Infer candidate continua, branch structure or directionality when sampling and biology support the assumptions. Pseudotime is a relative ordering, not measured chronological time, and requires orthogonal lineage or time-course validation.

Cell–Cell Communication

Prioritize expressed ligand–receptor hypotheses with source/receiver context and condition comparisons. Co-expression does not demonstrate secretion, receptor engagement or causal signaling.

Scientific Evidence

Cells are observations; biological replicates support inference

Single-cell datasets contain thousands of cells, but cells from one donor are not independent replacements for donor replication. The displayed open-access benchmark illustrates pseudobulk construction and compares differential-expression methods against ground-truth datasets. Aggregating counts within cell type and biological replicate allows established count models to represent between-replicate variability and helps prevent inflated discoveries caused by treating every cell as an independent sample.

We therefore distinguish marker discovery within a dataset from condition-level inference. Cluster markers may be calculated at cell level for descriptive annotation, while treatment comparisons use replicate-aware models whenever replication permits. For one sample per group, results are explicitly exploratory; adding more cells increases resolution but does not create missing biological replication.

Pseudobulk construction and differential expression benchmark in single-cell data
Pseudobulk aggregation by biological replicate and benchmark evidence on differential-expression behavior across single-cell datasets.1

Statistical and AI Strategy

Inference, validation and generalization

Sample size

Power depends primarily on independent samples, cell-type abundance, dispersion and target effect size. We assess whether each contrast has enough contributing replicates and cells; universal minimums are not promised.

Confounding

Donor, batch, tissue quality, sex, genotype, medication and processing enter the design when identifiable. Paired and longitudinal structures are preserved. Unidentifiable confounding is reported.

Multiplicity

Gene, pathway, population and interaction testing can create large hypothesis sets. We control false discovery rates within defined families and report effect magnitude beside significance.

Machine learning

Cell classifiers or latent models are trained with donor-aware splits. Normalization, feature selection and tuning remain inside training folds to avoid leakage; external cohorts test transportability.

The applicability domain includes species, tissue, disease state, platform, chemistry, annotation granularity and reference composition. A classifier trained on healthy blood may not label diseased tissue or tumors reliably. Cross-validation by randomly splitting cells can be misleading because cells from the same donor appear in both folds; donor-held-out or study-held-out validation is preferred. We report uncertain or out-of-reference cells rather than forcing every cell into a known label.

Six Deliverables

Outputs organized from provenance to decisions

01

Data audit

Sample sheet, design risks, reference and pipeline plan.

02

QC report

Library/cell metrics, thresholds, exclusions and retained counts.

03

Analysis object

Annotated H5AD, Seurat or agreed matrix and metadata.

04

Cell atlas

Clusters, labels, markers, uncertainty and composition.

05

Statistics

Replicate-aware contrasts, effect sizes and adjusted results.

06

Interpretation

Figures, candidate programs, limitations and validation plan.

Applications and Validation

From cell maps to testable hypotheses

Tumor and Immune Microenvironments

Resolve malignant, stromal and immune compartments; compare activation or exhaustion programs; and nominate interaction hypotheses for protein-level or spatial follow-up.

Development and Regeneration

Characterize changing states and candidate transitions across time or anatomy, with lineage tracing, time-series sampling or spatial localization recommended for directional claims.

Perturbation and Drug Response

Identify cell-type-specific response programs, resistant states and compositional shifts. Independent perturbations, dose response and orthogonal molecular assays strengthen causal interpretation.

Disease Cohorts and Biomarker Research

Compare cellular programs across donors while controlling clinical covariates. Candidate signatures require independent cohorts and fit-for-purpose assays before diagnostic or prognostic claims.

Validation options include flow cytometry or mass cytometry for population frequency, immunofluorescence or spatial transcriptomics for localization, RT-qPCR and protein assays for selected markers, targeted sequencing for rare states, and functional perturbation for candidate regulators. Replication in an independent cohort is preferred for high-value signatures. Reference mapping and concordance with public atlases provide supporting evidence but do not replace validation in the intended biological system.

Project Strategy and Boundaries

Match analysis scope to the claim

A basic atlas project may end with QC, clustering and carefully reviewed annotation. Comparative studies require sufficient biological replication and a contrast plan. Trajectory, communication, regulatory-network or machine-learning modules are added only when sampling, data quality and the research question support their assumptions. Integration is not automatically beneficial: aggressive correction can erase genuine condition-specific states, while inadequate correction can create technical clusters. We compare diagnostics and retain interpretable alternatives.

Single-cell RNA-seq measures captured RNA, not complete cellular function. Dropouts, dissociation bias, nuclei-versus-cell differences and finite sequencing depth shape the observed matrix. Cluster separation, correlations, enriched pathways and inferred networks are hypothesis-generating. We do not present computational cell states as clinically validated biomarkers or inferred regulators as proven therapeutic targets. Each major conclusion is paired with the evidence level and a practical next experiment.

References

Methods and evidence base

  1. Squair, J. W.; Gautier, M.; Kathe, C.; et al. Confronting false discoveries in single-cell differential expression. Nature Communications 2021, 12, 5692. https://doi.org/10.1038/s41467-021-25960-2. Distributed under the Creative Commons Attribution License (CC BY 4.0).
  2. Luecken, M. D.; Theis, F. J. Current best practices in single-cell RNA-seq analysis: a tutorial. Molecular Systems Biology 2019, 15, e8746. https://doi.org/10.15252/msb.20188746. Distributed under CC BY 4.0.
  3. Heumos, L.; Schaar, A. C.; Lance, C.; et al. Best practices for single-cell analysis across modalities. Nature Reviews Genetics 2023, 24, 550–572. https://doi.org/10.1038/s41576-023-00586-w. Open Access.

Turn single-cell data into defensible biological evidence

Share your platform, species, sample design, biological question and current data format.

Contact Us

Online Inquiry

Submit your project details below, and our team will respond within 24 hours.

x
Need help getting the data you need?

Talk to our technical team about your project!

I Want To Talk