Study Design and Power Planning
Define contrasts, blocking, pairing, time-course terms, interaction effects, covariates, expected dispersion, detectable effect sizes, replicate needs, and validation cohorts.
Transform disease-relevant transcriptomes into prioritized, testable target hypotheses supported by expression, pathway, network, and intervention evidence.
Discuss Your RNA-Seq ProjectRNA-seq captures disease-associated expression programs, pathway activation, cell-state shifts, treatment responses, isoform usage, and regulatory changes. Our target discovery analysis connects these signals to biological mechanism and intervention feasibility so that candidates can be compared by disease relevance, direction of change, reproducibility, network position, tissue context, and supporting evidence.
We support human, mouse, rat, non-human primate, livestock, plant, microbial, and other species with suitable reference genomes and annotations. Projects may include disease versus control, responder versus non-responder, genotype, perturbation, dose, time course, factorial, paired, longitudinal, or multi-cohort designs. Common specimens include tissue, biopsy, blood, sorted cells, cultured cells, organoids, xenografts, model organisms, FFPE-derived RNA, and extracellular RNA where the assay is appropriate.
Illumina, MGI/DNBSEQ, and other short-read platforms are supported, along with transcript abundance from validated long-read workflows when relevant. We can start from FASTQ, BAM/CRAM, gene or transcript count matrices, Salmon/Kallisto abundance files, normalized matrices for exploratory work, or public accessions. Raw counts are preferred for count-based differential expression.
Define contrasts, blocking, pairing, time-course terms, interaction effects, covariates, expected dispersion, detectable effect sizes, replicate needs, and validation cohorts.
Assess read quality, adapters, contamination, duplication, strandedness, mapping, gene-body coverage, rRNA content, transcript assignment, and gene or isoform abundance using versioned references.
Use count-aware models such as DESeq2, edgeR, or limma-voom as design-appropriate; estimate contrasts, interaction terms, shrunken effect sizes, confidence measures, and FDR-adjusted significance.
Evaluate alternative splicing, differential transcript usage, gene fusions, lncRNA or other RNA classes when supported by library design, depth, read length, and annotation quality.
Apply ranked gene-set analysis, over-representation testing, co-expression modules, upstream regulator analysis, protein association networks, and tissue or cell-type context.
Integrate expression effect, consistency, pathway role, disease evidence, essentiality, tractability, modality, safety, tissue selectivity, known drugs, and competitive activity into a traceable shortlist.
| Stage | Key Activities | Decision Output |
|---|---|---|
| 1. Biological question and design | Define disease context, primary endpoints, contrasts, controls, biological replicates, pairing, batch structure, covariates, effect size, and target-selection criteria. | Analysis plan and design matrix |
| 2. Data and metadata audit | Verify FASTQ/BAM/count integrity, sample identity, platform, library type, strandedness, genome and annotation compatibility, completeness, and potential confounding. | Data-readiness and risk report |
| 3. Processing and QC | Perform read QC, trimming where justified, alignment or pseudoalignment, quantification, sample-level QC, outlier review, exploratory structure, and batch assessment. | Validated counts and QC dashboard |
| 4. Statistical analysis | Model the agreed design, estimate contrasts and effect sizes, control multiple testing, assess interaction or temporal effects, and test sensitivity to covariates and influential samples. | Robust expression signatures |
| 5. Mechanism and target ranking | Analyze pathways, modules, regulators, network position, tissue context, known disease links, tractability, modality, safety, and therapeutic precedence. | Prioritized targets with evidence cards |
| 6. Independent confirmation | Check public or client cohorts, define qPCR/protein/functional assays, select perturbation models, establish acceptance criteria, and document versions and parameters. | Validation roadmap and final report |
Read, mapping, assignment, complexity, strandedness, coverage, sample correlation, PCA, outlier, and batch summaries.
Versioned gene or transcript counts, analysis-ready transformations, annotation mapping, normalized values for visualization, and sample metadata.
Contrast-specific effect sizes, shrunken fold changes where applicable, standard errors, test statistics, raw P values, adjusted P values, and diagnostic plots.
Ranked pathways, leading-edge genes, co-expression modules, regulators, network communities, tissue context, and interpretable visualizations.
Expression direction, reproducibility, disease and pathway evidence, tractability, modality, safety signals, existing agents, and prioritization rationale.
Methods, software and database versions, reference build, parameters, scripts or notebooks as scoped, target shortlist, uncertainties, and validation plan.
Identify reproducible disease-associated genes, pathways, and network modules across tissues or cohorts.
Compare responders and non-responders to nominate targets linked to sensitivity, resistance, or relapse.
Profile genetic or pharmacological perturbations to identify downstream mechanisms and compensatory targets.
Separate early from late programs and quantify dose-dependent expression trajectories.
Connect aberrant expression, splicing, and pathway effects to candidate disease mechanisms and therapeutic opportunities.
Reprocess compatible GEO, SRA, ArrayExpress, recount, GTEx, or other public datasets for replication and context.
Reliable target discovery depends on an analysis chain that preserves the experimental design from expression quantification through statistical testing and pathway interpretation. Integrated RNA-seq workflows combine gene identifier mapping, low-expression filtering, exploratory analysis, differential expression, clustering, network analysis, and gene-set enrichment to turn expression matrices into testable biological hypotheses.1
For differential expression, DESeq2 uses shrinkage estimation for dispersions and fold changes to improve the stability and interpretability of count-based estimates.2 Ranked pathway analysis and network visualization then help reveal coordinated biological programs that may be missed when genes are considered independently.3
We use these principles to build target shortlists that retain the observed direction and magnitude of change, disease and tissue context, pathway role, supporting evidence, and recommended confirmation assays.

1 Ge, S.X.; Son, E.W.; Yao, R. iDEP: an integrated web application for differential expression and pathway analysis of RNA-Seq data. BMC Bioinformatics 2018, 19, 534. https://doi.org/10.1186/s12859-018-2486-6. Open Access, CC BY 4.0.
2 Love, M.I.; Huber, W.; Anders, S. Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2. Genome Biology 2014, 15, 550. https://doi.org/10.1186/s13059-014-0550-8. Open Access, CC BY 4.0.
3 Reimand, J. et al. Pathway enrichment analysis and visualization of omics data using g:Profiler, GSEA, Cytoscape and EnrichmentMap. Nature Protocols 2019, 14, 482–517. https://doi.org/10.1038/s41596-018-0103-9.
Replicate needs are planned from variability, effect size, design complexity, and intended validation. Three biological replicates per group may support a simple exploratory comparison, but powered discovery and heterogeneous clinical cohorts commonly require more.
Controls, pairing, batch, donor, sex, age, site, cell composition, treatment, and other measured confounders are encoded where relevant. Results include effect size and uncertainty alongside FDR control, not significance alone.
When ML is justified, donors, sites, time points, or related samples are grouped to prevent leakage. Feature selection and tuning occur within training folds, with nested or repeated cross-validation and independent or temporal validation when available. Applicability domain, calibration, rank stability, and uncertainty accompany predictions.
RNA-seq associations nominate targets for follow-up; priority candidates are advanced through an independent cohort and orthogonal assays such as RT-qPCR, protein measurement, targeted genetic perturbation, rescue experiments, and disease-relevant cellular or in vivo models.
To scope an RNA-seq target discovery project, please Contact Us or submit the Online Inquiry below.
Submit your project details below, and our team will respond within 24 hours.
Talk to our technical team about your project!
I Want To Talk