Case Study
siRNA Off-Target Prediction Service

Inquiry
siRNA Off-Target Prediction Service - CD ComputaBio

Bioinformatics Services

siRNA Off-Target Prediction Service

Transcript-aware, evidence-linked assessment of sequence-dependent siRNA liabilities—from candidate design through RNA-seq confirmation.

Discuss Your Project

Overview

Specificity is more than a short-sequence similarity search

A useful siRNA off-target assessment must separate several biological routes: unintended near-perfect pairing that could support AGO2-mediated cleavage; microRNA-like recognition driven by the guide seed; passenger-strand activity; transcript-isoform and allele-specific matches; and sequence-independent responses that may appear in expression data. CD ComputaBio integrates these routes into a traceable risk assessment that helps research and development teams compare candidates, define controls, and plan validation.

The central question is not simply whether another transcript contains similar bases. We ask which unintended sites exist in the relevant transcriptome, where mismatches and gaps occur, whether the site lies in a biologically interpretable transcript region, whether that transcript is expressed in the intended tissue or model, and whether experimental expression changes are consistent with the predicted sequence mechanism. Each result remains a computational candidate until supported by orthogonal evidence. Our Bioinformatics Services team can work from sequence-only designs, candidate panels, or completed perturbation studies.

Project Inputs

Data and metadata required for a decision-ready analysis

Input classAccepted materialWhy it matters
siRNA designGuide and passenger sequences in CSV, XLSX, FASTA, or plain text; duplex orientation, length, overhangs, chemical modifications, and candidate identifiers.Correct strand definition and modification-aware interpretation prevent seed reversal and coordinate errors.
Target definitionGene symbol, stable gene/transcript accession, target sequence, intended isoform, species and genome/annotation build.Transcript versions change target coverage, UTR boundaries, paralog matches, and variant coordinates.
Biological contextHuman, mouse, rat, non-human primate, livestock, plant, pathogen, or another species with an adequate reference; tissue, cell type, disease state, delivery compartment, dose, and time point.Expression context and transcript availability affect which sequence matches deserve priority.
Expression evidenceBulk RNA-seq FASTQ or count matrices, microarray matrices, targeted expression data, public tissue atlases, or user-provided single-cell summaries. Illumina short-read data are common; platform, library layout and strandedness are required.Expression filters reduce irrelevant hits and enable empirical testing of seed-associated shifts.
Study metadataReplicate identity, negative and positive controls, dose, time, batch, sex, genotype, treatment history, sample exclusions and pairing structure.These variables define the statistical model and reveal batch–treatment confounding.

For organisms without a mature reference annotation, we can use a client-supplied transcriptome or construct a project-specific reference, but completeness must be reported as an applicability limitation. RNA-seq is not mandatory for sequence screening; it becomes especially valuable when the goal is to test whether predicted seed-bearing transcripts show a coordinated response.

Core Services

Mechanism-aware analysis modules

Transcriptome-wide alignment

Search guide and passenger strands against current transcript sequences using short-read-aware alignment or exact/approximate matching. We enumerate perfect and near-perfect sites, retain mismatch positions and indels where appropriate, distinguish coding sequence and UTR context, and map transcript hits back to genes without discarding isoform detail.

Seed-mediated risk profiling

Scan 3′ UTRs and other requested regions for canonical and project-defined guide or passenger seed matches. Site multiplicity, seed class, local sequence context and seed-duplex thermodynamics can be combined rather than reduced to one arbitrary homology cutoff.

Variant and population review

When relevant, common variants or client genotypes are intersected with target and off-target sites. We flag alleles that create, remove, or alter pairing and keep reference-build provenance. Population frequency informs risk stratification but does not prove an in vivo effect.

Expression-context prioritization

Candidate sites are annotated with tissue, cell-model, or study-specific expression. A strong match in a highly expressed, safety-relevant transcript may outrank numerous weak matches in genes not detected under the intended exposure context.

RNA-seq off-target diagnosis

For perturbation studies, we perform sample-level quality control, count modeling, effect-size estimation and seed-match enrichment analysis. We compare expression distributions for genes with and without predicted sites while accounting for the experimental design.

Candidate comparison and redesign

We build transparent multi-criteria rankings across on-target coverage, high-risk alignments, seed burden, passenger liability, expression relevance and evidence uncertainty. Alternative candidates are proposed only within stated design constraints and remain subject to experimental testing.

Six-Stage Workflow

From sequence audit to validation plan

Every stage records software, reference files, parameters and exclusions so that a future transcriptome release or candidate redesign can be compared reproducibly.

Six-stage siRNA off-target prediction workflow from sequence audit to candidate validation
Integrated workflow linking siRNA sequence review, transcript references, candidate-site searching, biological filtering, statistical assessment, and validation prioritization.

1. Scope and sequence audit

Confirm species, intended target isoform, strand orientation, chemistry, model system, development question and acceptable risk criteria. Normalize bases and identifiers, and resolve ambiguous or conflicting sequences before searching.

2. Reference construction

Select genome and transcript annotation releases; retain transcript isoforms, UTRs, coding regions and noncoding RNAs as appropriate. Add pathogen, vector, custom construct, or allele sequences when exposure makes them relevant.

3. Candidate-site enumeration

Search full guide and passenger sequences for near-perfect complementarity and scan defined seed windows. Record mismatch number, position and type, alignment span, transcript feature, site count and competing annotations.

4. Biological annotation

Integrate transcript abundance, tissue relevance, gene function, paralogy, essentiality or safety context, and optional variation. Missing expression is treated as uncertainty, not evidence of absence.

5. Statistical evidence integration

If expression data are supplied, run quality control and a prespecified contrast model, estimate fold changes and confidence intervals, adjust multiple tests, and evaluate enrichment or cumulative shifts among predicted target groups.

6. Ranking and validation design

Deliver candidate-level and site-level risk tiers with reasons, sensitivity analyses and proposed experiments. Conclusions distinguish sequence plausibility, expression association and experimentally demonstrated activity.

Scientific Evidence

Seed matches can create a transcriptome-wide signature

MicroRNA-like off-target regulation is often driven by guide positions near 2–8, yet shared seed complementarity does not imply equal repression. Experimental studies show that non-seed sequence features and target context can modulate the response. The displayed open-access study compared expression profiles of transcripts with perfect 3′ UTR seed matches against genes without those matches, then evaluated auxiliary sequence features. This supports a layered model rather than a binary BLAST-style rule.

In a project with RNA-seq, we can test whether seed-bearing genes tend to shift downward as a group and whether the pattern persists across doses, times, or independent siRNAs. Such a distributional association strengthens a sequence-mediated hypothesis but cannot by itself identify direct binding or causality for an individual gene.

Expression distributions and sequence-feature analyses for siRNA seed-matched off-target transcripts
Expression distributions of seed-matched transcripts and auxiliary non-seed sequence features associated with siRNA off-target effects.1

Statistical and AI Strategy

Controls, uncertainty, and generalization are designed into the analysis

Design adequacy

Sequence-only screening needs no biological sample size, but empirical confirmation does. RNA-seq studies should include independent biological replicates, matched negative controls and, where feasible, multiple siRNAs targeting the same gene. Technical replicates do not replace biological replication. Power depends on dispersion, effect size and multiplicity; we assess feasibility from pilot or comparable data rather than promise a universal minimum.

Confounding control

Batch, dose, time, sex, donor, cell passage and library preparation can mimic or mask an off-target signature. Covariates and paired structures enter the model when identifiable. If treatment is completely confounded with batch, computation cannot recover the missing comparison; the result is labeled exploratory and redesign is recommended.

Multiplicity and magnitude

Transcript-level testing uses false-discovery-rate control, with adjusted values reported beside effect sizes, uncertainty intervals and expression abundance. Seed enrichment across candidate definitions also creates multiple hypotheses, so the tested windows and contrasts are documented. Statistical significance alone does not establish biological importance.

Responsible prediction

Machine-learning scores may complement mechanistic features when an appropriate model and training domain are documented. Candidate-related sequences, duplicate experiments and measurements from the same study must remain in the same fold to prevent leakage. Feature selection, scaling and tuning occur inside training folds; nested cross-validation is preferred for honest internal estimates.

External validation should use sequences, laboratories, tissues or chemistries not represented during training. We report the model’s applicability domain—species, oligo length, chemistry, assay, target class and sequence space—and flag extrapolation. Calibrated probabilities or score intervals are preferred where supported; disagreement among alignment, seed, thermodynamic and model-based signals is retained as uncertainty rather than hidden in a single composite score.

Deliverables

Six outputs built for candidate decisions

1. Sequence and reference audit

Validated guide/passenger table, target transcript coverage, species/build information, reference checksums and unresolved assumptions.

2. Off-target site atlas

Searchable CSV/XLSX with gene, transcript, region, strand, alignment, mismatch positions, seed class, site multiplicity and annotation provenance.

3. Candidate risk dashboard

Readable candidate comparison with transparent tiering, major risk drivers, passenger-strand findings and sensitivity to alternative thresholds.

4. Expression evidence report

RNA-seq QC, model specification, contrasts, effect sizes, adjusted P values, seed-associated distribution tests and diagnostic plots when data are provided.

5. Biological interpretation package

Expression-context and functional annotations for prioritized genes, clearly separating predicted binding, observed association and confirmed evidence.

6. Validation and redesign plan

Recommended orthogonal assays, controls, concentration ranges, alternative siRNAs and decision gates, plus reproducible methods and parameter records.

Applications and Validation

Appropriate use across research and preclinical programs

Research reagent selection

Compare siRNAs before purchase or screening, identify likely paralog and seed liabilities, and select non-overlapping reagents so that a phenotype reproduced by independent sequences is more interpretable.

Therapeutic lead prioritization

Rank candidate duplexes in a tissue-relevant transcriptome, flag safety-relevant transcripts and common variants, and define which computational risks require dose-ranging or tissue-specific experiments.

Unexpected phenotype diagnosis

Use transcriptome data and sequence features to assess whether an observed response is compatible with on-target biology, a shared seed signature, passenger activity, innate immune activation, delivery toxicity or experimental confounding.

High-priority candidates should be tested with concentration–response experiments and at least one independently designed siRNA. Recommended orthogonal methods may include RT-qPCR or digital PCR for selected transcripts, protein-level assays, reporter constructs containing the predicted site, AGO2 immunoprecipitation or CLIP-based evidence, and rescue experiments with an siRNA-resistant on-target construct. RNA-seq in an independent experiment or relevant cell type can evaluate reproducibility. A mismatch control can probe sequence dependence, while chemistry-matched non-targeting controls help distinguish oligo and delivery effects. These experiments turn a ranked computational hypothesis into evidence; they are not interchangeable, and the final design depends on the proposed mechanism.

Project Strategy and Boundaries

Match the analysis depth to the development decision

For early candidate triage, a reference transcriptome, sequences and intended tissue may be enough to produce a specificity comparison. For lead selection, we recommend explicit passenger-strand analysis, transcript isoforms, population variants, tissue expression and a predefined risk rubric. For post-treatment diagnosis, raw counts or FASTQ files, complete sample metadata, suitable controls and sufficient biological replication are needed. Public expression atlases can inform context, but they do not substitute for the exposed model because tissue composition, disease state, dose and time may differ.

Predicted sites indicate sequence compatibility, not confirmed engagement. Absence of a predicted site does not exclude immune stimulation, saturation of RNAi machinery, delivery-related effects, unannotated transcripts, RNA editing, structural accessibility effects, or interactions unique to a chemical modification. Expression changes can be downstream consequences of intended target knockdown rather than direct off-target events. We therefore use calibrated language—candidate, predicted, associated, or consistent with—and keep direct binding, functional consequence and causality as separate evidence levels.

Before work begins, we agree on the decision threshold: eliminating candidates with close paralog matches, selecting the lowest relative seed burden, explaining an RNA-seq phenotype, or preparing a focused validation panel. This prevents retrospective threshold changes and makes the report actionable. Contact us to discuss species coverage, reference quality, candidate numbers and available experimental data.

References

Methods and evidence base

  1. Kamola, P. J.; Nakano, Y.; Takahashi, T.; Wilson, P. A.; Ui-Tei, K. The siRNA Non-seed Region and Its Target Sequences Are Auxiliary Determinants of Off-Target Effects. PLOS Computational Biology 2015, 11, e1004656. https://doi.org/10.1371/journal.pcbi.1004656. Distributed under the Creative Commons Attribution License (CC BY).
  2. Cazares, T.; et al. SeedMatchR: identify off-target effects mediated by siRNA seed regions in RNA-seq data. Bioinformatics 2024, 40, btae011. https://doi.org/10.1093/bioinformatics/btae011. Distributed under CC0.
  3. Naito, Y.; Yoshimura, J.; Morishita, S.; Ui-Tei, K. siDirect 2.0: updated software for designing functional siRNA with reduced seed-dependent off-target effect. BMC Bioinformatics 2009, 10, 392. https://doi.org/10.1186/1471-2105-10-392. Open Access.
  4. Das, S.; Ghosal, S.; Sen, R.; Chakrabarti, J. lnCeDB: database of human long noncoding RNA acting as competing endogenous RNA. Seed-related transcript resources and sequence-context methods inform transcriptome annotation; reference selection is project-specific.

Build a defensible siRNA specificity strategy

Share your candidate sequences, species, target isoform, biological model and available expression data.

Contact Us

Online Inquiry

Submit your project details below, and our team will respond within 24 hours.

x
Need help getting the data you need?

Talk to our technical team about your project!

I Want To Talk