Differential expression in RNA-seq: a matter of depth.
Next-generation sequencing (NGS) technologies are revolutionizing genome research, and in particular, their application to transcriptomics (RNA-seq) is increasingly being used for gene expression profiling as a replacement for microarrays. However, the properties of RNA-seq data have not been yet fully established, and additional research is needed for understanding how these data respond to differential expression analysis. In this work, we set out to gain insights into the characteristics of RNA-seq data analysis by studying an important parameter of this technology: the sequencing depth. We have analyzed how sequencing depth affects the detection of transcripts and their identification as differentially expressed, looking at aspects such as transcript biotype, length, expression level, and fold-change. We have evaluated different algorithms available for the analysis of RNA-seq and proposed a novel approach--NOISeq--that differs from existing methods in that it is data-adaptive and nonparametric. Our results reveal that most existing methodologies suffer from a strong dependency on sequencing depth for their differential expression calls and that this results in a considerable number of false positives that increases as the number of reads grows. In contrast, our proposed method models the noise distribution from the actual data, can therefore better adapt to the size of the data set, and is more effective in controlling the rate of false discoveries. This work discusses the true potential of RNA-seq for studying regulation at low expression ranges, the noise within RNA-seq data, and the issue of replication.
- Research Article
175
- 10.14806/ej.17.b.265
- Feb 28, 2012
- EMBnet.journal
http://bioinfo.cipf.es/aconesaIntroductionNext Generation Sequencing (NGS) technologies have brought a revolution to research in genome and genome regulation. One of the most breaking applications of NGS is in transcriptome analysis. RNA-seq has revealed exciting new data on gene models, alternative splicing and extra-genic expression. Also RNA-seq permits the quantification of gene expression across a large dynamic range and with more reproducibility than microarrays. Several methods for the assessment of differential expression from count data have been proposed but biases associated to transcript length and transcript frequency distributions have been reported. It is still not clear how much sequencing reads should be generated in a RNA-seq experiment to obtain reliable results and what’s exactly being detected. In general we observed that many RNA-seq datasets have not reached saturation for detection of expressed genes and that the relative proportion of different transcript biotypes changes with increasing sequencing depth. In this work we investigate the effect that library size has on the assessment of differential expression on different aspects of the selected genes. We show that current statistical methods suffer from a strong dependency of their significant calls on the number of mapped reads considered and proposed a novel differential expression methodology – NOISeq1- that is robust to the amount of reads.ResultsNOISeq is a non-parametric approach for the differential expression analysis of RNseq-data. NOISeq creates a null or noise distribution of count changes by comparing the number of reads of each gene in samples within the same condition. This reference distribution is then used to assess whether the change in count number between two conditions for a given gene is likely to be part of the noise or represents a true differential expression. Two variants of the method are implemented: NOISeq-real uses replicates, when available, to compute the noise distribution and, NOISeq-sim simulates them in absence of replication. We compared our method with edgeR2, DESeq3, baySeq4 and Fisher Exact Test (FET) using three different experimental datasets. Results are presented for MAQC experiment where the transcriptome of brain and Universal Human Reference (HUR) samples were sequenced at about 45 million Solexa reads each.We first determined that although protein-coding gene is the most abundant transcript type within differential expression calls for all methodologies, other RNA types, such as processed-transcript, pseudogenes and lincRNAs are readily detected. NOISeq dected comparatively more protein-coding genes than other methods that called significant a considerable number of non-coding and small RNA transcripts. Additionally, all comparing methods except FET greatly increased the number of detected (non-coding) genes as sequencing depth raised while NOISeq showed a constant pattern. Also these other methods tend to select shorter genes and smaller fold change differences with the increasing amounts of reads. In general, parametric approaches selected much more genes than NOISeq, specially at high sequencing depth rates. When analyzing the functional content of these genes by functional enrichment analysis, we observed that the pool of genes detected both by NOISeq and the parametric methods where highly enriched in functional categories, while genes selected only by parametric methods did not. To check whether this differences were indicative of different false calls between methods, we used the RT-PCR data available at the MAQC project that contains 330 true positive and 83 true negative differentially expressed genes. Performance plots indicate that edgeR, DESeq, baySeq strongly increased the number of false calls with sequencing depth, while NOISeq was constant and low. On the contrary true discoveries were slightly better for these methods, presumably consequence of their large number of selected genes. FET showed in low false and true discovery rates, due to its general lower detection power.ConclusionsWe showed that most current RNA-seq statistical analysis methods fail to control the number of false discoveries as the size of the sequenced library increases. These false positive are mainly short, non- coding genes and/or genes with small fold changes. NOISeq, but adopting an empirical approach to model the null distribution of differential expression captures better the shape of noise in RNA-seq data, resulting in a sequencing-depth robust method for differential expression analysis.
- Research Article
4
- 10.14806/ej.18.b.552
- Nov 9, 2012
- EMBnet.journal
A strategy to reduce technical variability and bias in RNA sequencing data
- Research Article
4
- 10.1007/s12561-016-9182-8
- Jun 1, 2017
- Statistics in Biosciences
Ultra high-throughput sequencing of transcriptomes (RNA-Seq) is a widely used method for quantifying gene expression levels due to its low cost, high accuracy, and wide dynamic range for detection. However, the nature of RNA-Seq makes it nearly impossible to provide absolute measurements of transcript abundances. Several units or data summarization methods for transcript quantification have been proposed in the past to account for differences in transcript lengths and sequencing depths across different genes and different samples. Nevertheless, further between-sample normalization is still needed for reliable detection of differentially expressed genes. In this paper, we propose a unified statistical model for joint detection of differential gene expression and between-sample normalization. Our method is independent of the unit in which gene expression levels are summarized. We also introduce an efficient algorithm for model fitting. Due to the L0-penalized likelihood used in our model, it is able to reliably normalize the data and detect differential gene expression in some cases when more than 50% of the genes are differentially expressed in an asymmetric manner. We compare our method with existing methods using simulated and real data sets.
- Book Chapter
8
- 10.1002/9780470015902.a0022508
- Apr 19, 2010
- Encyclopedia of Life Sciences
The advances in next generation sequencing (NGS) technologies have tremendous impacts on the studies of structural and functional genomics. Sequencing‐based approaches like ChIP‐Seq and RNA‐Seq have started taking the place of microarray experiments to study protein–DNA ( deoxyribonucleic acid ) interactions and transcriptomic profiling, respectively. The arrival of NGS technologies has also enabled several whole human genome resequencing studies to be completed efficiently at an affordable price. The major strengths of NGS technologies are their ultra high‐throughput production, characterized by their ability to generate several hundred megabases to tens of gigabases of sequencing data per instrument run, and more importantly, the steep reduction in cost compared to the traditional Sanger sequencing method. Hence, NGS technologies have rapidly become the primary choice for large scale as well as genome‐wide sequencing studies. The new sequencing‐based approaches to explore structural and functional genomics have produced important information and significantly expanded our knowledge in these areas. Key concepts The rapid developments in sequencing technologies have transformed the approaches in the studies of structural and functional genomics. The arrival of next generation sequencing (NGS) technologies has started substituting traditional Sanger sequencing method in many large‐scale or genome‐wide sequencing studies. Shortly after the first next generation sequencer was introduced by Roche® 454 Life Science, the Genome Sequencer 20 (GS 20) System (it was subsequently replaced by GS FLX System), another two biotechnology companies also marketed their sequencing platforms: Illumina® Genome Analyzer (GA) and Applied Biosystems® (ABI) Supported Oligonucleotide Ligation Detection System (SOLiD). The major attractions of NGS technologies are their ultra high‐throughput production, characterized by their ability to produce gigabases of sequencing data per instrument run, and more importantly, the steep reduction in cost compared to the traditional sequencing method. Previously, the molecular genomics studies mainly relied on microarray technologies such as gene expression microarrays and the ChIP‐chip method (i.e. chromatin immunoprecipitation coupled with microarray) for genome‐wide interrogation. The NGS technologies have been used in various research areas besides the standard sequencing applications such as whole genome sequencing; they have also been increasingly applied in detecting structural variations (paired‐end mapping), studies of protein–DNA interactions and histone modifications (ChIP‐Seq) and transcriptomic profiling of messenger RNAs (mRNAs) and noncoding RNAs (RNA‐Seq). Sequencing‐based approaches have already yielded numerous novel and important findings in research areas like genome‐wide mapping of histone modifications and protein–DNA interactions, discovery of genetic variations and transcriptomics studies even though the approaches are still new and maturing. The NGS technologies have shown their potential of being dominant in future genomics studies. This is evident from several international projects using NGS technologies like the ENCODE Project, 1000 Genomes Project and cancers sequencing project by the International Cancer Genome Consortium. It is only a matter of time before achieving the goal of $1000 per whole genome sequencing. This should not be too far from now given the progresses in the development of third generation sequencing technologies. Although the $1000 genome will technically make sequencing of thousands of human genomes a reality, the substantial cost that will be incurred for data storage, powerful computational packages and analytical softwares has to be borne in mind. However, beyond affordability, what are left behind are the bioinformatics challenges in processing and analysing the huge amount of sequencing data.
- Research Article
110
- 10.1016/j.celrep.2015.09.070
- Oct 29, 2015
- Cell Reports
Transcriptomics Identify CD9 as a Marker of Murine IL-10-Competent Regulatory B Cells.
- Research Article
117
- 10.1111/j.1469-8137.2011.03755.x
- May 9, 2011
- New Phytologist
Towards standardization of the description and publication of next‐generation sequencing datasets of fungal communities
- Research Article
58
- 10.1093/bioinformatics/btr005
- Jan 19, 2011
- Bioinformatics
Next-generation sequencing technologies are being rapidly applied to quantifying transcripts (RNA-seq). However, due to the unique properties of the RNA-seq data, the differential expression of longer transcripts is more likely to be identified than that of shorter transcripts with the same effect size. This bias complicates the downstream gene set analysis (GSA) because the methods for GSA previously developed for microarray data are based on the assumption that genes with same effect size have equal probability (power) to be identified as significantly differentially expressed. Since transcript length is not related to gene expression, adjusting for such length dependency in GSA becomes necessary. In this article, we proposed two approaches for transcript-length adjustment for analyses based on Poisson models: (i) At individual gene level, we adjusted each gene's test statistic using the square root of transcript length followed by testing for gene set using the Wilcoxon rank-sum test. (ii) At gene set level, we adjusted the null distribution for the Fisher's exact test by weighting the identification probability of each gene using the square root of its transcript length. We evaluated these two approaches using simulations and a real dataset, and showed that these methods can effectively reduce the transcript-length biases. The top-ranked GO terms obtained from the proposed adjustments show more overlaps with the microarray results. R scripts are at http://www.soph.uab.edu/Statgenetics/People/XCui/r-codes/.
- Research Article
203
- 10.1186/1471-2164-13-484
- Sep 17, 2012
- BMC Genomics
BackgroundRNA sequencing (RNA-Seq) has emerged as a powerful approach for the detection of differential gene expression with both high-throughput and high resolution capabilities possible depending upon the experimental design chosen. Multiplex experimental designs are now readily available, these can be utilised to increase the numbers of samples or replicates profiled at the cost of decreased sequencing depth generated per sample. These strategies impact on the power of the approach to accurately identify differential expression. This study presents a detailed analysis of the power to detect differential expression in a range of scenarios including simulated null and differential expression distributions with varying numbers of biological or technical replicates, sequencing depths and analysis methods.ResultsDifferential and non-differential expression datasets were simulated using a combination of negative binomial and exponential distributions derived from real RNA-Seq data. These datasets were used to evaluate the performance of three commonly used differential expression analysis algorithms and to quantify the changes in power with respect to true and false positive rates when simulating variations in sequencing depth, biological replication and multiplex experimental design choices.ConclusionsThis work quantitatively explores comparisons between contemporary analysis tools and experimental design choices for the detection of differential expression using RNA-Seq. We found that the DESeq algorithm performs more conservatively than edgeR and NBPSeq. With regard to testing of various experimental designs, this work strongly suggests that greater power is gained through the use of biological replicates relative to library (technical) replicates and sequencing depth. Strikingly, sequencing depth could be reduced as low as 15% without substantial impacts on false positive or true positive rates.
- Research Article
- 10.3760/cma.j.issn.1001-7097.2019.02.008
- Feb 15, 2019
- Chin J Nephrol
Objective To find the differentially expressed long non-coding RNA (lncRNA) between db/db mice that with nephropathy (DN) or not (DM). Methods In this study, 3 DM db/db mice and 2 DN db/db mice proven by renal biopsy were randomly selected, and 3 healthy db/m mice as normal control group. Then, differentially expressed lncRNAs, mRNAs and their fragments per kilobase million (FPKM) in kidney samples were detected by high-throughput next generation sequencing technology. Sequencing data were analyzed to filter out the differentially expressed lncRNA, and their function was preliminarily investigated by bioinformatics analysis and functional enrichment analysis to predict their target genes. Total RNAs of kidneys from these 8 mice were extracted to run real time PCR (RT-qPCR) for verifying the outcomes of the high-throughput sequencing. Results The urinary microalbumin/creatinine ratio (UACR), serum creatinine, and glomerular basement membrane thickness of DN db/db mice were higher than those of DM db/db mice (all P<0.05), while there was not significant difference in glucose between DM and DN mice. Totally 160 lncRNAs were up-regulated and 99 lncRNAs were down-regulated in kidneys of DN mice compared with those of DM mice, in which the differentially expressed lncRNAs with FPKM value≥2 and differential expression≥1 fold between groups were screened and verified by RT-qPCR. Finally three lncRNAs whose variation trend were consistent with the outcomes of the high-throughput sequencing were obtained. Conclusion There was a significantly different expression pattern of lncRNA between the kidneys of DN and DM mice, which may be involved in the progress of DN. Key words: Diabetes; Diabetic nephropathy; Real time PCR; Long non-coding RNA; High-throughput sequencing
- Research Article
348
- 10.1371/journal.pone.0103207
- Aug 13, 2014
- PLoS ONE
Recent advances in next-generation sequencing technology allow high-throughput cDNA sequencing (RNA-Seq) to be widely applied in transcriptomic studies, in particular for detecting differentially expressed genes between groups. Many software packages have been developed for the identification of differentially expressed genes (DEGs) between treatment groups based on RNA-Seq data. However, there is a lack of consensus on how to approach an optimal study design and choice of suitable software for the analysis. In this comparative study we evaluate the performance of three of the most frequently used software tools: Cufflinks-Cuffdiff2, DESeq and edgeR. A number of important parameters of RNA-Seq technology were taken into consideration, including the number of replicates, sequencing depth, and balanced vs. unbalanced sequencing depth within and between groups. We benchmarked results relative to sets of DEGs identified through either quantitative RT-PCR or microarray. We observed that edgeR performs slightly better than DESeq and Cuffdiff2 in terms of the ability to uncover true positives. Overall, DESeq or taking the intersection of DEGs from two or more tools is recommended if the number of false positives is a major concern in the study. In other circumstances, edgeR is slightly preferable for differential expression analysis at the expense of potentially introducing more false positives.
- Research Article
21
- 10.1016/j.cels.2020.01.002
- Feb 1, 2020
- Cell Systems
Differential Allele-Specific Expression Uncovers Breast Cancer Genes Dysregulated by Cis Noncoding Mutations.
- Research Article
- 10.69750/dmls.02.08.0143
- Sep 26, 2025
- DEVELOPMENTAL MEDICO-LIFE-SCIENCES
The last two decades have witnessed an unprecedented transformation in genomics research, driven largely by the development and maturation of next-generation sequencing (NGS) technologies. Among them, RNA sequencing (RNA-Seq) has emerged as a cornerstone technique for exploring the transcriptome, enabling researchers to characterize gene expression profiles, alternative splicing events, non-coding RNAs, and gene fusions with remarkable depth and resolution [1]. What began as a powerful research tool is now advancing steadily towards clinical utility, with applications ranging from precision oncology to rare disease diagnostics. Yet, the journey from bench to bedside is far from straightforward. The evolving landscape of RNA-Seq data analysis offers both opportunities and challenges that must be addressed to ensure reliable clinical translation [2]. Evolution of RNA-Seq Analytical Pipelines: Early RNA-Seq studies were primarily focused on counting reads aligned to known genes. Over time, analysis pipelines have expanded in scope and sophistication [3]. Current workflows integrate multiple layers quality control, alignment or pseudoalignment, normalization, statistical modeling, and downstream interpretation each supported by specialized software tools. Algorithms such as STAR, HISAT2, Salmon, and Kallisto have improved speed and accuracy in alignment, while normalization methods like DESeq2 and edgeR address inherent biases in sequencing depth and sample composition. The refinement of these pipelines has enabled researchers to handle increasingly complex datasets, including single-cell RNA-Seq and spatial transcriptomics [4]. Addressing Data Variability and Standardization: One of the most pressing challenges in translating RNA-Seq into clinical practice is variability across platforms, laboratories, and computational pipelines. Batch effects, differences in library preparation, and sequencing depth can profoundly influence results, raising concerns about reproducibility [5]. International initiatives such as the Genotype-Tissue Expression (GTEx) project and consortia like ENCODE have highlighted the necessity of standardized protocols. More recently, bioinformatics efforts have begun to converge on best practices, incorporating machine learning approaches for robust normalization and artifact correction. Such standardization is essential if RNA-Seq is to achieve regulatory acceptance for clinical decision-making [6]. Integration with Clinical Workflows: For RNA-Seq to impact patient care meaningfully, it must integrate seamlessly into clinical workflows [7]. This requires not only rapid turnaround times but also clear interpretability of results. Advances in automated pipelines and cloud-based computational platforms are helping reduce analysis bottlenecks. Additionally, user-friendly visualization tools are enabling clinicians to navigate complex datasets with greater ease. The development of clinical reporting standards, particularly in oncology where RNA-Seq can reveal actionable mutations or expression signatures, is bridging the gap between raw data and meaningful therapeutic recommendations [8]. Emerging Frontiers: Single-Cell and Multi-Omics: The shift towards single-cell RNA-Seq has opened new dimensions of biological discovery, allowing for the resolution of cellular heterogeneity within tissues and tumors. When integrated with proteomics, metabolomics, and epigenomics, RNA-Seq forms part of a powerful multi-omics toolkit [9]. These integrative approaches promise a holistic view of disease mechanisms, offering unprecedented potential for biomarker discovery and therapeutic targeting. However, the complexity of such datasets necessitates advanced computational models, often leveraging artificial intelligence, to extract clinically relevant insights [10]. Ethical and Regulatory Considerations: Clinical translation of RNA-Seq also raises ethical and regulatory questions. Patient privacy, informed consent, and the management of incidental findings are critical concerns [11]. Moreover, regulatory agencies require rigorous demonstration of analytic validity, clinical validity, and clinical utility before approving RNA-Seq-based diagnostics. Collaborative frameworks between researchers, clinicians, bioinformaticians, and regulators will be essential in establishing RNA-Seq as a reliable clinical tool [12]. CONCLUSION RNA-Seq has evolved from an exploratory research technology to a promising instrument for precision medicine. Advances in computational analysis, standardization efforts, and multi-omics integration are bringing the field closer to clinical translation. However, significant challenges remain in ensuring reproducibility, interpretability, and regulatory compliance. The future of RNA-Seq lies in collaborative innovation, where cutting-edge computational methods, rigorous standardization, and ethical foresight converge to transform transcriptomic insights into tangible clinical benefits. If these hurdles can be overcome, RNA-Seq has the potential to redefine diagnostics and therapeutics, marking a new era in personalized medicine.
- Research Article
- 10.1093/bioadv/vbag134
- May 9, 2026
- Bioinformatics Advances
PurposeIn this study, we evaluated 90 bioinformatics pipelines using RNA-Seq datasets from Rats, Zebrafish and Mice. The analysis was conducted in the context of weak signals, including exposure to metallic particles (tungsten), low-dose radiation or medical treatment. RNA-Seq data analysis involves several critical steps, from quality control to differential expression analysis, each offering multiple algorithmic options. Selecting the optimal pipeline is particularly challenging in complex scenarios with weak signals.We applied a dual strategy based on two complementary approaches to rank and evaluate the performance of these pipelines. The first approach, a widely used method, is based on the correlation between RNA-Seq and qRT-PCR expression data to ensure the direct validation of RNA-Seq results. The second approach leverages machine learning classifiers to rank the pipelines based on their ability to distinguish between exposure groups. This dual strategy was designed to identify the most reliable pipelines capable of providing accurate biological insights, with the top-performing pipeline highlighting key biological processes linked to a weak signal.ResultsOur results highlight the crucial role of pipeline selection in RNA-Seq studies, as it influences both analysis efficiency and biological insights. While our findings are particularly relevant for studies with weak RNA-Seq signals, the ranking methods we employed can be applied in other fields to identify the most appropriate pipeline for generating biologically meaningful data. We provide practical recommendations for bioinformaticians to select robust pipelines, ensuring reliable and insightful outcomes across various research contexts, including environmental exposures.ConclusionsRNA-Seq pipelines that are effective for strong signals may fail with weak signals or noisy RNA-seq data. Our classifier-based ranking approach is still particularly useful, even when the sequencing depth is lower. Pipelines' sensitivity has a greater impact on counting and normalization than on trimming and mapping. StringTie should be prioritized as a counting method for data with weak signals.
- Research Article
1
- 10.14806/ej.18.a.441
- Apr 29, 2012
- EMBnet.journal
Motivations. In recent years, RNA sequencing (RNA-seq) has rapidly become the method of choice for measuring and comparing gene transcription levels. Despite its wide application, it is now clear that this methodology is not free from biases and that a careful normalization procedure is the basis for a correct data interpretation. The most common normalization techniques account for: library size, gene or transcript length and sequence-specific biases such as GC-content effects. The aim of the present work is to investigate biases affecting RNA seq data and their effect on differential expression analysis. In order to reduce biases due to over-simplification of gene transcription models, we consider exon-based counts. Methods. We two used publicly available RNA-seq data sets from two-group comparison studies which are characterized by multiple technical replicates. We summarized read counts at exon level and investigated their dependence on sequence-specific covariates: GC-content and exon length. In addition, we considered the effect of library size correction on between-groups comparison and the impact of the above mentioned biases on the detection of differentially expressed exons. The assessment is performed on raw data, as well as on data normalized with different approaches: RPKM [1], library size scaling, based on Trimmed Mean of M-values (TMM) [2] and on Poisson goodness-of-fit statistic applied to non differentially expressed genes [3], and within-lane normalization based on loess regression of log-counts on GC-content and exon length [4]. We selected differentially expressed exons using the GLM-based version of edgeR [5] as it can consider an offset matrix which codifies counts normalization, that can be computed with the desired approach, and library size scaling factors specified by the user. Results. In our study, read counts show a significant dependence on exon length and a moderate dependence on GC-content. Exon length bias also affects differential expression analysis: longer exons tend to have lower P-values and to be selected as differentially expressed more frequently than shorter exons. The tested normalization techniques do not completely remove biases and, in particular, RPKM approach over-corrects for exon length bias. Moreover, the choice of the strategy for library size adjustment has a great impact on the direction of the detected differential expression. The results obtained on these data sets demonstrate that RNA-seq data normalization is still an open issue. Further efforts should be directed towards the clarification of the relationship between read counts and sequence-specific biases, which are, in turn, correlated to each other, and the definition of new models for their correction.
- Research Article
72
- 10.1371/journal.pone.0066883
- Jun 24, 2013
- PLoS ONE
Recent advances in RNA sequencing (RNA-Seq) have enabled the discovery of novel transcriptomic variations that are not possible with traditional microarray-based methods. Tissue and cell specific transcriptome changes during pathophysiological stress in disease cases versus controls and in response to therapies are of particular interest to investigators studying cardiometabolic diseases. Thus, knowledge on the relationships between sequencing depth and detection of transcriptomic variation is needed for designing RNA-Seq experiments and for interpreting results of analyses. Using deeply sequenced Illumina HiSeq 2000 101 bp paired-end RNA-Seq data derived from adipose of a healthy individual before and after systemic administration of endotoxin (LPS), we investigated the sequencing depths needed for studies of gene expression and alternative splicing (AS). In order to detect expressed genes and AS events, we found that ∼100 to 150 million (M) filtered reads were needed. However, the requirement on sequencing depth for the detection of LPS modulated differential expression (DE) and differential alternative splicing (DAS) was much higher. To detect 80% of events, ∼300 M filtered reads were needed for DE analysis whereas at least 400 M filtered reads were necessary for detecting DAS. Although the majority of expressed genes and AS events can be detected with modest sequencing depths (∼100 M filtered reads), the estimated gene expression levels and exon/intron inclusion levels were less accurate. We report the first study that evaluates the relationship between RNA-Seq depth and the ability to detect DE and DAS in human adipose. Our results suggest that a much higher sequencing depth is needed to reliably identify DAS events than for DE genes.