Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

Microbiome Datasets Are Compositional: And This Is Not Optional

  • Abstract
  • Highlights & Summary
  • PDF
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

Datasets collected by high-throughput sequencing (HTS) of 16S rRNA gene amplimers, metagenomes or metatranscriptomes are commonplace and being used to study human disease states, ecological differences between sites, and the built environment. There is increasing awareness that microbiome datasets generated by HTS are compositional because they have an arbitrary total imposed by the instrument. However, many investigators are either unaware of this or assume specific properties of the compositional data. The purpose of this review is to alert investigators to the dangers inherent in ignoring the compositional nature of the data, and point out that HTS datasets derived from microbiome studies can and should be treated as compositions at all stages of analysis. We briefly introduce compositional data, illustrate the pathologies that occur when compositional data are analyzed inappropriately, and finally give guidance and point to resources and examples for the analysis of microbiome datasets using compositional data analysis.

Similar Papers
  • Research Article
  • Cite Count Icon 378
  • 10.1139/cjm-2015-0821
Compositional analysis: a valid approach to analyze microbiome high-throughput sequencing data.
  • Apr 12, 2016
  • Canadian Journal of Microbiology
  • Gregory B Gloor + 1 more

A workshop held at the 2015 annual meeting of the Canadian Society of Microbiologists highlighted compositional data analysis methods and the importance of exploratory data analysis for the analysis of microbiome data sets generated by high-throughput DNA sequencing. A summary of the content of that workshop, a review of new methods of analysis, and information on the importance of careful analyses are presented herein. The workshop focussed on explaining the rationale behind the use of compositional data analysis, and a demonstration of these methods for the examination of 2 microbiome data sets. A clear understanding of bioinformatics methodologies and the type of data being analyzed is essential, given the growing number of studies uncovering the critical role of the microbiome in health and disease and the need to understand alterations to its composition and function following intervention with fecal transplant, probiotics, diet, and pharmaceutical agents.

  • Discussion
  • Cite Count Icon 53
  • 10.1093/annweh/wxaa056
Time-Based Data in Occupational Studies: The Whys, the Hows, and Some Remaining Challenges in Compositional Data Analysis (CoDA)
  • Jul 1, 2020
  • Annals of Work Exposures and Health
  • Nidhi Gupta + 3 more

Data on the use of time in different exposures, behaviors, and work tasks are common in occupational research. Such data are most often expressed in hours, minutes, or percentage of work time. Thus, they are constrained or ‘compositional’, in that they add up to a finite sum (e.g. 8 h of work or 100% work time). Due to their properties, compositional data need to be processed and analyzed using specifically adapted methods. Compositional data analysis (CoDA) has become a particularly established framework to handle such data in various scientific fields such as nutritional epidemiology, geology, and chemistry, but has only recently gained attention in public and occupational health sciences. In this paper, we introduce the reader to CoDA by explaining why CoDA should be used when dealing with compositional time-use data, showing how to perform CoDA, including a worked example, and pointing at some remaining challenges in CoDA. The paper concludes by emphasizing that CoDA in occupational research is still in its infancy, and stresses the need for further development and experience in the use of CoDA for time-based occupational exposures. We hope that the paper will encourage researchers to adopt and apply CoDA in studies of work exposures and health.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 111
  • 10.3389/fmicb.2021.727398
Compositional Data Analysis of Microbiome and Any-Omics Datasets: A Validation of the Additive Logratio Transformation.
  • Oct 11, 2021
  • Frontiers in Microbiology
  • Michael Greenacre + 2 more

Microbiome and omics datasets are, by their intrinsic biological nature, of high dimensionality, characterized by counts of large numbers of components (microbial genes, operational taxonomic units, RNA transcripts, etc.). These data are generally regarded as compositional since the total number of counts identified within a sample is irrelevant. The central concept in compositional data analysis is the logratio transformation, the simplest being the additive logratios with respect to a fixed reference component. A full set of additive logratios is not isometric, that is they do not reproduce the geometry of all pairwise logratios exactly, but their lack of isometry can be measured by the Procrustes correlation. The reference component can be chosen to maximize the Procrustes correlation between the additive logratio geometry and the exact logratio geometry, and for high-dimensional data there are many potential references. As a secondary criterion, minimizing the variance of the reference component's log-transformed relative abundance values makes the subsequent interpretation of the logratios even easier. On each of three high-dimensional omics datasets the additive logratio transformation was performed, using references that were identified according to the abovementioned criteria. For each dataset the compositional data structure was successfully reproduced, that is the additive logratios were very close to being isometric. The Procrustes correlations achieved for these datasets were 0.9991, 0.9974, and 0.9902, respectively. We thus demonstrate, for high-dimensional compositional data, that additive logratios can provide a valid choice as transformed variables, which (a) are subcompositionally coherent, (b) explain 100% of the total logratio variance and (c) come measurably very close to being isometric. The interpretation of additive logratios is much simpler than the complex isometric alternatives and, when the variance of the log-transformed reference is very low, it is even simpler since each additive logratio can be identified with a corresponding compositional component.

  • Research Article
  • Cite Count Icon 1
  • 10.1360/n012017-00147
High-dimensional count and compositional data analysis in\\ microbiome studies
  • Nov 16, 2017
  • SCIENTIA SINICA Mathematica
  • DENG MingHua + 2 more

The human microbiome plays an important role in human health and disease. The development of high-throughput sequencing technologies makes it possible to quantify all microbes constituting the microbiome. In this paper, we give a review of recent advances in high-dimensional count and compositional data analysis in microbiome studies. It includes the Dirichlet-multinomial model and its extensions, composition estimation from large sparse count matrix, high-dimensional regression with compositional covariates, and statistical inference for log-basis-based compositional data.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 65
  • 10.17713/ajs.v45i4.122
Compositional uncertainty should not be ignored in high-throughput sequencing data analysis
  • Jul 28, 2016
  • Austrian Journal of Statistics
  • Gregory Brian Gloor + 3 more

High throughput sequencing generates sparse compositional data, yet these datasets are rarely analyzed using a compositional approach. In addition, the variation inherent in these datasets is rarely acknowledged, but ignoring it can result in many false positive inferences. We demonstrate that examination of point estimates of the data can result in false positive results, even with appropriate zero replacement approaches, using an in vitro selection dataset with an outside standard of truth. The variation inherent in real high-throughput sequencing datasets is demonstrated, and we show that this varia- tion can be approximated, and hence accounted for, by Monte-Carlo sampling from the Dirichlet distribution. This approximation when used by itself is itself problematic, but becomes useful when coupled with a log-ratio approach commonly used in compositional data analysis. Thus, the approach illustrated here that merges Bayesian estimation with principles of compositional data analysis should be generally useful for high-dimensional count compositional data of the type generated by high throughput sequencing.

  • Research Article
  • Cite Count Icon 3
  • 10.1093/bioinformatics/btad700
Gmcoda: Graphical model for multiple compositional vectors in microbiome studies.
  • Nov 1, 2023
  • Bioinformatics
  • Huaying Fang

Microbes are essential components in the ecosystem and participate in most biological procedures in environments. The high-throughput sequencing technologies help researchers directly quantify the abundance of microbes in a natural environment. Microbiome studies explore the construction, stability, and function of microbial communities with the aid of sequencing technology. However, sequencing technologies only provide relative abundances of microbes, and this kind of data is called compositional data in statistics. The constraint of the constant-sum requires flexible statistical methods for analyzing microbiome data. Current statistical analysis of compositional data mainly focuses on one compositional vector such as bacterial communities. The fungi are also an important component in microbial communities and are always measured by sequencing internal transcribed spacer instead of 16S rRNA genes for bacteria. The different sequencing methods between fungi and bacteria bring two compositional vectors in microbiome studies. We propose a novel statistical method, called gmcoda, based on an additive logistic normal distribution for estimating the partial correlation matrix for cross-domain interactions. A majorization-minimization algorithm is proposed to solve the optimization problem involved in gmcoda. Through simulation studies, gmcoda is demonstrated to work well in estimating partial correlations between two compositional vectors. Gmcoda is also applied to infer cross-domain interactions in a real microbiome dataset and finds potential interactions between bacteria and fungi. Gmcoda is open source and freely available from https://github.com/huayingfang/gmcoda under GNU LGPL v3.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 98
  • 10.1186/s12966-018-0685-1
A comparison of standard and compositional data analysis in studies addressing group differences in sedentary behavior and physical activity
  • Jun 15, 2018
  • The International Journal of Behavioral Nutrition and Physical Activity
  • Nidhi Gupta + 6 more

BackgroundData on time spent in physical activity, sedentary behavior and sleep during a day is compositional in nature, i.e. they add up to a constant value. Compositional data have fundamentally different properties from unconstrained data in real space, and require other analytical procedures, referred to as compositional data analysis (CoDA). Most physical activity and sedentary behavior studies, however, still apply analytical procedures adapted to data in real space, which can lead to misleading results. The present study describes a comparison of time spent sedentary and in physical activity between age groups and sexes, and investigates the extent to which results obtained by CoDA differ from those obtained using standard analytical procedures.MethodsTime spent sedentary, standing, and in physical activity (walking/running/stair climbing/cycling) during work and leisure was determined for 1–4 days among 677 blue-collar workers using accelerometry. Differences between sexes and age groups were tested using MANOVA, using both a standard and a CoDA approach based on isometric log-ratio transformed data.ResultsWhen determining differences between sexes for different activities time at work, the effect size using standard analysis (η2 = 0.045, p < 0.001) was 15% smaller than that obtained with CoDA (η2 = 0.052, p < 0.001), although both approaches suggested a statistically significant difference. When determining corresponding differences between age groups, CoDA resulted in a 60% larger, and significant, effect size (η2 = 0.012, p = 0.02) than that obtained with the standard approach (η2 = 0.008, p = 0.07). During leisure, results based on standard (age; η2 = 0.007, p = 0.09; sex; η2 = 0.052, p < 0.001) and CoDA (age; η2 = 0.007, p = 0.09; sex; η2 = 0.051, p < 0.001) analyses were similar.ConclusionResults and, hence, inferences concerning age and sex-based differences in time spent sedentary and in physical activity at work differed between CoDA and standard analysis. We encourage researchers to use CoDA in similar studies, to adequately account for the compositional nature of data on physical activity and sedentary behavior.

  • Research Article
  • Cite Count Icon 5
  • 10.1186/s12874-025-02509-1
A comparison of methods for analysing compositional data with fixed and variable totals: a simulation study using the examples of time-use and dietary data
  • Apr 17, 2025
  • BMC Medical Research Methodology
  • Georgia D Tomova + 4 more

BackgroundCompositional data comprise the parts of a ‘whole’ (or ‘total’), which sum to that ‘whole’. The ‘whole’ may vary between units of analyses, or it may be fixed (constant). For example, total energy intake (a variable total) is the sum of intake from all foods or macronutrients. Total time in a day (a fixed total) is the sum of time spent engaging in various activities. There exist different approaches to analysing compositional data, such as the isocaloric or isotemporal model, ratio variables, and compositional data analysis (CoDA). Although the performance of the different approaches has been compared previously, this has only been conducted in real data. Since the true relationships are unknown in real data, it is difficult to compare model performance in estimating a known effect. We use data simulations of different parametric relationships, to explore and demonstrate the performance of each approach under various possible conditions.MethodsWe simulated physical activity time-use and dietary data as examples of compositional data with fixed and variable totals, respectively, using different parametric relationships between the compositional components and the outcome (fasting plasma glucose): linear, log2, and isometric log-ratios. We evaluated the performance of a range of generalised linear and additive models as well as CoDA, in estimating a 1-unit and either 10-unit (for physical activity) or 100-unit (for dietary data) reallocations under each parametric scenario. We simulated 10,000 datasets with 1,000 observations in each.ResultsThe performance of each approach to analysing compositional data depends on how closely its parameterisation matches the true data generating process. Overall, we demonstrated that the consequences of using an incorrect parameterisation (e.g. using CoDA when the true relationship is linear) are more severe for larger reallocations (e.g. 10-min or 100-kcal) than for 1-unit reallocations. The implications of choosing an unsuitable approach may be starker in compositional data with variable totals. For example, while models with ratio variables are mathematically equivalent to linear models in compositional data with fixed totals, their estimates may be radically different for variable totals.ConclusionsCompositional data with fixed and variable totals behave differently. All existing approaches to analysing such data have utility but need to be carefully selected. Investigators should explore the shape of the relationships between the compositional components and the outcome and chose an approach that matches it best.

  • Research Article
  • Cite Count Icon 7
  • 10.1016/j.sedgeo.2015.01.012
Compositional Data Analysis (CoDA) as a tool to study the (paleo)ecology of coccolithophores from coastal-neritic settings off central Portugal
  • Feb 13, 2015
  • Sedimentary Geology
  • Catarina Guerreiro + 4 more

Compositional Data Analysis (CoDA) as a tool to study the (paleo)ecology of coccolithophores from coastal-neritic settings off central Portugal

  • Research Article
  • Cite Count Icon 344
  • 10.1016/j.annepidem.2016.03.002
Compositional data analysis of the microbiome: fundamentals, tools, and challenges
  • Mar 31, 2016
  • Annals of Epidemiology
  • Matthew C.B Tsilimigras + 1 more

Compositional data analysis of the microbiome: fundamentals, tools, and challenges

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 17
  • 10.1186/s12966-020-00985-w
Can we walk away from cardiovascular disease risk or do we have to \u2018huff and puff\u2019? A cross-sectional compositional accelerometer data analysis among adults and older adults in the Copenhagen City Heart Study
  • Jul 6, 2020
  • The International Journal of Behavioral Nutrition and Physical Activity
  • Melker Staffan Johansson + 6 more

BackgroundIt is unclear whether walking can decrease cardiovascular disease (CVD) risk or if high intensity physical activity (HIPA) is needed, and whether the association is modified by age. We investigated how sedentary behaviour, walking, and HIPA, were associated with systolic blood pressure (SBP), waist circumference (WC), and low-density lipoprotein cholesterol (LDL-C) among adults and older adults in a general population sample using compositional data analysis. Specifically, the measure of association was quantified by reallocating time between sedentary behaviour and 1) walking, and 2) HIPA.MethodsCross-sectional data from the fifth examination of the Copenhagen City Heart Study was used. Using the software Acti4, we estimated daily time spent in physical behaviours from accelerometer data worn 24 h/day for 7 days (i.e., right frontal thigh and iliac crest; median wear time: 6 days, 23.8 h/day). SBP, WC, and LDL-C were measured during a physical examination. Inclusion criteria were ≥ 5 days with ≥16 h of accelerometer recordings per day, and no use of antihypertensives, diuretics or cholesterol lowering medicine. The 24-h physical behaviour composition consisted of sedentary behaviour, standing, moving, walking, HIPA (i.e., sum of climbing stairs, running, cycling, and rowing), and time in bed. We used fitted values from linear regression models to predict the difference in outcome given the investigated time reallocations relative to the group-specific mean composition.ResultsAmong 1053 eligible participants, we found an interaction between the physical behaviour composition and age. Age-stratified analyses (i.e., </≥65 years; 773 adults, 280 older adults) indicated that less sedentary behaviour and more walking was associated with lower SBP among older adults only. For less sedentary behaviour and more HIPA, the results i) indicated an association with a lower SBP irrespective of age, ii) showed an association with a smaller WC among adults, and iii) showed an association with a lower LDL-C in both age groups.ConclusionsLess sedentary behaviour and more walking seems to be associated with lower CVD risk among older adults, while HIPA types are associated with lower risk among adults. Therefore, to reduce CVD risk, the modifying effect of age should be considered in future physical activity-promoting initiatives.

  • Research Article
  • Cite Count Icon 20
  • 10.1161/circulationaha.124.069820
Device-Measured 24-Hour Movement Behaviors and Blood Pressure: A 6-Part Compositional Individual Participant Data Analysis in the ProPASS Consortium
  • Nov 6, 2024
  • Circulation
  • Joanna M Blodgett + 25 more

BACKGROUND:Blood pressure (BP)–lowering effects of structured exercise are well-established. Effects of 24-hour movement behaviors captured in free-living settings have received less attention. This cross-sectional study investigated associations between a 24-hour behavior composition comprising 6 parts (sleeping, sedentary behavior, standing, slow walking, fast walking, and combined exercise-like activity [eg, running and cycling]) and systolic BP (SBP) and diastolic BP (DBP).METHODS:Data from thigh-worn accelerometers and BP measurements were collected from 6 cohorts in the Prospective Physical Activity, Sitting and Sleep consortium (ProPASS) (n=14 761; mean±SD, 54.2±9.6 years). Individual participant analysis using compositional data analysis was conducted with adjustments for relevant harmonized covariates. Based on the average sample composition, reallocation plots examined estimated BP reductions through behavioral replacement; the theoretical benefits of optimal (ie, clinically meaningful improvement in SBP [2 mm Hg] or DBP [1 mm Hg]) and minimal (ie, 5-minute reallocation) behavioral replacements were identified.RESULTS:The average 24-hour composition consisted of sleeping (7.13±1.19 hours), sedentary behavior (10.7±1.9 hours), standing (3.2±1.1 hours), slow walking (1.6±0.6 hours), fast walking (1.1±0.5 hours), and exercise-like activity (16.0±16.3 minutes). More time spent exercising or sleeping, relative to other behaviors, was associated with lower BP. An additional 5 minutes of exercise-like activity was associated with estimated reductions of –0.68 mm Hg (95% CI, –0.15, –1.21) SBP and –0.54 mm Hg (95% CI, –0.19, 0.89) DBP. Clinically meaningful improvements in SBP and DBP were estimated after 20 to 27 minutes and 10 to 15 minutes of reallocation of time in other behaviors into additional exercise. Although more time spent being sedentary was adversely associated with SBP and DBP, there was minimal impact of standing or walking.CONCLUSIONS:Study findings reiterate the importance of exercise for BP control, suggesting that small additional amounts of exercise are associated with lower BP in a free-living setting.

  • Research Article
  • Cite Count Icon 30
  • 10.1016/j.gexplo.2022.107112
Identification of Rare Earth Elements (REEs) distribution patterns in the soils of Campania region (Italy) using compositional and multivariate data analysis
  • Oct 20, 2022
  • Journal of Geochemical Exploration
  • Maurizio Ambrosino + 6 more

Identification of Rare Earth Elements (REEs) distribution patterns in the soils of Campania region (Italy) using compositional and multivariate data analysis

  • Research Article
  • Cite Count Icon 1308
  • 10.1186/2049-2618-2-15
Unifying the analysis of high-throughput sequencing datasets: characterizing RNA-seq, 16S rRNA gene sequencing and selective growth experiments by compositional data analysis
  • May 5, 2014
  • Microbiome
  • Andrew D Fernandes + 5 more

BackgroundExperimental designs that take advantage of high-throughput sequencing to generate datasets include RNA sequencing (RNA-seq), chromatin immunoprecipitation sequencing (ChIP-seq), sequencing of 16S rRNA gene fragments, metagenomic analysis and selective growth experiments. In each case the underlying data are similar and are composed of counts of sequencing reads mapped to a large number of features in each sample. Despite this underlying similarity, the data analysis methods used for these experimental designs are all different, and do not translate across experiments. Alternative methods have been developed in the physical and geological sciences that treat similar data as compositions. Compositional data analysis methods transform the data to relative abundances with the result that the analyses are more robust and reproducible.ResultsData from an in vitro selective growth experiment, an RNA-seq experiment and the Human Microbiome Project 16S rRNA gene abundance dataset were examined by ALDEx2, a compositional data analysis tool that uses Bayesian methods to infer technical and statistical error. The ALDEx2 approach is shown to be suitable for all three types of data: it correctly identifies both the direction and differential abundance of features in the differential growth experiment, it identifies a substantially similar set of differentially expressed genes in the RNA-seq dataset as the leading tools and it identifies as differential the taxa that distinguish the tongue dorsum and buccal mucosa in the Human Microbiome Project dataset. The design of ALDEx2 reduces the number of false positive identifications that result from datasets composed of many features in few samples.ConclusionStatistical analysis of high-throughput sequencing datasets composed of per feature counts showed that the ALDEx2 R package is a simple and robust tool, which can be applied to RNA-seq, 16S rRNA gene sequencing and differential growth datasets, and by extension to other techniques that use a similar approach.

  • Abstract
  • 10.1136/jech-2023-ssmabstracts.253
OP123 DAG-informed regression-based modelling is a valid approach to estimating causal effects in compositional data with fixed totals: a simulation study
  • Aug 1, 2023
  • Journal of Epidemiology and Community Health
  • Georgia D Tomova + 1 more

BackgroundCompositional data comprise ‘parts’ of a ‘whole’ (or total), where the parts sum to the whole. In compositional data with fixed totals (e.g. hours within a day), only relative causal...

Save Icon
Up Arrow
Open/Close
Setting-up Chat
Loading Interface