Mass spectral databases for LC/MS- and GC/MS-based metabolomics: State of the field and future prospects
Mass spectral databases for LC/MS- and GC/MS-based metabolomics: State of the field and future prospects
- Research Article
77
- 10.3390/metabo8030051
- Sep 15, 2018
- Metabolites
The use of mass spectrometry-based metabolomics to study human, plant and microbial biochemistry and their interactions with the environment largely depends on the ability to annotate metabolite structures by matching mass spectral features of the measured metabolites to curated spectra of reference standards. While reference databases for metabolomics now provide information for hundreds of thousands of compounds, barely 5% of these known small molecules have experimental data from pure standards. Remarkably, it is still unknown how well existing mass spectral libraries cover the biochemical landscape of prokaryotic and eukaryotic organisms. To address this issue, we have investigated the coverage of 38 genome-scale metabolic networks by public and commercial mass spectral databases, and found that on average only 40% of nodes in metabolic networks could be mapped by mass spectral information from standards. Next, we deciphered computationally which parts of the human metabolic network are poorly covered by mass spectral libraries, revealing gaps in the eicosanoids, vitamins and bile acid metabolism. Finally, our network topology analysis based on the betweenness centrality of metabolites revealed the top 20 most important metabolites that, if added to MS databases, may facilitate human metabolome characterization in the future.
- Research Article
14
- 10.6028/jres.094.005
- Jan 1, 1989
- Journal of Research of the National Institute of Standards and Technology
Databases for use with analytical chemistry instrumental techniques are surveyed, with attention to existing databases and collection efforts now underway, as well as needs for new data-bases. Collections of spectra for use in NMR, infrared spectroscopy, and mass spectroscopy are described. Using mass spectral databases as an example, a critique is presented of automated quality control procedures used to evaluate individual spectra in large collections; the kinds of problems which have been en-countered in using these procedures are discussed. Finally, a brief critical review is presented covering the application of computers to the identification of unknown compounds using spectral data-bases; again, algorithms used with mass spectrometry are taken as the example. Ongoing work at NIST with the NIST/EPA/MSDC Mass Spectral Database is concerned with many of these problems; recent developments are described.
- Research Article
13
- 10.1007/978-1-0716-0239-3_8
- Jan 1, 2020
- Methods in molecular biology (Clifton, N.J.)
Liquid chromatography-mass spectrometry (LC-MS) is one of the most popular technologies in metabolomics. The large-scale and unambiguous identification of metabolite structures remains a challenging task in LC-MS based metabolomics. Tandem mass spectral databases provide experimental and in silico MS/MS spectra to facilitate the identification of both known and unknown metabolites, which has become a gold standard method in metabolomics. In addition, metabolite knowledge databases offer valuable biological and pathway information of metabolites. In this chapter, we have briefly reviewed the most common and important tandem mass spectral and metabolite databases, and illustrated how they could be used for metabolite identification.
- Research Article
69
- 10.1016/j.biotechadv.2007.11.002
- Nov 19, 2007
- Biotechnology Advances
Novel omics technologies in nutrition research
- Research Article
14
- 10.1002/bmc.6019
- Oct 7, 2024
- Biomedical chromatography : BMC
Mass spectrometry (MS) plays a crucial role in metabolomics, especially in the discovery of disease biomarkers. This review outlines strategies for identifying metabolites, emphasizing precise and detailed use of MS techniques. It explores various methods for quantification, discusses challenges encountered, and examines recent breakthroughs in biomarker discovery. In the field of diagnostics, MS has revolutionized approaches by enabling a deeper understanding of tissue-specific metabolic changes associated with disease. The reliability of results is ensured through robust experimental design and stringent system suitability criteria. In the past, data quality, standardization, and reproducibility were often overlooked despite their significant impact on MS-based metabolomics. Progress in this field heavily depends on continuous training and education. The review also highlights the emergence of innovative MS technologies and methodologies. MS has the potential to transform our understanding of metabolic landscapes, which is crucial for disease biomarker discovery. This article serves as an invaluable resource for researchers in metabolomics, presenting fresh perspectives and advancements that propels the field forward.
- Book Chapter
344
- 10.1007/978-1-4939-1258-2_1
- Jan 1, 2014
The field of metabolomics has witnessed an exponential growth in the last decade driven by important applications spanning a wide range of areas in the basic and life sciences and beyond. Mass spectrometry in combination with chromatography and nuclear magnetic resonance are the two major analytical avenues for the analysis of metabolic species in complex biological mixtures. Owing to its inherent significantly higher sensitivity and fast data acquisition, MS plays an increasingly dominant role in the metabolomics field. Propelled by the need to develop simple methods to diagnose and manage the numerous and widespread human diseases, mass spectrometry has witnessed tremendous growth with advances in instrumentation, experimental methods, software, and databases. In response, the metabolomics field has moved far beyond qualitative methods and simple pattern recognition approaches to a range of global and targeted quantitative approaches that are now routinely used and provide reliable data, which instill greater confidence in the derived inferences. Powerful isotope labeling and tracing methods have become very popular. The newly emerging ambient ionization techniques such as desorption ionization and rapid evaporative ionization have allowed direct MS analysis in real time, as well as new MS imaging approaches. While the MS-based metabolomics has provided insights into metabolic pathways and fluxes, and metabolite biomarkers associated with numerous diseases, the increasing realization of the extremely high complexity of biological mixtures underscores numerous challenges including unknown metabolite identification, biomarker validation, and interlaboratory reproducibility that need to be dealt with for realization of the full potential of MS-based metabolomics. This chapter provides a glimpse at the current status of the mass spectrometry-based metabolomics field highlighting the opportunities and challenges.
- Research Article
297
- 10.1016/s0031-9422(02)00703-3
- Feb 8, 2003
- Phytochemistry
Construction and application of a mass spectral and retention time index database generated from plant GC/EI-TOF-MS metabolite profiles
- Research Article
1
- 10.1016/j.aca.2005.04.066
- May 23, 2005
- Analytica Chimica Acta
Estimating the degree of similarity overlap for structural data in spectral databases
- Research Article
5
- 10.3389/ftox.2023.1216802
- Oct 16, 2023
- Frontiers in Toxicology
Introduction: The positive identification of xenobiotics and their metabolites in human biosamples is an integral aspect of exposomics research, yet challenges in compound annotation and identification continue to limit the feasibility of comprehensive identification of total chemical exposure. Nonetheless, the adoption of in silico tools such as metabolite prediction software, QSAR-ready structural conversion workflows, and molecular standards databases can aid in identifying novel compounds in untargeted mass spectral investigations, permitting the assessment of a more expansive pool of compounds for human health hazard. This strategy is particularly applicable when it comes to flame retardant chemicals. The population is ubiquitously exposed to flame retardants, and evidence implicates some of these compounds as developmental neurotoxicants, endocrine disruptors, reproductive toxicants, immunotoxicants, and carcinogens. However, many flame retardants are poorly characterized, have not been linked to a definitive mode of toxic action, and are known to share metabolic breakdown products which may themselves harbor toxicity. As U.S. regulatory bodies begin to pursue a subclass- based risk assessment of organohalogen flame retardants, little consideration has been paid to the role of potentially toxic metabolites, or to expanding the identification of parent flame retardants and their metabolic breakdown products in human biosamples to better inform the human health hazards imposed by these compounds. Methods: The purpose of this study is to utilize publicly available in silico tools to 1) characterize the structural and metabolic fates of proposed flame retardant classes, 2) predict first pass metabolites, 3) ascertain whether metabolic products segregate among parent flame retardant classification patterns, and 4) assess the existing coverage in of these compounds in mass spectral database. Results: We found that flame retardant classes as currently defined by the National Academies of Science, Engineering and Medicine (NASEM) are structurally diverse, with highly variable predicted pharmacokinetic properties and metabolic fates among member compounds. The vast majority of flame retardants (96%) and their predicted metabolites (99%) are not present in spectral databases, posing a challenge for identifying these compounds in human biosamples. However, we also demonstrate the utility of publicly available in silico methods in generating a fit for purpose synthetic spectral library for flame retardants and their metabolites that have yet to be identified in human biosamples. Discussion: In conclusion, exposomics studies making use of fit-for-purpose synthetic spectral databases will better resolve internal exposure and windows of vulnerability associated with complex exposures to flame retardant chemicals and perturbed neurodevelopmental, reproductive, and other associated apical human health impacts.
- Research Article
- 10.3389/fmed.2026.1791030
- Apr 22, 2026
- Frontiers in medicine
Current multiple myeloma (MM) risk stratification, anchored on the Revised International Staging System (R-ISS), provides a static snapshot of disease but fails to capture its dynamic biological evolution, functional tumor-microenvironment crosstalk, and real-time treatment response. Mass spectrometry (MS)-based proteomic and metabolomic profiling has emerged as a high-sensitivity tool for both novel biomarker discovery and minimal residual disease (MRD) monitoring. This systematic review evaluates the independent prognostic value and clinical utility of quantitative MS-based proteomics and metabolomics compared to standard-of-care risk models and traditional disease monitoring techniques. Following PRISMA 2020 guidelines, a systematic search of PubMed, Embase, and Web of Science was conducted. Inclusion required quantitative MS-based proteomics or metabolomics in MM cohorts with outcomes compared to ISS/R-ISS or traditional MRD detection methods. Data analysis was performed with a focus on overall survival (OS), progression-free survival (PFS), hazard ratios (HR), and MRD sensitivity thresholds. From 1,077 records, 19 studies met the inclusion criteria. Eleven discovery-focused studies identified specific MS-derived signatures, such as microenvironmental proteins (e.g., MTA2, CD44) and dysregulated lipid metabolites (e.g., acylcarnitines, LysoPE) that were consistently associated with PFS and OS. Crucially, MS-based biomarkers retained independent prognostic significance in multivariate models adjusted for R-ISS. Furthermore, 8 studies tackling blood-based MS-MRD detection demonstrated up to 1,000-fold higher sensitivity than traditional immunofixation electrophoresis, identified biochemical relapse 2-11 months earlier, and achieved high concordance with bone marrow-based assays (NGS/NGF). In conclusion, quantitative MS profiling provides a high-resolution molecular lens that significantly refines MM risk stratification beyond static staging. By enabling non-invasive, longitudinal MRD monitoring with superior lead times, MS integration facilitates a shift from reactive to proactive intervention. Standardization of bioinformatics pipelines and MS methodologies is now the final barrier to implementing MS-guided treatment adjustments in routine clinical practice.
- Research Article
11
- 10.1053/j.gastro.2006.09.025
- Nov 1, 2006
- Gastroenterology
Unraveling the Complex Proteome for Biomarker Discovery in Gastrointestinal and Liver Diseases
- Research Article
42
- 10.1002/pca.2763
- Apr 23, 2018
- Phytochemical Analysis
Medicinal plants are gaining increasing attention worldwide due to their empirical therapeutic efficacy and being a huge natural compound pool for new drug discovery and development. The efficacy, safety and quality of medicinal plants are the main concerns, which are highly dependent on the comprehensive analysis of chemical components in the medicinal plants. With the advances in mass spectrometry (MS) techniques, comprehensive analysis and fast identification of complex phytochemical components have become feasible, and may meet the needs, for the analysis of medicinal plants. Our aim is to provide an overview on the latest developments in MS and its hyphenated technique and their applications for the comprehensive analysis of medicinal plants. Application of various MS and its hyphenated techniques for the analysis of medicinal plants, including but not limited to one-dimensional chromatography, multiple-dimensional chromatography coupled to MS, ambient ionisation MS, and mass spectral database, have been reviewed and compared in this work. Recent advancs in MS and its hyphenated techniques have made MS one of the most powerful tools for the analysis of complex extracts from medicinal plants due to its excellent separation and identification ability, high sensitivity and resolution, and wide detection dynamic range. To achieve high-throughput or multi-dimensional analysis of medicinal plants, the state-of-the-art MS and its hyphenated techniques have played, and will continue to play a great role in being the major platform for their further research in order to obtain insight into both their empirical therapeutic efficacy and quality control.
- Research Article
- 10.1145/1217728.1217739
- Dec 1, 2006
- XRDS: Crossroads, The ACM Magazine for Students
Research in proteomics has created two significant needs: the need for an accurate public database of empirically derived mass spectrum information and the need for managing the I/O and organization of mass spectrometry data in the form of files and structures. Lack of an empirically derived database limits the ability of proteomic researchers to identify and study proteins. Managing the I/O and organization of mass spectrometry data is often time-consuming due to the many fields that need to be set and retrieved. As a result, incompatibilities and inefficiencies are created by each programmer handling this in his or her own way. Until recently, storage space and computing power has been the limiting factor in developing tools to handle the vast amount of mass spectrometry information. Now the resources are available to store, organize, and analyze mass spectrometry information.The Illinois Bio-Grid Mass Spectrometry Database is a database of empirically derived tandem mass spectra of peptides created to provide researchers with an organized and searchable database of curated spectrum information to allow more accurate protein identification. The Mass Spectrometry I/O Project creates a framework that handles mass spectrometry data I/O and data organization, allowing researchers to concentrate on data analysis rather than I/O. In addition, the Mass Spectrometry I/O Project leverages several cross-platform and portability-enhancing technologies, allowing it to be utilized on a variety of hardware and operating systems.
- Research Article
4
- 10.1515/psr-2018-0126
- Sep 25, 2019
- Physical Sciences Reviews
One major challenge in natural product (NP) discovery is the determination of the chemical structure of unknown metabolites using automated software tools from either GC–mass spectrometry (MS) or liquid chromatography–MS/MS data only. This chapter reviews the existing spectral libraries and predictive computational tools used in MS-based untargeted metabolomics, which is currently a hot topic in NP structure elucidation. We begin by focusing on spectral databases and the general workflow of MS annotation. We then describe software and tools used in MS, particularly those used to predict fragmentation patterns, mass spectral classifiers, and tools for fragmentation trees analysis. We then round up the chapter by looking at more advanced approaches implemented in tools for competitive fragmentation modeling and quantum chemical approaches.
- Book Chapter
9
- 10.1016/b978-0-444-64057-4.00007-7
- Jan 1, 2018
- Studies in Natural Products Chemistry
Chapter 7 - Isolation and Characterization of Phenolic Compounds From Selected Foods of Plant Origin Using Modern Spectroscopic Approaches