Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

NCBI Taxonomy: a comprehensive update on curation, resources and tools.

  • Abstract
  • Highlights & Summary
  • PDF
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

The National Center for Biotechnology Information (NCBI) Taxonomy includes organism names and classifications for every sequence in the nucleotide and protein sequence databases of the International Nucleotide Sequence Database Collaboration. Since the last review of this resource in 2012, it has undergone several improvements. Most notable is the shift from a single SQL database to a series of linked databases tied to a framework of data called NameBank. This means that relations among data elements can be adjusted in more detail, resulting in expanded annotation of synonyms, the ability to flag names with specific nomenclatural properties, enhanced tracking of publications tied to names and improved annotation of scientific authorities and types. Additionally, practices utilized by NCBI Taxonomy curators specific to major taxonomic groups are described, terms peculiar to NCBI Taxonomy are explained, external resources are acknowledged and updates to tools and other resources are documented. Database URL: https://www.ncbi.nlm.nih.gov/taxonomy.

Similar Papers
  • Research Article
  • 10.1172/jci26755
Nucleic acid sequence data turns 100,000,000,000 and looks to the future
  • Oct 1, 2005
  • Journal of Clinical Investigation
  • S Bloom

The 3 members of the International Nucleotide Sequence Database Collaboration (INSDC) — the European Molecular Biology Laboratory (EMBL) Bank, GenBank, and the DNA Data Bank of Japan (DDBJ) — have reached a milestone. Owing in large part to their daily exchange policies, these public databases for DNA and RNA sequences have reached 100 gigabases of information. These 100,000,000,000 bases of genetic code, collected since 1982, comprise over 55 million sequence entries from more than 200,000 different organisms. This collaborative effort ensures that information gleaned from molecular biology and genetic research is placed in the public domain where the scientific community can use it to push science forward. “Today’s nucleotide sequence databases allow researchers to share completed genomes, the genetic makeup of entire ecosystems, and sequences associated with patents,” said David Lipman, director of the National Center for Biotechnology Information. “The INSDC has realized the vision of the researchers who initiated the sequence database projects by making the global sharing of nucleotide sequence information possible.” The repositories got started in the 1970s, when researchers suggested a public storehouse be made available for the massive amounts of genetic code sequence information that were being generated. Two of the databases – the EMBL Data Library and GenBank – were launched in the early 1980s. The European Bioinformatics Institute (EBI) manages the EMBL database, while GenBank is the NIH’s National Center for Biotechnology Information genetic sequence database. Both EMBL and GenBank offer an annotated collection of all publicly available DNA sequences, and they were formed as nonprofit entities that collaborated from the beginning. By 1987, the INSDC was formed and included a third collaborator — DDBJ, launched at the National Institute of Genetics in Mishima. DDBJ is also an international nucleotide sequence database freely accessible online. Early on, staffers searched published journal articles for sequence data and entered it manually into the repository. But times have changed, and the modern sequencing centers at universities today have come a long way. New automated technology, robotics, and bioinformatics, combined with decreased cost, have fostered faster data collection. “The technology has come so fast that it blows my mind,” said Richard Wilson, director of the Genome Sequencing Center at Washington University School of Medicine. Wilson explained that in the mid-1980s his center was able to turn out several hundred bases of sequence per month, while today they are generating about 4.2 billion. A boost to the number of collected sequences is also due to the National Human Genome Research Institute (NHGRI) at the NIH, which now provides genome-sequencing grants. These funds support research aimed at sequencing a human-sized genome at a cost 100 times lower than is possible now — it presently costs nearly 10 million dollars to sequence the 3 billion base pairs of DNA found in humans. The immediate goal of the NHGRI is to lower the cost of these projects by tens of thousands of dollars in order to allow scientists to sequence genomes of human subjects involved in studies to find genes relevant for disease. The longer-term NHGRI goal is to reduce whole-genome sequencing to only 1,000 dollars so that this process can be used in routine medical tests and allow physicians to tailor diagnosis, prevention, and treatment to a patient’s individual genetic makeup. So far, the genetic information available has taught us a lot about the evolutionary relationships among different species. But according to Richard Gibbs, director of the Human Genome Sequencing Center at Baylor College of Medicine, “This will pale into insignificance once we find the full repertoire of alleles that cause different human genetic disease.” These are the tools, he explained, that would unlock our understanding of gene function and ultimately provide diagnostic and prognostic indicators. The benefit to human health will be immeasurable.

  • Research Article
  • Cite Count Icon 30
  • 10.1111/nph.13180
Towards the unification of sequence-based classification and sequence-based identification of host-associated microorganisms.
  • Nov 26, 2014
  • New Phytologist
  • Joshua R Herr + 2 more

Towards the unification of sequence-based classification and sequence-based identification of host-associated microorganisms.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 3
  • 10.3897/biss.6.93595
Enabling Community Curation of Biological Source Annotations of Molecular Data Through PlutoF and the ELIXIR Contextual Data Clearinghouse
  • Aug 23, 2022
  • Biodiversity Information Science and Standards
  • Vishnukumar Balavenkataraman Kadhirvelu + 11 more

The advancements in sequencing technologies have greatly contributed to the documentation of Earth’s biodiversity. However, for exploring the full potential of molecular resources for biodiversity, there needs to be a good linkage between sequence data and its biological source, contributing to a network of connected data in the biodiversity research cycle. This requires a foundation of well-structured and accessible annotations in the molecular sequence repositories. The International Nucleotide Sequence Database Collaboration (INSDC), of which the European Nucleotide Archive (ENA) is its European node, holds a large amount of annotations associated with sequence data, relating to its biological source (e.g., specimens in natural history collections). However, for a number of records, these annotations may be incomplete (e.g., missing voucher information), ambiguous or even inaccurate. Therefore, we have implemented a workflow that allows third-party annotations to be attached to sequence and sample records using two existing services, the PlutoF platform and the ELIXIR Contextual Data ClearingHouse. This work was developed within the scope of the BiCIKL (Biodiversity Community Integrated Knowledge Library) project, which aims to establish open science practices in the biodiversity domain. PlutoF is an online data management platform that also provides computing services for biology-related research. PlutoF features allow registered users to enter their own data and access public data at INSDC. Users can enter and manage a range of data, as taxonomic classifications, occurrences, etc. This platform also includes a module that allows the addition of third-party annotations (on material source, taxonomic identification, etc.) linked to specimens or sequence records. This module was already in use by the UNITE community for annotation of INSDC rDNA Internal Transcribed Spacer sequence datasets (Abarenkov et al. 2021). These UNITE annotations are displayed in the National Centre for Biotechnology Information (NCBI) records through links to the PlutoF platform. However, there was the need for an automated solution that allowed third-party annotations to any sequence or sample record at INSDC. This was implemented through the operation of the ELIXIR Contextual Data ClearingHouse (hereafter as Clearinghouse). The Clearinghouse holds a simple RESTful Application Programming Interface (API) to support the submission of additions and improvements to current metadata attributes, such as information on material sources, on records publicly available in the ELIXIR data resources. The Clearinghouse enables the submission of these corrected metadata from databases (such as the PlutoF platform) to the primary data repositories. The workflow developed is shown in Fig. 1 and consists of the following steps: i) users annotate sequence metadata that is regularly downloaded from INSDC using NCBI’s E-utilities; ii) an annotation proposal is created and a verification notification is sent to an assigned reviewer; iii) the reviewer evaluates the annotation proposal and accepts it or rejects it with comments; iv) if the annotation proposal is accepted, the annotated fields that may be mapped to ENA fields are then pushed to the Clearinghouse using their RESTful API. The annotations when received at ENA are then reviewed before being displayed. This workflow is implemented through a web interface in PlutoF, which allows user-friendly and effortless reporting of corrections or additions to biological source metadata in sequence records. Overall, we expect this tool to contribute to the enrichment of metadata associated with sequence records, and therefore increase the links between the molecular and biodiversity resources, and enable sequencing data to deliver their full potential for biodiversity conservation.

  • Research Article
  • Cite Count Icon 1500
  • 10.1093/nar/gkr1178
The NCBI Taxonomy database
  • Dec 1, 2011
  • Nucleic Acids Research
  • S Federhen

The NCBI Taxonomy database (http://www.ncbi.nlm.nih.gov/taxonomy) is the standard nomenclature and classification repository for the International Nucleotide Sequence Database Collaboration (INSDC), comprising the GenBank, ENA (EMBL) and DDBJ databases. It includes organism names and taxonomic lineages for each of the sequences represented in the INSDC’s nucleotide and protein sequence databases. The taxonomy database is manually curated by a small group of scientists at the NCBI who use the current taxonomic literature to maintain a phylogenetic taxonomy for the source organisms represented in the sequence databases. The taxonomy database is a central organizing hub for many of the resources at the NCBI, and provides a means for clustering elements within other domains of NCBI web site, for internal linking between domains of the Entrez system and for linking out to taxon-specific external resources on the web. Our primary purpose is to index the domain of sequences as conveniently as possible for our user community.

  • Research Article
  • Cite Count Icon 31
  • 10.1126/science.aaf7672
Reminder to deposit DNA sequences
  • May 11, 2016
  • Science
  • Mark Blaxter + 8 more

As members of the Advisory Committee to the International Nucleotide Sequence Database Collaboration (INSDC), which includes the DNA Data Bank of Japan (DDBJ), ENA, and GenBank databases, we wish to remind the research community of the importance of depositing complete DNA-sequence data in these databases upon publication of their results [see also S. L. Salzberg et al., Nature , (2016)]. Indeed, most journals demand a database accession number as a condition of publication. Access to the INSDC's databases is free and unrestricted ([ 1 ][1]), enabling researchers to plan experiments and to analyze existing data. As original contributions, deposited data form part of the scientific record and are citable in the literature. Authors can also correct and update their data. These amended records may be removed from the next database release but still remain permanently available by accession number. INSDC has also created major new repositories for large data collections, notably the Sequence Read Archive at the National Center for Biotechnology Information (NCBI) ([ 2 ][2]), the DDBJ Sequence Read Archive ([ 3 ][3]), and the European Nucleotide Archive at the European Molecular Biology Laboratory's European Bioinformatics Institute (EMBL-EBI) ([ 4 ][4]). These archive raw data from sequencing experiments, a crucial facility for reproducibility and reuse. For papers dependent on sequence data from human subjects, unrestricted data release may not be possible. In these cases, we would encourage journal editors to insist on data sharing through other repositories that are not part of INSDC, such as NCBI's Database of Genotype and Phenotype ([ 5 ][5]), EMBL-EBI's European Genome-phenome Archive ([ 6 ][6]), or DDBJ's Japanese Genotype-phenotype Archive ([ 7 ][7]). ![Figure][8] IMAGE: BEHOLDINGEYE/[ISTOCKPHOTO.COM][9] 1. [↵][10]1. G. Cochrane, 2. I. Karsch-Mizrachi, 3. T. Takagi , International Nucleotide Sequence Database Collaboration, Nucleic Acids Res 44, D48 (2016). [OpenUrl][11] 2. [↵][12][www.ncbi.nlm.nih.gov/sra][13]. 3. [↵][14] . 4. [↵][15][www.ebi.ac.uk/ena][16]. 5. [↵][17][www.ncbi.nlm.nih.gov/gap][18]. 6. [↵][19][www.ebi.ac.uk/ega/home][20]. 7. [↵][21] . [1]: #ref-1 [2]: #ref-2 [3]: #ref-3 [4]: #ref-4 [5]: #ref-5 [6]: #ref-6 [7]: #ref-7 [8]: pending:yes [9]: http://ISTOCKPHOTO.COM [10]: #xref-ref-1-1 View reference 1 in text [11]: {openurl}?query=rft.jtitle%253DInternational%2BNucleotide%2BSequence%2BDatabase%2BCollaboration%26rft.volume%253D44%26rft.spage%253DD48%26rft.genre%253Darticle%26rft_val_fmt%253Dinfo%253Aofi%252Ffmt%253Akev%253Amtx%253Ajournal%26ctx_ver%253DZ39.88-2004%26url_ver%253DZ39.88-2004%26url_ctx_fmt%253Dinfo%253Aofi%252Ffmt%253Akev%253Amtx%253Actx [12]: #xref-ref-2-1 View reference 2 in text [13]: http://www.ncbi.nlm.nih.gov/sra [14]: #xref-ref-3-1 View reference 3 in text [15]: #xref-ref-4-1 View reference 4 in text [16]: http://www.ebi.ac.uk/ena [17]: #xref-ref-5-1 View reference 5 in text [18]: http://www.ncbi.nlm.nih.gov/gap [19]: #xref-ref-6-1 View reference 6 in text [20]: http://www.ebi.ac.uk/ega/home [21]: #xref-ref-7-1 View reference 7 in text

  • Research Article
  • Cite Count Icon 39
  • 10.1093/nar/gks1152
DDBJ new system and service refactoring
  • Nov 23, 2012
  • Nucleic Acids Research
  • Osamu Ogasawara + 6 more

The DNA data bank of Japan (DDBJ, http://www.ddbj.nig.ac.jp) maintains a primary nucleotide sequence database and provides analytical resources for biological information to researchers. This database content is exchanged with the US National Center for Biotechnology Information (NCBI) and the European Bioinformatics Institute (EBI) within the framework of the International Nucleotide Sequence Database Collaboration (INSDC). Resources provided by the DDBJ include traditional nucleotide sequence data released in the form of 27 316 452 entries or 16 876 791 557 base pairs (as of June 2012), and raw reads of new generation sequencers in the sequence read archive (SRA). A Japanese researcher published his own genome sequence via DDBJ-SRA on 31 July 2012. To cope with the ongoing genomic data deluge, in March 2012, our computer previous system was totally replaced by a commodity cluster-based system that boasts 122.5 TFlops of CPU capacity and 5 PB of storage space. During this upgrade, it was considered crucial to replace and refactor substantial portions of the DDBJ software systems as well. As a result of the replacement process, which took more than 2 years to perform, we have achieved significant improvements in system performance.

  • PDF Download Icon
  • Research Article
  • 10.3897/biss.6.91118
The ENA Source Attribute Helper: An API for improved biological source data
  • Aug 2, 2022
  • Biodiversity Information Science and Standards
  • Vikas Gupta + 4 more

Metadata management for sequence data is essential for the accurate description of Earth’s biodiversity. Within metadata attributes, those that reference the biological sources of sequences and samples and allow linking to the specimen or sample of origin are fundamental for facilitating connections between molecular biology, taxonomy, systematic biology and biodiversity research, increasing the discoverability and usability of data by researchers worldwide. Sequence data is publicly archived at the International Nucleotide Sequence Database Collaboration (INSDC) that includes the National Centre for Biotechnology Information (NCBI), the DNA Data Bank of Japan (DDBJ) and the European Nucleotide Archive (ENA). Sequences stored at INSDC have associated a considerable range of metadata, including attributes related to its biological source, such as references to natural history collections or culture collections. But, these source attributes are not always submitted or may be incomplete, limiting the association of the sequence records to the original source material, hampering further data connections (e.g., biological data associated with the voucher or species distribution data). Therefore, we have developed the ENA Source Attribute Helper API, a tool that aims to assist users on the submission of accurate attributes referring to the biological source of samples and sequence data. This tool was developed within the scope of BiCIKL (Biodiversity Community Integrated Knowledge Library) (Penev et al. 2022), a Horizon 2020 project which targets building a wide, biodiversity related community for connecting data along the different axes of biodiversity research. The first version of the tool was designed to support correct annotation of the attributes that identify the source material from which the sample or sequence were obtained, namely /specimen_voucher, /culture_collection, and /biomaterial (INSDC 2021). These attributes follow a Darwin Core Triplet format (Wieczorek et al. 2012), composed of institution code, collection code and the specimen, culture, or material identifier, accordingly. Since the submission of the biological source attributes to the INSDC may be performed both when data is initially uploaded or on following updates using a variety of tools, we developed the API as an open source tool that is publicly accessible and may be used as a free-standing service. The API is built using Representational State Transfer (REST) API Architecture and it is designed to use the data available in the NCBI BioCollections (Sharma et al. 2018). NCBI Biocollections is a curated database of metadata for natural history collections, associated with records in INSDC, that includes the institution and collection codes. The API main functions include the querying of the metadata (the API presents both exact matches and similar matches) for the institutions and collections based on the user input, validation of institution and collection codes in the attribute strings provided by the user, and the construction of the attribute string based on the user-provided information. The API does not include the search or validation of the voucher specimen codes. The API is designed in a way that it can be extended easily for any future enhancements and initially expected to promote and support the submission and any subsequent curation of better structured and more richly described source data. We expect this tool to contribute to better connected biodiversity data and hence provide a stronger foundation to strengthen the value of natural history collections, taxonomic expertise, and biodiversity knowledge.

  • Research Article
  • Cite Count Icon 195
  • 10.1093/nar/gkx1097
The international nucleotide sequence database collaboration
  • Nov 28, 2017
  • Nucleic Acids Research
  • Ilene Karsch-Mizrachi + 2 more

For more than 30 years, the International Nucleotide Sequence Database Collaboration (INSDC; http://www.insdc.org/) has been committed to capturing, preserving and providing access to comprehensive public domain nucleotide sequence and associated metadata which enables discovery in biomedicine, biodiversity and biological sciences. Since 1987, the DNA Data Bank of Japan (DDBJ) at the National Institute for Genetics in Mishima, Japan; the European Nucleotide Archive (ENA) at the European Molecular Biology Laboratory's European Bioinformatics Institute (EMBL-EBI) in Hinxton, UK; and GenBank at National Center for Biotechnology Information (NCBI), National Library of Medicine, National Institutes of Health in Bethesda, Maryland, USA have worked collaboratively to enable access to nucleotide sequence data in standardized formats for the worldwide scientific community. In this article, we reiterate the principles of the INSDC collaboration and briefly summarize the trends of the archival content.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 59
  • 10.1093/nar/gkz982
DDBJ Database updates and computational infrastructure enhancement.
  • Nov 14, 2019
  • Nucleic acids research
  • Osamu Ogasawara + 4 more

The Bioinformation and DDBJ Center (https://www.ddbj.nig.ac.jp) in the National Institute of Genetics (NIG) maintains a primary nucleotide sequence database as a member of the International Nucleotide Sequence Database Collaboration (INSDC) in partnership with the US National Center for Biotechnology Information and the European Bioinformatics Institute. The NIG operates the NIG supercomputer as a computational basis for the construction of DDBJ databases and as a large-scale computational resource for Japanese biologists and medical researchers. In order to accommodate the rapidly growing amount of deoxyribonucleic acid (DNA) nucleotide sequence data, NIG replaced its supercomputer system, which is designed for big data analysis of genome data, in early 2019. The new system is equipped with 30 PB of DNA data archiving storage; large-scale parallel distributed file systems (13.8 PB in total) and 1.1 PFLOPS computation nodes and graphics processing units (GPUs). Moreover, as a starting point of developing multi-cloud infrastructure of bioinformatics, we have also installed an automatic file transfer system that allows users to prevent data lock-in and to achieve cost/performance balance by exploiting the most suitable environment from among the supercomputer and public clouds for different workloads.

  • Research Article
  • Cite Count Icon 24
  • 10.1093/sysbio/syad068
Improving the gold standard in NCBI GenBank and related databases: DNA sequences from type specimens and type strains.
  • Nov 13, 2023
  • Systematic biology
  • Susanne S Renner + 4 more

Scientific names permit humans and search engines to access knowledge about the biodiversity that surrounds us, and names linked to DNA sequences are playing an ever-greater role in search-and-match identification procedures. Here, we analyze how users and curators of the National Center for Biotechnology Information (NCBI) are flagging and curating sequences derived from nomenclatural type material, which is the only way to improve the quality of DNA-based identification in the long run. For prokaryotes, 18,281 genome assemblies from type strains have been curated by NCBI staff and improve the quality of prokaryote naming. For Fungi, type-derived sequences representing over 21,000 species are now essential for fungus naming and identification. For the remaining eukaryotes, however, the numbers of sequences identifiable as type-derived are minuscule, representing only 739 species of arthropods, 1542 vertebrates, and 125 embryophytes. An increase in the production and curation of such sequences will come from (i) sequencing of types or topotypic specimens in museum collections, (ii) the March 2023 rule changes at the International Nucleotide Sequence Database Collaboration requiring more metadata for specimens, and (iii) efforts by data submitters to facilitate curation, including informing NCBI curators about a specimen's type status. We illustrate different type-data submission journeys and provide best-practice examples from a range of organisms. Expanding the number of type-derived sequences in DNA databases, especially of eukaryotes, is crucial for capturing, documenting, and protecting biodiversity.

  • Research Article
  • Cite Count Icon 250
  • 10.1093/nar/gkaa967
The international nucleotide sequence database collaboration
  • Nov 9, 2020
  • Nucleic Acids Research
  • Masanori Arita + 2 more

The International Nucleotide Sequence Database Collaboration (INSDC; http://www.insdc.org/) has been the core infrastructure for collecting and providing nucleotide sequence data and metadata for >30 years. Three partner organizations, the DNA Data Bank of Japan (DDBJ) at the National Institute of Genetics in Mishima, Japan; the European Nucleotide Archive (ENA) at the European Molecular Biology Laboratory's European Bioinformatics Institute (EMBL-EBI) in Hinxton, UK; and GenBank at National Center for Biotechnology Information (NCBI), National Library of Medicine, National Institutes of Health in Bethesda, Maryland, USA have been collaboratively maintaining the INSDC for the benefit of not only science but all types of community worldwide.

  • Research Article
  • Cite Count Icon 958
  • 10.1093/nar/gkr854
The sequence read archive: explosive growth of sequencing data
  • Oct 18, 2011
  • Nucleic Acids Research
  • Yuichi Kodama + 2 more

New generation sequencing platforms are producing data with significantly higher throughput and lower cost. A portion of this capacity is devoted to individual and community scientific projects. As these projects reach publication, raw sequencing datasets are submitted into the primary next-generation sequence data archive, the Sequence Read Archive (SRA). Archiving experimental data is the key to the progress of reproducible science. The SRA was established as a public repository for next-generation sequence data as a part of the International Nucleotide Sequence Database Collaboration (INSDC). INSDC is composed of the National Center for Biotechnology Information (NCBI), the European Bioinformatics Institute (EBI) and the DNA Data Bank of Japan (DDBJ). The SRA is accessible at www.ncbi.nlm.nih.gov/sra from NCBI, at www.ebi.ac.uk/ena from EBI and at trace.ddbj.nig.ac.jp from DDBJ. In this article, we present the content and structure of the SRA and report on updated metadata structures, submission file formats and supported sequencing platforms. We also briefly outline our various responses to the challenge of explosive data growth.

  • Research Article
  • Cite Count Icon 6688
  • 10.1093/nar/gkv1189
Reference sequence (RefSeq) database at NCBI: current status, taxonomic expansion, and functional annotation.
  • Nov 8, 2015
  • Nucleic Acids Research
  • Nuala O’Leary + 53 more

The RefSeq project at the National Center for Biotechnology Information (NCBI) maintains and curates a publicly available database of annotated genomic, transcript, and protein sequence records (http://www.ncbi.nlm.nih.gov/refseq/). The RefSeq project leverages the data submitted to the International Nucleotide Sequence Database Collaboration (INSDC) against a combination of computation, manual curation, and collaboration to produce a standard set of stable, non-redundant reference sequences. The RefSeq project augments these reference sequences with current knowledge including publications, functional features and informative nomenclature. The database currently represents sequences from more than 55 000 organisms (>4800 viruses, >40 000 prokaryotes and >10 000 eukaryotes; RefSeq release 71), ranging from a single record to complete genomes. This paper summarizes the current status of the viral, prokaryotic, and eukaryotic branches of the RefSeq project, reports on improvements to data access and details efforts to further expand the taxonomic representation of the collection. We also highlight diverse functional curation initiatives that support multiple uses of RefSeq data including taxonomic validation, genome annotation, comparative genomics, and clinical testing. We summarize our approach to utilizing available RNA-Seq and other data types in our manual curation process for vertebrate, plant, and other species, and describe a new direction for prokaryotic genomes and protein name management.

  • Research Article
  • 10.3897/biss.8.137771
Aspects of NCBI GenBank as a Biodiversity Information Resource
  • Sep 24, 2024
  • Biodiversity Information Science and Standards
  • Takeru Nakazato

DNA sequencing of museum specimens, also known as museomics, provides new insights into the study of biodiversity, including taxonomy, phylogeny, and environmental studies. Also, sequencing specimens have led to the rediscovery of extinct species (Suzuki et al. 2016), identification of related species (Waku et al. 2016), and analysis of ancient DNA (Kanzawa-Kiriyama et al. 2016). Nucleotide sequence data have been collected for more than 30 years under the framework of the International Nucleotide Sequence Database Collaboration (INSDC) by three institutes, namely, National Center for Biotechnology Information, US (NCBI), European Bioinformatics Institute (EBI), and DNA Data Bank of Japan (DDBJ) (Arita et al. 2020). NCBI has collated a database of sequence data, GenBank, which contains approximately 494 million sequences as of April 2022 (Sayers et al. 2021). In fact, GenBank is designed with qualifiers to describe various types of biodiversity information such as "/specimen_voucher", "/lat_lon" (latitude and longitude) and "/collection_date". Also, INSDC now requires that all submissions include the sampling location and date (INSDC 2023). I surveyed the biodiversity information assigned to GenBank records to determine the potential of GenBank as a biodiversity resource. I downloaded all GenBank data as of August 2023 from the FTP site. The “/specimen_voucher” qualifier was introduced to describe specimen ID in Release 104 in December 1997. This qualifier was designed to fill the value in free text: for example, /specimen_voucher="Smith s. n. 4-IV-1995 (U. S. Natl. Herbarium)". After Release 162 in October 2007, a method of writing with a structured value of "[<institution-code>: [<collection-code>:]] <specimen_id>" was added (institution-code and collection-code are optional). There are 527,215 records (37.8%) with "/specimen_voucher" qualifier for fish, 3,096,112 records (40.3%) for insects, 1,505,556 records (39.0%) for flowering plants. But fewer than 10% of records have specimen IDs listed using this structured description. To utilize these ambiguous specimen IDs in GenBank, these IDs may need to be cleansed using databases such as NCBI BioCollections, GRSciColl (Global Registry of Scientific Collections) or AI to map them to IDs in databases rich in specimen information such as those of the Global Biodiveristy Information Facility (GBIF) and Barcode of Life System (BOLD). In GenBank, the BOLD ID is listed in the /db_xref qualifier in the “Features” field as the ID of the external database. The 70% of insect sequence data with a specimen ID in the /specimen_voucher qualifier are also assigned a BOLD ID (Nakazato and Jinbo 2022). The correspondence between specimen IDs in biodiversity information databases such as GBIF and specimen IDs in GenBank is expected to further enhance the value of museum specimens. In addition, GenBank provides the /type_material qualifier for describing the type of voucher (e.g., holotype of Asphondylia bursicola). In GenBank insect data, there were over 2,000 records for type material, and approximately 450 species were mentioned, including 269 for holotypes. We found approximately 3,000 records with type information by including “/notes” and “/specimens_voucher” qualifiers in addition to “/type_material”. Thus, GenBank has potential as a biodiversity information resource, but for more effective use, data mining and linkage with other specimen-based biodiversity databases are essential.

  • Research Article
  • Cite Count Icon 33
  • 10.1093/nar/gkae1058
The international nucleotide sequence database collaboration (INSDC): enhancing global participation.
  • Nov 13, 2024
  • Nucleic acids research
  • Ilene Karsch-Mizrachi + 7 more

The members of the International Nucleotide Sequence Database Collaboration (INSDC; https://insdc.org) have built systems to collect, archive and disseminate sequence data for more than four decades. The three collaborating organizations, the National Library of Medicine, National Center for Biotechnology Information (NLM-NCBI) in the UnitedStates, Research Organization of Information and Systems, National Institute of Genetics (ROIS-NIG) in Japan; and the European Molecular Biology Laboratory-European Bioinformatics Institute (EMBL-EBI) formalized their relationship through the adoption of an arrangement which documents their commitment to free and open access to genomic sequences. The INSDC is committed to expand the collaboration to be more representative of the global community of sequences and users. Diversifying participation through new membership will advance open science and data sharing and, in turn, drive innovation. This expansion will additionally benefit the INSDC and its broad user base by providing additional diverse perspectives as it explores emerging areas of data management, including federation, attribution and management.

Save Icon
Up Arrow
Open/Close
Setting-up Chat
Loading Interface