PEMA v2: addressing metabarcoding bioinformatics analysis challenges

Haris Zafeiropoulos,Christina Pavloudi,Evangelos Pafilis

doi:10.3897/aca.4.e64902

Abstract

Environmental DNA (eDNA) and metabarcoding have launched a new era in bio- and eco-assessment over the last years (Ruppert et al. 2019). The simultaneous identification, at the lowest taxonomic level possible, of a mixture of taxa from a great range of samples is now feasible; thus, the number of eDNA metabarcoding studies has increased radically (Deiner and 2017). While the experimental part of eDNA metabarcoding can be rather challenging depending on the special characteristics of the different studies, computational issues are considered to be its major bottlenecks. Among the latter, the bioinformatics analysis of metabarcoding data and especially the taxonomy assignment of the sequences are fundamental challenges. Many steps are required to obtain taxonomically assigned matrices from raw data. For most of these, a plethora of tools are available. However, each tool's execution parameters need to be tailored to reflect each experiment's idiosyncrasy; thus, tuning bioinformatics analysis has proved itself fundamental (Kamenova 2020). The computation capacity of high-performance computing systems (HPC) is frequently required for such analyses. On top of that, the non perfect completeness and correctness of the reference taxonomy databases is another important issue (Loos et al. 2020). Based on third-party tools, we have developed the Pipeline for Environmental Metabarcoding Analysis (PEMA), a HPC-centered, containerized assembly of key metabarcoding analysis tools. PEMA combines state-of-the art technologies and algorithms with an easy to get-set-use framework, allowing researchers to tune thoroughly each study thanks to roll-back checkpoints and on-demand partial pipeline execution features (Zafeiropoulos 2020). Once PEMA was released, there were two main pitfalls soon to be highlighted by users. PEMA supported 4 marker genes and was bounded by specific reference databases. In this new version of PEMA the analysis of any marker gene is now available since a new feature was added, allowing classifiers to train a user-provided reference database and use it for taxonomic assignment. Fig. 1 shows the taxonomy assignment related PEMA modules; all those out of the dashed box have been developed for this new PEMA release. As shown, the RDPClassifier has been trained with Midori reference 2 and has been added as an option, classifying not only metazoans but sequences from all taxonomic groups of Eukaryotes for the case of the COI marker gene. A PEMA documentation site is now also available. PEMA.v2 containers are available via the DockerHub and SingularityHub as well as through the Elixir Greece AAI Service. It has also been selected to be part of the LifeWatch ERIC Internal Joint Initiative for the analysis of ARMS data and soon will be available through the Tesseract VRE.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

PEMA v2: addressing metabarcoding bioinformatics analysis challenges

Abstract

Talk to us

Similar Papers

More From: ARPHA Conference Abstracts

Lead the way for us

Journal: ARPHA Conference Abstracts	Publication Date: Mar 4, 2021
License type: CC BY 4.0

Similar Papers

PEMA: a flexible Pipeline for Environmental DNA Metabarcoding Analysis of the 16S/18S ribosomal RNA, ITS, and COI marker genes.
Haris Zafeiropoulos ... Antonis Potirakis
GigaScience | VOL. 9
Haris Zafeiropoulos, et. al.Haris Zafeiropoulos ... Antonis Potirakis
01 Mar 2020
GigaScience | VOL. 9

EnsembleTax: an R package for determinations of ensemble taxonomic assignments of phylogenetically-informative marker gene sequences.
Dylan Catlett ... Kevin Son
PeerJ | VOL. 9
Dylan Catlett, et. al.Dylan Catlett ... Kevin Son
26 Jul 2021
PeerJ | VOL. 9

Direct-to-consumer raw genetic data and third-party interpretation services: more burden than bargain?
Tia Moscarello ... Erin Demo
Genetics in Medicine | VOL. 21
Tia Moscarello, et. al.Tia Moscarello ... Erin Demo
01 Mar 2019
Genetics in Medicine | VOL. 21

DNA metabarcoding of littoral hard-bottom communities: high diversity and database gaps revealed by two molecular markers.
Owen S Wangensteen ... Magdalena Guardiola
PeerJ | VOL. 6
Owen S Wangensteen, et. al.Owen S Wangensteen ... Magdalena Guardiola
04 May 2018
PeerJ | VOL. 6

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

PEMA v2: addressing metabarcoding bioinformatics analysis challenges

Abstract

Talk to us

Similar Papers

More From: ARPHA Conference Abstracts