Modeling coral reef biological indicators using passive acoustic monitoring and machine learning.
This study develops a machine learning framework integrating passive acoustic monitoring data with ecological surveys to predict coral reef indicators such as fish abundance, species richness, and coral cover. Evaluated across ten sites, LightGBM achieved the highest accuracy, offering an efficient, scalable, non-invasive tool to support marine ecosystem management decisions.
This study systematically integrates acoustic methods and machine learning (ML) into marine ecosystem management, developing a comprehensive ML framework that combines passive acoustic monitoring (PAM) data with ecological survey observations to predict key coral reef ecological indicators, including fish abundance, fish species richness, and live coral cover. The framework extracts features from multiple acoustic frequency bands and deploys a complete ML workflow covering seven algorithms across three categories: tree-based models (Random Forest, LightGBM, Gradient Boosting), neural networks (Multilayer Perceptron, Recurrent Neural Networks, Bayesian Neural Networks), and an ensemble strategy (Voting Regressor). Evaluated on ten coral reef sites in Sanya, China, the framework was comprehensively compared in terms of predictive accuracy and computational efficiency. Results indicate that the LightGBM model achieves the highest predictive performance of these biological indicators, providing a more efficient, scalable, and non-invasive solution for marine fish monitoring, which can support decision-making in marine ecosystem management. The proposed machine learning-based framework has the potential to be integrated into decision-support tools for ecosystem management, enabling more efficient monitoring of coral reef ecosystems worldwide.
- Research Article
4
- 10.1111/ddi.13790
- Nov 22, 2023
- Diversity and Distributions
AimSpecies distribution models (SDMs) are essential tools in ecology and conservation. However, the scarcity of visual sightings of marine mammals in remote polar areas hinders the effective application of SDMs there. Passive acoustic monitoring (PAM) data provide year‐round information and overcome foul weather limitations faced by visual surveys. However, the use of PAM data in SDMs has been sparse so far. Here, we use PAM‐based SDMs to investigate the spatiotemporal distribution of the critically endangered Antarctic blue whale in the Weddell Sea.LocationThe Weddell Sea.MethodsWe used presence‐only dynamic SDMs employing visual sightings and PAM detections in independent models. We compared the two independent models with a third combined model that integrated both visual and PAM data, aiming at leveraging the advantages of each data type: the extensive spatial extent of visual data and the broader temporal/environmental range of PAM data.ResultsVisual and PAM data prove complementary, as indicated by a low spatial overlap between daily predictions and the low predictability of each model at detections of other data types. Combined data models reproduced suitable habitats as given by both independent models. Visual data models indicate areas close to the sea ice edge (SIE) and with low‐to‐moderate sea ice concentrations (SIC) as suitable, while PAM data models identified suitable habitats at a broader range of distances to SIE and relatively higher SIC.Main ConclusionsThe results demonstrate the potential of PAM data to predict year‐round marine mammal habitat suitability at large spatial scales. We provide reasons for discrepancies between SDMs based on either data type and give methodological recommendations on using PAM data in SDMs. Combining visual and PAM data in future SDMs is promising for studying vocalized animals, particularly when using recent advances in integrated distribution modelling methods.
- Research Article
11
- 10.7717/peerj.1459
- Dec 1, 2015
- PeerJ
Large disturbances can cause rapid degradation of coral reef communities, but what baseline changes in species assemblages occur on undisturbed reefs through time? We surveyed live coral cover, reef fish abundance and fish species richness in 1997 and again in 2007 on 47 fringing patch reefs of varying size and depth at Mersa Bareika, Ras Mohammed National Park, Egypt. No major human or natural disturbance event occurred between these two survey periods in this remote protected area. In the absence of large disturbances, we found that live coral cover, reef fish abundance and fish species richness did not differ in 1997 compared to 2007. Fish abundance and species richness on patches was largely related to the presence of shelters (caves and/or holes), live coral cover and patch size (volume). The presence of the ectoparasite-eating cleaner wrasse, Labroides dimidiatus, was also positively related to fish species richness. Our results underscore the importance of physical reef characteristics, such as patch size and shelter availability, in addition to biotic characteristics, such as live coral cover and cleaner wrasse abundance, in supporting reef fish species richness and abundance through time in a relatively undisturbed and understudied region.
- Conference Article
186
- 10.1145/3357223.3362711
- Nov 20, 2019
Machine learning (ML) workflows are extremely complex. The typical workflow consists of distinct stages of user interaction, such as preprocessing, training, and tuning, that are repeatedly executed by users but have heterogeneous computational requirements. This complexity makes it challenging for ML users to correctly provision and manage resources and, in practice, constitutes a significant burden that frequently causes over-provisioning and impairs user productivity. Serverless computing is a compelling model to address the resource management problem, in general, but there are numerous challenges to adopt it for existing ML frameworks due to significant restrictions on local resources. This work proposes Cirrus---an ML framework that automates the end-to-end management of datacenter resources for ML workflows by efficiently taking advantage of serverless infrastructures. Cirrus combines the simplicity of the serverless interface and the scalability of the serverless infrastructure (AWS Lambdas and S3) to minimize user effort. We show a design specialized for both serverless computation and iterative ML training is needed for robust and efficient ML training on serverless infrastructure. Our evaluation shows that Cirrus outperforms frameworks specialized along a single dimension: Cirrus is 100x faster than a general purpose serverless system [36] and 3.75x faster than specialized ML frameworks for traditional infrastructures [49].
- Research Article
15
- 10.1002/ajp.23599
- Jan 20, 2024
- American journal of primatology
The urgent need for effective wildlife monitoring solutions in the face of global biodiversity loss has resulted in the emergence of conservation technologies such as passive acoustic monitoring (PAM). While PAM has been extensively used for marine mammals, birds, and bats, its application to primates is limited. Black-and-white ruffed lemurs (Varecia variegata) are a promising species to test PAM with due to their distinctive and loud roar-shrieks. Furthermore, these lemurs are challenging to monitor via traditional methods due to their fragmented and often unpredictable distribution in Madagascar's dense eastern rainforests. Our goal in this study was to develop a machine learning pipeline for automated call detection from PAM data, compare the effectiveness of PAM versus in-person observations, and investigate diel patterns in lemur vocal behavior. We did this study at Mangevo, Ranomafana National Park by concurrently conducting focal follows and deploying autonomous recorders in May-July 2019. We used transfer learning to build a convolutional neural network (optimized for recall) that automated the detection of lemur calls (57-h runtime; recall = 0.94, F1 = 0.70). We found that PAM outperformed in-person observations, saving time, money, and labor while also providing re-analyzable data. Using PAM yielded novel insights into V. variegata diel vocal patterns; we present the first published evidence of nocturnal calling. We developed a graphic user interface and open-sourced data and code, to serve as a resource for primatologists interested in implementing PAM and machine learning. By leveraging the potential of this pipeline, we can address the urgent need for effective primate population surveys to inform conservation strategies.
- Conference Article
4
- 10.2118/208125-ms
- Dec 9, 2021
Estimation of petrophysical properties is essential for accurate reservoir predictions. In recent years, extensive work has been dedicated into training different machine-learning (ML) models to predict petrophysical properties of digital rock using dry rock images along with data from single-phase direct simulations, such as lattice Boltzmann method (LBM) and finite volume method (FVM). The objective of this paper is to present a comprehensive literature review on petrophysical properties estimation from dry rock images using different ML workflows and direct simulation methods. The review provides detailed comparison between different ML algorithms that have been used in the literature to estimate porosity, permeability, tortuosity, and effective diffusivity. In this paper, various ML workflows from the literature are screened and compared in terms of the training data set, the testing data set, the extracted features, the algorithms employed as well as their accuracy. A thorough description of the most commonly used algorithms is also provided to better understand the functionality of these algorithms to encode the relationship between the rock images and their respective petrophysical properties. The review of various ML workflows for estimating rock petrophysical properties from dry images shows that models trained using features extracted from the image (physics-informed models) outperformed models trained on the dry images directly. In addition, certain tree-based ML algorithms, such as random forest, gradient boosting, and extreme gradient boosting can produce accurate predictions that are comparable to deep learning algorithms such as deep neural networks (DNNs) and convolutional neural networks (CNNs). To the best of our knowledge, this is the first work dedicated to exploring and comparing between different ML frameworks that have recently been used to accurately and efficiently estimate rock petrophysical properties from images. This work will enable other researchers to have a broad understanding about the topic and help in developing new ML workflows or further modifying exiting ones in order to improve the characterization of rock properties. Also, this comparison represents a guide to understand the performance and applicability of different ML algorithms. Moreover, the review helps the researchers in this area to cope with digital innovations in porous media characterization in this fourth industrial age – oil and gas 4.0.
- Research Article
6
- 10.1145/3773084
- Oct 27, 2025
- ACM Transactions on Software Engineering and Methodology
Machine Learning (ML) workflows—spanning data preprocessing and feature engineering, model selection and hyperparameter optimization, and workflow evaluation—are increasingly embedded in complex software systems. Building these workflows manually demands substantial ML expertise, domain knowledge, and engineering effort. Automated ML (AutoML) frameworks address parts of this challenge but often suffer from constrained search spaces, limited adaptability, and low interpretability. Recent advances in Large Language Models (LLMs) have opened new opportunities to automate and enhance ML workflows by leveraging their capabilities in language understanding, reasoning, interaction, and code generation, posing new practical and theoretical challenges for software engineering (SE). This survey provides the first SE-oriented, stage-wise review of LLM-based ML workflow automation. We introduce a taxonomy covering all three workflow stages, systematically compare and analyze state-of-the-art methods, and synthesize both stage-specific and cross-stage trends. Our analysis yields SE-oriented implications, including the need for robust verification, quality management, context-aware deployment, and risk mitigation, alongside ensuring key quality attributes such as usability, modularity, traceability, and performance. The findings also call for adapting development models, rethinking lifecycle boundaries, and formalizing uncertainty handling to address the probabilistic and collaborative nature of LLM-assisted workflow generation. We further identify major open challenges and outline future research directions to guide the reliable and effective adoption of LLMs in ML workflow development. Our artifacts are publicly available at https://github.com/t-harden/LLM4AutoML .
- Conference Article
11
- 10.1109/ccgrid54584.2022.00047
- May 1, 2022
Machine Learning (ML) projects are currently heavily based on workflows composed of some reproducible steps and executed as containerized pipelines to build or deploy ML models efficiently because of the flexibility, portability, and fast delivery they provide to the ML life-cycle. However, deployed models need to be watched and constantly managed, supervised, and debugged to guarantee their availability, validity, and robustness in unexpected situations. Therefore, containerized ML workflows would benefit from leveraging flexible and diverse autonomic capabilities. This work presents an architecture for autonomic ML workflows with abilities for multi-layered control, based on an agent-based approach that enables autonomic management and supervision of ML workflows at the application layer and the infrastructure layer (by collaborating with the orchestrator). We redesign the Scanflow ML framework to support such multi-agent approach by using triggers, primitives, and strategies. We also implement a practical platform, so-called Scanflow-K8s, that enables autonomic ML workflows on Kubernetes clusters based on the Scanflow agents. MNIST image classification and MLPerf ImageNet classification benchmarks are used as case studies to show the capabilities of Scanflow-K8s under different scenarios. The experimental results demonstrate the feasibility and effectiveness of our proposed agent approach and the Scanflow-K8s platform for the autonomic management of ML workflows in Kubernetes clusters at multiple layers.
- Conference Article
1
- 10.1145/3447545.3451185
- Apr 19, 2021
Today, machine learning (ML) workloads are nearly ubiquitous. Over the past decade, much effort has been put into making ML model-training fast and efficient, e.g., by proposing new ML frameworks (such as TensorFlow, PyTorch), leveraging hardware support (TPUs, GPUs, FPGAs), and implementing new execution models (pipelines, distributed training). Matching this trend, considerable effort has also been put into performance analysis tools focusing on ML model-training. However, as we identify in this work, ML model training rarely happens in isolation and is instead one step in a larger ML workflow. Therefore, it is surprising that there exists no performance analysis tool that covers the entire life-cycle of ML workflows. Addressing this large conceptual gap, we envision in this work a holistic performance analysis tool for ML workflows. We analyze the state-of-practice and the state-of-the-art, presenting quantitative evidence about the performance of existing performance tools. We formulate our vision for holistic performance analysis of ML workflows along four design pillars: a unified execution model, lightweight collection of performance data, efficient data aggregation and presentation, and close integration in ML systems. Finally, we propose first steps towards implementing our vision as GradeML, a holistic performance analysis tool for ML workflows. Our preliminary work and experiments are open source at https://github.com/atlarge-research/grademl.
- Research Article
5
- 10.1080/14634980903140364
- Sep 24, 2009
- Aquatic Ecosystem Health & Management
This study is aimed at analyzing key indicators for the evaluation of changes and tendencies of sustainable utilization and management in some marine ecosystems in the coastal waters of Hai Phong - Quang Ninh area. To complete this study, the methods employed consisted of the application of remote sensing data and geographic information systems to extract information on spatial distribution of mangrove and tidal flat ecosystems, as well as on the reclamation area for human development and indicator investigation and analysis, using remote sensing extracted and field survey monitored data with Microsoft Excel (MS Excel). Applying indicators for sustainable utilization and management of marine ecosystems, developed in recent studies for the Hai Phong - Quang Ninh area, demonstrated changes and trends in significant ecosystems, such as coral reefs and mangroves. The outcomes showed that from 1998 to 2003, living coral cover was reduced by 20% on average in the Ha Long Bay area and by 13% in Cat Ba (Hang Trai - Dau Be), and projected that living coral cover would be reduced by a further 10% to 50% by 2010. The number of coral species was much reduced by 15% to 72% in the Ha Long—Cat Ba area. Mangrove area also decreased between 1995 and 2004, particularly after 2000, and is projected to be only 10,000 ha in 2010 in Quang Ninh. Marine ecosystems in the coastal waters of Hai Phong - Quang Ninh area have degraded significantly, particularly coral reefs, mangroves and tidal flats.
- Research Article
15
- 10.1016/j.ecoinf.2013.12.004
- Dec 16, 2013
- Ecological Informatics
Integration of passive acoustic monitoring data into OBIS-SEAMAP, a global biogeographic database, to advance spatially-explicit ecological assessments
- Research Article
1
- 10.1121/10.0026935
- Mar 1, 2024
- The Journal of the Acoustical Society of America
Passive acoustic monitoring (PAM) data collection has been growing exponentially, resulting in petabytes of data that document ocean soundscapes, how they change over time, and what animals use these ecosystems at varying timescales. Efficiently extracting this critical information and comparing it to other datasets in the context of ecosystem-based management is a Big Data challenge that traditional desktop processing methods cannot address. The curation, management, and dissemination of PAM datasets is another challenge in need of collaborative progress. To meet these exigencies, a multi-agency funded Sound Cooperative (SoundCoop) project is building community-focused, national cyberinfrastructure capability for PAM data to promote improved, scalable and sustainable accessibility and applications for management and science. Driven by partnerships and framed by four case studies, the SoundCoop has established guidance on the standardized processing of sound level metrics using free software toolkits and begun developing core cyberinfrastructure components that future PAM projects can leverage. U.S. and international scientists contributed PAM data collected across 10 long-term monitoring projects to operationalize the production of hybrid-millidecade spectra across a diversity of labs/instruments. Collectively, the contributed data demonstrate the value of standardized processing that enables the creation of comparable results from disparate monitoring efforts.
- Research Article
21
- 10.7717/peerj.10761
- Feb 8, 2021
- PeerJ
BackgroundProviding coral reef systems with the greatest chance of survival requires effective assessment and monitoring to guide management at a range of scales from community to government. The development of rapid monitoring approaches amenable to collection at community level, yet recognised by policymakers, remains a challenge. Technologies can increase the scope of data collection. Two promising visual and audio approaches are (i) 3D habitat models, generated through photogrammetry from video footage, providing assessment of coral cover structural metrics and (ii) audio, from which acoustic indices shown to correlate to vertebrate and invertebrate diversity, can be extracted.MethodsWe collected audio and video imagery using low cost underwater cameras (GoPro Hero7™) from 34 reef samples from West Papua (Indonesia). Using photogrammetry one camera was used to generate 3D models of 4 m2 reef, the other was used to estimate fish abundance and collect audio to generate acoustic indices. We investigated relationships between acoustic metrics, fish abundance/diversity/functional groups, live coral cover and reef structural metrics.ResultsGeneralized linear modelling identified significant but weak correlations between live coral cover and structural metrics extracted from 3D models and stronger relationships between live coral and fish abundance. Acoustic indices correlated to fish abundance, species richness and reef functional metrics associated with overfishing and algal control. Acoustic Evenness (1,200–11,000 Hz) and Root Mean Square RMS (100–1,200 Hz) were the best individual predictors overall suggesting traditional bioacoustic indices, providing information on sound energy and the variability in sound levels in specific frequency bands, can contribute to reef assessment.ConclusionAcoustics and 3D modelling contribute to low-cost, rapid reef assessment tools, amenable to community-level data collection, and generate information for coral reef management. Future work should explore whether 3D models of standardised transects and acoustic indices generated from low cost underwater cameras can replicate or support ‘gold standard’ reef assessment methodologies recognised by policy makers in marine management.
- Research Article
2
- 10.47941/ijce.1714
- Mar 2, 2024
- International Journal of Computing and Engineering
Purpose: This paper addresses the comprehensive security challenges inherent in the lifecycle of machine learning (ML) systems, including data collection, processing, model training, evaluation, and deployment. The imperative for robust security mechanisms within ML workflows has become increasingly paramount in the rapidly advancing field of ML, as these challenges encompass data privacy breaches, unauthorized access, model theft, adversarial attacks, and vulnerabilities within the computational infrastructure.
 Methodology: To counteract these threats, we propose a holistic suite of strategies designed to enhance the security of ML workflows. These strategies include advanced data protection techniques like anonymization and encryption, model security enhancements through adversarial training and hardening, and the fortification of infrastructure security via secure computing environments and continuous monitoring.
 Findings: The multifaceted nature of security challenges in ML workflows poses significant risks to the confidentiality, integrity, and availability of ML systems, potentially leading to severe consequences such as financial loss, erosion of trust, and misuse of sensitive information.
 Unique Contribution to Theory, Policy and Practice: Additionally, this paper advocates for the integration of legal and ethical considerations into a proactive and layered security approach, aiming to mitigate the risks associated with ML workflows effectively. By implementing these comprehensive security measures, stakeholders can significantly reinforce the trustworthiness and efficacy of ML applications across sensitive and critical sectors, ensuring their resilience against an evolving landscape of threats.
- Research Article
- 10.55041/ijsrem46241
- Apr 27, 2025
- INTERNATIONAL JOURNAL OF SCIENTIFIC RESEARCH IN ENGINEERING AND MANAGEMENT
Abstract—Access to clean water and affordable electricity is vital for sustainable living. This study introduces a Machine Learning (ML) framework to locate groundwater concurrently. The ML framework integrates diverse data modalities into a unified hypersurface, accommodating numeric and categorical field observations. These modalities include field measurements and data-driven machine learning outputs. Applied to various locations, the ML workflow predicts geophysical, geologic, and hydrogeologic features. Despite challenges like data disparity and spatial limitations, our model accurately identifies hidden groundwater. The study contributes insights/into sustainable resource/management/and demonstrates the applicability of/the ML framework to address local challenges. This research presents a promising approach to address water and energy resource needs in India, aligning with the country's sustainability goals. Keywords—"Groundwater localization, machine learning framework, data-driven predictions, water resource management, spatial data analysis."
- Research Article
20
- 10.1088/1755-1315/473/1/012058
- Mar 1, 2020
- IOP Conference Series: Earth and Environmental Science
Coral reefs are currently suffering from serious degradation due to human activities. In 2015, the condition of coral reefs in Kapoposang Island has been very poor with the live coral cover only 16%. Therefore, the coral reef ecosystem on this island needs to be rehabilitated. This study aims to assess coral cover based on the age of transplantation and examine abundance of reef fish in relation to age of transplant module at Kapoposang Island which is in the Wallacea region. Coral transplant was carried out from 2014 to 2018. The transplanted corals were corals of the genus Acropora. Transplants were carried out at a depth 3 to 4 m. The determination of the transplant module as the reef fish’s observation was based on the age of the transplant module, i.e. 1, 2, 3 and 4 years old. Data collection was carried out using the UVC (Underwater Visual Census) method. Data collection was done by using the UPT (Underwater Photo Transect) method. Photograph data was processed using CPCe (Coral Point Count with Excel extension) software using 30 random points for each frame. The significant relationship between live coral cover and reef fish shows that the coral transplantation was successful. There was linear relation between coral habitat cover and the reef fish. The difference in abundance in each transplant module shows the linier relation between the increase of reef fishes and live coral cover. The live coral cover was higher at the two and three years old of the transplant module. During the study, it was found 13 families and 56 species of reef fish. Planktivorous group was the most dominant of reef fish.