Identification of herbarium specimen sheet components from high-resolution images using deep learning.

Karen M. Thompson,Joanne L. Birch,Emily Fitzgerald,Robert Turnbull

doi:10.1002/ece3.10395

Abstract

Advanced computer vision techniques hold the potential to mobilise vast quantities of biodiversity data by facilitating the rapid extraction of text- and trait-based data from herbarium specimen digital images, and to increase the efficiency and accuracy of downstream data capture during digitisation. This investigation developed an object detection model using YOLOv5 and digitised collection images from the University of Melbourne Herbarium (MELU). The MELU-trained 'sheet-component' model-trained on 3371 annotated images, validated on 1000 annotated images, run using 'large' model type, at 640 pixels, for 200 epochs-successfully identified most of the 11 component types of the digital specimen images, with an overall model precision measure of 0.983, recall of 0.969 and moving average precision (mAP0.5-0.95) of 0.847. Specifically, 'institutional' and 'annotation' labels were predicted with mAP0.5-0.95 of 0.970 and 0.878 respectively. It was found that annotating at least 2000 images was required to train an adequate model, likely due to the heterogeneity of specimen sheets. The full model was then applied to selected specimens from nine global herbaria (Biodiversity Data Journal, 7, 2019), quantifying its generalisability: for example, the 'institutional label' was identified with mAP0.5-0.95 of between 0.68 and 0.89 across the various herbaria. Further detailed study demonstrated that starting with the MELU-model weights and retraining for as few as 50 epochs on 30 additional annotated images was sufficient to enable the prediction of a previously unseen component. As many herbaria are resource-constrained, the MELU-trained 'sheet-component' model weights are made available and application encouraged.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Journal: Ecology and evolution	Publication Date: Aug 1, 2023
Citations: 5	License type: CC BY 4.0

R Discovery Prime

R Discovery Prime

Identification of herbarium specimen sheet components from high-resolution images using deep learning.

Abstract

Talk to us

Similar Papers

More From: Ecology and evolution

Lead the way for us

Similar Papers

Plants meet machines: Prospects in machine learning for plant biology
Pamela S Soltis ... Alina Zare
Applications in Plant Sciences | VOL. 8
Pamela S Soltis, et. al.Pamela S Soltis ... Alina Zare
01 Jun 2020
Applications in Plant Sciences | VOL. 8

Green digitization: Online botanical collections data answering real‐world questions
Pamela S Soltis ... Shelley A James
Applications in Plant Sciences | VOL. 6
Pamela S Soltis, et. al.Pamela S Soltis ... Shelley A James
01 Feb 2018
Applications in Plant Sciences | VOL. 6

Beyond dead trees: integrating the scientific process in the Biodiversity Data Journal.
Vincent Smith ... Jeremy Miller
Biodiversity Data Journal | VOL. 1
Vincent Smith, et. al.Vincent Smith ... Jeremy Miller
16 Sep 2013
Biodiversity Data Journal | VOL. 1

Data ownership and data publishing
Lyubomir Penev
ARPHA Conference Abstracts | VOL. 2
Lyubomir PenevLyubomir Penev
20 Aug 2019
ARPHA Conference Abstracts | VOL. 2

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Identification of herbarium specimen sheet components from high-resolution images using deep learning.

Abstract

Talk to us

Similar Papers

More From: Ecology and evolution