Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

A Fine-Tuned RetinaNet for Real-Time Lettuce Detection

  • Abstract
  • Highlights & Summary
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

The agricultural industry plays a vital role in the global demand for food production. Along with population growth, there is an increasing need for efficient farming practices that can maximize crop yields. Conventional methods of harvesting lettuce often rely on manual labor, which can be time-consuming, labor-intensive, and prone to human error. These challenges lead to research into automation technology, such as robotics, to improve harvest efficiency and reduce reliance on human intervention. Deep learning-based object detection models have shown impressive success in various computer vision tasks, such as object recognition. RetinaNet model can be trained to identify and localize lettuce accurately. However, the pre-trained models must be fine-tuned to adapt to the specific characteristics of lettuce, such as shape, size, and occlusion, to deploy object recognition models in real-world agricultural scenarios. Fine-tuning the models using lettuce-specific datasets can improve their accuracy and robustness for detecting and localizing lettuce. The data acquired for RetinaNet has the highest accuracy of 0.782, recall of 0.844, f1-score of 0.875, and mAP of 0,962. Metrics evaluate that the higher the score, the better the model performs.

Similar Papers
  • PDF Download Icon
  • Research Article
  • Cite Count Icon 1
  • 10.1017/nlp.2024.16
A case study on decompounding in Indian language IR
  • Jun 3, 2024
  • Natural Language Processing
  • Siba Sankar Sahu + 1 more

Decompounding is an essential preprocessing step in text-processing tasks such as machine translation, speech recognition, and information retrieval (IR). Here, the IR issues are explored from five viewpoints. (A) Does word decompounding impact the Indian language IR? If yes, to what extent? (B) Can corpus-based decompounding models be used in the Indian language IR? If yes, how? (C) Can machine learning and deep learning-based decompounding models be applied in the Indian language IR? If yes, how? (D) Among the different decompounding models (corpus-based, hybrid machine learning-based, and deep learning-based), which provides the best effectiveness in the IR domain? (E) Among the different IR models, which provides the best effectiveness from the IR perspective? This study proposes different corpus-based, hybrid machine learning-based, and deep learning-based decompounding models in Indian languages (Marathi, Hindi, and Sanskrit). Moreover, we evaluate the effectiveness of each activity from an IR perspective only. It is observed that the different decompounding models improve IR effectiveness. The deep learning-based decompounding models outperform the corpus-based and hybrid machine learning-based models in Indian language IR. Among the different deep learning-based models, the Bi-LSTM-A model performs best and improves mean average precision (MAP) by 28.02% in Marathi. Similarly, the Bi-RNN-A model improves MAP by 18.18% and 6.1% in Hindi and Sanskrit, respectively. Among the retrieval models, the In_expC2 model outperforms others in Marathi and Hindi, and the BB2 model outperforms others in Sanskrit.

  • Research Article
  • 10.64615/fjes...2025.72
A Deep Learning Ensemble Model for Flood Image Classification
  • Nov 10, 2025
  • Fusion Journal of Engineering and Sciences
  • Asghar Ali Chandio + 2 more

Flood is a type of natural disaster that leads to a widespread devastation. The increasing amount of rain specifically in the Urban regions of Sindh province causes several issues, whereas the drainage system is not very efficient to handle the large amount of water in a short period of time. Identification of floods is essential for disaster response, as it helps locate areas which need immediate help. Recently, the deep learning-based models have shown the best performance for image classification tasks. In this paper, a deep learning-based ensemble model has been developed where four state-of-the-art deep learning models are combined to classify flood from the images either captured with the mobile camera or other image capturing devices. The deep learning ensemble model has been trained and tested on the two publicly available datasets labelled with flood and non-flood images. To enhance the efficacy of the deep learning-based ensemble model, the hyper-parameters of the four models are fine-tuned. The results obtained show that the deep learning-based ensemble model outperforms than the individual models.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 34
  • 10.1186/s12859-020-3393-1
DTranNER: biomedical named entity recognition with deep learning-based label-label transition model
  • Feb 11, 2020
  • BMC Bioinformatics
  • S K Hong + 1 more

BackgroundBiomedical named-entity recognition (BioNER) is widely modeled with conditional random fields (CRF) by regarding it as a sequence labeling problem. The CRF-based methods yield structured outputs of labels by imposing connectivity between the labels. Recent studies for BioNER have reported state-of-the-art performance by combining deep learning-based models (e.g., bidirectional Long Short-Term Memory) and CRF. The deep learning-based models in the CRF-based methods are dedicated to estimating individual labels, whereas the relationships between connected labels are described as static numbers; thereby, it is not allowed to timely reflect the context in generating the most plausible label-label transitions for a given input sentence. Regardless, correctly segmenting entity mentions in biomedical texts is challenging because the biomedical terms are often descriptive and long compared with general terms. Therefore, limiting the label-label transitions as static numbers is a bottleneck in the performance improvement of BioNER.ResultsWe introduce DTranNER, a novel CRF-based framework incorporating a deep learning-based label-label transition model into BioNER. DTranNER uses two separate deep learning-based networks: Unary-Network and Pairwise-Network. The former is to model the input for determining individual labels, and the latter is to explore the context of the input for describing the label-label transitions. We performed experiments on five benchmark BioNER corpora. Compared with current state-of-the-art methods, DTranNER achieves the best F1-score of 84.56% beyond 84.40% on the BioCreative II gene mention (BC2GM) corpus, the best F1-score of 91.99% beyond 91.41% on the BioCreative IV chemical and drug (BC4CHEMD) corpus, the best F1-score of 94.16% beyond 93.44% on the chemical NER, the best F1-score of 87.22% beyond 86.56% on the disease NER of the BioCreative V chemical disease relation (BC5CDR) corpus, and a near-best F1-score of 88.62% on the NCBI-Disease corpus.ConclusionsOur results indicate that the incorporation of the deep learning-based label-label transition model provides distinctive contextual clues to enhance BioNER over the static transition model. We demonstrate that the proposed framework enables the dynamic transition model to adaptively explore the contextual relations between adjacent labels in a fine-grained way. We expect that our study can be a stepping stone for further prosperity of biomedical literature mining.

  • Research Article
  • 10.1016/j.jgar.2025.08.012
Optimising personalised antibiotic treatment for methicillin-resistant Staphylococcus aureus bloodstream infections in ICU patients using a deep learning-based causal inference approach.
  • Dec 1, 2025
  • Journal of global antimicrobial resistance
  • Min Woo Kang + 1 more

Methicillin‑resistant Staphylococcus aureus (MRSA) bloodstream infections (BSIs) in intensive care units (ICUs) carry high mortality, and although vancomycin remains standard treatment, daptomycin and linezolid may benefit specific subgroups. This study evaluates the mortality reduction associated with vancomycin, daptomycin, and linezolid using a deep learning-based causal inference model. Data were extracted from the Medical Information Mart for Intensive Care (MIMIC)-III and MIMIC-IV databases, including 270 ICU patients with MRSA BSI. A deep learning-based causal inference model was used to assess the treatment effect of linezolid, daptomycin, and vancomycin on in-hospital mortality. Multivariable logistic regression was employed to identify patient characteristics associated with the effectiveness of each antibiotic. The deep learning-based model predicted that vancomycin, daptomycin, and linezolid reduced mortality by 15.86% (17.90% to 13.82%), 9.68% (11.83% to 7.53%), and 10.74% (12.64% to 8.84%), respectively, with vancomycin showing the greatest reduction. The average treatment effect for in-hospital mortality reduction with vancomycin was significantly greater than that with linezolid and daptomycin (both P < 0.001). Multivariable logistic regression for treatment effects revealed that vancomycin was particularly effective in patients of advanced age, those with chronic liver disease, and those with end-stage kidney disease, while it was less effective in patients with congestive heart failure or cancer. Daptomycin exhibited superior efficacy over vancomycin in patients with cancer, and linezolid was more effective in patients with cancer, hypertension, and congestive heart failure. This study highlights linezolid and daptomycin treatment in select subgroups, while a deep learning-based model enables personalised antibiotic recommendations for ICU treatment strategies.

  • Research Article
  • Cite Count Icon 9
  • 10.1080/15481603.2024.2325720
Bridging satellite missions: deep transfer learning for enhanced tropical cyclone intensity estimation
  • Mar 11, 2024
  • GIScience & Remote Sensing
  • Minki Choo + 4 more

Geostationary satellites are valuable tools for monitoring the entire lifetime of tropical cyclones (TCs). Although the most widely used method for TC intensity estimation is manual, several automatic methods, particularly artificial intelligence (AI)-based algorithms, have been proposed and have achieved significant performance. However, AI-based techniques often require large amounts of input data, making it challenging to adopt newly introduced data such as those from recently launched satellites. This study proposed a transfer-learning-based TC intensity estimation method to combine different source data. The pre-trained model was built using the Swin Transformer (Swin-T) model, utilizing data from the Communication Ocean and Meteorological Satellite Meteorological Imager sensor, which has been in operation for an extensive period (2011–2021) and provides a large dataset. Subsequently, a transfer learning model was developed by fine-tuning the pre-trained model using the GEO-KOMPSAT-2A Advanced Meteorological Imager, which has been operational since 2019. The transfer learning approach was tested in three different ways depending on the fine-tuning ratio, with the optimal performance achieved when all layers were fine-tuned. The pre-trained model employed TC observations from 2011 to 2017 for training and 2018 for testing, whereas the transfer learning model utilized data from 2019 and 2020 for training and 2021 for testing to evaluate the model performance. The best pre-trained and transfer learning models achieved mean absolute error of 6.46 kts and 6.48 kts, respectively. Our proposed model showed a 7–52% improvement compared to the control models without transfer learning. This implies that the transfer learning approach for TC intensity estimation using different satellite observations is significant. Moreover, by employing a deep learning model visualization approach known as Eigen-class activation map, the spatial characteristics of the developed model were validated according to the intensity levels. This analysis revealed features corresponding to the Dvorak technique, demonstrating the interpretability of the Swin-T-based TC intensity estimation algorithm. This study successfully demonstrated the effectiveness of transfer learning in developing a deep learning-based TC intensity estimation model for newly acquired data.

  • Research Article
  • Cite Count Icon 5
  • 10.11591/ijeecs.v30.i1.pp491-500
An abbreviated review of deep learning-based image classification models
  • Apr 1, 2023
  • Indonesian Journal of Electrical Engineering and Computer Science
  • Zaman Talal Abbood + 2 more

Image classification is an extensively researched sub-fields of computer vision implemented in face recognition, self-driving, medical image segmentation, biological identification, and others. Traditional models of image classification require manual construction of feature extraction techniques and classification accuracy which are closely associated with these utilized techniques. During the rapid progress of multimedia technologies, the number of images that require classification got bigger, and this led to making image classification more complicated, hence, the manual construction of feature extraction techniques consumes more time and provides lower accuracy. In the recent decade, deep learning-based models have appeared in various applications. These models hold the merits of an effective extraction of image features, low-weight features filtering, a large capacity for processing, and higher classification speed and accuracy. Thus, lots of researchers have attempted to utilize deep learning algorithms, especially convolutional neural networks (CNNs) for image classification. Therefore, this paper concentrates on providing an abbreviated review of deep learning-based image classification models, by covering the recently utilized deep learning algorithms, comparing various related works and benchmark datasets mentioned in this paper, and summarizing the fundamental analysis and discussion.

  • Research Article
  • Cite Count Icon 1
  • 10.1007/s10278-025-01800-3
Customized CNN Architectures Outperform Pre-Trained Models in Differentiating Normal Brain Tissues, Glioma, Meningioma, and Pituitary Tumors.
  • Jan 20, 2026
  • Journal of imaging informatics in medicine
  • Ahmed M Taha + 3 more

Early and accurate detection of brain tumors is essential for improving treatment outcomes and patient survival. While pre-trained deep learning models such as ResNet, VGG, and MobileNet have achieved notable success in medical image classification, their generalized architectures often struggle to capture the intricate heterogeneity of brain tissues. This study introduces a customized Convolutional Neural Network (CNN) specifically designed for brain tumor classification, demonstrating superior performance over widely used pre-trained models. The primary objective of this research is to evaluate the diagnostic performance of the customized CNN and pre-trained models in real-world scenarios and establish benchmarks for their accuracy, reliability, and computational efficiency. Furthermore, this study aims to explore and implement optimization techniques that enhance the diagnostic capabilities of the models under investigation. The proposed CNN incorporates optimized convolutional blocks, adaptive fine-tuning, and an advanced data augmentation pipeline to enhance feature extraction and minimize overfitting. When evaluated on the CE-MRI Figshare dataset containing 3064 T1-weighted contrast-enhanced images, the model achieved a validation accuracy of 97.01%, outperforming ResNet50 (89.15%), MobileNetV2 (92.89%), and VGG16 (96.76%). Furthermore, the CNN exhibited strong consistency across all tumor categories-glioma, meningioma, and pituitary tumors-proving its robustness in real-world diagnostic scenarios. These findings confirm that a well-optimized CNN architecture can outperform generic pre-trained models, underscoring the importance of task-specific deep learning designs for medical imaging applications.

  • Conference Article
  • 10.26868/25222708.2025.1893
Fast prediction of urban wind distribution with deep learning-based surrogate models
  • Aug 24, 2025
  • Houzhi Wang + 3 more

Urban wind environment assessment is crucial for enhancing pedestrian wind comfort and managing pollutant dispersion at the planning and design stage. Deep learning-based models have shown potential to replace computationally intensive Computational Fluid Dynamics (CFD) simulations for accelerated assessment. However, the effectiveness and characteristics of the models in predicting wind distribution for urban microclimate under different building configurations remain unclear. Moreover, there is a lack of knowledge about the influence of the training dataset composition on the model performance. This study aims to comprehensively evaluate a deep learning-based surrogate model for fast prediction of urban wind distribution. To this end, a substantial training dataset comprising 4,000 CFD simulations was created. A surrogate model based on U-net architecture was then developed and trained to predict wind distribution for urban microclimate. The trained model achieved mean absolute percentage errors (MAPE) ranging from 1.74% to 12.49% for unseen configurations with 1 to 4 buildings, offering a speed-up of 3–4 orders of magnitude over traditional CFD methods thus enabling near real-time wind distribution assessments. The model exhibited limited domain transferability, as it can learn transferable wind flow patterns across different building configurations. From the model development perspective, integrating diverse building configurations into the training dataset proved effective in improving the model's robustness and generalization capabilities, with cases including multiple buildings yielding more substantial improvements in predictive performance. This study shows that the deep learning-based surrogate model has demonstrated significant potential for accelerating the assessment of wind distribution for urban microclimate, thus greatly benefitting the early stages of urban planning and design.

  • Conference Article
  • Cite Count Icon 29
  • 10.2118/212690-ms
Deep Learning-Based and Kernel-Based Proxy Models for Nonlinearly Constrained Life-Cycle Production Optimization
  • Jan 24, 2023
  • Aykut Atadeger + 3 more

In this study, we investigate the use of deep learning-based and kernel-based proxy models in nonlinearly constrained production optimization and compare their performances with directly using the high-fidelity simulators (HFS) for such optimization in terms of computational cost and optimal results obtained. One of the proxy models is embed to control and observe (E2CO), a deep learning-based model, and the other model is a kernel-based proxy, least-squares support-vector regression (LS-SVR). Both proxies have the capability of predicting well outputs. The sequential quadratic programming (SQP) method is used to perform nonlinearly constrained production optimization. The objective function considered here is the net present value (NPV), and the nonlinear state constraints are field liquid production rate (FLPR) and field water production rate (FWPR). NPV, FLPR, and FWPR are constructed by using two different types of proxy models. The gradient of the objective function as well as the Jacobian matrix of constraints are computed analytically for the LS-SVR, whereas the method of stochastic simplex approximated gradient (StoSAG) is used for optimization with E2CO and HFS. The reservoir model considered in this study is a two-phase, three-dimensional reservoir with heterogeneous permeability which is taken from the SPE10 benchmark case. Well controls are optimized to maximize the NPV in an oil-water waterflooding scenario. It is observed that all proxy models can find optimal NPV results like optimal NPV obtained by HFS with much less computational effort. Among proxy models, LS-SVR is found to be less computationally demanding in the training process. Overall, both proxy models are orders of magnitude faster than numerical models in the prediction. We provide new insights into the accuracy and prediction performances of these machine learning-based proxy models for 3D oil-water systems as well as their efficiency in nonlinearly constrained production optimization for waterflooding applications.

  • Research Article
  • Cite Count Icon 19
  • 10.1016/j.egyai.2022.100140
Comparison between conventional and deep learning-based surrogate models in predicting convective heat transfer performance of U-bend channels
  • Jan 25, 2022
  • Energy and AI
  • Qi Wang + 3 more

Comparison between conventional and deep learning-based surrogate models in predicting convective heat transfer performance of U-bend channels

  • Research Article
  • Cite Count Icon 6
  • 10.32604/cmc.2023.035063
OffSig-SinGAN: A Deep Learning-Based Image Augmentation Model for Offline Signature Verification
  • Jan 1, 2023
  • Computers, Materials &amp; Continua
  • M Muzaffar Hameed + 4 more

Offline signature verification (OfSV) is essential in preventing the falsification of documents. Deep learning (DL) based OfSVs require a high number of signature images to attain acceptable performance. However, a limited number of signature samples are available to train these models in a real-world scenario. Several researchers have proposed models to augment new signature images by applying various transformations. Others, on the other hand, have used human neuromotor and cognitive-inspired augmentation models to address the demand for more signature samples. Hence, augmenting a sufficient number of signatures with variations is still a challenging task. This study proposed OffSig-SinGAN: a deep learning-based image augmentation model to address the limited number of signatures problem on offline signature verification. The proposed model is capable of augmenting better quality signatures with diversity from a single signature image only. It is empirically evaluated on widely used public datasets; GPDSsyntheticSignature. The quality of augmented signature images is assessed using four metrics like pixel-by-pixel difference, peak signal-to-noise ratio (PSNR), structural similarity index measure (SSIM), and frechet inception distance (FID). Furthermore, various experiments were organised to evaluate the proposed image augmentation model’s performance on selected DL-based OfSV systems and to prove whether it helped to improve the verification accuracy rate. Experiment results showed that the proposed augmentation model performed better on the GPDSsyntheticSignature dataset than other augmentation methods. The improved verification accuracy rate of the selected DL-based OfSV system proved the effectiveness of the proposed augmentation model.

  • Research Article
  • Cite Count Icon 2
  • 10.1109/ojemb.2025.3610160
Enhancing Super-Resolution Network Efficacy in CT Imaging: Cost-Effective Simulation of Training Data
  • Jan 1, 2025
  • IEEE Open Journal of Engineering in Medicine and Biology
  • Zeyu Tang + 3 more

Deep learning-based Generative Models have the potential to convert low-resolution CT images into high-resolution counterparts without long acquisition times and increased radiation exposure in thin-slice CT imaging. However, procuring appropriate training data for these Super-Resolution (SR) models is challenging. Previous SR research has simulated thick-slice CT images from thin-slice CT images to create training pairs. However, these methods either rely on simplistic interpolation techniques that lack realism or on sinogram reconstruction, which requires the release of raw data and complex reconstruction algorithms. Thus, we introduce a simple yet realistic method to generate thick CT images from thin-slice CT images, facilitating the creation of training pairs for SR algorithms. The training pairs produced by our method closely resemble real data distributions (PSNR = 49.74 vs. 40.66, p < 0.05). A multivariate Cox regression analysis involving thick slice CT images with lung fibrosis revealed that only the radiomics features extracted using our method demonstrated a significant correlation with mortality (HR = 1.19 and HR = 1.14, p < 0.005). This paper represents the first to identify and address the challenge of generating appropriate paired training data for Deep Learning-based CT SR models, which enhances the efficacy and applicability of SR models in real-world scenarios.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 26
  • 10.3390/s23010311
Unusual Driver Behavior Detection in Videos Using Deep Learning Models.
  • Dec 28, 2022
  • Sensors
  • Hamad Ali Abosaq + 10 more

Anomalous driving behavior detection is becoming more popular since it is vital in ensuring the safety of drivers and passengers in vehicles. Road accidents happen for various reasons, including health, mental stress, and fatigue. It is critical to monitor abnormal driving behaviors in real time to improve driving safety, raise driver awareness of their driving patterns, and minimize future road accidents. Many symptoms appear to show this condition in the driver, such as facial expressions or abnormal actions. The abnormal activity was among the most common causes of road accidents, accounting for nearly 20% of all accidents, according to international data on accident causes. To avoid serious consequences, abnormal driving behaviors must be identified and avoided. As it is difficult to monitor anyone continuously, automated detection of this condition is more effective and quicker. To increase drivers' recognition of their driving behaviors and prevent potential accidents, a precise monitoring approach that detects abnormal driving behaviors and identifies abnormal driving behaviors is required. The most common activities performed by the driver while driving is drinking, eating, smoking, and calling. These types of driver activities are considered in this work, along with normal driving. This study proposed deep learning-based detection models for recognizing abnormal driver actions. This system is trained and tested using a newly created dataset, including five classes. The main classes include Driver-smoking, Driver-eating, Driver-drinking, Driver-calling, and Driver-normal. For the analysis of results, pre-trained and fine-tuned CNN models are considered. The proposed CNN-based model and pre-trained models ResNet101, VGG-16, VGG-19, and Inception-v3 are used. The results are compared by using the performance measures. The results are obtained 89%, 93%, 93%, 94% for pre-trained models and 95% by using the proposed CNN-based model. Our analysis and results revealed that our proposed CNN base model performed well and could effectively classify the driver's abnormal behavior.

  • Research Article
  • Cite Count Icon 216
  • 10.1007/s10462-022-10237-x
Biometrics recognition using deep learning: a survey
  • Jan 13, 2023
  • Artificial Intelligence Review
  • Shervin Minaee + 4 more

In the past few years, deep learning-based models have been very successful in achieving state-of-the-art results in many tasks in computer vision, speech recognition, and natural language processing. These models seem to be a natural fit for handling the ever-increasing scale of biometric recognition problems, from cellphone authentication to airport security systems. Deep learning-based models have increasingly been leveraged to improve the accuracy of different biometric recognition systems in recent years. In this work, we provide a comprehensive survey of more than 150 promising works on biometric recognition (including face, fingerprint, iris, palmprint, ear, voice, signature, and gait recognition), which deploy deep learning models, and show their strengths and potentials in different applications. For each biometric, we first introduce the available datasets that are widely used in the literature and their characteristics. We will then talk about several promising deep learning works developed for that biometric, and show their performance on popular public benchmarks. We will also discuss some of the main challenges while using these models for biometric recognition, and possible future directions to which research in this area is headed.

  • Research Article
  • Cite Count Icon 25
  • 10.1016/j.cose.2022.102789
Threat classification model for security information event management focusing on model efficiency
  • Jun 6, 2022
  • Computers &amp; Security
  • Jae-Yeol Kim + 1 more

Threat classification model for security information event management focusing on model efficiency

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant