YOLOv8m-segmentation for detecting cervical burnout and caries in bitewing radiographs: A deep learning approach
PurposeThis study evaluated the performance of the YOLOv8m-seg model in detecting and delineating interproximal caries and cervical burnout on bitewing radiographs and examined whether increasing the number of training epochs improved segmentation accuracy and consistency.Materials and MethodsIn total, 1,410 bitewing radiographs were annotated using polygon-based masks by a trained dental clinician. The YOLOv8m-seg model was trained for 50, 100, and 150 epochs on 1,128 images and validated on 282 images using the Ultralytics segmentation framework. Model performance was assessed using precision, recall, and mean average precision at intersection-over-union thresholds of 0.5 and 0.5 to 0.95 (mAP0.5, mAP0.5-0.95) for both bounding box and mask outputs. Additional evaluation was conducted on a non-augmented validation subset.ResultsExtended training duration was associated with improved segmentation performance. The highest mask mAP0.5-0.95 value was 0.828 at epoch 150. Both box-based precision and recall increased with longer training, whereas mask-based evaluation more accurately reflected the model’s ability to delineate the boundaries of caries and cervical burnout. Performance appeared consistent across both classes in the augmented validation split but was reduced in the non-augmented validation subset.ConclusionThe YOLOv8m-seg model demonstrated high diagnostic accuracy in distinguishing proximal caries from cervical burnout on bitewing radiographs. Its mask-based outputs may assist clinicians in early lesion recognition and support improved diagnostic decision-making. Future studies should evaluate model generalizability across broader populations and diverse clinical environments and should prioritize assessment using non-augmented validation sets and independent test datasets.
- Research Article
15
- 10.3390/jcm12093058
- Apr 23, 2023
- Journal of clinical medicine
Supervised deep learning requires labelled data. On medical images, data is often labelled inconsistently (e.g., too large) with varying accuracies. We aimed to assess the impact of such label noise on dental calculus detection on bitewing radiographs. On 2584 bitewings calculus was accurately labeled using bounding boxes (BBs) and artificially increased and decreased stepwise, resulting in 30 consistently and 9 inconsistently noisy datasets. An object detection network (YOLOv5) was trained on each dataset and evaluated on noisy and accurate test data. Training on accurately labeled data yielded an mAP50: 0.77 (SD: 0.01). When trained on consistently too small BBs model performance significantly decreased on accurate and noisy test data. Model performance trained on consistently too large BBs decreased immediately on accurate test data (e.g., 200% BBs: mAP50: 0.24; SD: 0.05; p < 0.05), but only after drastically increasing BBs on noisy test data (e.g., 70,000%: mAP50: 0.75; SD: 0.01; p < 0.05). Models trained on inconsistent BB sizes showed a significant decrease of performance when deviating 20% or more from the original when tested on noisy data (mAP50: 0.74; SD: 0.02; p < 0.05), or 30% or more when tested on accurate data (mAP50: 0.76; SD: 0.01; p < 0.05). In conclusion, accurate predictions need accurate labeled data in the training process. Testing on noisy data may disguise the effects of noisy training data. Researchers should be aware of the relevance of accurately annotated data, especially when testing model performances.
- Research Article
23
- 10.1007/s10278-023-00871-4
- Aug 28, 2023
- Journal of digital imaging
The study aimed to evaluate the impact of image size, area of detection (IoU) thresholds and confidence thresholds on the performance of the YOLO models in the detection of dental caries in bitewing radiographs. A total of 2575 bitewing radiographs were annotated with seven classes according to the ICCMS™ radiographic scoring system. YOLOv3 and YOLOv7 modelswere employed with different configurations, and their performances were evaluated based on precision, recall, F1-score and mean average precision (mAP). Results showed that YOLOv7 with 640 × 640pixel images exhibitedsignificantly superior performance compared to YOLOv3 in terms of precision (0.557 vs. 0.268), F1-score (0.555 vs. 0.375) and mAP (0.562 vs. 0.458), while the recall was significantly lower (0.552 vs. 0.697). The following experiment found that the overall mAPs did not significantly differ between 640 × 640 pixel and 1280 × 1280 pixelimages, for YOLOv7 with an IoU of 50% and a confidence threshold of 0.001 (p = 0.866). The last experiment revealed that the precision significantly increased from 0.570 to 0.593 for YOLOv7 with an IoU of 75% and a confidence threshold of 0.5, but the mean-recall significantly decreased and led to lower mAPs in both IoUs. In conclusion, YOLOv7 outperformed YOLOv3 in caries detection and increasing the image size did not enhance the model's performance. Elevating the IoU from 50% to 75% and confidence threshold from 0.001 to 0.5 led to a reduction of the model's performance, while simultaneously improving precision and reducing recall (minimizing false positives and negatives) for carious lesiondetection in bitewing radiographs.
- Research Article
21
- 10.3290/j.qi.b1244461
- Jun 9, 2021
- Quintessence international (Berlin, Germany : 1985)
The aim of this study was to examine the success of deep learning-based convolutional neural networks (CNN) in the detection and differentiation of amalgam, composite resin, and metal-ceramic restorations from bitewing and periapical radiographs. Five hundred and fifty bitewing and periapical radiographs were used. Eighty percent of the images were used for training, and 20% were left for testing. Twenty percent of the images allocated for training were then used for validation during learning. The image classification model was based on the application of CNN. The model used Resnet34 architecture, which is pre-trained on the ImageNet dataset. Average sensitivity, receiver operating characteristic (ROC) curve, and area under the curve (AUC) were calculated for performance evaluation of the model. The model training loss was 0.13, and the validation loss was 0.63. The independent test group result was 0.67. Amalgam AUC was 0.95, composite AUC was 0.95, and metal-ceramic AUC was 1.00. The average AUC was 0.97. The false positive rate in the validation set was 18, the false negative rate was 18, the true positive rate was 60, and the true negative rate was 138. The true positive rate was 0.82 for amalgam, 0.75 for composite, and 0.73 for metal-ceramic. Deep learning-based CNNs from periapical and bitewing radiographs appear to be a promising technique for the detection and differentiation of restorations.
- Research Article
5
- 10.1155/ijod/6644310
- Jan 1, 2025
- International Journal of Dentistry
Background: Dental caries is considered a public health issue, with early detection being crucial for effective management. Traditional diagnostic methods, including visual examination and bitewing radiographs, are prone to interpretation variability. Artificial intelligence (AI), particularly deep learning (DL), has shown promise in improving diagnostic accuracy. This study evaluates the YOLOv11 model for dental caries detection and segmentation in bitewing radiographs, using the standardized International Caries Classification and Management System (ICCMS) framework.Methods: A dataset of 730 bitewing radiographs, containing 1115 annotated carious lesions, was used for training and validation. Annotation was performed by experienced dentists using the Roboflow platform. To evaluate annotation consistency, a subset of 10 images was independently annotated by both dentists. Agreement was assessed using Intersection over Union (IoU) and Dice similarity coefficient (DSC). The YOLOv11 model was trained for 50 epochs with data augmentation techniques. Performance was assessed using precision (P), recall (R), and mean average precision at 50% IoU (mAP50).Results: The reliability analysis showed strong agreement, with an average interrater IoU of 0.82 and DSC of 0.85, and intrarater IoU of 0.84 and DSC of 0.87 across the 10 images. The YOLOv11 model excelled in detecting and segmenting advanced carious lesions, achieving high mAP50 values of 0.74 and 0.80 for RB4 + RC5 and RC6 classes, respectively. However, it showed moderate performance for early-stage lesions (RA1 + RA2 and RA3), with mAP50 scores of 0.61 and 0.52, respectively. This disparity highlights areas for potential enhancement through additional data augmentation and model fine-tuning.Conclusion: The YOLOv11 model is highly effective in identifying dental caries, especially advanced lesions, but struggles with detecting early stages of caries. AI enhancements could improve diagnostic accuracy, enable better early interventions and improve patient outcomes. The research supports incorporating AI technologies into dental radiographic evaluations to improve diagnostics and clinical results.
- Research Article
17
- 10.1093/dmfr/twae036
- Jul 18, 2024
- Dentomaxillofacial Radiology
ObjectivesThis study aimed to assess the effectiveness of deep convolutional neural network (CNN) algorithms for the detecting and segmentation of overhanging dental restorations in bitewing radiographs.MethodsA total of 1160 anonymized bitewing radiographs were used to progress the artificial intelligence (AI) system for the detection and segmentation of overhanging restorations. The data were then divided into three groups: 80% for training (930 images, 2399 labels), 10% for validation (115 images, 273 labels), and 10% for testing (115 images, 306 labels). A CNN model known as You Only Look Once (YOLOv5) was trained to detect overhanging restorations in bitewing radiographs. After utilizing the remaining 115 radiographs to evaluate the efficacy of the proposed CNN model, the accuracy, sensitivity, precision, F1 score, and area under the receiver operating characteristic curve (AUC) were computed.ResultsThe model demonstrated a precision of 90.9%, a sensitivity of 85.3%, and an F1 score of 88.0%. Furthermore, the model achieved an AUC of 0.859 on the receiver operating characteristic (ROC) curve. The mean average precision (mAP) at an intersection over a union (IoU) threshold of 0.5 was notably high at 0.87.ConclusionsThe findings suggest that deep CNN algorithms are highly effective in the detection and diagnosis of overhanging dental restorations in bitewing radiographs. The high levels of precision, sensitivity, and F1 score, along with the significant AUC and mAP values, underscore the potential of these advanced deep learning techniques in revolutionizing dental diagnostic procedures.
- Research Article
22
- 10.1007/s00784-023-05335-1
- Nov 16, 2023
- Clinical Oral Investigations
The aim of this work was to assemble alarge annotated dataset of bitewing radiographs and to use convolutional neural networks to automate the detection of dental caries in bitewing radiographs with human-level performance. A dataset of 3989 bitewing radiographs was created, and 7257 carious lesions were annotated using minimal bounding boxes. The dataset was then divided into 3 parts for the training (70%), validation (15%), and testing (15%) of multiple object detection convolutional neural networks (CNN). The tested CNN architectures included YOLOv5, Faster R-CNN, RetinaNet, and EfficientDet. To further improve the detection performance, model ensembling was used, and nested predictions were removed during post-processing. The models were compared in terms of the [Formula: see text] score and average precision (AP) with various thresholds of the intersection over union (IoU). The twelve tested architectures had [Formula: see text] scores of 0.72-0.76. Their performance was improved by ensembling which increased the [Formula: see text] score to 0.79-0.80. The best-performing ensemble detected caries with the precision of 0.83, recall of 0.77, [Formula: see text], and AP of 0.86 at IoU=0.5. Small carious lesions were predicted with slightly lower accuracy (AP 0.82) than medium or large lesions (AP 0.88). The trained ensemble of object detection CNNs detected caries with satisfactory accuracy and performed at least as well as experienced dentists (see companion paper, Part II). The performance on small lesions was likely limited by inconsistencies in the training dataset. Caries can be automatically detected using convolutional neural networks. However, detecting incipient carious lesions remains challenging.
- Research Article
69
- 10.1016/j.jag.2014.08.008
- Sep 6, 2014
- International Journal of Applied Earth Observation and Geoinformation
Evaluating the robustness of models developed from field spectral data in predicting African grass foliar nitrogen concentration using WorldView-2 image as an independent test dataset
- Research Article
- 10.1038/s41598-025-30774-z
- Dec 4, 2025
- Scientific Reports
This study presents a proof-of-concept deep learning approach for automated detection and classification of restorative dental instruments on standardized trays, aiming to support workflow automation and infection control in dental supply units. A dataset comprising 1,000 images and 14,000 annotated instances of restorative dental instruments across 14 categories was developed. The YOLOv8 model was trained and evaluated on this dataset using standard object detection metrics, including precision, recall, and mean average precision at IoU thresholds 0.5 (mAP@0.5) and 0.5:0.95 (mAP@[0.5:0.95]). To assess model advancement, YOLOv8 performance was compared against its predecessors, YOLOv5, YOLOv6, and YOLOv7, under identical experimental settings. A session-level data split was implemented as the primary evaluation to minimize data leakage and provide a realistic estimate of generalization across unseen tray configurations. The YOLOv8 model achieved highest mean average precision mAP@0.5 of 95.9% and mAP@[0.5:0.95] of 80.9%, demonstrating robust detection capability under both standard and stringent evaluation thresholds. Across instrument categories, YOLOv8 demonstrated precision ranging from 90.3% to 100% and recall from 80.6 to 98.5%. The findings demonstrate the feasibility of using YOLOv8 for automated restorative dental instrument detection as an early-stage tool for improving supply unit efficiency. While results indicate high detection accuracy and robustness, further validation in diverse clinical environments is needed. Future deployment should incorporate human-in-the-loop verification, audit trails, and error escalation mechanisms to ensure safe and accountable AI-assisted workflows.
- Research Article
6
- 10.1016/j.compbiomed.2024.109139
- Sep 12, 2024
- Computers in Biology and Medicine
Automatic motion artifact detection in electrodermal activity signals using 1D U-net architecture
- Research Article
9
- 10.1016/j.compbiomed.2024.109262
- Oct 30, 2024
- Computers in Biology and Medicine
Automated detection and labeling of posterior teeth in dental bitewing X-rays using deep learning
- Discussion
4
- 10.2174/0929866526666191112150636
- Mar 17, 2020
- Protein & Peptide Letters
Neuropeptides are a class of bioactive peptides produced from neuropeptide precursors through a series of extremely complex processes, mediating neuronal regulations in many aspects. Accurate identification of cleavage sites of neuropeptide precursors is of great significance for the development of neuroscience and brain science. With the explosive growth of neuropeptide precursor data, it is pretty much needed to develop bioinformatics methods for predicting neuropeptide precursors' cleavage sites quickly and efficiently. We started with processing the neuropeptide precursor data from SwissProt and NueoPedia into two sets of data, training dataset and testing dataset. Subsequently, six feature extraction schemes were applied to generate different feature sets and then feature selection methods were used to find the optimal feature subset of each. Thereafter the support vector machine was utilized to build models for different feature types. Finally, the performance of models were evaluated with the independent testing dataset. Six models are built through support vector machine. Among them the enhanced amino acid composition-based model reaches the highest accuracy of 91.60% in the 5-fold cross validation. When evaluated with independent testing dataset, it also showed an excellent performance with a high accuracy of 90.37% and Area under Receiver Operating Characteristic curve up to 0.9576. The performance of the developed model was decent. Moreover, for users' convenience, an online web server called NeuroCS is built, which is freely available at http://i.uestc.edu.cn/NeuroCS/dist/index.html#/. NeuroCS can be used to predict neuropeptide precursors' cleavage sites effectively.
- Research Article
9
- 10.1016/j.jdent.2025.105992
- Oct 1, 2025
- Journal of dentistry
The detection and classification of oral mucosal lesions is a challenging task due to high heterogeneity and overlap in clinical appearance. Nevertheless, differentiating benign from potentially malignant lesions is essential for appropriate management. This study evaluated whether a deep learning model trained to discriminate 11 classes of oral mucosal lesions could exceed the performance of general dentists. 4079 intraoral photographs of benign, potentially malignant and malignant oral lesions were labeled using bounding boxes and classified into 11 classes. The data were split 80:20 for training (n = 3031) and validation (n = 766), keeping an independent test set (n = 282). The YOLOv8 computer vision model was implemented for image classification and object detection. Model performance was evaluated on the test set which was also assessed by six general dentists and three specialists in oral surgery. Evaluation metrics included sensitivity, specificity, F1-score, precision, area under the receiver operating characteristic curve (AUROC), and average precision (AP) at multiple thresholds of intersection over union. In terms of classification, the highest F1-score (0.80) and AUROC (0.96) were observed for human papillomavirus (HPV)-related lesions, whereas the lowest F1-score (0.43) and AUROC (0.78) were obtained for keratosis. In terms of object detection, the best results were achieved for HPV-related lesions (AP25 = 0.82) and proliferative verrucous leukoplakia (AP25 = 0.80; AP50 = 0.76), while the lowest values were noted for leukoplakia (AP25 = 0.36; AP50 = 0.20). Overall, the model performed comparable to specialists (p = 0.93) and significantly better than general dentists (p < 0.01). The developed model performed as well as specialists in oral surgery, highlighting its potential as a valuable tool for oral lesion assessment. By providing performance comparable to oral surgeons and superior to general dentists, the developed multi-class model could support the clinical evaluation of oral lesions, potentially enabling earlier diagnosis of potentially malignant disorders, enhancing patient management and improving patient prognosis.
- Research Article
10
- 10.2174/0115734056344235241217155930
- Jan 2, 2025
- Current Medical Imaging
Background:The BraTS Generalizability Across Tumors (BraTS-GoAT) initiative addresses the critical need for robust and generalizable models in brain tumor segmentation. Despite advancements in automated segmentation techniques, the variability in tumor characteristics and imaging modalities across clinical settings presents a significant challenge.Objective:This study aims to develop an advanced CNN-based model for brain tumor segmentation that enhances consistency and utility across diverse clinical environments. The objective is to improve the generalizability of CNN models by applying them to large-scale datasets and integrating robust preprocessing techniques.Methods:The proposed approach involves the application of advanced CNN models to the BraTS 2024 challenge dataset, incorporating preprocessing techniques such as standardization, feature extraction, and segmentation. The model's performance was evaluated based on accuracy, mean Intersection over Union (IOU), average Dice coefficient, Hausdorff 95 score, precision, sensitivity, and specificity.Results:The model achieved an accuracy of 98.47%, a mean IOU of 0.8185, an average Dice coefficient of 0.7, an average Hausdorff 95 score of 1.66, a precision of 98.55%, a sensitivity of 98.40%, and a specificity of 99.52%. These results demonstrate a significant improvement over the current gold standard in brain tumor segmentation.Conclusion:The findings of this study contribute to establishing benchmarks for generalizability in medical imaging, promoting the adoption of CNN-based brain tumor segmentation models in diverse clinical environments. This work has the potential to improve outcomes for patients with brain tumors by enhancing the reliability and effectiveness of automated segmentation techniques.
- Research Article
- 10.1016/j.jdent.2025.106252
- Feb 1, 2026
- Journal of dentistry
AI-based technologies are increasingly integrated into dental diagnostics, showing promising performance in supporting radiographic interpretation and pathology detection. However, few studies have studied the evolution of these tools over time. Thus, understanding their longitudinal development is crucial to ensure clinical reliability and use. This longitudinal observational study reassessed the dataset evaluated in 2021 and reanalyzed it in 2025 using the latest version of an AI-Based software. A total of 300 digital bitewing radiographs were examined, with clinical validation serving as the reference standard. Carious lesions were classified following a modified ICDAS criteria. Diagnostic accuracy metrics (Sensitivity, Specificity, F1-score, and Area Under the Receiver Operating Characteristic Curve and McNemar's test) were calculated. A total of 2646 interproximal surfaces were evaluated, of which 579 were carious and 2067 were sound. The latest version showed lower sensitivity (80 % vs. 87 %) but improved specificity (93 %), PPV (56 %), and AUC (0.794), with fewer error-free surfaces (65.7 %) and more bounding box and omission errors. McNemar's test for paired classifications found a significant difference between versions (p < 0.001). The longitudinal assessment of the AI software reveals a clear shift toward higher specificity and reduced overdiagnosis, albeit with lower sensitivity and an increase in false negatives. Ongoing longitudinal evaluations are crucial for ensuring the safe and effective integration of AI tools in real-world dental settings. AI-based diagnostic tools evolve with data exposure, requiring clinicians to account for changing performance over time. While these systems can sharpen detection decisions from bitewing radiographs, successful use depends on continuous, longitudinal performance monitoring, tracking accuracy, and adjusting thresholds as needed.
- Research Article
10
- 10.1007/s00784-024-05528-2
- Jan 1, 2024
- Clinical Oral Investigations
ObjectiveThe objective of this study was to compare the detection of caries in bitewing radiographs by multiple dentists with an automatic method and to evaluate the detection performance in the absence of a reliable ground truth.Materials and methodsFour experts and three novices marked caries using bounding boxes in 100 bitewing radiographs. The same dataset was processed by an automatic object detection deep learning method. All annotators were compared in terms of the number of errors and intersection over union (IoU) using pairwise comparisons, with respect to the consensus standard, and with respect to the annotator of the training dataset of the automatic method.ResultsThe number of lesions marked by experts in 100 images varied between 241 and 425. Pairwise comparisons showed that the automatic method outperformed all dentists except the original annotator in the mean number of errors, while being among the best in terms of IoU. With respect to a consensus standard, the performance of the automatic method was best in terms of the number of errors and slightly below average in terms of IoU. Compared with the original annotator, the automatic method had the highest IoU and only one expert made fewer errors.ConclusionsThe automatic method consistently outperformed novices and performed as well as highly experienced dentists.Clinical significanceThe consensus in caries detection between experts is low. An automatic method based on deep learning can improve both the accuracy and repeatability of caries detection, providing a useful second opinion even for very experienced dentists.