Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

HISTAI: a valuable dataset with a valuable lesson.

  • Abstract
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

The application of artificial intelligence in computational pathology depends on both robust algorithms and high-quality, clinically reliable data. Progress in this field has been limited by the scarcity of large, diverse, and well-validated whole slide image (WSI) datasets. To address this gap, HISTAI introduced an open-source resource comprising over 112,000 WSIs across multiple organ systems with associated clinical metadata. Here, we present a pathologist-led evaluation of label accuracy, metadata completeness, and dataset composition across 328 selected cases from this resource. Although HISTAI reports 47,279 cases, we identified only 44,564 unique cases after accounting for missing entries and duplicate records. Basic demographic information, including age and sex, was available for only 55% of cases. Dataset composition was uneven, with dermatopathology accounting for 47.1% of cases and gastrointestinal pathology for 24.0%; however, primary specialty was explicitly reported for only 39.6% of cases, obscuring this imbalance within the provided metadata. Notably, clinical ground truth is recorded in the Conclusion column. Concordance between the dataset's Conclusion and Diagnosis fields was observed in only 20.7% of cases, while 27.1% contained conflicting diagnoses. In a focused review of 198 cases, 30.3% were found to contain unclear or ambiguous diagnostic conclusions, including eight cases in which the diagnosis was incorrect. Assessment of molecular annotation revealed that only 18.9% of analyzed lung and colorectal cancer cases included molecular information. Furthermore, among adult-type diffuse gliomas, none of the 55 cases met current World Health Organisation Classification of Tumors of the Central Nervous System 5th Edition (WHO CNS5) diagnostic criteria, with IDH mutation status reported in only 15 cases. Together, these findings highlight substantial ambiguities in ground-truth labeling, incomplete molecular annotation, and limited documentation of dataset provenance and ethical oversight. While HISTAI represents a valuable open-source resource, its effective and responsible use requires careful clinical validation and close collaboration between computational researchers and pathologists.

Similar Papers
  • Discussion
  • Cite Count Icon 4
  • 10.1016/s2589-7500(23)00070-5
Clinical ground truth in machine learning for early sepsis diagnosis
  • May 24, 2023
  • The Lancet Digital Health
  • Holger A Lindner + 2 more

Clinical ground truth in machine learning for early sepsis diagnosis

  • Research Article
  • Cite Count Icon 1
  • 10.3390/jmp7010002
The Role of Whole Slide Imaging in AI-Based Digital Pathology: Current Challenges and Future Directions—An Updated Literature Review
  • Jan 1, 2026
  • Journal of Molecular Pathology
  • Samya A Omoush + 3 more

Background/Objectives: Combining Whole Slide Imaging (WSI) and Artificial Intelligence (AI) in digital pathology (DP) is accelerating the field of diagnostic pathology by improving analysis metrics accuracy, reproducibility, and speed. AI applications in pathology include automated image capture, assessment and analysis, risk stratification, and prognostic prediction. This integration introduces significant challenges, including data quality, high computational demands, the ability to generalize across different settings, and a range of ethical considerations. This review provides an end-to-end roadmap covering WSI acquisition, preprocessing, and deep learning (DL) channels through tumor recognition, biomarker prediction, and evolving computational methods such as original models and combined learning, highlighting the specific challenges and opportunities of WSI-attached AI in pathology. Methods: This review provides a WSI-centric analysis that examines AI and DL applications specifically as they overlap with the acquisition, processing, and computational analysis of WSI. Therefore, this review aims to comprehensively examine the challenges and pitfalls associated with the use of WSI in AI-Based Digital Pathology. Results: Pre-analytical factors like how the tissue is prepared, staining, and scanning artifacts affect AI and contain possible post-analytical barriers such as the range of colors used, color standardization, and algorithm transparency. Furthermore, there may be bias found in the training datasets that can blur the ethical and legal boundaries alongside regulatory uncertainty. Conclusions: Even though there is an array of challenges, AI applied in DP can enhance the accuracy of medical diagnosis, encourage workflow efficiency, facilitate cross-collaboration for pediatric research, and enable research into rare diseases. Further development on the topic needs to focus on defining standard operating procedures and guidelines alongside dependable datasets through teamwork from various scientific fields.

  • Research Article
  • Cite Count Icon 3
  • 10.1158/1538-7445.sabcs21-pd11-01
Abstract PD11-01: An artificial intelligence-based predictor of CDH1 biallelic mutations and invasive lobular carcinoma
  • Feb 15, 2022
  • Cancer Research
  • Jorge S Reis‐Filho + 28 more

Introduction: Invasive lobular carcinoma (ILC) is the most frequent special histologic subtype of breast cancer (BC). ILC is identifiable by pathologic assessment given its distinctive discohesive growth pattern, largely caused by CDH1 inactivation. Compared to common forms of BC, ILCs display lower responses to chemotherapy and selective estrogen receptor modulators. The low interobserver agreement for the diagnosis of ILC, however, renders the inclusion of histologic subtyping in therapeutic decision-making challenging. Artificial intelligence (AI)-based algorithms hold promise for improving pathologic diagnosis; their performance, however, depends on the ground truth labeling used. Here, we seek to develop an AI-based methodology for detection of ILC using ‘CDH1 biallelic mutations’ (i.e., mutation + loss-of-heterozygosity of the wild-type allele or two pathogenic somatic mutations) as ground truth, reasoning that in BC, >95% of CDH1 bi-allelic inactivation is found in ILCs.Materials and methods: We developed a convolutional neural network system to detect CDH1 biallelic genetic inactivation (AI-CDH1) using whole slide images (WSI) of 1,100 primary BCs with available targeted sequencing data. The model was trained using a 10-fold cross-validation method to detect biallelic mutations. The mean number of positive and negative samples in the training set was 85.2 (SD=2.57) and 562.8 (SD=10.51) per fold, respectively. The evaluation set consisted of a mean of 14.2 (SD=2.04) positive and 93.8 (SD=9.13) negative samples. We evaluated the performance of the AI-CDH1 classifier to predict the lobular phenotype and CDH1 status using original and revised labels, following a histopathologic re-review of the histologic type and CDH1 status curation. The latter was conducted by incorporating information on biallelic CDH1 inactivation beyond CDH1 mutations (homozygous deletions, deleterious structural rearrangements, and loss-of-heterozygosity and gene promoter methylation).Results: The AI-CDH1 classifier predicted biallelic CDH1 mutations with an area under the curve (AUC)=0.944 (95 CI: 0.925-0.963), sensitivity=91.6% and specificity=85.9%, PPV=49.8%, NPV=98.5% and accuracy=86.7%, and the original ‘lobular phenotype’ with an AUC=0.941 (95 CI: 0.922-0.960), sensitivity=89%, specificity=86.7%, PPV=55.6%, NPV=97.7% and accuracy=87.1%. Review of the CDH1 gene status revealed that 7/957 BCs lacking CDH1 biallelic mutations harbored biallelic CDH1 inactivation by promoter methylation, homozygous deletions or structural rearrangements. The AI-CDH1 classifier detected all seven reclassified BCs and predicted the revised CDH1 biallelic inactivation with an AUC=0.948 (95 CI: 0.930-0.966), sensitivity=92%, specificity=86.5%, PPV=52.3%, NPV=98.5% and accuracy=87.2%. Upon histologic re-review, which resulted in reclassification of 36/927 non-lobular BCs as ‘lobular’ and 5/173 ‘lobular’ BCs as ‘non-lobular’, the AI-CDH1 classifier detected the ‘lobular phenotype’ with an AUC=0.953 (95 CI: 0.935-0.971), sensitivity=90.7%, specificity=89.7%, PPV=66.8%, NPV=97.7% and accuracy=89.9%. Using the revised histologic re-classification and CDH1 biallelic inactivation status labels, the AI-CDH1 classifier predicted the lobular phenotype irrespective of CDH1 status (P>0.05).Conclusions: By training a machine learning system to detect ‘CDH1 biallelic mutations’, as ground truth rather than histologic diagnosis of lobular carcinoma, which might be confounded by human subjectivity, we developed an AI-based system that can detect ILCs accurately, providing a new paradigm for the development of AI-based cancer classification systems. Citation Format: Jorge S Reis-Filho, Fresia Pareja, Fatemeh Derakhshan, David N Brown, Jillian Sue, Pier Selenica, Yi Kan Wang, Arnaud Da Cruz Paula, Monami Banerjee, Zahra Ebrahimzadeh, Manuel Isava, Matthew Lee, Ran Godrich, Adam Casson, Ruben Padron, George Shaikovski, Alexander van Eck, Antonio Marra, Higinio Dopeso, Hannah Y Wen, Edi Brogi, Matthew G Hanna, Chris Kanan, Jeremy D Kunz, Felipe C Geyer, Carla Leibowitz, David Klimstra, Leo Grady, Thomas J Fuchs. An artificial intelligence-based predictor of CDH1 biallelic mutations and invasive lobular carcinoma [abstract]. In: Proceedings of the 2021 San Antonio Breast Cancer Symposium; 2021 Dec 7-10; San Antonio, TX. Philadelphia (PA): AACR; Cancer Res 2022;82(4 Suppl):Abstract nr PD11-01.

  • Supplementary Content
  • Cite Count Icon 29
  • 10.1002/mp.14508
Deep learning‐based digitization of prostate brachytherapy needles in ultrasound images
  • Oct 27, 2020
  • Medical Physics
  • Christoffer Andersén + 3 more

PurposeTo develop, and evaluate the performance of, a deep learning‐based three‐dimensional (3D) convolutional neural network (CNN) artificial intelligence (AI) algorithm aimed at finding needles in ultrasound images used in prostate brachytherapy.MethodsTransrectal ultrasound (TRUS) image volumes from 1102 treatments were used to create a clinical ground truth (CGT) including 24422 individual needles that had been manually digitized by medical physicists during brachytherapy procedures. A 3D CNN U‐net with 128 × 128 × 128 TRUS image volumes as input was trained using 17215 needle examples. Predictions of voxels constituting a needle were combined to yield a 3D linear function describing the localization of each needle in a TRUS volume. Manual and AI digitizations were compared in terms of the root‐mean‐square distance (RMSD) along each needle, expressed as median and interquartile range (IQR). The method was evaluated on a data set including 7207 needle examples. A subgroup of the evaluation data set (n = 188) was created, where the needles were digitized once more by a medical physicist (G1) trained in brachytherapy. The digitization procedure was timed.ResultsThe RMSD between the AI and CGT was 0.55 (IQR: 0.35–0.86) mm. In the smaller subset, the RMSD between AI and CGT was similar (0.52 [IQR: 0.33–0.79] mm) but significantly smaller (P < 0.001) than the difference of 0.75 (IQR: 0.49–1.20) mm between AI and G1. The difference between CGT and G1 was 0.80 (IQR: 0.48–1.18) mm, implying that the AI performed as well as the CGT in relation to G1. The mean time needed for human digitization was 10 min 11 sec, while the time needed for the AI was negligible.ConclusionsA 3D CNN can be trained to identify needles in TRUS images. The performance of the network was similar to that of a medical physicist trained in brachytherapy. Incorporating a CNN for needle identification can shorten brachytherapy treatment procedures substantially.

  • Research Article
  • 10.56922/mchc.v4i8.1745
AI-based early detection of cervical cancer: A new hope for cancer prevention in indonesia
  • Nov 28, 2025
  • THE JOURNAL OF Mother and Child Health Concerns
  • Taufiq Qul Hidayat + 1 more

Background: Cervical cancer is one of the leading causes of death among women in Indonesia, mostly due to delayed diagnosis. Early detection is a key step in reducing mortality rates, but conventional methods such as Pap smears and visual inspection with acetic acid (VIA) still face obstacles such as limited medical personnel, subjectivity of results, and low screening coverage in remote areas. Purpose: to analyze the potential application of artificial intelligence (AI) as an innovative solution to improve the effectiveness and efficiency of early detection of cervical cancer in Indonesia. Method: The study used a descriptive qualitative approach through a literature review of scientific journals, international agency reports (WHO, GLOBOCAN), and national policies related to digital health transformation. Results: The study shows that the application of deep learning-based AI can increase the sensitivity and specificity of detection to over 90%, speed up the analysis process from day to minute, and reduce the operational costs of examinations by up to 40%. In addition, AI has the potential to expand the scope of screening and strengthen the national health referral system through digital integration and cloud-based telemedicine. However, the main challenges faced include data privacy issues, algorithmic bias, legal liability, and digital infrastructure gaps that must be addressed with strong ethical policies and oversight. Conclusion: The application of AI in early detection of cervical cancer is a strategic step towards a more equitable, efficient, and sustainable healthcare system in Indonesia, provided that it is implemented responsibly, transparently, and with a focus on patient safety.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 10
  • 10.3390/biomedinformatics4010028
Whole Slide Image Understanding in Pathology: What Is the Salient Scale of Analysis?
  • Feb 14, 2024
  • BioMedInformatics
  • Eleanor Jenkinson + 1 more

Background: In recent years, there has been increasing research in the applications of Artificial Intelligence in the medical industry. Digital pathology has seen great success in introducing the use of technology in the digitisation and analysis of pathology slides to ease the burden of work on pathologists. Digitised pathology slides, otherwise known as whole slide images, can be analysed by pathologists with the same methods used to analyse traditional glass slides. Methods: The digitisation of pathology slides has also led to the possibility of using these whole slide images to train machine learning models to detect tumours. Patch-based methods are common in the analysis of whole slide images as these images are too large to be processed using normal machine learning methods. However, there is little work exploring the effect that the size of the patches has on the analysis. A patch-based whole slide image analysis method was implemented and then used to evaluate and compare the accuracy of the analysis using patches of different sizes. In addition, two different patch sampling methods are used to test if the optimal patch size is the same for both methods, as well as a downsampling method where whole slide images of low resolution images are used to train an analysis model. Results: It was discovered that the most successful method uses a patch size of 256 × 256 pixels with the informed sampling method, using the location of tumour regions to sample a balanced dataset. Conclusion: Future work on batch-based analysis of whole slide images in pathology should take into account our findings when designing new models.

  • Research Article
  • Cite Count Icon 1
  • 10.48129/kjs.12929
Monocular weakly supervised depth and pose estimation method based on multi-information fusion
  • Dec 9, 2021
  • Kuwait Journal of Science
  • Zhimin Zhang + 2 more

Monocular weakly supervised depth and pose estimation method based on multi-information fusion

  • Book Chapter
  • 10.1016/b978-0-443-15688-5.00050-4
Chapter 15 - Artificial intelligence in dermatopathology
  • Sep 15, 2023
  • Artificial Intelligence in Clinical Practice
  • Puneet K Bhullar + 5 more

Chapter 15 - Artificial intelligence in dermatopathology

  • Conference Article
  • Cite Count Icon 2
  • 10.1117/12.2300262
SlideSeg: a Python module for the creation of annotated image repositories from whole slide images
  • Mar 6, 2018
  • Brendan Crabb + 1 more

Machine learning methods are being widely used in medicine to aid cancer diagnosis and detection. In the area of digital pathology, prediction heat maps produced by convolutional neural networks (CNN) have already exceeded the performance of a trained pathologist with no time constraints. To train deep learning networks, large datasets of accurately labeled ground truth data are required; however, whole slide images are often on the scale of 10p gigapixels when digitized at 40X magnification, contain multiple magnification levels, and have unstandardized formats. Due to these characteristics, traditional techniques for the production of training and validation data cannot be used, resulting in the limited availability of annotated datasets. This research presents a Python module and method to rapidly produce accurately annotated image patches from whole slide images. This module is built on OpenCV, an open source computer vision library, OpenSlide, an open source library for reading virtual slide images, and NumPy, a library for scientific computing with Python. These Python scripts successfully produce 'ground truth' image patches and will help transfer advances in research laboratories into clinical application by addressing many of the challenges associated with the development of annotated datasets for machine learning in histopathology.

  • Preprint Article
  • 10.31219/osf.io/5zbv2_v1
A Resilient HER2 Grading System for Brest Cancer using CISH and FISH Slides
  • Apr 30, 2025
  • Sayeed Sohail

In this study, we propose a robust textitHER2 grading system for digital pathology using CISH whole slide image (WSI). The proposed system utilizes deep learning technology to achieve robustness for nuclei and bio-marker detection. Then, the CISH and FISH results of the proposed system were compared with pathologists’ manual FISH counts which is the clinical ground truth. The HER2 to CEP17 ratio and average copy number per nucleus were estimated by the Pearson correlation coefficient. The correlation of proposed system with the clinical results was 0.993 and 0.991, accordingly in terms ofHER2-to-CEP17 ratio. In the experiment, we have found high concordances for the proposed method with clinical scores for both FISH and CISH.

  • Research Article
  • Cite Count Icon 16
  • 10.5146/tjpath.2023.01601
Whole Slide Images in Artificial Intelligence Applications in Digital Pathology: Challenges and Pitfalls.
  • Jan 1, 2023
  • Turk patoloji dergisi
  • Kayhan Basak + 2 more

The use of digitized data in pathology research is rapidly increasing. The whole slide image (WSI) is an indispensable part of the visual examination of slides in digital pathology and artificial intelligence applications; therefore, the acquisition of WSI with the highest quality is essential. Unlike the conventional routine of pathology, the digital conversion of tissue slides and the differences in its use pose difficulties for pathologists. We categorized these challenges into three groups: before, during, and after the WSI acquisition. The problems before WSI acquisition are usually related to the quality of the glass slide and reflect all existing problems in the analytical process in pathology laboratories. WSI acquisition problems are dependent on the device used to produce the final image file. They may be related to the parts of the device that create an optical image or the hardware and software that enable digitization. Post-WSI acquisition issues are related to the final image file itself, which is the final form of this data, or the software and hardware that will use this file. Because of the digital nature of the data, most of the difficulties are related to the capabilities of the hardware or software. Being aware of the challenges and pitfalls of using digital pathology and AI will make pathologists' integration to the new technologies easier in their daily practice or research.

  • Research Article
  • Cite Count Icon 1
  • 10.46792/fuoyejet.v7i2.800
Clustering Based Approach for Ground Truth Inference in Crowdsourced Data
  • Jun 30, 2022
  • FUOYE Journal of Engineering and Technology
  • Victor T Odumuyiwa + 5 more

Crowdsourcing provides a means of gathering data from the public in order to infer what the ground truth label of an unfamiliar entity is. Such data are not used for decision making in their raw form until further processing is done to infer ground truth from the crowdsourced data. This paper presents a detailed comparative analysis of the ground truth inference ability of three clustering algorithms on crowd sourced datasets with different experimental scenarios (Initializing centroids and extracting class labels). The algorithms include, the self-organizing maps, the k-means and the expectation maximization clustering algorithm. The three algorithms were experimented on different datasets. The datasets used are Adult2, weather sentiments, emotion, valence5 and employee review dataset Four possible experimental scenarios for inferring the ground truth label from the curated dataset were analysed. The first scenario makes use of the clustering algorithm alone relying on the inner workings of the algorithm to predict the ground truth, while the second scenario makes use of an extract class label mechanism where the ground truth label was inferred by performing a further analysis on the clusters provided by the algorithm. In the third scenario, the centroids of the clustering algorithm were pre-initialized by setting the maximum value in each class from the curated data as a centroid, where centroid might mean something different relative to the particular algorithm. The fourth experimental scenario is a combination of the second and third scenario. Experimental results show that the self-organizing map (SOM) performs best across all the datasets when the weights of the units in the SOM are pre-initialized. SOM had the best performance on the weather sentiments dataset recording 92.49% accuracy and ROC AUC score of 0.88. It also recorded the best overall average accuracy of 50.2% and ROC AUC score of 0.59365 across all the datasets.

  • Research Article
  • Cite Count Icon 38
  • 10.1016/j.isci.2023.107407
Artificial intelligence and machine learning in prehospital emergency care: A scoping review
  • Jul 17, 2023
  • iScience
  • Marcel Lucas Chee + 10 more

SummaryOur scoping review provides a comprehensive analysis of the landscape of artificial intelligence (AI) applications in prehospital emergency care (PEC). It contributes to the field by highlighting the most studied AI applications and identifying the most common methodological approaches across 106 included studies. The findings indicate a promising future for AI in PEC, with many unique use cases, such as prognostication, demand prediction, resource optimization, and the Internet of Things continuous monitoring systems. Comparisons with other approaches showed AI outperforming clinicians and non-AI algorithms in most cases. However, most studies were internally validated and retrospective, highlighting the need for rigorous prospective validation of AI applications before implementation in clinical settings. We identified knowledge and methodological gaps using an evidence map, offering a roadmap for future investigators. We also discussed the significance of explainable AI for establishing trust in AI systems among clinicians and facilitating real-world validation of AI models.

  • Research Article
  • Cite Count Icon 14
  • 10.1016/j.prp.2020.153233
Whole slide imaging and colorectal carcinoma: A validation study for tumor budding and stromal differentiation
  • Sep 28, 2020
  • Pathology - Research and Practice
  • Sean Hacking + 6 more

Whole slide imaging and colorectal carcinoma: A validation study for tumor budding and stromal differentiation

  • Research Article
  • Cite Count Icon 5
  • 10.1111/epi.18141
Electroencephalographic source imaging of spikes with concurrent high-frequency oscillations is concordant with the clinical ground truth.
  • Oct 10, 2024
  • Epilepsia
  • Colton B Gonsisko + 5 more

Epilepsy raises critical challenges to accurately localize the epileptogenic zone (EZ) to guide presurgical planning. Previous research has suggested that interictal spikes overlapping with high-frequency oscillations, referred to here as pSpikes, serve as a reliable biomarker for EZ estimation, but there remains a question as to whether and to how pSpikes perform as compared to other types of epileptic spikes. This study aims to address this question by investigating the source imaging capabilities of pSpikes alongside other spike types. A total of 2819 interictal spikes from 76-channel scalp electroencephalography (EEG) were analyzed in a cohort of 24 drug-resistant focal epilepsy patients. All patients received surgical resection, and 16 were declared seizure-free based on at least 1 year of postoperative follow-up. A recently developed electrophysiological source imaging algorithm-fast spatiotemporal iteratively reweighted edge sparsity (FAST-IRES)-was used for source imaging of the detected interictal spikes. The performance of 217 pSpikes was compared with 772 nSpikes (spikes with irregular high-frequency activations), 1830 rSpikes (spikes with no high-frequency activity), and all 2819 aSpikes (all interictal spikes). The localization and extent estimation using pSpikes are concordant with the clinical ground truth; using pSpikes yields the best performance compared with nSpikes, rSpikes, and conventional spike imaging (aSpikes). For multiple spike type seizure-free patients, the mean localization error for pSpike imaging was 6.8 mm, compared with 15.0 mm for aSpikes. The sensitivity, precision, and specificity were .41, .67, and .93 for pSpikes compared with .32, .48, and .93 for aSpikes. These results demonstrate the merits of noninvasive EEG source localization, and that (1) pSpike is a superior biomarker, outperforming conventional spike imaging for the localization of epileptic sources, and especially those with multiple irritative zones; and (2) FAST-IRES provides accurate source estimation that is highly concordant with clinical ground truth, even insituations of single spike analysis with low signal-to-noise ratio.

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant