KOG: A secret sharing-based scalable privacy-preserving training framework for decision trees
KOG: A secret sharing-based scalable privacy-preserving training framework for decision trees
- Research Article
74
- 10.1155/2019/5627156
- Jan 21, 2019
- Computational Intelligence and Neuroscience
This paper proposes a novel classification framework and a novel data reduction method to distinguish multiclass motor imagery (MI) electroencephalography (EEG) for brain computer interface (BCI) based on the manifold of covariance matrices in a Riemannian perspective. For method 1, a subject-specific decision tree (SSDT) framework with filter geodesic minimum distance to Riemannian mean (FGMDRM) is designed to identify MI tasks and reduce the classification error in the nonseparable region of FGMDRM. Method 2 includes a feature extraction algorithm and a classification algorithm. The feature extraction algorithm combines semisupervised joint mutual information (semi-JMI) with general discriminate analysis (GDA), namely, SJGDA, to reduce the dimension of vectors in the Riemannian tangent plane. And the classification algorithm replaces the FGMDRM in method 1 with k-nearest neighbor (KNN), named SSDT-KNN. By applying method 2 on BCI competition IV dataset 2a, the kappa value has been improved from 0.57 to 0.607 compared to the winner of dataset 2a. And method 2 also obtains high recognition rate on the other two datasets.
- Research Article
29
- 10.1109/tnet.2020.3048666
- Jan 25, 2021
- IEEE/ACM Transactions on Networking
Major commercial client-side video players employ adaptive bitrate (ABR) algorithms to improve the user quality of experience (QoE). With the evolvement of ABR algorithms, increasingly complex methods such as neural networks have been adopted to pursue better performance. However, these complex methods are too heavyweight to be directly deployed in client devices with limited resources, such as mobile phones. Existing solutions suffer from a trade-off between algorithm performance and deployment overhead. To make the deployment of sophisticated ABR algorithms practical, we propose PiTree, a general, high-performance, and scalable framework that can faithfully convert sophisticated ABR algorithms into decision trees with teacher-student learning. In this way, network operators can train complex models offline and deploy converted lightweight decision trees online. We also present theoretical analysis on the conversion and provide two upper bounds of the prediction error during the conversion and the generalization loss after conversion. Evaluation on three representative ABR algorithms with both trace-driven emulation and real-world experiments demonstrates that PiTree could convert ABR algorithms into decision trees with <; 3% average performance degradation. Moreover, compared to original deployment solutions, PiTree could save considerable operating expenses for content providers.
- Research Article
- 10.1007/s44163-026-01152-z
- Mar 27, 2026
- Discover Artificial Intelligence
Early detection of diabetes is crucial for effective intervention and management. This study presents a scalable and interpretable machine learning framework using a real-world dataset obtained from Kaggle consisting of 96,146 valid records after removal of duplicate entries. The dataset contains clinically relevant and demographically diverse features, including age, gender, BMI, hypertension, heart disease, smoking history, HbA1c level, and blood glucose level. The framework includes enhanced preprocessing and advanced class balancing using Synthetic Minority Oversampling Technique (SMOTE). Two evaluation approaches, one using the full feature set and another using a reduced subset of top features identified through feature importance analysis and Recursive Feature Elimination (RFE) with Random Forest (RF) as the base estimator. In the full-feature pipeline, classifiers including Logistic Regression (LR), Decision Tree (DT), RF, and XGBoost are trained and evaluated. The RF model is fine-tuned using a randomized search, and the performance of the XGBoost model is improved by using the grid search. Model interpretability is assessed using LIME analysis, which asserts the role of smoking history and hypertension on the full and reduced feature set using RF. The tuned XGBoost model achieves the highest accuracy of 96.8% and a ROC-AUC of 0.973. In the reduced-feature pipeline, an RF trained on the top five features achieves an ROC-AUC of 0.958 with a lower computational cost. The results demonstrate that high predictive accuracy can be maintained with a minimal and interpretable feature set, making the proposed framework suitable for deployment in real-world, resource-constrained environments.
- Conference Article
27
- 10.1145/3343031.3350866
- Oct 15, 2019
Major commercial client-side video players employ adaptive bitrate (ABR) algorithms to improve user quality of experience (QoE). With the evolvement of ABR algorithms, increasingly complex methods such as neural networks have been adopted to pursue better performance. However, these complex methods are too heavyweight to be directly implemented in client devices, especially mobile phones with very limited resources. Existing solutions suffer from a trade-off between algorithm performance and deployment overhead. To make the implementation of sophisticated ABR algorithms practical, we propose PiTree, a general, high-performance and scalable framework that can faithfully convert sophisticated ABR algorithms into lightweight decision trees to reduce deployment overhead. We also provide a theoretical upper bound on the optimization loss during the conversion. Evaluation results on three representative ABR algorithms demonstrate that PiTree could faithfully convert ABR algorithms into decision trees with <3% average performance degradation. Moreover, comparing to original implementation solutions, PiTree could save operating expenses for large content providers.
- Research Article
30
- 10.1111/1752-1688.12701
- Nov 29, 2018
- JAWRA Journal of the American Water Resources Association
There has recently been a return in climate change risk management practice to bottom‐up, robustness‐based planning paradigms introduced 40 years ago. The World Bank's decision tree framework (DTF) for “confronting climate uncertainty” is one incarnation of those paradigms. In order to better represent the state of the art in climate change risk assessment and evaluation techniques, this paper proposes: (1) an update to the DTF, replacing its “climate change stress test” with a multidimensional stress test; and (2) the addition of a Bayesian network framework that represents joint probabilistic behavior of uncertain parameters as sensitivity factors to aid in the weighting of scenarios of concern (the combination of conditions under which a water system fails to meet its performance targets). Using the updated DTF, water system planners and project managers would be better able to understand the relative magnitudes of the varied risks they face, and target investments in adaptation measures to best reduce their vulnerabilities to change. Next steps for the DTF include enhancements in: modeling of extreme event risks; coupling of human‐hydrologic systems; integration of surface water and groundwater systems; the generation of tradeoffs between economic, social, and ecological factors; incorporation of water quality considerations; and interactive data visualization.
- Single Book
151
- 10.1596/978-1-4648-0477-9
- Aug 20, 2015
Water infrastructure projects are a significant portion of the World Bank’s lending portfolio and a major need for developing countries throughout the world. Many water resources projects have long periods of economic return, with significant uncertainties in the behavior of the natural system as well as that of human factors (technology, population dynamics, economic development, and the like). The uncertainties associated with climate change, however, have led to a reconsideration of whether the water development community is adequately taking into account the uncertainties that characterize the future. The goal of this book is to outline a pragmatic process for risk assessment of Bank water resources projects that can serve as a decision support framework - a decision tree to assist project planning under uncertainty. The decision tree described in this book is based on the growing consensus that robustness-based approaches are needed to address uncertainty and its potential impacts on infrastructure planning. These approaches emphasize assessment of individual projects and their ability to perform well over a wide range of future uncertainty, including climate and other uncertainties. The decision tree is designed to address the fundamental issues to provide a path forward for project planners who face decisions potentially affected by climate change uncertainty. The decision tree was also designed so that human and financial resources can be used economically. The goal of this work was to develop a tool that will be applicable to all water resources projects, but that will also allocate climate risk assessment effort in a way that is consistent with each project’s potential sensitivity to that climate risk. Report is organized as follows: chapter one gives introduction. Chapter two provides background on the risks relevant to water systems planning, describes the different approaches to scenario definition in water systems planning, and introduces the decision scaling methodology upon which the general structure of the decision tree framework is based. Chapter three describes the decision tree tool and explains each of the steps and processes that make up the tool. Chapter four focuses on a case study of a small hydropower project as an illustration of the decision tree procedure. Chapter five describes some of the tools available for decision making under uncertainty and methods available for climate risk management. Concluding thoughts on implementation are presented in chapter six.
- Research Article
- 10.1142/s1469026826500033
- Mar 12, 2026
- International Journal of Computational Intelligence and Applications
The growing global emphasis on multilingual communication has brought English listening and speaking training to the forefront of research in computer science and human–computer interaction. Traditional language training approaches, often based on static audio content and scripted dialogues, lack adaptability, provide limited feedback, and fail to reflect real-world contexts, ultimately impeding personalized and effective skill development. To address these challenges, we propose an immersive English training system that integrates advanced speech synthesis (TTS) and automatic speech recognition (ASR) technologies within the audio-interactive linguistic enhancement network (AILEN) architecture, supported by a progressive contextual refinement scheme (PCRS). AILEN employs dual-stream encoder–decoder networks and cross-modal attention bridges to simultaneously enhance auditory comprehension and spoken language production. PCRS introduces a dynamic curriculum learning strategy, adapting task difficulty and feedback granularity in response to individual learner progress. Key innovations include the use of bidirectional cycle-consistency loss, phoneme-level refinement mechanisms, domain adaptation for accent robustness, and self-supervised pretraining, all contributing to improved user engagement and learning outcomes. Experimental evaluations show that the system significantly outperforms traditional baselines in terms of comprehension accuracy, pronunciation fluency, and cross-domain generalization. This paper presents a scalable and adaptive framework for English language training, effectively combining advanced computational techniques with practical language learning needs. It offers strong potential for the development of next-generation intelligent educational systems and more natural, effective human–computer interactions.
- Research Article
5
- 10.1016/j.jclinepi.2024.111641
- Mar 1, 2025
- Journal of clinical epidemiology
The aim of this study was to highlight the effects of entering duplicated or overlapping data from published studies using the same data registries into a meta-analysis, including its identification and management using a novel structured framework. Secondary analysis of data from a proportional meta-analysis of 30-day cumulative incidence of venous thromboembolic events (VTE) after metabolic and bariatric surgery was performed. Sensitivity analysis was conducted a) including all studies regardless of duplication (uncorrected sample) and b) comparing it to a corrected sample of studies. We developed a decision tree framework to identify duplicated data from prospective studies and data registries. We demonstrated that biasing from duplicated data, primarily from data registries, underestimated the incidence of VTE in the literature by 0.15% of the patient population (an erroneous difference equivalent to 22.06% of total VTE). This error persisted at 8.16% of total VTE when limiting to studies using a primarily laparoscopic approach. The decision tree framework used a comparison of the data source (country and hospital or registry), sampling time frame (dates/years of included data) and inclusion characteristics (included procedures/diagnoses or inclusion criteria) to identify potentially duplicated data. Inter-rater reliability was excellent (κ=1.00, P<.001), although only 17.86% of studies coded as containing data duplication were verified by the authors while the remaining studies could not be verified. Lastly, we identified a strong lack of diversity in the geographical origins of the data from the included studies. We demonstrated that inadvertently including duplicated data in a meta-analysis can result in substantially inaccurate pooled estimates. We outlined a comprehensive decision tree framework that future researchers can apply to assist with decision making when identifying and managing duplicated data, including that from prospective trials and data registries or other publicly accessible datasets. We explored the effects of entering duplicated or overlapping data from published studies using the same data registries into a meta-analysis; and developed a decision tree framework to identify such duplicated data from prospective studies and data registries. We analyzed data of 30-day incidence of venous thromboembolic events after metabolic and bariatric surgery. We demonstrated that including duplicated data, mainly from data registries, in a meta-analysis can result in substantially inaccurate pooled estimates, underestimating the incidence of total venous thromboembolic events by 22.06%. We also found a lack of diversity in the geographical origins of the data. The decision tree compared data source (country and hospital/registry), sampling time frame (dates/years of included data) and inclusion characteristics (inclusion criteria/procedures/diagnoses) to identify potentially duplicated data. Future researchers can apply the framework to make decisions when identifying and managing duplicated data from data registries or other publicly accessible datasets.
- Conference Article
2
- 10.1109/icde53745.2022.00213
- May 1, 2022
Decision trees and tree ensembles are popular supervised learning models on tabular data. Two recent research trends on tree models stand out: (1) bigger and deeper models with many trees, and (2) scalable distributed training frameworks. However, existing implementations on distributed systems are IO-bound leaving CPU cores underutilized. They also only find best node-splitting conditions approximately due to row-based data partitioning scheme. In this paper, we target the exact training of tree models by effectively utilizing the available CPU cores. The resulting system called TreeServer adopts a column-based data partitioning scheme to minimize communication, and a node-centric task-based engine to fully explore the CPU parallelism. Experiments show that TreeServer is up to 10× faster than models in Spark MLlib. We also showcase TreeServer's high training throughput by using it to build big “deep forest” models.
- Research Article
- 10.4018/ijdldc.399172
- Jan 22, 2026
- International Journal of Digital Literacy and Digital Competence
This article presents a structured framework developed for the Information and Communication Technology Authority of Kenya's training of foundational digital literacy skills for citizens. The study aims to evaluate the effectiveness of the training framework in narrowing the digital divide and empowering communities. Using a mixed-methods approach, including baseline and endline surveys and statistical analyses (paired t-tests and effect sizes), the article assesses the impact of digital training on more than 600,000 citizens in Mandera and Busia counties. Results demonstrate significant improvements in digital skills, with large effect sizes (Cohen's d ranging from 1.14 to 1.25) and consistently high post-training scores across modules. The findings underscore the framework's effectiveness and highlight gaps in reaching persons with disabilities. It concludes that a standardized, inclusive, and scalable framework is essential for national digital empowerment. It recommends policy support, public-private partnerships, and targeted interventions to enhance the framework's reach and sustainability.
- Research Article
1
- 10.71097/ijsat.v9.i3.5273
- Jul 6, 2018
- International Journal on Science and Technology
The fast pace of deep learning requires efficient and scalable training frameworks to support large data and intricate models. This paper presents design patterns for provisioning and managing multi-GPU clusters, specifically using platforms like AWS EC2 P3 instances, to enable training large convolutional neural networks (CNNs) and recurrent neural networks (RNNs) at scale. Important strategies are presented, including multi-source streaming broadcast, GPU-specialized parameter servers, distributed training frameworks, and scalable scheduling systems to maximize resource utilization and performance. Emphasis is placed on efficient data sharding techniques to enable load balancing and minimize communication overhead, thereby enabling accelerated convergence and improved throughput. Fault tolerance techniques like check pointing and dynamic resource management are outlined to ensure training continuity in case of hardware or network failure. Comparative analysis of frameworks like GeePS, CNTK, Nexus, and DeCUVE demonstrate the practical trade-offs between latency, scalability, and energy efficiency across various cluster configurations. Cost-effectiveness strategies for using cross-region GPU spot instances are also analyzed for deep learning applications. Topology-aware scheduling and edge-cloud distributed training paradigms are also explored to further improve system resilience and training effectiveness. This paper presents actionable insights and best practices for researchers and practitioners to deploy resilient, scalable deep learning architectures in modern cloud environments.
- Research Article
7
- 10.1016/j.peva.2024.102451
- Nov 6, 2024
- Performance Evaluation
Enabling scalable and adaptive machine learning training via serverless computing on public cloud
- Book Chapter
39
- 10.1007/978-3-319-58943-5_64
- Jan 1, 2017
We develop a scalable and extendable training framework that can utilize GPUs across nodes in a cluster and accelerate the training of deep learning models based on data parallelism. Both synchronous and asynchronous training are implemented in our framework, where parameter exchange among GPUs is based on CUDA-aware MPI. In this report, we analyze the convergence and capability of the framework to reduce training time when scaling the synchronous training of AlexNet and GoogLeNet from 2 GPUs to 8 GPUs. In addition, we explore novel ways to reduce the communication overhead caused by exchanging parameters. Finally, we release the framework as open-source for further research on distributed deep learning (https://github.com/uoguelph-mlrg/Theano-MPI).
- Research Article
- 10.51519/journalisi.v7i3.1215
- Sep 30, 2025
- Journal of Information Systems and Informatics
Data leakage in cloud storage systems poses a significant security threat, potentially leading to unauthorized access, loss of sensitive information, and operational disruptions. This research proposes a classification model for detecting potential data leakage incidents using the Decision Tree algorithm. The dataset, obtained from the Kaggle public repository, contains user activity logs representing both normal and anomalous behaviors in cloud storage environments. Several preprocessing steps were applied to improve model quality, including handling missing values, removing outliers, and converting categorical data into numerical form. Hyperparameter optimization was performed using GridSearchCV to determine the best configuration for the Decision Tree classifier. Experimental results demonstrate that the optimized model achieved high classification performance, with an accuracy of 70,84%, a precision of 55% for the data leakage class, and an F1-score of 40%. The analysis also highlights the significance of certain features, such as multi-factor authentication usage and access to confidential data, in predicting potential leakage events. This study provides a theoretical contribution by \establishing a robust methodology for applying Decision Tree algorithms to a novel cloud security dataset, offering a scalable and interpretable framework for automated threat detection.
- Research Article
1
- 10.1609/aaai.v37i8.26180
- Jun 26, 2023
- Proceedings of the AAAI Conference on Artificial Intelligence
There has been a surge of interest in learning optimal decision trees using mixed-integer programs (MIP) in recent years, as heuristic-based methods do not guarantee optimality and find it challenging to incorporate constraints that are critical for many practical applications. However, existing MIP methods that build on an arc-based formulation do not scale well as the number of binary variables is in the order of 2 to the power of the depth of the tree and the size of the dataset. Moreover, they can only handle sample-level constraints and linear metrics. In this paper, we propose a novel path-based MIP formulation where the number of decision variables is independent of dataset size. We present a scalable column generation framework to solve the MIP. Our framework produces a multiway-split tree which is more interpretable than the typical binary-split trees due to its shorter rules. Our framework is more general as it can handle nonlinear metrics such as F1 score, and incorporate a broader class of constraints. We demonstrate its efficacy with extensive experiments. We present results on datasets containing up to 1,008,372 samples while existing MIP-based decision tree models do not scale well on data beyond a few thousand points. We report superior or competitive results compared to the state-of-art MIP-based methods with up to a 24X reduction in runtime.