Related Topics
Articles published on Educational data mining
Authors
Select Authors
Journals
Select Journals
Duration
Select Duration
1350 Search results
Sort by Recency
- New
- Research Article
- 10.1038/s41598-026-58086-w
- Jun 24, 2026
- Scientific reports
- Savitha Acharya + 3 more
Educational Data Mining (EDM) techniques are increasingly employed to analyze student data for predicting optimal career paths and providing tailored recommendations. A major challenge, however, is the lack of a benchmark dataset that effectively supports this objective, along with the difficulty of identifying the most relevant student attributes for career growth decision support. This study addresses the need for a comprehensive and well-structured dataset to facilitate research on personalized career growth recommendations for engineering students. It presents the methodology used to curate and preprocess a novel benchmark dataset encompassing student demographics, academic background, technical and soft skills, and stress-related factors. Challenges such as data heterogeneity, sparsity, and noise were managed through rigorous data cleaning, feature engineering, and dimensionality reduction techniques.
- New
- Research Article
- 10.1038/s41598-026-59158-7
- Jun 22, 2026
- Scientific reports
- Huanhuan Zheng + 2 more
Accurately predicting student employment outcomes remains a significant challenge in educational data mining, particularly given the increasing diversity of student backgrounds and the dynamic nature of labor market demands. This study proposes an explainable hybrid deep learning framework that integrates multi-source heterogeneous data, including academic records, demographic profiles, financial attributes, and engagement in scientific or organizational activities, to perform multi-class employment prediction. The framework employs recursive feature elimination for precise feature selection, a bi-directional long short-term memory network with attention mechanisms to capture temporal academic patterns, and a tree-structured Parzen estimator-optimized XGBoost classifier to model complex feature interactions. To enhance model interpretability, SHAP values are utilized to quantify the contribution of each feature to the final prediction. Extensive experiments on a real-world vocational college dataset demonstrate that the proposed model consistently surpasses competitive baselines across multiple evaluation metrics, including accuracy, macro-F1, AUC, and Cohen's kappa. The results confirm the framework's capability to effectively leverage both longitudinal educational trajectories and static characteristics for accurate and interpretable prediction of employment outcomes.
- New
- Research Article
- 10.1038/s41598-026-58289-1
- Jun 18, 2026
- Scientific reports
- Junjie Liu + 1 more
With the growing demand for Educational Data Mining (EDM), addressing challenges in educational practice through advanced algorithms remains complex. Key issues include the heterogeneity of educational data, the difficulty in quantifying students' cognitive and behavioral traits, the absence of robust multimodal data fusion strategies, and the limited generalizability of traditional models in small-sample contexts. This study investigates the application of Artificial Intelligence (AI), specifically Convolutional Neural Network (CNN), in predicting student academic performance in colleges and universities to optimize resource allocation and improve learning outcomes. The dataset incorporates students' demographic information, academic records, and learning behavior, with comprehensive preprocessing to ensure data integrity. Experimental results indicate that the proposed model achieves a prediction accuracy of 90.7%, precision of 86.4%, recall of 85.2%, and an F1 score of 86.1%. For a test group of 500 students, the model predicts an average score of 82.5 and a median score of 83. These findings demonstrate the feasibility of integrating AI into educational systems for accurate performance forecasting and personalized learning, offering a strong foundation for future research and practical implementation.
- Research Article
- 10.1038/s41598-026-55514-9
- Jun 5, 2026
- Scientific reports
- Sanjay Agal
The accurate and early prediction of student academic outcomes is fundamental to enabling timely interventions and personalized support in higher education. However, progress in educational data mining is severely constrained by the scarcity of publicly available data sets, as student records are protected by stringent privacy regulations that prohibit their distribution. This paper introduces a comprehensive machine learning framework for early student profiling and outcome prediction, underpinned by a novel public synthetic data set of over 100,000 student records designed specifically for machine learning research in educational contexts. The data set encompasses 28 attributes spanning demographic characteristics, academic background, entrance examination scores, socio economic indicators, and behavioral metrics, incorporating realistic complexities including non linear relationships, systematically introduced missing value patterns, and controlled class imbalance. Two complementary prediction tasks are defined: ordinal classification of student academic level into Beginner, Intermediate, Advanced, and Exceptional categories, and nominal classification of division allotment into Remedial, Regular, Advanced, and Honors divisions. A modular machine learning framework is developed encompassing comprehensive preprocessing, feature engineering including composite indices and interaction terms, implementation of eleven distinct algorithms ranging from interpretable baselines to ensemble methods and neural networks, rigorous cross validation, and multi faceted evaluation incorporating accuracy, per class metrics, macro averaged F1 scores, and ordinal specific measures. Experimental results demonstrate that ensemble methods substantially outperform simpler approaches, with LightGBM achieving macro F1 scores of 0.842 for student level prediction and 0.826 for division allotment. An ordinal neural network incorporating the ordinal nature of the student level target achieves the highest overall performance with macro F1 score of 0.846 and quadratic weighted kappa of 0.892, confirming the value of explicit ordinal modeling. Feature importance analysis reveals that prior academic achievement, particularly class 12 percentage, dominates predictions, while socio economic and behavioral factors provide meaningful secondary contributions. SHAP analysis uncovers important interaction effects, including the moderating role of prior achievement on the impact of entrance examination scores, and provides local explanations that enable targeted intervention design. Cross validation confirms stability with standard deviations below 0.011 across folds, and computational efficiency analysis identifies LightGBM as offering optimal balance between predictive performance and resource requirements. The complete framework, including all preprocessing modules, model implementations, evaluation protocols, and interpretation tools, is released as open source software alongside the validated synthetic data set, providing the research community with foundational resources to advance reproducible educational data mining research without privacy constraints. This work establishes a comprehensive benchmark for student outcome prediction and provides actionable insights for developing equitable, data driven student support systems.
- Research Article
2
- 10.1016/j.caeai.2026.100548
- Jun 1, 2026
- Computers and Education: Artificial Intelligence
- Salma Boujmiraz + 2 more
The application of Machine Learning (ML) and Deep Learning (DL) in Educational Data Mining (EDM) is revolutionizing the educational field. Researchers have been particularly interested in predicting student performance at an early stage. These early predictions can significantly benefit students’ learning experiences, allowing educators and other stakeholders to plan timely interventions. This review examines studies employing these technologies, respecting the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines to ensure the approach is transparent and replicable. This systematic review begins with preselecting a total of 281 records that go through a rigorous screening process. Then, 72 retained studies are analyzed to uncover the types of databases utilized, examine the most common features and labels in predictive models, and identify the prevalent ML and DL methods and the reasons for their selection. In addition, the review highlights the role of explainability in making complex models more interpretable and pedagogically meaningful. However, only a limited number of studies explicitly connect these predictive approaches to educational innovation. This review therefore emphasizes how predictive and explainable AI (XAI) can bridge that gap by supporting evidence-based teaching, adaptive learning, and more equitable decision-making in education. • A systematic review of ML/DL techniques for predicting student performance. • PRISMA methodology used to select 72 studies from an initial pool of 281. • OULAD and UCI datasets are the most commonly employed in this domain. • Tree-based models and neural networks dominate predictive approaches. • Explainable AI (XAI) is emerging as a key component to support pedagogical decisions, although its adoption remains limited.
- Research Article
- 10.65339/ijsair.v2.i2.527
- May 31, 2026
- International Journal of Sustainability and Advanced Integrated Research
- Arnold Fuentes
This study developed and evaluated supervised machine learning models for the early identification of academically at-risk students using selected academic performance indicators. Anchored on Educational Data Mining and supervised machine learning classification, the study employed a quantitative experimental research design using a synthetic dataset composed of 100 student records. The dataset included attendance rate, quiz average score, assignment completion rate, midterm examination score, final examination score, and class participation score. The students were classified as either at-risk or not at-risk based on patterns of attendance and academic performance. Data preprocessing involved mean imputation, Min-Max Scaling, categorical encoding, and an 80% training and 20% testing split. Logistic Regression, Decision Tree, and Random Forest models were implemented and evaluated using accuracy, precision, recall, F1-score, and ROC-AUC. Results showed that Random Forest achieved the highest predictive performance, with 93% accuracy, 0.92 precision, 0.91 recall, 0.91 F1-score, and 0.95 ROC-AUC. The model correctly classified 46 non-at-risk students and 47 at-risk students, with only 4 false positives and 3 false negatives. The study concludes that machine learning can support early warning systems and improve educational decision-making by helping institutions identify students who may need timely academic support. Future studies should use larger real-world institutional datasets and include broader learner-related variables to improve model generalization. The study supports SDG 4, Quality Education, and SDG 9, Industry, Innovation and Infrastructure, by promoting data-driven academic intervention and educational innovation. Its sustainability impact is linked to educational, institutional, and technological sustainability through improved student monitoring, support planning, and evidence-based decision-making.
- Research Article
- 10.1016/j.dib.2026.112843
- May 12, 2026
- Data in Brief
- Nazira Jesmin Lina + 6 more
CodeStream: A dataset of iterative programming submissions with sequential verdict traces and attempt histories
- Research Article
- 10.56778/jdlde.v4i11.681
- Apr 30, 2026
- JOURNAL OF DIGITAL LEARNING AND DISTANCE EDUCATION
- Georgios Alexandropoulos
In the current digital era, data science has emerged as a transformative discipline with profound potential for reshaping the educational landscape. This paper explores the multifaceted role of data science—specifically through educational data mining (EDM) and learning analytics—in enhancing teaching and learning processes across various platforms. Through a critical literature review, the study examines the dual nature of data utilization. On one hand, it highlights significant benefits such as personalized learning, early detection of student behavioral patterns, and evidence-based decision-making. On the other hand, it addresses critical risks, including privacy concerns, ethical violations, social stereotyping (labeling), and the potential commodification of education by corporate interests. The analysis further demonstrates that an overreliance on quantitative metrics risks neglecting the psychological dimensions and sociocultural contexts inherent in human learning. To mitigate these imbalances, the paper proposes the application of the DELICATE framework (determination, explain, legitimate, involve, consent, anonymize, technical aspects, and external partners) to ensure transparency and data protection. The study concludes by emphasizing a necessary shift from a purely technocratic perspective to a human-centered design approach. The authors argue that data science should serve as a pedagogical support tool rather than a substitute for teacher intuition. By integrating quantitative methods with qualitative-ethnographic approaches, a more just, innovative, and humane educational environment can be achieved.
- Research Article
- 10.65102/is2026293
- Apr 30, 2026
- Ingegneria Sismica
- Honglian Bian
Under the background of artificial intelligence, learning analytics and educational data mining continuing to enter the foreign language teaching scene, the reform of private undergraduate German curriculum has the basis of process data collection, feature modeling and quantitative evaluation. Focusing on the evaluation needs of course goal achievement, classroom interaction quality, language ability growth and teaching support adaptation, this paper constructs an evaluation model for the effect of private undergraduate German course reform in the AI-enabled context. The classroom behavior records, assignment texts, test scores, platform access trajectories and feedback data in 4120 effective samples are uniformly cleaned, coded and associated mapped. The model consists of three parts: multi-source data representation, key feature modeling and evaluation output mechanism, and combines attention weighting, gated fusion and hierarchical scoring methods to complete reform effect identification and difference analysis. The experimental results on the validation set show that the evaluation accuracy of the model is 92.7%, the Recall is 91.4%, the F1 score is 90.9%, and the average output delay is 1.6 seconds. The model can reflect the effect and change characteristics of the German curriculum reform more stably, and can simultaneously show the association changes between vocabulary training, oral interaction, writing revision and stage evaluation. It provides continuous calculation basis and quantitative reference for course content adjustment, teaching method revision and learning support configuration.
- Research Article
- 10.1142/s0218213026500132
- Apr 22, 2026
- International Journal on Artificial Intelligence Tools
- Ji Hongzheng
Predicting student academic performance has become increasingly vital in the field of educational data mining, as institutions seek data-driven strategies to enhance learning outcomes. However, many existing models rely solely on behavioral indicators or static features, often overlooking the role of time and context in shaping learning behavior. This limitation reduces predictive accuracy and adaptability in academic environments. To address this challenge, this study introduces EduFuseNet, a hybrid deep learning framework that integrates behavioral and spatiotemporal data for accurate classification of student performance. The workflow begins with data collection from a Student Academic Performance dataset, comprising both behavioral metrics and spatiotemporal information. The raw data undergoes preprocessing, including missing value imputation, one-hot encoding of categorical variables, and min-max scaling of numerical features. The processed data is then passed through two specialized branches: a Tabular Neural Structure-Aware (TabNSA) module that captures complex interdependencies within behavioral data, and a Spatiotemporal Transformer module that models temporal and sequential patterns in learning activities. The feature embeddings from both branches are fused and passed through fully connected layers to generate predictions across five academic performance bands, enabling precise classification and early risk identification. EduFuseNet achieved an accuracy of 99.00%, with a precision of 99.04%, recall of 99.00%, and F1-score of 99.01%, reflecting strong and reliable predictive performance. By leveraging both behavioral and temporal learning indicators, the model serves as an effective tool for early academic monitoring and intervention.
- Research Article
- 10.62762/tedm.2026.988161
- Apr 11, 2026
- ICCK Transactions on Educational Data Mining
- Farshid Keivanian + 1 more
Educational Data Mining (EDM) has achieved substantial gains in predictive performance, yet many existing approaches remain centered on single-objective optimization, most often accuracy. This does not adequately reflect the multi-dimensional nature of real-world educational decision-making, which requires balancing interpretability, fairness, robustness, efficiency, and timeliness. This perspective advocates a shift toward multi-objective, interpretable, and trustworthy EDM frameworks. We highlight the role of multi-objective optimization in modeling trade-offs through Pareto-optimal solutions and address the challenge of actionable decision-making through bargaining-based mechanisms, such as Nash bargaining, to select balanced and transparent outcomes. In addition, we discuss the value of fuzzy logic and adaptive methods for handling uncertainty and supporting interpretable reasoning in dynamic learning environments. Finally, we emphasize the importance of governance, accountability, and rigorous evaluation, and argue that emerging technologies should be assessed not only by performance gains but also by their practical and educational relevance. Overall, this perspective outlines a human-centered research agenda for the development of trustworthy, interpretable, and context-aware EDM systems.
- Research Article
- 10.70593/deepsci.0202045
- Apr 5, 2026
- International Journal of Applied Resilience and Sustainability
- Taibat Bolarinwa
The adoption of Artificial Intelligence (AI) in educational fields has developed a great potential and problem related to the academic progress, critical thinking, mental abilities, and final student results. The increased application of generative AI, intelligent tutoring machines, adaptive learning systems, predictive analytics, and AI-assisted learning tools have altered the traditional learning environment, although issues have been raised about over-reliance on technology, lower-level thinking, algorithm biases and academic dishonesty. This literature review was a systematic investigation of recent articles relevant to the topic of Artificial Intelligence in Education (AIEd) and its effects on student engagement, academic achievement, cognitive growth, and learning customization. The focus was on emerging trends as ChatGPT in education, machine learning in education, educational data mining, personalized feedback, and smart classrooms. The review established that AI-assisted learning and individualized instructional setting have a positive effect on knowledge retention, student motivation, self-regulated learning, digital literacy, and problem-solving abilities. Learning analytics and adaptive learning systems enhance the quality of academic outcomes by providing personalized learning channels and real-time feedback. The results suggest that over-dependence on AI tools can negatively affect critical thinking, creativity, metacognition, and independent reasoning in case the pedagogy does not carefully incorporate it. The concern increasing on algorithmic bias, cognitive load, ethical issues, and the impact of large language models on academic integrity were also identified as a major concern in the review.
- Research Article
1
- 10.3390/data11040075
- Apr 3, 2026
- Data
- Erika María López-López + 2 more
Student attrition remains a persistent challenge in higher education and is shaped by interacting socioeconomic, academic, institutional, and wellbeing-related mechanisms. Although learning analytics and educational data mining increasingly support early-warning and intervention workflows, dataset reuse is often limited by incomplete documentation and inconsistent variable definitions. This Data Descriptor presents a structured cross-sectional survey dataset on factors influencing student persistence at a Colombian public university campus (La Paz). Data were collected between August and December 2025 through an online questionnaire and subsequently cleaned to remove duplicate entries and personally identifiable information. The released dataset contains 333 student records and 33 variables covering demographics (e.g., age, gender, first-generation status), socioeconomic conditions (e.g., residential stratum, housing, financial aid), academic experience and satisfaction (multiple 1–5 Likert items), perceived dropout intention across personal/socioeconomic/academic domains, thematically coded open-ended items describing challenges and motives, and a self-allocation of 0–100 weights across three dropout-factor domains. We provide a machine-readable codebook, a transparent preprocessing description, and technical validation checks (value ranges, category consistency, and composite-score integrity). The dataset is intended to support reproducible retention research, equity-oriented analyses, and benchmarking of predictive models, while encouraging responsible reuse through privacy-preserving release practices and FAIR-aligned metadata, repository deposition, and versioning.
- Research Article
- 10.37134/jsml.vol14.2.2.2026
- Apr 1, 2026
- Journal of Science and Mathematics Letters
- Nurulhuda Ramli
Learning Management Systems (LMS) have become integral tools in higher education, generating vast amounts of data that can be leveraged to analyze and enhance academic performance. Despite the abundance of this data, effectively harnessing it to understand complex relationships between learning activities and student outcomes remains a challenge. This paper explores the application of Bayesian Network (BN), a powerful technique in Educational Data Mining (EDM) to model and predict student outcomes using LMS data. BN provides a probabilistic framework to explore how various learning analytics variables influence academic success. Using LMS data from an online undergraduate Mathematics course, the model investigated the impact of student engagement, resource utilization, and participation on exam grades. The results show that consistent attendance (88%), active participation in lecturing sessions (85%), and involvement in online mathematical laboratory activities (62%), despite lower engagement in other areas such as assessments and gamification, are strongly associated with favourable final exam outcomes (62% achieving ‘Good’ or ‘Excellent’ grades). Numerical simulations were conducted to explore future student outcomes by manipulating key variables, demonstrating the potential of improved learning strategies such as full participation, improved prior knowledge and complete utilization of digital resources. This study highlights the utility of BN in analyzing LMS data to inform educational practices and ultimately enhance academic performance in higher education.
- Research Article
- 10.1016/j.dib.2026.112512
- Apr 1, 2026
- Data in brief
- Mehedi Hasan
Educational data mining and learning analytics have become important research areas for supporting pedagogical analysis, algorithm development, and privacy-preserving educational research. The advancement of natural language processing (NLP) methods in educational contexts depends on the availability of structured and well-documented textual datasets; however, access to real student data is often restricted due to ethical, legal, and privacy concerns. This article presents a fully synthetic textual dataset of student learning habits and preferences generated using a large language model (LLM). The dataset contains 10,000 CSV-formatted records representing fictional students and includes attributes such as education level, study hours, preferred learning methods, learning challenges, motivation levels, opinions on online learning, and primary devices used for study. Data generation was performed using structured prompting strategies with explicitly defined controlled vocabularies to ensure internal consistency and reproducibility while avoiding the use of any real personal information. The resulting dataset follows intentionally controlled and near-uniform distributions, with variables generated under independent constraints. This design limits its suitability for modelling real-world stochastic behaviour or discovering natural correlations but makes it appropriate for benchmarking educational NLP pipelines, evaluating synthetic data generation techniques, and conducting privacy-preserving survey and machine learning experiments.
- Research Article
- 10.22214/ijraset.2026.78840
- Mar 31, 2026
- International Journal for Research in Applied Science and Engineering Technology
- Dr P C Khanzode
Student academic performance prediction is a crucial topic in Educational Data Mining (EDM) and Learning Analytics, which can be used to help at-risk students by undertaking timely actions. The given paper is a systematic review of machine learning methods used in this field. It analyzes a range of approaches, starting with interpretable models such as Multiple Linear Regression and Decision Trees to ensemble and deep learning high-performance models such as Random Forest and Neural Networks. The review highlights the central role of feature engineering and is discussing predictors of academic and behavioral data, social-economic and psychological conditions. One of the broad implications of this paper is providing a comparative analysis of these methods with an emphasis on the continuing trade-off between predictive accuracy and model inter-pretability. Moreover, the disconnect between theory and real-world, full-stack deployment systems, which are more and more critical when it comes to actual usability, is also critically discussed in this review. Major gaps in the research, such as excessive use of synthetic data, lack of practical testing, and ethics, are determined. Lastly, the paper presents future directions which include the use of Explainable AI (XAI), federated learning in privacy and creation of real-time adaptive feedback systems
- Research Article
- 10.3390/informatics13040050
- Mar 27, 2026
- Informatics
- Yuri Reina Marín + 6 more
Student retention has become a major challenge for higher education institutions due to the influence that academic, socioeconomic, family, and motivational factors exert on students’ academic continuity. In this context, understanding the determinants that explain university persistence is essential for designing effective retention strategies. Based on the analysis of factors related to motivation, commitment, attitude, academic integration, and social and economic conditions, retention patterns were examined in a population of 532 university students, of whom 57.7% showed high retention, 38.2% medium retention, and 4.1% low retention. To identify the factors with the greatest influence on academic continuity, educational data mining techniques and supervised classification models were applied and evaluated using stratified 10-fold cross-validation. Tree-based ensemble models showed the most consistent predictive performance, with Random Forest achieving the best results (accuracy = 0.729 ± 0.058; F1-macro = 0.636 ± 0.136). Model interpretability was examined through SHAP analysis, which revealed that transportation conditions (0.249), task completion (0.170), absence of work obligations (0.168), and course completion (0.164) were the most influential predictors in the classification of retention levels. In addition, sensitivity analysis indicated that academic commitment accounts for 41.6% of the predictive impact, followed by motivation (23.5%). These findings demonstrate that student retention is shaped by the interaction of academic, motivational, and contextual factors and provide practical implications for the development of **early warning systems, personalized tutoring programs, psychosocial support initiatives, and financial assistance policies aimed at strengthening university retention.
- Research Article
- 10.1038/s41598-026-40502-w
- Mar 26, 2026
- Scientific reports
- Yongkang Duan + 1 more
Accurately predicting student dropout in Massive Open Online Courses (MOOCs) remains a critical challenge in educational data mining. While Spatio-Temporal Graph Neural Networks (STGNNs) have shown promise, established frameworks typically rely on first-order temporal dependencies, recursively deriving the current state solely from its immediate predecessor. We argue that such recursive compression fails to capture complex student behaviors, which are driven by the interplay between immediate short-term shocks and accumulated long-term patterns. To address this, we propose the Multi-Scale Spatio-Temporal Graph Network (MST-GCN). The core of our framework is a novel MST-RGCN layer featuring a Spatially-Conditioned Adaptive Gate. This mechanism dynamically modulates the fusion of short-term and long-term memories by explicitly conditioning on the evolving heterogeneous graph context. Comprehensive experiments on two large-scale benchmarks, KDD Cup 2015 and XuetangX, demonstrate that MST-GCN yields superior predictive performance compared to established baselines. Notably, our model exhibits remarkable robustness in unstructured, self-paced learning environments. Furthermore, qualitative analysis reveals that the model learns an interpretable policy: prioritizing long-term history to identify at-risk students while leveraging short-term momentum to predict successful learners. Our source code is publicly available at https://github.com/wudongze9/MST-GCN .
- Research Article
- 10.55041/isjem05836
- Mar 24, 2026
- International Scientific Journal of Engineering & Management
- Dr Satyam K + 1 more
Gathering and evaluating student input is essential to raising academic achievement and teaching quality in contemporary educational institutions. However, conventional feedback systems are frequently labour-intensive, manual, and incapable of drawing significant conclusions from massive amounts of data. This research proposes an intelligent academic feedback analysis system that combines ensemble machine learning and deep learning methods to address these issues. Students' textual feedback is processed by the suggested system, which also preprocesses the data and uses feature extraction techniques to transform unstructured data into a format that can be analysed. To improve forecast accuracy and robustness, ensemble techniques are used with deep learning models, such as neural networks. Institutions can make data-driven decisions thanks to the system's ability to automatically classify input into several categories and spot sentiment patterns. According to experimental findings, the hybrid model performs more accurately and efficiently than conventional machine learning techniques. In the end, this method improves educational results by lowering human labour and offering a scalable solution for real-time academic feedback evaluation. Keywords:Academic Feedback Analysis, Deep Learning, Ensemble Learning, Natural Language Processing, Sentiment Analysis,Educational Data Mining, Text Classification, Machine Learning
- Research Article
- 10.66104/374y8m47
- Mar 24, 2026
- Journal International Review of Research Studies
- Joelson Lopes Da Paixão
The incorporation of artificial intelligence (AI) into educational systems has expanded the possibilities for monitoring learning, producing feedback, and personalizing instruction. In school assessment, however, the adoption of these technologies requires more precise conceptual distinctions between different AI paradigms and a critical analysis of their pedagogical and ethical effects. This study aims to analyze the impacts of AI on school assessment processes by examining its formative potential, epistemological limits, and the ethical challenges involved in its use. Methodologically, this is a qualitative bibliographic study with an analytical and interpretive orientation. Rather than presenting itself as an exhaustive systematic review, the study explicitly adopts the format of an analytical bibliographic review, organized around academic literature and institutional documents relevant to the topic. Theanalysis is guided by the articulation between formative assessment theory and a critical sociotechnical reading of educational datafication. The study shows that the effects of AI on assessment are not homogeneous: rule-based systems tend to operate better in structured tasks; models supported by learning analytics and educational data mining expand monitoring and diagnostic capacity; and generative systems open new possibilities for open-ended tasks, but still show instability, opacity, and a persistent need for human oversight. The article concludes that AI can contribute to more continuous, responsive, and formative assessment practices, provided that its use remains subordinated to teachers' pedagogical judgment, data protection, algorithmic transparency, and principles of equity.