Articles published on Internet traffic classification
Authors
Select Authors
Journals
Select Journals
Duration
Select Duration
91 Search results
Sort by Recency
- Research Article
37
- 10.1109/comst.2024.3504955
- Oct 1, 2025
- IEEE Communications Surveys & Tutorials
- Alfredo Nascita + 5 more
With the increasing complexity and scale of modern networks, the demand for transparent and interpretable Artificial Intelligence (AI) models has surged. This survey comprehensively reviews the current state of eXplainable Artificial Intelligence (XAI) methodologies in the context of Network Traffic Analysis (NTA) (including tasks such as traffic classification, intrusion detection, attack classification, and traffic prediction), encompassing various aspects such as techniques, applications, requirements, challenges, and ongoing projects. It explores the vital role of XAI in enhancing network security, performance optimization, and reliability. Additionally, this survey underscores the importance of understanding why AI-driven decisions are made, emphasizing the need for explainability in critical network environments. By providing a holistic perspective on XAI for Internet NTA, this survey aims to guide researchers and practitioners in harnessing the potential of transparent AI models to address the intricate challenges of modern network management and security.
- Research Article
- 10.3390/a18080457
- Jul 23, 2025
- Algorithms
- Ramazan Enisoglu + 1 more
Accurate and real-time classification of low-latency Internet traffic is critical for applications such as video conferencing, online gaming, financial trading, and autonomous systems, where millisecond-level delays can degrade user experience. Existing methods for low-latency traffic classification, reliant on raw temporal features or static statistical analyses, fail to capture dynamic frequency patterns inherent to real-time applications. These limitations hinder accurate resource allocation in heterogeneous networks. This paper proposes a novel framework integrating wavelet transform (WT) and artificial neural networks (ANNs) to address this gap. Unlike prior works, we systematically apply WT to commonly used temporal features—such as throughput, slope, ratio, and moving averages—transforming them into frequency-domain representations. This approach reveals hidden multi-scale patterns in low-latency traffic, akin to structured noise in signal processing, which traditional time-domain analyses often overlook. These wavelet-enhanced features train a multilayer perceptron (MLP) ANN, enabling dual-domain (time–frequency) analysis. We evaluate our approach on a dataset comprising FTP, video streaming, and low-latency traffic, including mixed scenarios with up to four concurrent traffic types. Experiments demonstrate 99.56% accuracy in distinguishing low-latency traffic (e.g., video conferencing) from FTP and streaming, outperforming k-NN, CNNs, and LSTMs. Notably, our method eliminates reliance on deep packet inspection (DPI), offering ISPs a privacy-preserving and scalable solution for prioritizing time-sensitive traffic. In mixed-traffic scenarios, the model achieves 74.2–92.8% accuracy, offering ISPs a scalable solution for prioritizing time-sensitive traffic without deep packet inspection. By bridging signal processing and deep learning, this work advances efficient bandwidth allocation and enables Internet Service Providers to prioritize time-sensitive flows without deep packet inspection, improving quality of service in heterogeneous network environments.
- Research Article
- 10.58190/imiens.2025.119
- Apr 30, 2025
- Intelligent Methods in Engineering Sciences
- Poonam B Lohiya + 1 more
The increasing complexity and volume of internet traffic have led researchers to explore machine learning as an effective approach for traffic classification. By integrating intelligence into network processes, machine learning enhances network management and optimization. This study investigates four supervised learning techniques—Support Vector Machine (SVM), Random Forest (RF), K-Nearest Neighbors (KNN), and Decision Tree (DT)—to forecast network traffic categorization. Through a comparative analysis, we evaluate the performance of these algorithms in terms of accuracy, precision, recall, and computational efficiency using a standardized dataset. The results demonstrate that while each algorithm has its strengths and weaknesses, our findings indicate that Random Forest outperforms the other algorithms in most metrics, providing valuable insights for future applications in network management. This study provides valuable insights into the applicability of these algorithms for real-time internet traffic management.
- Research Article
2
- 10.1016/j.comcom.2025.108068
- Mar 1, 2025
- Computer Communications
- Jacek Krupski + 2 more
On the right choice of data from popular datasets for Internet traffic classification
- Research Article
30
- 10.11591/ijece.v14i6.pp6958-6968
- Dec 1, 2024
- International Journal of Electrical and Computer Engineering (IJECE)
- Alhamza Munther + 3 more
The pursuit of effective models with high detection accuracy has sparked great interest in anomaly detection of internet traffic. The issue still lies in creating a trustworthy and effective anomaly detection system that can handle massive data volumes and patterns that change in real-time. The detection techniques used, especially the feature selection methods and machine learning algorithms, are crucial to the design of such a system. The fundamental difficulty in feature selection is selecting a smaller subset of features that are more related to the class but are less numerous. To reduce the dimensionality of the dataset, this research offered a multi-feature selection technique (MFST) using four filter techniques: fast correlation-based filter, significance feature evaluator, chi-square, and gain ratio. Each technique's output vector is put via ranker and Borda voting filters. The feature with the highest number of votes and rank values will be selected from the dataset. The performance of the given MFST framework was the best when compared to the four strategies listed above functioning alone; three different classifiers were employed to test the accuracy. C4.5, nave Bayes, and support vector machine. The experiment outcomes employed ten datasets of different sizes with 10,000-300,000 instances. Only 8 out of 248 characteristics were chosen, with classifiers percentages averaging 65%, 93.8%, and 95.5%.
- Research Article
1
- 10.22266/ijies2024.1031.72
- Oct 31, 2024
- International Journal of Intelligent Engineering and Systems
Network traffic classification has become more important with the rapid growth of the Internet and online applications.The rapid development of the Internet has enabled explosive growth of various network traffic.The challenge lies in how to classify and identify different categories of network traffic among these huge network traffic.The classification with the massive data network traffic suffers from noise and imbalanced data.Traditional classification algorithms are becoming less effective in handling these issues of the large number of traffic generated by these technologies.This paper proposes an advanced clustering model to enhance network traffic classification and improve the quality of services based on Advanced Density-Based Spatial Clustering of Applications with Noise (A-DBSCAN) with similarity and probability distance.A-DBSCAN with adaptive parameters are applied to identify clusters.The similarity distance is utilized to distinguish between clusters to identify the quality of clusters, where the value of similarity between (-1,1).Moreover, the cluster with a value similarity of more than 0 is identified as a highquality cluster.The probability distance is used to re-evolve the instances of negative clusters to suitable positive clusters.This stage results in consolidated optimal clusters to overcome the problem of imbalances data in the dynamic network efficiently.Additionally, the standard classifiers, such as the Random Forest (RF), K Nearest Neighbours (KNN), Decision Trees (DT), and Nave Bayes (NB) classifier are utilized to classify data network traffic.Finally, the ISCX VPN-nonVPN dataset remarks as a benchmark to evaluate the proposed solution.The experiment results show that the performance evaluation achieves higher accuracy 81.9% compared to the standard classifiers and related works.
- Research Article
26
- 10.1109/mcom.001.2300361
- Sep 1, 2024
- IEEE Communications Magazine
- Giuseppe Aceto + 4 more
Traffic classification (TC) is pivotal for network traffic management and security. Over time, TC solutions leveraging Artificial Intelligence (AI) have undergone significant advancements, primarily fueled by Machine Learning (ML). This paper analyzes the history and current state of AI-powered TC on the Internet, highlighting unresolved research questions. Indeed, despite extensive research, key desiderata goals to product-line implementations remain. AI presents untapped potential for addressing the complex and evolving challenges of TC, drawing from successful applications in other domains. We identify novel ML topics and solutions that address unmet TC requirements, shaping a comprehensive research landscape for the TC future. We also discuss the interdependence of TC desiderata and identify obstacles hindering AI-powered next-generation solutions. Overcoming these roadblocks will unlock two intertwined visions for future networks: self-managed and human-centered networks.
- Research Article
- 10.32620/reks.2024.3.09
- Aug 28, 2024
- Radioelectronic and Computer Systems
- Arkadii Kravchuk + 1 more
The subject matter of this article is the methods to detect distributed denial-of-service (DDoS) attacks at the Hypertext Transfer Protocol (HTTP) level with the purpose of justifying the requirements for creating software capable of identifying malicious web server clients. The goal of this article is to develop an information technology to evaluate the efficiency of DDoS attack detection methods, which will quantify their operating time, memory consumption, and approximate classification accuracy. In addition, this paper proposes hypotheses and a potential approach to improve existing application-layer DDoS attack detection methods with the intention of increasing their accuracy and identification speed. The tasks of this study are as follows: to analyse modern methods for detecting application-layer DDoS attacks; to investigate their features and shortcomings; to develop a software system to assess DDoS attack detection methods; to programmatically implement these methods and experimentally measure their performance indicators, specifically: classification accuracy, operating time, and memory usage; to compare the efficiency of the investigated methods; to formulate hypotheses and propose an approach to improve existing methods and/or develop new methods based on the results obtained. The methods employed are abstraction, analysis, systematic approach, and empirical research. In particular, the datasets generated by DDoS utilities were processed using the synthetic minority oversampling technique (SMOTE) to balance them. Furthermore, the studied DDoS attack detection methods were implemented, including fitting the required parameters and training artificial neural network models for evaluation. The following results were obtained. The average classification accuracy, operating time, and random-access memory (RAM) consumption during Internet traffic classification were determined for six DDoS attack detection methods under the same conditions. This study has demonstrated that the development of a novel method to detect DDoS attacks at the HTTP level with enhanced accuracy and classification speed is strongly required. The experimental results demonstrate that the time series-based method exhibited the shortest operating time (1.33 ms for 5000 vectors), whereas the deep neural network-based method exhibited the highest average classification accuracy (ranging from 99.07% to 99.97%) and the lowest memory consumption (39.09 KB for 5000 vectors). Conclusions. In this study, a software system was developed to assess the average accuracy of DDoS attack classification methods and measure the computational resources utilized. The scientific novelty of the obtained results lies in the formulation of two hypotheses and a potential approach to the creation of a novel method for detecting DDoS attacks at the HTTP level, which will have both high classification accuracy and a short operating time to surpass previously studied analogues in these respects. The first hypothesis is based on the additional usage of HTTP request attributes during Internet traffic classification. The second hypothesis is to analyse a graph of user transitions between website pages. The article also superficially describes a potential approach that involves the implementation of the described hypotheses as well as the proposed software architecture of an application-layer DDoS attack detection system for the Kubernetes platform and the Istio framework, which addresses the issue of collecting web request parameter values for websites that use the cryptographically secured HTTPS protocol.
- Research Article
22
- 10.1109/tnsm.2024.3366848
- Jun 1, 2024
- IEEE Transactions on Network and Service Management
- Eyal Horowicz + 2 more
Internet traffic classification has been intensively studied over the past decade due to its importance for traffic engineering and cyber security. A promising approach to several traffic classification problems is the FlowPic approach, where histograms of packet sizes in consecutive time slices are transformed into a picture that is fed into a Convolution Neural Network (CNN) model for classification. However, CNNs (and the FlowPic approach included) require a relatively large labeled flow dataset, which is not always easy to obtain. In this paper, we show that we can overcome this obstacle by using Contrastive Representation Learning in order to learn from an unlabeled flow dataset a flow representation that can be embedded in a latent space, enabling clustering of flows belonging to the same class together. We then show that by using just a few labeled flows (a few shots) from each class, we can achieve high accuracy in flow classification. We show that common picture augmentation techniques can help, but accuracy improves further when introducing augmentation techniques that mimic network behavior, such as changes in the RTT (Round-trip time). Finally, we show that we can replace the large FlowPics suggested in the past with much smaller mini-FlowPics and achieve two advantages: improved model performance and easier engineering. Interestingly, this even improves accuracy in some cases.
- Research Article
1
- 10.55041/isjem01429
- Mar 23, 2024
- International Scientific Journal of Engineering and Management
- Debmalya Ray
Internet traffic classification is a fundamental task for network services and management. There are good machine learning models to identify the class of traffic. However, finding the most discriminating features to have efficient models remains essential. In this paper, we use interpretable machine learning algorithms such as random forest and gradient boosting to find the most discriminating features for internet traffic classification. This paper aims to overcome these challenges by proposing machine learning classification mechanism.
- Research Article
5
- 10.1016/j.comcom.2023.10.011
- Nov 10, 2023
- Computer Communications
- Ofek Bader + 4 more
OSF-EIMTC: An open-source framework for standardized encrypted internet traffic classification
- Research Article
1
- 10.54254/2755-2721/6/20230939
- Jun 14, 2023
- Applied and Computational Engineering
- Tianhao Fu + 2 more
Network traffic classification is significant due to the fast growth of the number of internet users. The traditional way of classifying the large number of traffic generated by these users is becoming less effective. Therefore, many researchers made a network traffic classifier based on deep learning. However, those classifiers do not provide far better results and perform poorly when dealing with encrypted information. This paper tries to approach highly accurate and robust results in both encrypted and unencrypted networks by using machine learning algorithms. The algorithm used is the convolutional neural network (CNN). The performance of the proposed CNN is compared with that of the classical LeNet-5 network. Experimental results show that the classifier based on the proposed CNN performed better when dealing with both encrypted and unencrypted datasets, achieving a maximum average accuracy of 83.55%. Moreover, it is not sensitive to hyper-parameter choices, indicating its superiority in robustness. Compared with traditional network classifiers, the network classifier based on CNN can improve accuracy and improve stability.
- Research Article
5
- 10.48047/nq.2021.19.5.nq21080
- May 20, 2023
- neuroquantology
- Saumitra Chattopadhyay + 1 more
In this article, we offer a method to successfully enhance a traffic classifier based on Naive Bayes Prediction with a limited collection of training samples. Here, you should put forth a fresh traffic classification strategy to make use of data from correlated traffic patterns produced by an application. Traffic patterns are represented by statistical characteristics that are extracted. To implement Correlation based feature selection to eliminate pointless and superfluous features from the feature collection with low intercorrelation and high class-specific correlation. The classification ability of different traffic classification methods can be successfully demonstrated by Naive Bayes prediction. In comparison with current state-of-the-art traffic classification methods, suggested strategycan produce significantly improved classification results. Getting a high-performance analytic characteristic is one of the system's goals.
- Research Article
2
- 10.1093/comjnl/bxad022
- Mar 13, 2023
- The Computer Journal
- Oğuzhan Erdem + 2 more
Abstract Decision tree (DT)-based machine learning (ML) algorithms are one of the preferred solutions for real-time internet traffic classification in terms of their easy implementation on hardware. However, the rapid increase in today’s newly developed applications and the resulting diversity in internet traffic greatly increases the size of DTs. Therefore, the tree-based hardware classifiers cannot keep up with this growth in terms of resource usage and classification speed. To alleviate the problem, we propose to group application classes by certain rules and create an individual small DT per each group. In this article, a pipelined organization of multiple DT data structures, called pipelined decision trees, is proposed as a scalable solution to tree-based traffic classification. We also propose two distinct algorithms, namely confusion matrix-based class aggregation and leaf count-based class aggregation algorithms, to set group creation rules that allows traffic classification on pipelined smaller DTs in a hierarchical order. We further designed an hardware engine on field programmable gate arrays, which can search those pipelined trees within a single clock cycle by transforming them into bit vectors and implementing multiple range comparisons in parallel. Our architecture with 12 classes can run in 928.88 giga bit per second and achieve 96.04% accuracy.
- Research Article
46
- 10.1016/j.cose.2022.103000
- Nov 2, 2022
- Computers & Security
- Adi Lichy + 4 more
When a RF beats a CNN and GRU, together—A comparison of deep learning and classical machine learning approaches for encrypted malware traffic classification
- Research Article
2
- 10.11591/ijai.v11.i3.pp1175-1183
- Sep 1, 2022
- IAES International Journal of Artificial Intelligence (IJ-AI)
- Erick A Adje + 2 more
Internet traffic classification is a fundamental task for network services and management. There are good machine learning models to identify the class of traffic. However, finding the most discriminating features to have efficient models remains essential. In this paper, we use interpretable machine learning algorithms such as decision tree, random forest and eXtreme gradient boosting (XGBoost) to find the most discriminating features for internet traffic classification. The dataset used contains 377,526 traffics. Each traffic is described by 248 features. From these features, we propose a 12-feature model with an accuracy of up to 99.76%. We tested it on another dataset with 19626 flows and obtained 98.40% of accuracy. This shows the efficiency and stability of our model. Also, we identify a set of 14 important features for internet traffic classification, including two that are crucial: port number (server) and minimum segment size (client to server).
- Research Article
27
- 10.1016/j.neucom.2022.06.055
- Jul 1, 2022
- Neurocomputing
- Ali Rasteh + 5 more
Encrypted internet traffic classification using a supervised spiking neural network
- Research Article
4
- 10.7717/peerj-cs.860
- Feb 28, 2022
- PeerJ Computer Science
- Meijia Zhang + 4 more
Internet traffic classification is fundamental to network monitoring, service quality and security. In this paper, we propose an internet traffic classification method based on the Echo State Network (ESN). To enhance the identification performance, we improve the Salp Swarm Algorithm (SSA) to optimize the ESN. At first, Tent mapping with reversal learning, polynomial operator and dynamic mutation strategy are introduced to improve the SSA, which enhances its optimization performance. Then, the advanced SSA are utilized to optimize the hyperparameters of the ESN, including the size of the reservoir, sparse degree, spectral radius and input scale. Finally, the optimized ESN is adopted to classify Internet traffic. The simulation results show that the proposed ESN-based method performs much better than other traditional machine learning algorithms in terms of per-class metrics and overall accuracy.
- Research Article
42
- 10.1016/j.comcom.2022.02.003
- Feb 11, 2022
- Computer Communications
- Sangita Roy + 2 more
Fast and lean encrypted Internet traffic classification
- Research Article
12
- 10.1109/access.2022.3227073
- Jan 1, 2022
- IEEE Access
- Vanice Canuto Cunha + 4 more
Internet traffic classification aims to identify the kind of Internet traffic. With the rise of traffic encryption and multi-layer data encapsulation, some classic classification methods have lost their strength. In an attempt to increase classification performance, Machine Learning (ML) strategies have gained the scientific community interest and have shown themselves promising in the future of traffic classification, mainly in the recognition of encrypted traffic. However, some of these methods have a high computational resource consumption, which make them unfeasible for classification of large traffic flows or in real-time. Methods using statistical analysis have been used to classify real-time traffic or large traffic flows, where the main objective is to find statistical differences among flows or find a pattern in traffic characteristics through statistical properties that allow traffic classification. The purpose of this work is to address statistical methods to classify Internet traffic that were little or unexplored in the literature. This work is not generally focused on discussing statistical methodology. It focuses on discussing statistical tools applied to Internet traffic classification Thus, we provide an overview on statistical distances and divergences previously used or with potential to be used in the classification of Internet traffic. Then, we review previous works about Internet traffic classification using statistical methods, namely Euclidean, Bhattacharyya, and Hellinger distances, Jensen-Shannon and Kullback–Leibler (KL) divergences, Support Vector Machines (SVM), Correlation Information (Pearson Correlation), Kolmogorov-Smirnov and Chi-Square tests, and Entropy. We also discuss some open issues and future research directions on Internet traffic classification using statistical methods.