Font Size: a A A

Video Traffic Classification Based On Deep Neural Networks And Semi-Supervised Learning

Posted on:2023-06-14Degree:MasterType:Thesis
Country:ChinaCandidate:M R ZhouFull Text:PDF
GTID:2568306836472054Subject:Multimedia Communications and Video Business Streams (Professional Degree)
Abstract/Summary:
With the diversified development of network service types and the rapid growth of network resource deployment scale,traffic classification technology plays a crucial role in network resource allocation and network security prevention.Traffic classification methods based on deep learning often require a large number of labeled datasets to achieve excellent performance.However,collecting and labeling a large number of datasets requires huge time and labor costs,and the continuous innovation of Internet technology has accelerated the speed of change in the network environment.,resulting in time-consuming and laborious sample collection at risk of becoming obsolete.In view of the above problems,this paper proposes a traffic classification method based on deep neural network and semi-supervised learning with the help of the idea of transfer learning.The Pearson correlation coefficient is used to screen out 90 traffic statistical features with less correlation,so as to reduce the probability of misclassification.Through the method of model transfer and semi-supervised training,the model is first pre-trained on a large set of unlabeled old samples,and after the training is completed,it is transferred to a new model with more linear layers,and a part of the new model has observed a large amount of traffic The input-output mapping relationship of the category,so only a small amount of labeled data is required to retrain it.The experimental results show that the proposed method utilizes 40% of the labeled samples in the entire dataset,and achieves a classification accuracy of 98.3%,which effectively reduces the need for the number of labeled samples for the deep learning model and avoids the waste of outdated and old samples.This paper compares the characteristic probability distribution in different data sets,analyzes the concept drift phenomenon between different data sets,blends the labeled data into the unlabeled data sets according to the proportion of 20%,40% and 60%,and explores the impact of the concept drift amplitude on the classification accuracy of the model.The experimental results show that the overall accuracy of the proposed method is 98.5%,98.6% and 99.2% respectively when the concept drift of different ranges occurs in the dataset.When new classes(classes that do not exist in the pre training unlabeled dataset)appear in the labeled dataset,the classification accuracy of the proposed method is 93.8%,which is significantly improved compared with other literature methods.This paper also explores the impact of different packet length sequences on the performance of the model.The results show that short packet sequences are more difficult to classify than long packet sequences.
Keywords/Search Tags:Semi-supervised learning, deep learning, network traffic classification
Related items