Font Size: a A A

Data Mining On The Classification Of Unipolar And Bipolar Depression

Posted on:2022-03-22Degree:MasterType:Thesis
Country:ChinaCandidate:M Y GaoFull Text:PDF
GTID:2504306338473844Subject:Master of Applied Statistics
Abstract/Summary:
With the increasing social pressure,the incidence rate of depression at home and abroad is increasing year by year.This has aroused widespread concern in the society,and depression has also become a hot issue in today’s society.In the field of depression,some scholars have made some explorations,such as using traditional statistical methods to test the significance of the factors affecting the severity of depression,continuously optimizing the scale to assess the degree of depression of patients,and so on,There is no objective and clear quantitative indicators,and it is particularly difficult for doctors to diagnose depression types of new patients.With the rapid development of computer technology and big data,the combination of machine learning and medical data has made great progress in the information construction of medical field,especially the application of some classification algorithms has played an important role in promoting the medical cause.Based on the real data of patients with depression in a hospital in Sichuan Province,this paper established a binary classification model for the diagnosis of depression.Firstly,the original data is preprocessed.including original data integration,data cleaning,data type conversion and so on.Then,the traditional statistical methods are used to test the significance of discrete variables and continuous variables respectively,and remove their collinearity.Then,the random forest,xgboost and lightgbm models are established by removing all the collinearity variables.Then,the characteristics with high contribution to the model are selected to model again,The results of the two modeling were analyzed and compared after 5000 times of simulation,so as to verify whether the influence factors we screened have reference value.At the same time,we will also explain the advantages and disadvantages of each model according to the differences of AUC,F1,recall and running time of the three models.Finally,the classification effect of the three models on different real datasets is compared,and the influencing factors of depression types screened out by lightgbm model with the best performance are analyzed,so as to assist clinical diagnosis and provide support and help for the medical industry.
Keywords/Search Tags:Depression, Data mining, Statistical test, LightGBM, Classification
Related items