Dissertation > Excellent graduate degree dissertation topics show
Classification Algorithm and Evaluation on Imbalanced Datasets
Author: LiMingFang
Tutor: ZhangHuaXiang
School: Shandong Normal University
Course: Computer Software and Theory
Keywords: Classification Unbalanced datasets BP algorithm Bagging algorithm Threshold evaluation criteria
CLC: TP181
Type: Master's thesis
Year: 2011
Downloads: 45
Quote: 0
Read: Download Dissertation
Abstract
|
The classification algorithm is one of the important research direction in the field of machine learning today, many classical algorithm has been widely used in the actual problem and achieved good results. Traditional classification algorithms assume that the data distribution balance, and accuracy was the main objective to improve the overall classification of the classification algorithm on the data set. However, the reality in many different categories of the data set contains the number of data is highly inclined, non-balanced, one or several categories of data quantity is far greater than the number of data in the other categories, such data sets are called class imbalance data. Number of dominant classes or categories known as the majority class, and other so-called minority class. Traditional classification algorithm classification imbalanced data, in order to improve the overall classification of the data set to the correct rate target, due to the small number of minority class sample, and thus to predict the tendency of the number of dominant majority minority class classification recognition rate; however, In many practical problems, to improve the recognition rate of minority class data is often more valuable than the majority class data recognition rate. However, in some extreme cases, the classification algorithm minority sample misclassification can still guarantee a higher overall accuracy. Therefore, how to improve the classification performance of the classification algorithm for a class imbalanced data, in particular, to improve the recognition rate of the classification algorithm for a minority class, becoming one of the hot research in machine learning. Unbalanced data classification study the following aspects: First, the data level, data pre-processing algorithm, usually by changing the distribution of the original data, to reduce the degree of imbalance, to achieve the purpose of the balance data, commonly used methods include sampling techniques, feature extraction techniques; followed by the algorithm level, by modifying the traditional classification algorithm to adapt to the imbalanced data problem, and thus tend to minority class classification to improve the recognition rate of minority class data commonly used techniques such as cost-sensitive learning threshold value method is an imbalance data classification performance evaluation criteria, due to the traditional classification algorithm performance evaluation criteria usually negligible the classification algorithm minority class data recognition rate, so how to choose the imbalance data classification evaluation criteria recent years; Last one of the hot G-Mean, F-Measure are unbalanced data commonly used standards. In this paper, the class imbalanced data classification problem, a study from the data level and the algorithm level and propose new evaluation criteria in the original performance evaluation criteria based on the combination of classic evaluation criteria. The main innovations are: (1) algorithms and data planes combined: the the traditional Bagging method and SMOTE algorithm combination, using the SMOTE algorithm for processing of sample concentration of minority class sample, Bagging algorithm based on category values ??and correct rate re-adjust the weight of each sample and base classifier. Typically, most algorithms for imbalanced data only applies to the two-class problem, the algorithm based on improved classification of multi-class imbalanced data. The experimental results show that the proposed algorithm can guarantee the overall classification accuracy achieved on the two types of unbalanced datasets and multi-class imbalanced data sets, but also improve minority class the purpose of classification accuracy. (2) the algorithm level: on the basis of the results of previous studies, a new application of the artificial neural network threshold criteria. And proved in theory: the standard can be free from the influence of the proportion of sample categories under the minority class and the majority class classification accuracy at the same time get the best. This standard back-propagation algorithm (BP algorithm) as base classifiers, combined with genetic algorithm to search the optimal threshold value, the experimental results show that: the new standard threshold artificial neural network can improve the classification accuracy of minority class data. In addition, this threshold criteria on the basis of preliminary propose a new evaluation criteria. Experimental results show that this standard focuses on the error rate of the sample of concern. (3) Data level: SMOTE only according to the linear rule the shortcomings of the new sample, a new over-sampling method. This new algorithm can be random walk in the minority class data based on the synthesis of new sample. Theory has proved, under the premise of certain assumptions established close to the expectation and variance of the mean and variance of the synthetic sample of the new algorithm with the original data in the minority class data. The experimental results showed that, from a statistical standpoint, when the group classification device C4.5, Naive Bayes (NB), k Nearest Neighbor (KNN) use a variety of sampling methods, the new algorithm is obviously superior to other sampling methods.
|
Related Dissertations
- K Company’s Improving Planning and Forecasting for the Reasonable Allocation of Inventory,F224
- Research on Decoding Algorithm for LDPC Codes,TN911.22
- The Simulation and Analysis of Spectrum Sensing Based on Higher Order Cumulants,TN911.23
- Research on the Classification Based on the Reconstruction of Solder Joint,TP391.41
- The Research and Implementation of Protein Classification Algorithm on the Basic of String Kernel,TP301.6
- Context-Dependent Lexical Paraphrasing,TP391.1
- Research on Text Classification Based on Biomimetic Pattern Recongnition,TP391.1
- Tourism Comments on the Internet’s Semantic Analysis and Usefulness Research,TP391.1
- Research on Classification Method of Tongue Substance Color and Tongue Coating Color Based on SVM,TP391.41
- Research and Application of Diverse Density Learning Algorithm,TP181
- Research on Target Tracking and System in Wireless Sensor Networks Based on Character of Movement and Terrain Restriction,TP212.9
- Study on the Preparation of the Low Molecular Weight Fuciodan from Sargassum Henslowianum (C.Agardh) and Anti-tumor Activity,TS254.9
- Uptake and Accumulation of Heavy Metals of Pistia Stratiotes and Eicharnia Crassipes under Combinated Pollution Condition,X173
- Application of Numerical Designing Cigarette Blending Formula on Computational Intelligence,TS44
- Establishment of the Molecular Identifying System of Curvualria and Application in Difficult Species,Q949.32
- Research on Nondestructive Detection Technology for External Qualities of Papayas Based-on Vision,S667.9
- Research on Design and Practice of Educational Video Co-Construction and Sharing Platform for the Public,G434
- Studies on Effects of Soybean Species on North Tofu Quality and Methods of Evaluation in North Tofu,TS214.2
- Comparison of Gene Expression Data Cluster Methods and Gene Network Construction for Phytophthora Sojae Genes,S435.651
- Study on Growth Monitoring Technique Based on Pixel Un-Mixing Method and HJ Remote Sensing Images in Paddy Rice,S511
- Methodology for Linkage-map-based Quantitative Trait Loci Synthesis Analysis,S562
CLC: > Industrial Technology > Automation technology,computer technology > Automated basic theory > Artificial intelligence theory > Automated reasoning,machine learning
© 2012 www.DissertationTopic.Net Mobile
|