Dissertation > Excellent graduate degree dissertation topics show

Text Classification Algorithm Based on Imbalanced Data Sets

Author: XieNaNa
Tutor: FangBin
School: Chongqing University
Course: Applied Computer Technology
Keywords: Imbalanced data sets text classification CHI-square selection method data distribution re-sampling
CLC: TP391.1
Type: Master's thesis
Year: 2013
Downloads: 25
Quote: 0
Read: Download Dissertation

Abstract


As the fast development of computer network technology, electronic documentsgradually become a main style of text information. The diversity and bad organizationof network information make users have difficulty finding the exact information thatthey really want. Text classification which is considered to be the most importanttechnology in information retrieval plays a great role in organizing the documents. Thedata sets in library for text processing are relatively balanced. However, this is differentfrom text categorization collections in practical applications, especially that the texts onthe network are often marked incomplete or imbalanced. The data imbalance problemhas become a major problem of text classification technology due to the broadapplication and its importance in various fields. Besides, text classification onimbalanced data sets is becoming a focus in text mining.In this paper some research has been done on text categorization on imbalanceddata sets. A new text classification algorithm on imbalanced data sets is proposed basedon the improvement of the feature selection in text classification and the re-sampling indata set layer. The main contents of this paper are as follows:①A deep research has been made on the traditional CHI statistical featureselection method and the one-sided metric CHI-square which only considered thepositive feature. However, the experiment result shows that they both give poorperformances.②Based on the research and analysis of the imbalanced data sets, a newimprovement on the one-sided metric CHI-square method is proposed. A tendentiousfactor is introduced to preserve part of the negative feature which may have acontribution on the classification of small class. Besides, in order to overcome thedefects of the CHI-square, the ICF (Inverse categorization frequency) is also proposedas a factor of the new feature selection method. The features which can most respect thecategorizations are selected by using the new method. Then, the texts of corpus arequantified to the vector space mode.③In order to solve the inefficient classification result because of the imbalanceddata, a re-sampling process is made on the data layer after the quantification of the textcorpus. First, a re-sampling method which is based on the combination of randomoversampling and random under-sampling is applied. Though it has better achieved filtering the imbalance of the data distribution and give a relatively balanced data setwhich is used to train a classifier. The random oversampling always lead to over-fittingin classification, while the random under-sampling can’t avoid to delete some sampleswhich play important role in classification which may produce the reduction of theclassification result. So an improvement of the combined re-sampling method isproposed by using the SMOTE on oversampling which often behaves well and theunder-sampling method based on improved clustering algorithm. The experiment resultsshow that the new method has produced a better classification result.

Related Dissertations

  1. Research on Text Classification Based on Biomimetic Pattern Recongnition,TP391.1
  2. Tourism Comments on the Internet’s Semantic Analysis and Usefulness Research,TP391.1
  3. Development of the Digitizer Based on PCI-E Bus,TP274.2
  4. An Approach for Identifying a Plant Resistance Gene Based on the Random Forest,Q943
  5. Study of Acoustical Signal Transmitter Which Based on Sigma-Delta Modulation,TN761
  6. Design of Receiving Processor on Radio Signal and Multi-Rate,TN851
  7. Space Infrared Target Simulation and Application of Target Trace,TP391.41
  8. Bayesian methods under two ordinal value of multi-value data model with simultaneous identification of outliers,O212.1
  9. Based on Data Distribution Characteristics of Text Classification,TP391.1
  10. Research of Text Categorization on Food Complaint Documentation Based on Ontology,TP391.1
  11. Research on Ontology-based Scientific Papers of Chinese Classification,TP391.1
  12. Chinese Keyword Extraction Method Based on Word Span and Its Application in Text Classification,TP391.1
  13. Sampling Based on Stochastic Optimization,O212.2
  14. Research on Improved K Neighbor Support Vector Machine Algorithm Faced Text Classification,TP391.1
  15. The Research on Cllective Multi-Label Classification,TP391.1
  16. Researching of Scheduling and Control Co-design Based on Variable Sampling Period in Networked Control System,TP273.5
  17. Quasi-Monte Carlo Method for the Structured Stochastic Variational Inequalities,O22
  18. Text classification based on Bayesian filtering technology research and implementation,TP393.098
  19. The Research of Markov Switching Cointegration Model,F224
  20. Research on Content-Based Chinese Junk Short Messages Classification Technology,TP391.1
  21. Research for Event Extraction Method in Specific Domain Based on Tree Conditional Random Field,TP391.1

CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer applications > Information processing (information processing) > Text Processing
© 2012 www.DissertationTopic.Net  Mobile