Dissertation > Excellent graduate degree dissertation topics show

Research of Weight Algorithm in KNN Text Classification

Author: ZhaoXiaoHua
Tutor: MaJianFen
School: Taiyuan University of Technology
Course: Applied Computer Technology
Keywords: Text Classification Feature selection Feature weight KNN TFIDF
CLC: TP391.1
Type: Master's thesis
Year: 2010
Downloads: 227
Quote: 5
Read: Download Dissertation

Abstract


With the popularity of the Internet, and now the government, enterprises, families are using the online electronic resources, or to send a large number of electronic documents on the Internet, such as e-news, e-mail, e-books, and so on. Since the rapid expansion in the number of the electronic document on the Internet, how quickly within a short period of time, the effective processing of a huge number of text documents, has become a hot research. At present, the text categorization technology has been widespread concern because of its practicality. Text classification has called this a popular research topic. Text classification is the text to be classified according to their content automatically divided into predetermined good category. Text classification steps generally include the text pre-processing, feature selection and feature weight calculation, classification, performance evaluation. Feature weight calculation method is a key link in the text classification related to the classification of the final results. The traditional feature weight algorithm TFIDF use important consideration to appear more frequently in the text, and other text in the text set feature fewer items. Because the TFIDF no taking into account the relationship between categories of the feature items and text TFIDF algorithm accuracy is not high. In view of the the the TFIDF algorithm itself defects, we guess whether the calculation function based on the aggregation feature selection algorithm to form a new weighting function TFIDF function so that you can avoid the defects in the traditional weights. In this paper, re-calculation method of feature selection for text classification and rights discussed proposed new class-based distribution of information and the location information of the new term weight algorithm. In this paper, the Chinese Lexical Analysis System Institute of Computing ICTCLAS (Institute of Computing Technology, Chinese Lexical Analysis System). Test use of the system in the KNN classifier TFIDF improved algorithm, based on the experimental results the following conclusions can be drawn: (1) select the mutual information of the expected cross entropy, the different characteristics of the information gain, the text weight of evidence, and CHI Probability and Statistics The selection algorithm for the comparative experiments, experimental results show that the feature selection algorithm is relatively large impact on classification accuracy. In these types of algorithms, feature selection effect of the text weight of evidence is the best, followed by information gain, CHI, expect cross entropy, mutual information classification worst effect. (2) to compare the weights have calculated respectively selected as the TFIDF the TF * CHI, the TFIDF * CHI single category classification results by different types of micro-precision experiments feature weights formula. Through experiments, we can draw the following conclusions: TF * CHI improved formula weight category classification results either recall or check the precision is the worst TFIDF * CHI weights formula with the traditional TFIDF formula is not very different from the overall point of view, but improved weight calculation method can be found through careful study of the impact on the already high precision category is not very big, but it was originally classified effect The not-so-good category classification performance has greatly improved. (3) Select the mutual information, the expected cross entropy, information gain, the text weight of evidence, and CHI probability and statistics different feature selection algorithm, and then were to select the TFIDF TF * characteristic selection function, the TFIDF * Feature selection function different weights value calculating method, the use of KNN classification model, comparative experiment, comparison improved feature item weight calculation method of the final classification results. Experiments can be seen, no matter what kind of feature selection function, function selected as TF when the feature weights * Feature selection function, the classification results are lower than the weighting functions TFIDF TFIDF * Feature selection function situation. 5 kinds of feature selection in addition to the information gain function in any of the other kind of high classification accuracy than the original TFIDF the TFIDF * Feature selection function. This case illustrates, and can not arbitrarily say certain what kind of feature selection algorithm or weight calculation method is necessarily good, but depends on the combination of the feature selection algorithm and weights algorithm, only with the right to play to their strengths, optimizing the classification results .

Related Dissertations

  1. Research on Text Classification Based on Biomimetic Pattern Recongnition,TP391.1
  2. Tourism Comments on the Internet’s Semantic Analysis and Usefulness Research,TP391.1
  3. Feature Extraction, Selection and Combination in Lipreading,TP391.41
  4. Based on Data Distribution Characteristics of Text Classification,TP391.1
  5. Research on Feature Extraction, Selection and Classification Algorithms for Pulmonary CAD,TP391.41
  6. Research for Event Extraction Method in Specific Domain Based on Tree Conditional Random Field,TP391.1
  7. One kind of empirical data on the workload of a software bug fixes Prediction Model,TP311.53
  8. Ontology-based medicine named entity recognition technology research,TP391.1
  9. Classification model based monitoring of e-commerce Prohibited Research and Implementation,TP393.09
  10. Research on Web Content Filtering Based on Concepts of Collection,TP393.092
  11. Sentiment Classification by Combining Lexicon-based and Machine Learning Methods,TP391.1
  12. Research of Spatial Outliers Detection Algorithm,TP311.13
  13. The Research and Implementation of Web Page Classification in Enterprise Search Engine,TP393.092
  14. Text Categorization of Chinee Small Sample Based on Graph Model,TP391.1
  15. Research on the Keyphrase Extraction and Relevant Technology,TP391.1
  16. Research and Realization on Correlation Techniques of Topic Search-Specific Engine,TP391.3
  17. Production Forecast for Daliudi.S1 Gas Reservoir,P618.13
  18. Mining and Analysis of Relation Users in Social Network,TP391.3
  19. On Research for Chinese Automatic Text Categorization Technology Based on VSM Model and Feature Selection,TP391.1
  20. Lead-free Piezoelectric Ceramics (K_ (0.44)) Na_ (0.52) Li_ (0.04)) (Ta_ (0.2) Nb_ (0.8)) O_3-xCeO_2 Preparation and Properties,TM282
  21. A Study on GM and Its Functions in ESD,H05

CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer applications > Information processing (information processing) > Text Processing
© 2012 www.DissertationTopic.Net  Mobile