Dissertation > Excellent graduate degree dissertation topics show

Study on Chinese Text Classification Algorithm Based on Rough Set and It’s Application

Author: ZhangBaoFu
Tutor: ShiHuaJi
School: Jiangsu University
Course: Applied Computer Technology
Keywords: text classification TFIDF vector space model Rough sets attribute reduction
CLC: TP391.1
Type: Master's thesis
Year: 2010
Downloads: 155
Quote: 6
Read: Download Dissertation

Abstract


With the rapid development of the network technology, the amount of information on the network is dramatically increasing. It is how to effectively organize and manage these online documents that has become an urgent problem to be solved. Text classification has become a key solution to the problem. And it is an important branch of text mining, which has get more deeply researched because of its unique knowledge discovery function. Text classification has been in a wide range of application fields, such as information filtering, retrieval, digital library services and so on, and it,has broad application prospects.Rough set theory can deal with fuzzy and uncertain knowledge. it can effectively analyze and deal with incomplete, inconsistent, inaccurate data, without any prior information. Thus knowledge can be analyzed and dealt with using a mathematical method, and implicit knowledge can be discovered, and, potential rules can be revealed. The main idea of rough set theory is to lower the dimension of feature vectors without affecting the classification accuracy, and to obtain the simplest classification rules.This paper mainly researches the system of text classification based on the Rough sets theory systematically and deeply, and the algorithm were applied to the classification system in the Public Security Intelligence System. The main work of the paper is as follows:(1) This paper describes the relevant text classification techniques, and some commonly used text classification algorithm is a detailed analysis and comparison.(2) Aiming at the problem that document set is dealt with as a whole and distribution of features among and in classes is not taken into account when using traditional TFIDF method, an improved TFIDF method which is combined with information entropy is proposed. This method modifies the method of calculating weights of features of TFIDF by combining itself with information entropy of features among and in classes, which overcomes the defect that the features that made less contribution to the categorization are given greater weight, and calculates weights of text features more efficiently.(3)Aiming at the problem that the traditional feature selection method which filters features using frequency threshold would result in information loss and reduce the classification precision, a novel automatic text categorization method based on rough set is proposed. In the proposed method, the weighted attribute features discretization is carried out to form a decision table; then, selection of conditional attributes at the decision table is carried out on the basis of attribute significance which is based on dependency degree; finally, the reduction of text attribute features is performed by heuristic algorithm which is based on conditional information entropy.(4) The improved TFIDF method and a novel automatic text categorization method based on rough set which were proposed in this paper were applied to the public security intelligence classification subsystem. Practical application shows that the use of the system can obtain better results for text classification.

Related Dissertations

  1. Research on Text Classification Based on Biomimetic Pattern Recongnition,TP391.1
  2. Tourism Comments on the Internet’s Semantic Analysis and Usefulness Research,TP391.1
  3. Methodology for Linkage-map-based Quantitative Trait Loci Synthesis Analysis,S562
  4. Based on Rough Set of Urban Areas When Traffic Green Control System Research,TP18
  5. Incremental rough set attribute reduction,TP18
  6. Research on Ontology-based Scientific Papers of Chinese Classification,TP391.1
  7. The Forecast Methods of Mining Risk of the Coal Resources Affected by Dynamic Disaster,TD713
  8. Research on Improved K Neighbor Support Vector Machine Algorithm Faced Text Classification,TP391.1
  9. Public Research and implementation of information-gathering,TP311.13
  10. Research for Event Extraction Method in Specific Domain Based on Tree Conditional Random Field,TP391.1
  11. The Research of Text Classification Based on Hadoop,TP391.1
  12. The Implementation and Research of the Probabilistic Latent Semantic Analysis Model in the Search Engine’s Business Text Classification System,TP391.1
  13. The Research of Tibetan Text Categorization Base on N-Gram Information,TP391.1
  14. Design and Inplementation of a Data Classification System Based on Improved KNN Algorithm,TP391.1
  15. Sentiment Classification by Combining Lexicon-based and Machine Learning Methods,TP391.1
  16. Research on Focused Crawler Based on SVM Classification Algorithm,TP391.3
  17. Text based on SVM multi-class classification,TP391.1
  18. BBS Spam Filtering Model Based on Word Co-ocurrence,TP393.094
  19. The Research and Implementation of Web Page Classification in Enterprise Search Engine,TP393.092
  20. Text filtering isolation technology - based research,TP393.08
  21. Research on Adaptive Teaching and Research Information Recommendation Model,TP393.09

CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer applications > Information processing (information processing) > Text Processing
© 2012 www.DissertationTopic.Net  Mobile