Dissertation > Excellent graduate degree dissertation topics show

The Research on Text Classification Based on Clique Model

Author: HuXiaoHui
Tutor: WangMingWen
School: Jiangxi Normal University
Course: Computer Software and Theory
Keywords: Text Classification Text group Graph model Link Web page classification
CLC: TP391.1
Type: Master's thesis
Year: 2008
Downloads: 30
Quote: 4
Read: Download Dissertation

Abstract


In recent years, with the number of electronic documents on the Internet exponentially increase, text classification techniques in information retrieval, information filtering, and content management applications become more and more important, has become an information retrieval and machine learning the field of cutting-edge research. Automatic text categorization is a powerful means of information in the organization and management of the text, to make it easier for users to more accurately locate desired information can solve the messy problem of information to a greater extent. Automatic text classification given classification system for automatic processing of documents of unknown class, and in accordance with the process of document characteristics to determine its category; text classification techniques based on machine learning has become a mainstream technology. Currently, researchers have proposed many a mature text classification algorithm, most of these algorithms from pattern classification, such as KNN classification algorithm, support vector machine algorithm. Most of the existing text classification algorithm based on the vector space model, did not consider the semantic features of the document structure information. For the lack of the traditional classification of text classification and its related technologies, proposed two effective resolved or improved methods and techniques. Content and innovative work of this research include the following two points. 1) This paper studies a text classification method based on the text group on the training set by the textual similarity matrix constructed text similar to Figure extract text from the figure group (complete subgraph) to construct the information from each category group classifier, and then combined with SVM classifier. Experiments on Fudan University Chinese Text Classification Corpus and 20 Newsgroups corpus, and at the same time under the same pretreatment conditions, with the traditional classification methods of comparative experiments, experiments show that our proposed method on two data sets greater improvement the classification performance. 2) With the rapid growth of the Web information, especially the increase of online information on the Internet, no longer rely on artificial ways to process information is impractical. Therefore, the web page automatic classification has become a key technology has great practical value, is a powerful means of organization and management of data. In order to effectively organize the extremely rich source of information on the Internet, the Web page automatic classification has become an increasingly important area of ??research. The WEB document has its own characteristics, subject to a lot of the attention of scholars in recent years, traditional text WEB document classification, classifier has its own limitations, so the characteristics of the Web document, on the basis of the traditional classification use of WEB document link information. The experiment shows that the classification of this article on the data set provided by the North Skynet combined page link information to improve the classification effect.

Related Dissertations

  1. Research on Text Classification Based on Biomimetic Pattern Recongnition,TP391.1
  2. Modeling Internet Link Delay Based on Measurement,TP393.4
  3. Based on Data Distribution Characteristics of Text Classification,TP391.1
  4. Research on Ontology-based Scientific Papers of Chinese Classification,TP391.1
  5. The Procedure of the Tax Law Enforcement and the Procedure of the Tax Law Judicature,D922.22
  6. Internet phone service quality assurance mechanism,TP393.02
  7. Research on cross-language text categorization,TP391.1
  8. Classification model based monitoring of e-commerce Prohibited Research and Implementation,TP393.09
  9. Graph model based on statistical and machine learning algorithm and its application of certain,TP181
  10. DSR protocol in Ad Hoc Networks Research and Optimization,TN929.5
  11. Underactuated three-link robot control strategy,TP242
  12. Link Aggregation of asymmetric exchange scheduling algorithm,TP301.6
  13. Research on Estimation and Prediction of Urban Link Travel Time Based on Low-Frequency Floationg Car Data,U491
  14. High-end core routers distributed link aggregation technology research and implementation,TP393.05
  15. Spam Filtering System Study Based on Non-negative Matrix Factorization,TP393.098
  16. Research and Implementation on Key Technology of Web Text Collection and Analysis,TP391.1
  17. The Optimization Study on IEEE802.16 QoS Scheduling and Handover Mechanism,TN929.5
  18. Experimental Study on RC Frames Strengthened with Y-eccentrically Brace,TU392.1
  19. Application Study on Pre-stressed Oblique Lifting-Link in Girder Transfer Structure Construction,TU755.2
  20. Design and Implementation of Power Meter Reading System Based on ZigBee Wireless Network,TM764
  21. Development of New Kind of Steel for High Strength Round Chain for Mining of Ф34×126–C,TF762

CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer applications > Information processing (information processing) > Text Processing
© 2012 www.DissertationTopic.Net  Mobile