Dissertation > Excellent graduate degree dissertation topics show
Research of Domain Text Classification Based on Semantics
Author: ZhangGuiXi
Tutor: ZuoJianZhuo
School: Beijing University of Technology
Course: Information and Communication Engineering
Keywords: Text Classification Semantics Distance map Support Vector Machine
CLC: TP391.1
Type: Master's thesis
Year: 2011
Downloads: 106
Quote: 1
Read: Download Dissertation
Abstract
|
With the rapid development of information technology, especially the popularity of Internet applications, the text information on the Web explosive show in front of people. The proliferation of information we need to effectively classify it, filter efficiency to continue to raise awareness in the mass of information to find effective. The text classification technology is under a given classification system according to the contents of the text or attributes automatically by a computer will be a large amount of text attributed to the appropriate category. It is able to handle large amounts of text, Information disorders of the status quo can be solved to some extent, user-friendly positioning to find the desired information. Traditional text classification algorithm using keywords as features to build that keywords are independent of each other, there is no semantic association. Been rapid development, but still faces some problems, if not taking into account the structure of the text information and the rich semantics associated characteristics lead to classification effect failed to reach a satisfactory accuracy. In recent years, the network and rich semantic data available, such as WordNet, Wikipedia. This paper argues that the full use of these rich semantic data is one of the ideas to improve the text classification performance. Text categorization primary facing the problem of how to represent text data, text representation model. Most text classification methods are based on the vector space model. This text representation model method is relatively simple, but it triggered a high dimensional vector space model and the sparsity problem, on the one hand the text classification with a very high time complexity; the other hand, ignoring the semantic feature between correlation, resulting in the loss of a large number of semantic information, thus obtained feature vector is not a good representation of the text contents. All of these issues are interfering with the efficiency and accuracy of text classification, text classification performance declined. In order to solve the above problems, the references of semantic dictionary WordNet semantic information, the proposed new text representation model, the research prototype system to achieve a text classification. For different types of features of a document, select the model or based on the distance map of the text based on the text representation of the concept representation model, wherein in the concept vector space model, the concept as a text feature, will have synonymous keywords relationship mapped to The same concept; model analysis of the structure of the document, the text feature item indicates the distance node co-occurrence relations between feature items described as the edges of the graph structure, in turn mapped to the text graph structure in the distance graph structure . Finally, support vector machine (SVM) classification algorithm text representation model combination and use Fudan University Chinese corpus to test the system performance compared with traditional text classification system experiments from the recall rate, accuracy and F-Measure 3 results show that the method proposed in this paper system has improved the overall effect than traditional text classification system, which the average accuracy rate of 12.49% increase in average recall rate of 13.5%, the average F-Measure increased by 23.16%. In short, the article is intended for a few specific areas, characteristic extraction of text classification, text representation of the key technologies of the theoretical analysis and experimental verification of the model, text classification algorithm, proposed solutions, and experimental results prove that these programs effectiveness. These algorithms and models for future research text classification and other text processing problems will have a certain reference value and reference.
|
Related Dissertations
- Research on Automatic Detection Algorithm for Substructure Distress of Highway Pavement Based on SVM,U418.6
- Research on Autamatic Music Structrue Analysis,TN912.3
- Research on Transductive Support Vector Machine and Its Application in Image Retrieval,TP391.41
- Research on Text Classification Based on Biomimetic Pattern Recongnition,TP391.1
- Tourism Comments on the Internet’s Semantic Analysis and Usefulness Research,TP391.1
- Fault Diagnosis Method Based on Support Vector Machine,TP18
- Process Support Vector Machine and Its Application to Satellite Thermal Equilibrium Temperature Prediction,TP183
- Study of the Bank’s ATM Interface Design,TP11
- Research for Infrared Image Target Identification and Tracking Technology,TP391.41
- Study on the Road Condition Monitoring Based on Vehicular 3D Acceleration Sensor,TP274
- Research of Diagnosing Cucumber Diseases Based on Hyperspectral Imaging,S436.421
- Based on Data Distribution Characteristics of Text Classification,TP391.1
- A Study on the Predicate-complement Structure of Er Pai,H141
- The Summary of Redundancy in Chinese,H146
- The Research on Intrusion Detection System Based on Machine Learning,TP393.08
- Pragmatic logic true value of,H030
- \,H131
- The Study of Chinese Farmer’s Proverb,H136
- Comparative Study of Chinese and English parenthesis,H314
- On its post- semantic study the impact of general semantics,H030
- Cognitive Semantics Studies of Modern Chinese Touch Sensory Words,H136
CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer applications > Information processing (information processing) > Text Processing
© 2012 www.DissertationTopic.Net Mobile
|