Dissertation > Excellent graduate degree dissertation topics show
The Chinese Text Categorization Research Based on Support Vector Machine and Clustering Algorithm
Author: WuHuiXiong
Tutor: TanJunShan
School: Central South University of Forestry Science and Technology
Course: Applied Computer Technology
Keywords: Support Vector Machine Chinese text classification Clustering Reduce
CLC: TP391.1
Type: Master's thesis
Year: 2009
Downloads: 91
Quote: 0
Read: Download Dissertation
Abstract
|
With the rapid development of the Internet, the surge in the number of a variety of electronic text data, how to quickly and efficiently access, management and use of these text data, has become an urgent need to address the important issues of Information Systems Science. Over the past decade, as one of the basic tools to solve these problems, automatic text classification based on text content, an unprecedented development, has aroused widespread concern. Automatic Text Categorization is the process by a computer automatically determine the text category text categorization problems with text vector sparsity, dimension between the characteristics of greater relevance. Support vector machine is not sensitive features and sparse, and has more advantages than other methods in dealing with the problem of high-dimensional support vector machine is very suitable for text classification problems, and has great application in text classification potential. , Text classification for support vector machine proposed a number of challenging issues, support vector text classification training classification model is slower significant defects. This article there is training for text classification for support vector classification model slower defects, try some pretreatment of the support vector machine training samples, support vector machine in order to try to overcome such defects, better used text classification. Previous studies have shown the establishment of the support vector machine depends only on the support vectors and non-support vector in the training sample, so this article is to cut the number of the vector angle from the lower text classification process Chinese text vector concentration of non-support vector accelerate the speed of training support vector machine classifier. In this paper, the k-means clustering algorithm based on density concept, make full use of the known training text category clustering in the training sample pretreatment stage as much as possible to reduce training samples of non-support vector support vector reserved greatly reduced, making the final number of samples involved in training support vector machines, which generated one is both quite traditional support vector machine classification classification accuracy, but also in terms of time complexity greatly improved Chinese text classification model. Meanwhile, for some of the shortcomings exist in the k-means clustering algorithm targeted improvements involved in clustering data standardization, so involved in clustering data by the irregular distribution of cluster-like distribution, The clustering process so as to facilitate the introduction of density to determine the initial cluster centers, the introduction of a the purity metric to determine the appropriate density of the cluster center radius, and thereby ultimately to achieve the most effective reduction SVM training sample.
|
Related Dissertations
- Research on Automatic Detection Algorithm for Substructure Distress of Highway Pavement Based on SVM,U418.6
- Research on Autamatic Music Structrue Analysis,TN912.3
- Research on Transductive Support Vector Machine and Its Application in Image Retrieval,TP391.41
- Research and Implementation of Mining Implicit User Interest,TP311.13
- Establishment and Update of Similar Users’ Cluster in Personalized Information Retrieval,TP391.3
- Research on Removal Algorithm of Shadows in Image Segmentation,TP391.41
- The Research of the Text Extraction Method Based on Spectral Cut,TP391.41
- Fault Diagnosis Method Based on Support Vector Machine,TP18
- Process Support Vector Machine and Its Application to Satellite Thermal Equilibrium Temperature Prediction,TP183
- Gao Zhong-ying academic thought and experience and use of Bufei Decoction treatment of common diseases of the respiratory system drug law,R249.2
- Research and Improvement on K-Means Clustering Algorithm,TP311.13
- Research of Scheduling Algorithm Based on Hybrid Adaptive Genetic Algorithm in Computing Grid,TP393.09
- Evaluation of Photosynthetic Efficiancy of Seedlings of the Hybrid Progenies (F1) in Peach,S662.1
- Research for Infrared Image Target Identification and Tracking Technology,TP391.41
- The Application of Ant Colony Algorithm in Meteorological Satellite Cloud Pictures Segmentation,TP391.41
- Research on Clustering Algorithm Based on Mutation Particle Swarm Optimization,TP18
- Research on K-means Optimization Clustering Algorithm,TP311.13
- Research on Fuzzy C-Mean Clustering Algorithm Based on Particle Swarm Optimization and Shuffled Frog Leaping Algorithm,TP18
- Research on Clustering Algorithm Based on Genetic Algorithm and Rough Set Theory,TP18
- Study on Photosynthetic Characteristics of Peach Based on Heterosis of Assimilation Capacity,S662.1
- Study on the Road Condition Monitoring Based on Vehicular 3D Acceleration Sensor,TP274
CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer applications > Information processing (information processing) > Text Processing
© 2012 www.DissertationTopic.Net Mobile
|