Dissertation > Excellent graduate degree dissertation topics show
Research of Optimizing on KNN Algorithm Based on Clustering Concept in Text Classification
Author: LinFeng
Tutor: ChenQiuLian
School: Guangxi University
Course: Computer System Architecture
Keywords: Feature Bit-string Feature Multi-Class Matrix clusteringconcept KNN text classification
CLC: TP391.1
Type: Master's thesis
Year: 2013
Downloads: 1
Quote: 0
Read: Download Dissertation
Abstract
|
In the era of knowledge and economy, information has become one of the most important signs; more and more attentions focus on the management and obtaining of information. The structure of the information has changed, from structured, semi-structured to unstructured. The organization and processing of unstructured information have become increasingly important, text classification as one of the key technologies have been widely used in the field of information retrieval, knowledge discovery and management, but for amount of information in vast texts, the efficiency and accuracy of text classification severely restrict the application in the field of real-time.The existing classification methods are based on statistical theory and machine learning; well-known methods of text-classification are Bayes, k-nearest neighbors, support vector machine, neurons network. In these methods, KNN as a simple, effective, non-parametric method, widely used in text classification and achieved good results of classification. But KNN algorithm calculation limits its application in the field of real-time, so how to improve the efficiency of KNN has been widely concerned, the focus of this paper also gathered on improving the efficiency of text categorization without decreasing classification accuracy.In this paper, we proposed the clustering concept, and then base on this clustering, we present two methods, Feature bit-string, Feature Multi-Class Matrix, to improve the efficiency of text classification.The main innovation of this paper as follow:1. The clustering concept based on the semantic. In the text, there will be some different forms of each concept, and the similarity calculation of texts can not recognize the relationship between these kinds of word, and we often ignore the effects of them. In this paper, we would cluster these into a concept. Experimental results show that the clustering concept can effectively express the meaning of these features, and the similarity calculation of texts can also reflect its contribution, and improve the accuracy of text classification and reduce the dimension of the text vector.2. The amount of computation could be reduced by feature bit-string of text. To overcome the large calculation of KNN algorithm, we proposed feature bit-string of text, which could quickly filter out the texts that may be similar with test text. Thereby reducing training set of texts could decrease the amount of computation of KNN. Theoretical analysis and experiments show that feature bit-string of text could improve the efficiency of KNN classification algorithm without reducing the classification accuracy.3. The feature multi-class matrix also could reduce the amount of computation of KNN algorithm. After the analysis of KNN, we could not only filter the similar text, but also get rid of the impossible class, to reduce the amount of computation. Thereby reducing the possible category of the text could reduce the amount of calculation by feature multi-class matrix. The experiments show that feature multi-class matrix can effectively improve the speed of classification.
|
Related Dissertations
- Research on Text Classification Based on Biomimetic Pattern Recongnition,TP391.1
- Tourism Comments on the Internet’s Semantic Analysis and Usefulness Research,TP391.1
- Based on Data Distribution Characteristics of Text Classification,TP391.1
- Transportation network path and neighbor selection study,U116.2
- Research on Improved K Neighbor Support Vector Machine Algorithm Faced Text Classification,TP391.1
- The Research on Cllective Multi-Label Classification,TP391.1
- Financial Index Forecast Based on Data Mining,F830
- Feasibility Study on Discriminating Adulterated Milks by Near Infrared Spectroscopy,TS252.7
- Computer-aided Method of Detecting Lung Fissure Based on CT Images,TP391.41
- Ontology-based medicine named entity recognition technology research,TP391.1
- Research on Relation Extraction of Person Entity in News Webpage,TP393.092
- Subspace-based face recognition algorithm,TP391.41
- Research and Application on Outlier Data Mining Algorithm in Large Data Set,TP311.13
- Research on Web Content Filtering Based on Concepts of Collection,TP393.092
- Content-based spam filtering technology research,TP393.098
- Research and Implementation on Key Technologies of Web Text Mining Oriented to Enterprise Competitive Intelligence,TP391.1
- The Research and Implementation of Text Classification Based on Meta-Information and Optimization,TP391.1
- Research of Spatial Outliers Detection Algorithm,TP311.13
- The Research and Implementation of Web Page Classification in Enterprise Search Engine,TP393.092
- Text Categorization of Chinee Small Sample Based on Graph Model,TP391.1
- Research and Realization on Correlation Techniques of Topic Search-Specific Engine,TP391.3
CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer applications > Information processing (information processing) > Text Processing
© 2012 www.DissertationTopic.Net Mobile
|