Dissertation > Excellent graduate degree dissertation topics show
A Study on the Clustering of Chinese Web Text
Author: YuWenMing
Tutor: WanJian
School: Hangzhou University of Electronic Science and Technology
Course: Computer Software and Theory
Keywords: Text Clustering K- average Suffix tree Mapping reduction
CLC: TP391.1
Type: Master's thesis
Year: 2009
Downloads: 170
Quote: 1
Read: Download Dissertation
Abstract
|
With the popularity of the Internet, the text information on a web page showing a trend of explosive growth. How to indexing, retrieval, management, mining the mass text information on the Web has become a huge challenge facing the field of computational science. The emergence of text clustering technology provides an effective way for the mass text information classification management and visualization. Text clustering technique as an unsupervised machine learning methods in recent years in the field of information retrieval, multi-text automatic summarization Internet has been widely used. The object of this discussion is the Chinese Web text clustering, based on the inspection of existing domestic and foreign academic achievements as well as the latest research found, more in-depth study of two typical scenarios: (1) News Chinese text clustering massive text number in the portal; (2) Chinese search engine returns results scene in real-time clustering. In the first scene, in the MapReduce distributed parallel computing framework to design a series of distributed improvement of the text clustering algorithm. In the text pre-processing stage, in order to evaluate the importance of a text in a word of the text, in MapReduce designed and implemented a new iterative algorithm for a calculation words tfidf weight. Text clustering stage, first a rough distance measure individual text in the text collection is divided into overlapping subset. Then step clustering on the basis of further design distributed K-average text clustering algorithm, this step text more precise measure of distance, but the complexity of the algorithm is also quite high. Finally, a series of new improved algorithm given above in MapReduce platform, a distributed system of the Chinese text clustering, the efficient and stable system running massive text clustering task. The actual Chinese corpus on the experimental show that the proposed method can effectively deal with the a large text clustering problem sets, a linear relationship between the size of the text number and clustering time within a certain range, and relatively satisfied with the quality of clustering. In the second scene, combined with the re-calculation of the weights of the vector space model, a suffix tree-based clustering method for Chinese text fragments. First Chinese word segmentation tool focused on the Chinese text fragments in the text pre-processing stage, the use of each one pick out the meaning of words (usually a verb or noun). Filtered off after using a linear algorithm to construct a Chinese suffix tree node has too high a document frequency (phrase), and use of the proposed a formula for computing the score of the node (phrase), select a high score phrase as a final text feature . Then use the text feature to redefine the method of calculation of similarity between the Chinese text fragments. Combination of text features extracted from the suffix tree similarity calculation in the new text, this paper quasi real agglomerative hierarchical clustering algorithm. The experimental results show that the new method can improve the quality of clustering and clustering speed to meet the needs of real-time applications. This work provides a good example of the clustering problem for a particular field of Chinese studies, and the accumulated experience and results obtained in the experiment also has certain reference value and practical value.
|
Related Dissertations
- The Research and Implementation of Protein Classification Algorithm on the Basic of String Kernel,TP301.6
- Research and Implementation of Mining Implicit User Interest,TP311.13
- Research and Application on Short Message Text Clustering,TP391.1
- Web news hot discovery system design and implementation,TP393.09
- Research of Finding Maximal Unique Matches in Genome,TP301.6
- Optimization of SOM Algorithm and Application in Chinese Text Clustering,TP391.1
- Cloud storage system for mass data,TP333
- The Design and Implementation of the Hot Education News Topic Detection System,TP391.1
- Research of Text Clustering on Food Complaint Documents Based on Ontology,TP391.1
- Evolutionary Clustering Algorithm and Its Application,TP311.13
- The Research and Implementation of Biological Sequence Alignment,TP399-C8
- Based on latent semantic indexing text clustering algorithm,TP301.6
- Open electronic document plagiarism detection services to build technology research,TP309.7
- Chinese blog for information collection and retrieval bias,TP391.3
- Research of Text Clustering Technology Based on Colony Intelligence,TP391.1
- Data Feature Extraction of Blogs and Filtering of Splogs Based on Classification,TP393.092
- Research on Chinese Text Clustering of Neural Network of Support Vector Machine,TP391.1
- Research and Implementation of Clustering Systems of Web Search Results,TP391.3
- The Research of the Clustering Ensembles Based on SEAM Algorithm and It’s Application on Text,TP301.6
- Research on Text Clustering Algorithm Based on 2 Degree Frequent Word Sequence,TP301.6
CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer applications > Information processing (information processing) > Text Processing
© 2012 www.DissertationTopic.Net Mobile
|