Dissertation > Excellent graduate degree dissertation topics show
The Research of Text Clustering and Keywords Extraction Based on Complex Network Theory
Author: XieFengHong
Tutor: XieFuDing
School: Liaoning Normal University
Course: Computer Software and Theory
Keywords: text clustering keywords extraction weighted complex network density set multifeature value
CLC: TP391.1
Type: Master's thesis
Year: 2011
Downloads: 207
Quote: 0
Read: Download Dissertation
Abstract
|
With the rapid development of information technology, the number of text data is increasing amazingly. How to quickly access the useful text information in large text data, properly manage and use these text messages has become the urgent problem. Getting use of the data mining technology reasonable can efficiently help to solve this problem.Text clustering and text keyword extraction is an important field in text mining research. Text clustering divides the text of document into several clusters, which requires that the texts assigned to each cluster are more similar to each other than the texts assigned to different clusters. As an unsupervised machine learning method, text clustering doesn’t require the training set or need to know the number of clusters in advance. It has a great of flexibility and reality. Text keyword extraction is one of the important text information processing technology. It is the premise and foundation of information processing including automatic categorization, automatic clustering, automatic summary generation and so on.This thesis introduced the background of the text mining and text keywords extraction, research significance, research status and relevant theoretical knowledge. This thesis summarized domestic and foreign classics theoretical knowledge, proposed a new text clustering method and a new text keywords extraction. Main work includes the following two aspects:1. Based on partitioning community in complex network a text clustering method is proposed. Firstly, a new algorithm for detecting community structures in a weighted complex network is proposed. To partition the weighted complex network into groups, the algorithm looks for the density sets constantly and some proper operations are executed. Secondly, the proposal is applied to cluster text documents which are represented by the vector space model. A weighted complex network is constructed in terms of the similarity between two documents calculated by the cosine function. And then the community structure in this network is detected by the proposed algorithm. Finally, the experiment results show that the proposed algorithm has a good clustering efficiency by clustering some samples of Reuters-21578 data sets.2. Analyzed the characteristic and disadvantages of the existing keywords extraction algorithm based on complex network, a new keywords extraction algorithms based on weighted complex network is proposed. First of all, a weighted complex network model is constructed according to the relationship between the feature words of text. Secondly, the weighted clustering coefficient and betweenness are introduced to calculate the node’s multi-feature value. Finally, the keywords are extracted by the multi-feature value. The experiment results show that the keywords extracted in this algorithm have great contribution to the text subject, and the accuracy of keywords extraction is better than the existing algorithms.
|
Related Dissertations
- Research and Implementation of Mining Implicit User Interest,TP311.13
- Evolutionary Clustering Algorithm and Its Application,TP311.13
- A Novel Labels and Similarity Reconstruction Based on K-means Algorithm Application on Text Clustering,TP181
- Research and Application on Short Message Text Clustering,TP391.1
- Research on Antomatic Chinese Text Summarization of Web-oriented Text Mining,TP391.1
- Discovery of Texts’ Hot Topics Based on Improved TF-IDF,TP391.1
- The Research on the Constructing of Ontology Knowledge Base on Social Security Audit,TP391.1
- Research of the Influence of Node Failure to Network Status Based Directed Complex Networks,TP393.08
- Research on Summary Sentence Selection and Ordering in Query-focused Multi-document Summarization,TP391.1
- Research of Text Clustering on Food Complaint Documents Based on Ontology,TP391.1
- The Research on Methods of Text Mining Based on aiNet,TP311.13
- Based on frequent term sets text clustering,TP311.13
- Research on Some Field Text Information Processing Based on Latent Semantic Analysis,TP391.1
- Chinese clustering search engine major technology research,TP391.3
- Chinese blog for information collection and retrieval bias,TP391.3
- Research of Text Clustering Technology Based on Colony Intelligence,TP391.1
- Data Feature Extraction of Blogs and Filtering of Splogs Based on Classification,TP393.092
- Research on Chinese Text Clustering of Neural Network of Support Vector Machine,TP391.1
- Research and Implementation of Clustering Systems of Web Search Results,TP391.3
- The Research of the Clustering Ensembles Based on SEAM Algorithm and It’s Application on Text,TP301.6
CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer applications > Information processing (information processing) > Text Processing
© 2012 www.DissertationTopic.Net Mobile
|