Dissertation > Excellent graduate degree dissertation topics show

Study on Semi-supervised Text Clustering Based on Pairwise Constrains

Author: LiuLongHai
Tutor: ZhongJiang
School: Chongqing University
Course: Computer Software and Theory
Keywords: semi-supervised study pairwise constrain text clustering active study latent semantic analysis
CLC: TP391.1
Type: Master's thesis
Year: 2011
Downloads: 101
Quote: 0
Read: Download Dissertation

Abstract


With the development of Internet and information technology, more and more text needs addressing in daily life and work. How to process lots of text quickly and effectively comes to be a serious problem. Text clustering is an important text processing technology. Though traditional text clustering technology is with high flexibility and automated processing capability, requiring neither training process nor marking text category in advance, there still exist some problems. In many applications, we might get a large number of unlabeled samples as well as a few labeled samples. However, the traditional text clustering algorithm can not boost the clustering performance with these labeled samples. For semi-supervised learning can use both labeled and unlabeled data to build better learners, more and more researchers concern about it.In this thesis, the semi-supervised clustering algorithm is introduced to the text clustering to boost the text clustering performance. The improvement of the performance of the semi-supervised clustering algorithm mainly depends on the supervised information used by the semi-supervised clustering algorithm. In order to address this problem, a new selection method of supervised information was proposed. Firstly, the least well-defined cluster was identified with the fuzzy hyper volume method. Secondly, several data items were selected at the boundary of the identified cluster. The closest data item in its nearest cluster was found, which with the selected data item formed a pairwise constraint. The new selection method of supervised information can select most informative supervised information.In order to improve the results of text clustering with a few pairwise constrains, a semi-supervised text clustering algorithm based on pairwise constraints was proposed. At first, latent semantic analysis was used to descend the dimension of text features. In the clustering process, the proposed selection method of supervised information was used to actively select pairwise constraints by which to steer the clustering process towards an appropriate partition.To verify the effectiveness of the new selection method of supervised information and semi-supervise text clustering algorithm based on pairwise constrains, several experiments were carried out in the Chinese text clustering system. The results show that the new selection method of supervised information can select supervised information with rich information and the new semi-supervised text clustering algorithm could significantly improve the results of text clustering with a few constrains.

Related Dissertations

  1. Research and Implementation of Mining Implicit User Interest,TP311.13
  2. Evolutionary Clustering Algorithm and Its Application,TP311.13
  3. Research on Topic Clustering Model of Socail Tagging Based on Bayesian Theory,C93
  4. Research of Text Clustering on Food Complaint Documents Based on Ontology,TP391.1
  5. A Novel Labels and Similarity Reconstruction Based on K-means Algorithm Application on Text Clustering,TP181
  6. The Research on the Automatic Generation of Concept Weights and Optimal Composition for Test Item in E-learning System,TP391.6
  7. A Study of Based on LSA and Paragraph Clustering of Automatic Abstracting System,TP391.1
  8. Research and Application on Short Message Text Clustering,TP391.1
  9. The Design and Implementation of the Hot Education News Topic Detection System,TP391.1
  10. Research on Approaches of the Subjective Automated Assessment,TP391.1
  11. Finding Web Services Based on Clustering Probabilistic Semantic Approach,TP393.09
  12. Based on the model and features of the package ML_pLSA Target Recognition Algorithm,TP391.41
  13. The literature of resources for research and application clustering system,TP391.1
  14. Research on Data Mining Technologies Applied to Web Chinese Text,TP311.13
  15. Research of Patent Document Analysis and Retrieval Based on Latent Semantic Analysis,TP391.3
  16. The Research of Web Pages Information Extraction Based on Page Structure Analysis Technique,TP393.092
  17. The Research of Structured P2P Network Model Based on Semantic Search,TP393.02
  18. The Research of Modeling Multi-Networks Based on Unstructured Data,TP391.1
  19. Intent analysis and application of the search word,TP391.3
  20. Study on Exploration of Reputation Dimensions Based on Text-Clustering and Corpus,TP391.1

CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer applications > Information processing (information processing) > Text Processing
© 2012 www.DissertationTopic.Net  Mobile