Dissertation > Excellent graduate degree dissertation topics show
The Research of Keywords Extraction Algorithm in Text Mining
Author: WangLiangFang
Tutor: LiuDuanYang
School: Zhejiang University of Technology
Course: Applied Computer Technology
Keywords: Tongyici Cilin Lexical Chain Keyword Extraction Semantic Analysis Semantic Expansion
CLC: TP391.1
Type: Master's thesis
Year: 2013
Downloads: 1
Quote: 0
Read: Download Dissertation
Abstract
|
As the information technology develops continually, a large amount of text information stored in the form of computer readable and many areas of information emerged as the explosive growth. So how to extract useful information to readers quickly and accurately in the large amount of information will be an important issue. Keywords extraction is an effective means of solving the above problems. Keywords refine subject information of texts, to enable readers to grasp the important information about the text quickly, and to improve the efficiency of the access to information. So it has obvious practical significance.Keywords extraction is one of the core technologies in the field of text mining and plays a very important role. The main carrier of information is based on text. However the vast majority of text messages are also not yet provided keywords. Existing keyword extraction algorithms can not be a good solution to the problems about word sense disambiguation, synonym redundant expression, over-fitting in the classifier training process and lexical chain can not expresses the text semantic structure accurately and so on. So this thesis proposes two improved method based on semantic analysis. They mine the potential expression of text theme in the semantic level. The proposed methods are better to solve the problems about word sense disambiguation, lexical chains expresses the text semantic structure accurately and comprehensively, etc. While synonymous avoid redundant expression, especially articles have synonym the more the more obvious. The main work of this thesis are shown as followed.1. Keywords extraction algorithm based on semantic dictionary and Lexical ChainThe semantic dictionary of Tongyici Cilin not only coding is simple but also has the advantage of synonyms group that is more abundant and easy to semantic understanding than other knowledge bases, and the vocabulary chains express the semantic structure of the text and multiple topics excellently. Therefore, a complete keywords extraction algorithm based on semantic dictionary and Lexical Chain is proposed ultimately. Its name is KETCLC (Keyword Extraction based on Tongyici Cilin and Lexical Chain). It analysis the characteristics of Tongyici Cilin and lexical chains, combing both of them and according to the advantage of their combined, then makes the text processing to improve the quality of the keyword extraction in terms of preprocess, polysemy disambiguation, synonym mergence, the construction of lexical chains, feature selection and improvement of weights computation.2. Keywords Extraction Algorithm Based on Semantic Expansion Integrated With Lexical ChainThe building methods of lexical chain are based on semantic similarity values calculated or semantic relevancy values calculated independently at present. However, these lexical chains can not express the association and semantic relations between words accurately and fully. Accordingly, they affect the accuracy of the expression of the theme of the article and the quality of keywords extraction. Therefore, a complete keywords extraction algorithm based on semantic expansion integrated with lexical chain is proposed ultimately. Its name is KESELC (Keyword Extraction based on Semantic Expansion integrated with Lexical Chain). It calculates the semantic similarity and semantic correlation based on the semantic dictionary of Tongyici Cilin from the perspective of semantic analysis, then comes to a semantic expansion degree and its calculation method considering the both of the above. Finally, it integrated semantic extension with lexical chain to extract keywords. It is good to mine the vocabularies that their frequency is not high but having an important contribution to the article.The experimental results show that the above two methods both are able to take full account of the semantic knowledge. The extracted keywords not only avoid a redundant expression, but also cover the subjects of the article accurately and comprehensively. The above two methods of keyword extraction both have excellent performance to extract keywords from Chinese texts.
|
Related Dissertations
- Tourism Comments on the Internet’s Semantic Analysis and Usefulness Research,TP391.1
- A Study of English Neologisms in Semantics and Context,H313
- Research of Intelligent Search Engine Based on Semantic Web,TP391.3
- Research on Topic Clustering Model of Socail Tagging Based on Bayesian Theory,C93
- Ontology-Based Hazard Information Extraction from Chinese Food Complaint Documents,TP391.1
- Chinese Keyword Extraction Method Based on Word Span and Its Application in Text Classification,TP391.1
- English verbs hate comparative study psychology class,H314
- The Research on the Automatic Generation of Concept Weights and Optimal Composition for Test Item in E-learning System,TP391.6
- A Study of Based on LSA and Paragraph Clustering of Automatic Abstracting System,TP391.1
- Research on Approaches of the Subjective Automated Assessment,TP391.1
- Finding Web Services Based on Clustering Probabilistic Semantic Approach,TP393.09
- The Reasreach and Implementation of Semantic-Based Event Extraction Method for Chinese Text,TP391.1
- Emotional background music for movies Classification Method,TP391.1
- Image / video automatic replacement of the face,TP391.41
- Based on the model and features of the package ML_pLSA Target Recognition Algorithm,TP391.41
- Based on semantic analysis of text mining research,TP391.1
- Detection and Analysis of Malware Network Behavior Based on Isolation Environment,TP393.08
- Conceptual Integration Approach to Semantic Analysis of Chinese Noun-Noun Compounds,H136
- Chinese named entity recognition and disambiguation of,TP391.1
- Research of Keywords Extraction Algorithm for Chinese Text Based on Gene Expression Programming,TP391.1
- Research on Text Clustering Based on Latent Semantic Analysis and Self-organizing Maps,TP391.1
CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer applications > Information processing (information processing) > Text Processing
© 2012 www.DissertationTopic.Net Mobile
|