Dissertation > Excellent graduate degree dissertation topics show
The Research and Implementation of Keyword Extraction
Author: LuoZhunChen
Tutor: WangTing
School: National University of Defense Science and Technology
Course: Computer Science and Technology
Keywords: keyword extraction key phrases separating models mutual information parameter table of word-sequence boundary feature selection machine learning linguistic feature
CLC: TP391.1
Type: Master's thesis
Year: 2008
Downloads: 280
Quote: 8
Read: Download Dissertation
Abstract
|
Keywords are widely used in many applications such as Information Retrieval, Automatic Summarizing, Text Classification, Text Clustering and so on. Only a small minority of documents have author-assigned keywords, and manually assigning keywords to documents is very laborious. Therefore it is highly desirable to automate the keyword extraction process. Many academic journals require their authors to provide a list of about five to fifteen keywords on the first page of each article. Since these keywords are often phrases consisting of two or more words, we prefer to call them key phrases. Most of the keywords from other kinds of documents are also actually phrases, which make the task more difficult.This paper argues that the keywords extraction can be treated as two problems: extracting key words and extracting key phrases. A keywords extraction method based on separating models was proposed for extracting keywords from the documents. This method develops different features for the two mentioned problems in order to improve the accuracy. This paper also considers the problem of automatically extracting keywords from text as a supervised learning task. We treat a document as a set of words or phrases, which the learning algorithm must learn to classify as positive or negative examples of keywords. Based on the different structure of the key words and key phrases, we develop a set of features. For example using the features of mutual information and parameter table of word-sequence boundary can improve the phrases identification. We also use the part-of-speech rule of key words and key phrases to develop some linguistic features to improve the result of extracting key words and key phrases.Based on the above work, we run the experiment to evaluate the effect of the keyword extraction method based on separation model. The result shows that, using the same features, the performance of keyword extraction algorithm based on separation model is better than that based on integrated model. In addition we also evaluated the effect of features for key words and key phrases. At last, to compare the work with the famous keywords extractor KEA, we implemented an keywords extractor based on separation model adopting different key words’features and key phrases’features. The result shows that our extractor is better than KEA.
|
Related Dissertations
- Word Sense Disambiguation Corpus Automatic Acquisition,TP391.1
- Research on Text Classification Based on Biomimetic Pattern Recongnition,TP391.1
- Feature Extraction, Selection and Combination in Lipreading,TP391.41
- Research on Feature Selection and Construction in Emotion Speech Recognition,TP18
- Based on Data Distribution Characteristics of Text Classification,TP391.1
- Research and Implementation of a Dynamic Feature Selection Method for Vehicle Recognition System,TP391.41
- Algorithm and Implementation Based on Non-rigid Medical Image Registration,TP391.41
- Research on Face Recognition Based on AdaBoost Algorithm,TP391.41
- Research on Feature Extraction, Selection and Classification Algorithms for Pulmonary CAD,TP391.41
- Research and Implement of Chinese Word Segment Techniques Based on the Conditional Random Field,TP391.1
- Semi-supervised Learning and Active Learning of Sentiment Classification Coupled with Domain Knowledge,TP181
- A Static Behavior-Based Method to Detect Malware on Android,TP309
- Supervision topic model research and application,TP391.1
- Distortion effects on image quality evaluation and classification,TP391.41
- Learning-based human motion synthesis inverse kinematics,TP391.41
- Content-based malicious code detection Research and Implementation,TP393.092
- Home Academic Information Extraction System,TP393.092
- Attribute-based collection of product reviews mining research,F274
- Social network algorithm and its application in recommendation,TP391.3
- Application of Data Mining in Email Anti-Spam System,TP393.098
- Medical image registration platform and combines new gray with geometric information Registration Measures,TP391.41
CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer applications > Information processing (information processing) > Text Processing
© 2012 www.DissertationTopic.Net Mobile
|