Dissertation > Excellent graduate degree dissertation topics show

The Research on Topic Extraction from Web Pages Based on Semantic

Author: ZhangWei
Tutor: ShaoXiongKai
School: Hubei University of Technology
Course: Applied Computer Technology
Keywords: Subject extraction TF * IDF WordNet Semantic similarity
CLC: TP391.1
Type: Master's thesis
Year: 2011
Downloads: 75
Quote: 0
Read: Download Dissertation

Abstract


When people search via the Internet on the same theme, through the link showing many contain themes relevant pages, in order to enable users to more accurately find useful information, or faster and more easily transferred to the related topics of interest, good maintaining the correct identification of these closely linked page is very meaningful. Maintenance of numerous web pages set or by artificial means, therefore, more intelligent identification technology urgent needs. This paper focuses on the introduction of the ontology database, web theme extraction algorithm combines semantic information. WordNet ontology, for keyword semantic similarity calculation, and then carry out the merger of the keyword semantic level semantic abstract summarized the entire text information, so as to achieve the effect of the themes are more accurate. This method improves the existing manual methods or word frequency statistics for lack of subject extraction, the search becomes self-energy, more accurate search results. The implementation of this method is divided into the following steps. First, obtain a sample Web page filter, clear information unrelated to the topic, such as web pages contain pictures, advertising and navigation links and other information, the removal of this information makes during the subject extraction process will not be affected by these theme extracting contribute little interference of information more accurately, which makes the extraction of the final subject. And then converted to the text information, because in the subsequent operation, a direct page clustering calculation amount is large, and the efficiency is not high. Then, the initial treatment of text information coreference resolution and word processing. Introducing coreference the digestion system allows higher frequency pronouns resolve, a lot of help on the subject extraction accuracy. Using the Lucene segmentation operation, you can distinguish between noun phrases, such as names, terms, and also provides for the function of the word index construction, improve the the theme extraction efficiency and accuracy. Secondly, the introduction of WordNet, improved Lesk algorithm to calculate the semantic similarity, and clustering of web pages, gathered in a different class, so as to reduce the noise generated by the different theme pages so that the page has a different theme. Using semantic similarity between MeSH, not just a simple calculation of the inner product of the keyword can guarantee semantic information is preserved between the pages. Finally, the use of improved TF * IDF algorithm Web MeSH extraction, improved TF * IDF algorithm taking into account the distribution of feature items in between classes and within classes to make up for the lack of TF * IDF, the ultimate theme Words more accurate. Finally, the experiment proved that the proposed method is feasible and effective.

Related Dissertations

  1. Ontology -based Semantic Web service matching and composition method,TP393.09
  2. WordNet and the \,G254
  3. Based on community structure found in the data warehouse subject extraction studies,TP311.13
  4. The importance of link-based prediction algorithm for dynamic linking,TP393.03
  5. Ontology-based Knowledge Retrieval System Smart Grid,TM76
  6. Reasoning knowledge retrieval method based on ontology and SWRL,TP391.3
  7. Research on Information Retrieval Technology Based on Semantic Web,TP391.3
  8. Research on the Semantic Annotation of Domain with WordNet,TP391.1
  9. A Improved Text Similarity Model Based on PageRank Value,TP391.1
  10. Discovery of Texts’ Hot Topics Based on Improved TF-IDF,TP391.1
  11. Web Services Discovery Based on Kernel and WordNet,TP393.09
  12. On Research for Chinese Automatic Text Categorization Technology Based on VSM Model and Feature Selection,TP391.1
  13. SMS User Interest Hierarchy Algorithm Based on Text Classification Algorithm,TP391.1
  14. Research and Implementation of Method for Component Testing Information Semantic Retrieval System Based on Ontology,TP311.52
  15. Researching on Chinese Text Classification Based on Naive Bayes and BP Neural Network,TP391.1
  16. Research on the Algorithm for Chinese Duplicated Web Pages Detection,TP393.092
  17. Blog Distillation with Feed Similarity Distribution,TP393.092
  18. Chinese Spelling Correction Research in Search Engines Based on Statistical Model,TP391.3
  19. Based on semantic similarity analysis of the soft component clustering algorithm,TP311.13
  20. Research of Feature Selection and Weighting Algorithm in Text Classification System Based on SVM,TP391.1

CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer applications > Information processing (information processing) > Text Processing
© 2012 www.DissertationTopic.Net  Mobile