Dissertation > Excellent graduate degree dissertation topics show

Research of Web Information Extraction Based on DOM and Ontology

Author: WangYuanLong
Tutor: ZhangLi
School: Liaoning University
Course: Applied Computer Technology
Keywords: Web information extraction DOM Similarity ontology semanticsannotation
CLC: TP391.1
Type: Master's thesis
Year: 2012
Downloads: 92
Quote: 0
Read: Download Dissertation

Abstract


At present with the explosive growth of the Internet,it is more and moredifficult to find useful information for users. Under the circumstances, how toextract the useful information from Web page with large data and lacking structurehas become a research focus. A mass of methods for Web information extractionhave been proposed at home and abroad in recent years. These methods got a goodeffect as a whole. But, the research of Web automatic information extractionbecomes the main research direction in general. However, if the method of Webautomatic information extraction can get an effect, it needs a lot of training samplesto train. Meanwhile, it needs heavy work.Aiming to the above shortages, an automatic method for Web informationextraction is proposed. The main content is shown as follows.Firstly, because lots of HTML pages is not with standard grammatical structure,the pages serious impact on the Web information extraction effect. So this paperclears the HTML pages On the basis of predecessors by wiping off the uselessinformation such as advertisement and the content which is nothing to the extractingdata. The cleared pages become not only in XML standard, but also removeinterference of irrelevant content. Meanwhile, at the process of parser HTML pagesby DOM API, the path of DOM tree from root node to leaf node is got, and preparesfor the future works very well.Secondly, by the comparison of two DOM trees of Web pages which come fromthe same data source and is the same category, the set of different DOM trees withthe same DOM path and different entity information is obtained. Then, by gatheringand merge-ring the DOM trees in the set through calculating the Similarity of twoDOM trees, the data zone containing the extracted information can be obtained.Finally, the information extracted by the above method lacks semantics andcannot be helpful for the understanding of human and computer. Meanwhile, theextracted information lacked semantics cannot be deal with very well. The paperbuilds the ontology to give the extracted information the corresponding semantics. Firstly, this paper builds the ontology for the extracted information item on basis ofpredecessors. Then, it needs to enlarge the above data zone appropriately to containthe extracting information and prefix characters information. The informationsuitable for the concept of ontology will be extracted by the prefix charactersinformation to achieve the semantics annotation and information extraction.The method of this paper is easy to operate. The method only needs two goodtraining Web documents to train and automatically extract the entity information.The experiment demonstrates that the method has the higher precision for the webpages with the same data source and category. At the same time this method hascertain flexibility for different data source Web pages with the same category.

Related Dissertations

  1. Syntactic Features Based Pronoun Resolution,TP391.1
  2. Research of Multiple Emails Automatic Summarization,TP391.1
  3. Research of IRC Botnet Detection Based on Behavior,TP393.08
  4. Research on Auto-Evaluation Method of Programming Based on Similarity,TP312.1
  5. The Effect of Instruction for Middle School with Philosophy,G633.6
  6. Comprehensive Evaluation of Flue-cured Tobacco Quality in Pingdingshan and Comparative Analysis with American Tobacco,S572
  7. Murine Peritoneal Macrophages Transcriptional Responses Following in Vivo Infection with Streptococcus Suis Type 2,S858.91
  8. The Impact of Tourism on Typical Vegetation in Luya Mountain Nature Reserve, Shanxi Province,S759.9
  9. Ontology -based Distributed Description Logic Modular Construction Methods,TP391.1
  10. Ontology -based Semantic Web service matching and composition method,TP393.09
  11. WordNet and the \,G254
  12. Graph-based similarity search scientific literature key technology research,TP391.3
  13. In Four Regions of Qinghai Fritillaria.Unibracteata Hsiao et K.C.Hsia RAPD Analysis of Genetic Diversity,S567.231
  14. In Some Areas of Qinghai Cordyceps Sinensis RAPD Analysis of Genetic Diversity,S567.35
  15. Research on the Patent Map Based on Domain Ontology,TP391.1
  16. Research and System Implementation of Image Retrieval Method Based on Fuzzy Clustering,TP391.41
  17. A Comparative Study of Nominal Suffixes between Chinese and English,H314
  18. The Research and Implement of Chinese Input System Based on Dynamic Self-adaptation Language Model,TP391.14
  19. Anomaly Detection Research Based on Similarity Analysis of Time Series,TP311.13
  20. Similarity-preserving Linear Maps on B(x),O177.2
  21. Research of Text Clustering on Food Complaint Documents Based on Ontology,TP391.1

CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer applications > Information processing (information processing) > Text Processing
© 2012 www.DissertationTopic.Net  Mobile