Dissertation > Excellent graduate degree dissertation topics show

Research on Domain-oriented Deep Web Information Extraction

Author: GaoYuan
Tutor: GuYunHua
School: Nanjing University of Information Engineering
Course: Applied Computer Technology
Keywords: Deep Web domain ontology DIV blocks template table template template matching
CLC: TP391.1
Type: Master's thesis
Year: 2013
Downloads: 3
Quote: 0
Read: Download Dissertation

Abstract


The Deep Web contains more abundant and professional data resources compared to the Surface Web. With the rapid growth of information, research on the Deep Web has aroused increasingly more attention. The data of the Deep Web are semi-structured. How to extract the information and give them semantic information becomes one of the Deep Web widely concerned research topics.According to the application requirement of Deep Web information extraction, this paper uses several developing technologies, such as the Chinese word segmentation ontology modeling and machine learning, to do researches on the pretreatment of web pages、construction of domain ontology、construction of template and template matching, and carry on the Deep Web information extraction experiment on the field of weather and books. The main research work of this paper includes:(1) Research on pretreatment of web pages. Mainly studies how to present the HTML document as a hierarchical tree which has DIV block elements、attributes and text, and other process of pretreatment of web pages, which includes converting the DIV block to string flow、Chinese word segmentation and word frequency statistics. The target of these processes is to handle the HTML document into a data set which with the DIV block as basic unit and contains the segmentation results.(2) Research on construction of domain ontology. As a semantic foundation in communication between different subjects, domain ontology in the process of template construction can have the effect of optimization, reducing the unrelated content which would appear in the template.(3) Research on construction of dual template. In consideration of the features of the HTML pages which usually use "DIV+CSS" on the whole structure design and table layout in detail place, this paper uses the combinations of DIV block template and table template. Using the result of web page pretreatment, by the algorithm of C4.5decision tree to train the classifier which can select the number of extracted DIV blocks, building the template of DIV blocks which can locate the data area. Then using the technology of XML, constructing XSLT document under the help of machine, to form the table template which can extract the data fragment. Experiment results show that the accurate rate of the classifier trained by C4.5decision tree can reach95.2%, which can ensure that DIV blocks judgment will not be wrong. Eventually uses the dual template for extracting and average accuracy rate and recall rate can achieve95%above, better extraction effect is obtained. Dual template is more stable than single template and at the same time, more conducive to maintenance.(4) Research on template matching. Error caused by the traditional regular string matching method based on the URL is very big. This paper, on the basis of the traditional methods, combines web page similarity calculation method in template matching. Firstly use the URL string matching roughly and then use the web page similarity matching accurately. Experiment results show that, under the condition of maintaining the efficiency, template matching accuracy can achieve93%, increased by32.9%compared with the traditional method.

Related Dissertations

  1. Research on Method of Deep Web Schema Matching Based on Query Interface,TP311.13
  2. Research on the Patent Map Based on Domain Ontology,TP391.1
  3. Research of Text Categorization on Food Complaint Documentation Based on Ontology,TP391.1
  4. Tracking Events for Food Complaint Documents Based on Ontology,TP391.1
  5. Study on Data Extraction and Semantic Annotation for Specific Field Deep Web,TP311.13
  6. The Research of Incremental Deep Web Acquisition Technology,TP393.09
  7. Gesture Recognition for Traffic Control Based on Thinning Algorithm and Template Matching,TP391.41
  8. Based on the Deep Web of Books Information Integration and Inquires the System,TP311.52
  9. The Research on Deep Web Interfaces Integration and Query Results Ranking,TP274
  10. Study on Discovering the Relationships among Data Resources in DataSpace,TP311.13
  11. The Research on Key Technologies for Web Information Personalization Collection and Management,TP393.09
  12. Research on Vision Detection Algorithm for Tracing Printing System,TP391.41
  13. Visual Servoing Approaches Based on Parallel Mechanism,TP391.41
  14. Study on Requirement Group-based Web Service Scheduling Model,TP393.09
  15. Research on Crawling Deep Web Information,TP393.09
  16. Deep Web Interface Discovery Based on Domain Knowledge,TP393.09
  17. Study on Data Annotation of Deep Web Data Integration System,TP393.09
  18. The Research of Data Extraction and Semantic Annotation in Deep Web,TP393.09
  19. Design and Implementation of Warpper Generation System Based on Nested-Pattern in Web Pages,TP393.092
  20. Research and Implementation of DOM-Tree Based Entity Extraction for Deep Web,TP393.09
  21. The Research on Data Extraction Mechanism in Deep Web Based on Result Pattern,TP393.09

CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer applications > Information processing (information processing) > Text Processing
© 2012 www.DissertationTopic.Net  Mobile