Dissertation > Excellent graduate degree dissertation topics show

Research on Interactive Web Data Extraction Based on Tree Matching

Author: QuZhuWei
Tutor: ZhangHuiYing
School: Tianjin University
Course: Management Science and Engineering
Keywords: Web Data Extraction TOP-DOWN tree matching DOM Component system
CLC: TP393.092
Type: Master's thesis
Year: 2005
Downloads: 213
Quote: 1
Read: Download Dissertation

Abstract


Web browser and search engine keyword search is the traditional method for people to get information from the Internet , and its limitations : the web browser is difficult to locate specific knowledge on the Internet ; Search engine easily gain the required knowledge submerged in a large number of irrelevant pages . Present some new Web-based applications , such as monitoring the stock market in a timely manner Quotes , compare various sites in commodity prices , and so on . Alone search engines and Web query technology can not meet these emerging needs . Information retrieval technology , data extraction technology is more focused on the related documents found in the collection of documents from the web , and out from the data of interest to users , and the upcoming HTML semi-structured data into structured data . The online data is stored in the website 's back-end database , Web database based on the query special requirements or conditions for dynamic generation of Web pages , queries related Web . Between data records in the Web page having a high similarity of the code structure , naturally has a high structural similarity between the Web data corresponding to the record DOM subtree . The proposed based the the subtree matching interactive extraction method , XHTML corresponding DOM tree is the basis of Web data structure and analysis of the method extracted from the analysis of Web document DOM tree correspondence between . Method based on one-to-one relationship of Web data recording and DOM subtree , by analyzing the code structure of one or several data records , the discovery and extraction of all similar data records . Way to interact with the user generated data extraction rules , and multiple sequence alignment algorithm is introduced into the extraction rules generation process ; introduction of the sub- division of the tree to the data extraction process , combined with the TOP-DOWN tree matching algorithm realize the DOM tree the discovery and extraction of the data recording . Application of component-based model of the prototype system implementation, test proved that to extract high efficiency the subtree matching - based interactive extraction method , the extraction recall rate and the accuracy rate can reach a high level .

Related Dissertations

  1. Structure Biological Studies of Histone Demethylases Jhd2 & JMJD5 and Expression, Purification of Staphylococcus Aureus HssSR,Q93
  2. High-performed Kernel Classification Methods Based on Multi-kernel Learning,TP391.41
  3. The MObile Widget Engine Researching and Implementing Based on the Webkit,TP391.3
  4. The Research of Data Extraction and Semantic Annotation in Deep Web,TP393.09
  5. Classification for Webpage Trojan Detection Based on DOM Modeling,TP309.5
  6. Experimental study based POS system aerophotogrammetry,P231.2
  7. Removal of PAHs from Contaminated Soil by DOM Derived from Fructus Gleditsiae Sinensis and Its Influence Factors,X53
  8. Design and Implementation of Fingerprint Spectrum Differenece Analysis System of the Complex Multi-component System,O657
  9. The Graphic Editing System of the Distribution Running Management Information System Based on GIS Technology,TM769
  10. Research on the Methos of Producting High-Quality DOM in Urban Areas,TP751
  11. Information Extraction Based on Table Area Locating for E-Commerce Websites,TP393.092
  12. Research and Implementation of DOM-Tree Based Entity Extraction for Deep Web,TP393.09
  13. Design and implementation of vehicle insurance marketing support system based on software component technology,TP311.52
  14. Research and Implementation of Search Engine Prototype Based on Deep Web Crawler,TP391.3
  15. The Design and Implementation of RSS Technology-based Library Information Release and Subscribe System,TP311.52
  16. Studies on Characteristics of Dissolved Organic Matter and Organic Nitrogen in Lake Sediments,X524
  17. Photodegradation of dissolved organic matter in number average molecular weight and its related parameter changes,X703
  18. DOM presence of copper oxide nanoparticles in water bacteriostatic mechanism,TB383.1
  19. Design and Implementation of Budget Management System for the Power Equipment Maintenance,F426.61
  20. Model-based Web Pages Automatically Standardized Application Software Refactoring Tool Design and Implementation,TP393.092
  21. Research on Wrapper Adaptation in Web Data Integration,TP393.09

CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer applications > Computer network > General issues > The application of computer network > Web browser
© 2012 www.DissertationTopic.Net  Mobile