Dissertation > Excellent graduate degree dissertation topics show

Vertical Search Engine for Crawling the Web Page Design and Implementation

Author: ChenZhe
Tutor: YaoWenLin; QinYongSheng
School: Ocean University of China
Course: Computer technology
Keywords: DOM Web Crawler dynamic page crawling IE kernel Web Information Extraction DOM
CLC: TP393.092
Type: Master's thesis
Year: 2009
Downloads: 0
Quote: 0
Read: Download Dissertation

Abstract


Today’s society, Internet technology is moving very quickly, network information has been the rapid increase in search engine information on the network’s coverage on the capacity of the overall downward trend, but at the same time, people search engine search information to the quality requirements but more and more high, all kinds of user information on the request of english from the original quantity on the quality of the up to now has changed. In this context, how can the network at the growing amount of information quickly to find more accurate information became more valuable field of the current search engine a challenging research problem hot spots. In relation to all areas and topics for a comprehensive search engine, vertical search engines are often only for a particular discipline or field of information to meet the needs of the field or discipline-specific requirements, so information on the network for more in-depth and accurate extraction, make search more targeted, more targeted, accurate recall is higher. Vertical search engine to narrow your search at the premise, can quickly search for a more accurate information more valuable, but at the same time a vertical search engine web crawling depth and accuracy of data extracted from information in relation to integrated search engine is also made a higher requirement.In this thesis, how to design and implementation in accordance with the vertical search engine for crawling the web site needs browser of this topic in light of the current web design technology and the characteristics of browser technology to crawl the web page formulate the overall structure of system and system design of each module.In this, based on the requirement to construct a vertical search engine and combined with the technology of making web page and the characteristic of the web browser technology, architecture and modules design schemes are made. In the aspect of Information Extraction, researched the method of Information Extraction based on IE kernel. The method is based on the DOM' dissertation">DOM Tree Model of IE kernel and through DOM, as the input of the characterized information of the content of the tree nodes and the HTML labels, the regular express is created automatically by the program. Then combined with the created regular express and the index of the tree node in the whole DOM tree, the node containing search content is found and the result is extracted from the whole node.In the aspect of web crawling, researched the method based on the IE kernel to crawl the web pages . The method realize the function of the machine browser the web pages automatically, moreover, realize the function of crawling the dynamic pages and the hidden web data in the database of the website through the machine fill in the forms and simulate the user’s clicking on the web page.Through applying the above presented methods, the platform to assist the creation of web crawler is constructed and through using the configuration data created by the platform, the web page collecting system not depending on the structure of the concrete website and with the characteristics of universal purposes and lower implement and maintenance cost.

Related Dissertations

  1. Enhanced Coagulation and Characterization of Dissolved Organic Matter Fractions Removal in Drinking Water,TU991.2
  2. High-performed Kernel Classification Methods Based on Multi-kernel Learning,TP391.41
  3. The MObile Widget Engine Researching and Implementing Based on the Webkit,TP391.3
  4. Design and Implementation of News Web Data Extraction,TP393.092
  5. Study of Webpage-Trojan Detection Technology,TP393.08
  6. Hardware Accelerated Parallel Speculative XML Parser,TP311.10
  7. Application of Unmanned Aerial Vehicle (UAV) in the Dynamic Monitoring of Regional Land Use,F301
  8. Classification for Webpage Trojan Detection Based on DOM Modeling,TP309.5
  9. Experimental study based POS system aerophotogrammetry,P231.2
  10. Removal of PAHs from Contaminated Soil by DOM Derived from Fructus Gleditsiae Sinensis and Its Influence Factors,X53
  11. The Graphic Editing System of the Distribution Running Management Information System Based on GIS Technology,TM769
  12. Research on the Methos of Producting High-Quality DOM in Urban Areas,TP751
  13. Information Extraction Based on Table Area Locating for E-Commerce Websites,TP393.092
  14. Research and Implementation of DOM-Tree Based Entity Extraction for Deep Web,TP393.09
  15. Research and Implementation of Search Engine Prototype Based on Deep Web Crawler,TP391.3
  16. The Design and Implementation of RSS Technology-based Library Information Release and Subscribe System,TP311.52
  17. Studies on Characteristics of Dissolved Organic Matter and Organic Nitrogen in Lake Sediments,X524
  18. Photodegradation of dissolved organic matter in number average molecular weight and its related parameter changes,X703
  19. DOM presence of copper oxide nanoparticles in water bacteriostatic mechanism,TB383.1
  20. Design and Implementation of Budget Management System for the Power Equipment Maintenance,F426.61
  21. Model-based Web Pages Automatically Standardized Application Software Refactoring Tool Design and Implementation,TP393.092

CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer applications > Computer network > General issues > The application of computer network > Web browser
© 2012 www.DissertationTopic.Net  Mobile