Dissertation > Excellent graduate degree dissertation topics show
Design and Implementation of Web Information Extraction System for Evaluation of Search Engine
Author: LiuBao
Tutor: LuoJiZhou; WangYa
School: Harbin Institute of Technology
Course: Software Engineering
Keywords: Web information extraction Dom search engine evaluation onsatisfaction of users data of evaluation
CLC: TP391.3
Type: Master's thesis
Year: 2012
Downloads: 4
Quote: 0
Read: Download Dissertation
Abstract
|
With the rapid growth of the information on Internet, the web has graduallybecome the main platform for people to get information. Thus, the major searchengines come with the tide of fashion with intensive competitions. The searchresults as well as the user’s experience are two factors to evaluate the quality ofthe engine. Since the user’s experience can bring flow for the engine, so theevaluation that focus on the degree of user’s satisfaction about search engine hasbecome more and more important, and companies that has large search enginehave already set up a special team or departments for evaluation to lead thedirection for search engine by the evaluation data.The principle of the evaluation is to collect up information of search resultsof search engine for the users to score, and to make statistics for the index in orderto make a comparison. To get evaluation data successfully is the key to theevaluation task.In this thesis, the importance of information extraction is stressed becausethat the experimental data verifies that the accuracy of data extraction will directlyaffect the results of evaluation. The author compares some technologies of theexisted Web information’s extraction, and makes analyses according to thedemand of this system, and points out advantages and disadvantages with thecombination of the characteristics of the source code on the results page of thesearch engine. The author also proposes a new method that is to combine regularmatch with Dom resolve to extract and process the evaluation data. Based on thisidea, the author wants to achieve an extraction system of Web information withstrong applicability and higher automation to solve the issue which is to collect theevaluation data of evaluation system.The system mainly consists of the page download, the page filter, rules togenerate extraction, the information extraction, and the data storage, etc… Thethesis makes a more detail introduction for these elements, in which the rules togenerate extraction is more important in the realization of the system. By usingDom structure and the study of the sample, the system can automatically generatethe rules of extraction; find the maximum common path of the node, and recordcharacteristics of the sample node. At the same time, the system filters irrelevantinformation of the node by matching algorithm of similarity on node to realize thehigh automation on extraction of information for some products. Of course, therules of extraction can also be manually amended. In order to improve accuracy, some products use regular matching to extract information. To encode and enteredinto the base of rules artificially in advance, the system will call the templatematching module to allocate template for the extraction.Finally, the thesis introduces two indicators in the evaluation of extraction ofinformation: rate of precision and rate of recall, and makes tests on data downloadand analyses on results of information extraction. According to the indicators ofevaluation, the author proposed that the system has a good effect on the page ofresult set that is generated by the search engine, and also in solving the problemthat getting evaluation data efficiently and accurately for the reviewers.
|
Related Dissertations
- The Design and Implementation of Lucene-Based Network Literature Vertical Search Engine,TP391.3
- The Research of Link Structure in Tibetan Web Base on Social Network Analysis,TP393.09
- Enhanced Coagulation and Characterization of Dissolved Organic Matter Fractions Removal in Drinking Water,TU991.2
- Web Page Attribute Extraction Method Research,TP391.1
- High-performed Kernel Classification Methods Based on Multi-kernel Learning,TP391.41
- Research and Implementation of Vertical Search Engine Key Technology,TP311.52
- Based on analysis of public opinion judged the emergency warning platform for the design and implementation of,TP311.52
- The Research on Key Technologies for Web Information Personalization Collection and Management,TP393.09
- The MObile Widget Engine Researching and Implementing Based on the Webkit,TP391.3
- Design and Implementation of News Web Data Extraction,TP393.092
- Study of Webpage-Trojan Detection Technology,TP393.08
- Hardware Accelerated Parallel Speculative XML Parser,TP311.10
- Application of Unmanned Aerial Vehicle (UAV) in the Dynamic Monitoring of Regional Land Use,F301
- Research and Implementation of Test Data Management System for State of Railway Tunnel Lining,TP311.13
- Research on Related Theme of Search Engines,TP391.3
- Classification for Webpage Trojan Detection Based on DOM Modeling,TP309.5
- Research on Web-based Opinion Analysis for Stock Reviews,TP391.1
- The Study of Search Recommendation for Brand Information,F273.2
- Experimental study based POS system aerophotogrammetry,P231.2
- Removal of PAHs from Contaminated Soil by DOM Derived from Fructus Gleditsiae Sinensis and Its Influence Factors,X53
- The Graphic Editing System of the Distribution Running Management Information System Based on GIS Technology,TM769
CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer applications > Information processing (information processing) > Retrieval machine
© 2012 www.DissertationTopic.Net Mobile
|