Dissertation > Excellent graduate degree dissertation topics show
On Research of Deep Search and Information Extraction for E-commerce Websites
Author: Zhou
Tutor: WangHao
School: Jiangxi Normal University
Course: Management Science and Engineering
Keywords: Deep Search Deep Web Web Information Extraction URL Collection Regular expression
CLC: TP393.092
Type: Master's thesis
Year: 2011
Downloads: 91
Quote: 1
Read: Download Dissertation
Abstract
|
With the rapid development of network and database technology, Internet has become a carrier of a great deal of information, how to extract the meaningful information accurately and rapidly from the mass of web information has become an urgent problem which people want to solve. There is a lot of web information in the E-commerce websites, these websites’web pages are dynamic, and have high degree of structures. Actually, the E-commerce websites belongs to Deep Web, Deep Web is the information which can not access with hyperlinks and needs the dynamic web technology to access, users need to submit queries on the specific interface(the query interface) to access the data in the database. Deep search is the kind of search which aim at Deep Web. Deep Web has lots of information, high degree of structures and public access interface. So we choose the E-commerce websites to do the object of research of deep search, and get the purpose that extract the products’information. The search engines of E-commerce websites offer convenient for researchers to extracting information. Depending on some keywords that users input, these query interfaces can show users the information about keywords by web database dynamically. We can use these query interfaces to do some research into deep search, getting the web pages what we need by simulating the filling and submission of keywords.The webpages what get by deep search have a lot of products’information, we use these information source to extraction. The key point of web information extraction is how to generate extraction rules. There are two approaches to generating extraction rules, one is generating automatically, and another one is generating by manual work, these two approaches are all have advantages and shortcomings, they have their respective applicable scopes. The automatic approach is suitable for websites which have different structrue, but the precision rate of this approach is lower. Although the manual approach needs someone to write the regule rules, the precision rate of extraction results is higer. Due to the web pages’structures of E-commerce websites are mainly the same, and we want to extract the information includes: products’name, products’price, freight and other information about products, we choose the manual method what has more precise extraction results.The main works and innovations of this paper are as follows: 1、Design an interface of keywords’files, allowing the system to accept to the keywords’files(text files, there is a carriage return between two keywords), and putting these keywords in our system for filling and submitting query forms. Moreover, we consider the question of incremental keywords, the system of this paper do not accept the keywords what in the old keywords’library.2、Extracting the HTML codes of E-commerce websites. Extracting the part of query forms’HTML codes of E-commerce websites by analyzing the HTML codes of E-commerce websites. According to these HTML codes, using WebBrowser Control to simulate the filling and submission of keywords, and getting the initial pages about keywords.3、Extracting the hyperlinks with selectivity, only extract the hyperlinks of products’information, but not the hyperlinks of advertisement and any other unrelated hyperlinks. Moreover, we need get more comprehensive hyperlinks of products’information by getting the hyperlinks of“next-page”, because of the multi-page hyperlinks of products’information. In this paper, we introduce some approaches of getting the hyperlinks of“next-page”, and propose an approach that has good applicability.4、According to the structures of different websites, generating the extraction rules by regular expression, thus information extraction. The extraction results are saved as text files which is convenient for updating keywords’files.
|
Related Dissertations
- Study on Data Extraction and Semantic Annotation for Specific Field Deep Web,TP311.13
- The Research of Incremental Deep Web Acquisition Technology,TP393.09
- Analysis of Deep Web Page’s Structure and Its Rich-Content Extraction,TP393.092
- Research of Data Source Selection with Similar Theme in Deep Web Integrated System,TP311.13
- Deep Web Data Cleaning Method Research and Application,TP393.09
- Research on Data Acquisition and Topic Analysis of Online Public Opinion,TP393.09
- Research on Crawling Deep Web Information,TP393.09
- The Study on Deep Web Interface Integration and Search Strategy,TP393.09
- Deep Web Interface Discovery Based on Domain Knowledge,TP393.09
- Study on Data Annotation of Deep Web Data Integration System,TP393.09
- The Research of Data Extraction and Semantic Annotation in Deep Web,TP393.09
- Selection of Deep Web Database,TP311.13
- Design and Implementation of Warpper Generation System Based on Nested-Pattern in Web Pages,TP393.092
- Research of Query Interface Integration Mechanism in DWIIS System,TP393.09
- Study on Data Sources Discovery and Selection on Deep Web,TP393.09
- Research and Implementation of DOM-Tree Based Entity Extraction for Deep Web,TP393.09
- The Research on Data Extraction Mechanism in Deep Web Based on Result Pattern,TP393.09
- Study and Implementation of a Two-Phase Query Interface Extraction Technique Based on Domain Features,TP393.09
- Research on Data Extraction and Schema Recognition on Deep Web,TP393.09
- Study on Query Relaxation for the Deep Web,TP393.09
- Result Pattern Semantic Annotation Based on CPN Network in Deep Web Integrated System,TP393.09
CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer applications > Computer network > General issues > The application of computer network > Web browser
© 2012 www.DissertationTopic.Net Mobile
|