Dissertation > Excellent graduate degree dissertation topics show
Research of Data Source Selection with Similar Theme in Deep Web Integrated System
Author: SangYinBang
Tutor: WangChengLiang
School: Chongqing University
Course: Computer Software and Theory
Keywords: Repeatability Estimation Quality Indicators Deep Web Data Source Selection Quality Evaluation Model
CLC: TP311.13
Type: Master's thesis
Year: 2011
Downloads: 22
Quote: 0
Read: Download Dissertation
Abstract
|
The way that people get high quality data which is hidden in the database by using Deep Web integrated system is one of important ways for people to obtain information. Deep Web integrated system needs to get information from mass of data sources on the web when querying data, so the more of data sources on the web, the more cost that system obtains information, the quality of these data sources is uneven at same time, and there is shortage that it can’t be efficient for user to get high-quality data from the data sources.This paper presents a selection method of Deep Web data sources with similar theme, which is based on study of using data source quality indicator to select high-quality data sources. This method can effectively compute the repeatability of the content between new data source and integrated system by differential analysis of the data source, and uses quality indicators of accuracy, sequence, size of data source and authority that represent the characteristics of database to assess the quality of data source from different perspectives.The main contents of this paper can be summarized as follows:①Discusses and analyzes the research of Deep Web technology, research status at home and abroad, practical significance, domain knowledge and related technologies of Deep Web integrated framework.②It can obtain repeatability of three or more data sources with similar themes by using improved data sources estimation methods. Firstly, the paper chooses key attributes set for the record of data source, then use edit distance method to realize the comparison on value of corresponding attribute between single data source records with records of similar data source in integrated system, finally uses FR (Frequent Records) method to obtain the contents repeatability of single data source and data source sets in the integrated system.③It improves correlation discrimination method of record in the query result set. This method gets the frequency of record in the data source set with similar theme by the way of probing query, and the record is related to the query when the frequency of it exceeds a given threshold. This method can obtain a different number of relevant records with the diversification of threshold value, and eliminates the influence about type restrictions of property in the query interface of traditional correlation discrimination method of record.④For current deep web data source quality assessment exists shortages of poor objective truth and low accuracy , this paper uses quality indicators of accuracy, sequence, size of data source and authority to get the quality of Deep Web data source by establishing data source quality assessment model, so as to select the N-Deep Web data sources of highest quality for user query.The experiment results on mainstream book sites show that the proposed method not only can reduce the burden on the system, but also can assess quality of data sources with same theme effectively, then the system can obtain higher quality data sources.
|
Related Dissertations
- Studieson Effects of Soybean Species on Yuba and Initial Establishment of Quality Evalution System for Yuba,TS214.2
- Studies on the Preparation of 20% Chlorpyrifos CS and Its Analytic Methods,S482.3
- The Energy-saving and Pollution-reduction for Cotton Dyeing Industry,F206;X322
- Typical areas of Guizhou Tobacco and tobacco quality meteorological factors Relations,S572
- The Body of Appearance and Particularly Item the Character of Our Country Excellnet Changquan and Taiji Athlete Be Simple to Evaluate Model of,G852.11
- Deep Web Data Cleaning Method Research and Application,TP393.09
- Research on Crawling Deep Web Information,TP393.09
- The Study on Deep Web Interface Integration and Search Strategy,TP393.09
- Deep Web Interface Discovery Based on Domain Knowledge,TP393.09
- Study on Data Annotation of Deep Web Data Integration System,TP393.09
- The Research of Data Extraction and Semantic Annotation in Deep Web,TP393.09
- Selection of Deep Web Database,TP311.13
- Study on Population Quality Index of Super High-yield Cotton in Xinjiang,S562
- Design and Implementation of Warpper Generation System Based on Nested-Pattern in Web Pages,TP393.092
- Research of Query Interface Integration Mechanism in DWIIS System,TP393.09
- Study on Data Sources Discovery and Selection on Deep Web,TP393.09
- Research and Implementation of DOM-Tree Based Entity Extraction for Deep Web,TP393.09
- The Research on Data Extraction Mechanism in Deep Web Based on Result Pattern,TP393.09
- Study and Implementation of a Two-Phase Query Interface Extraction Technique Based on Domain Features,TP393.09
- Research on Data Extraction and Schema Recognition on Deep Web,TP393.09
CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer software > Program design,software engineering > Programming > Database theory and systems
© 2012 www.DissertationTopic.Net Mobile
|