Dissertation > Excellent graduate degree dissertation topics show
Study on Data Extraction and Semantic Annotation for Specific Field Deep Web
Author: YangZhou
Tutor: CuiZhiMing
School: Suzhou University
Course: Applied Computer Technology
Keywords: Deep Web Information Integration Page Purification Information Extraction Semantic Annotation
CLC: TP311.13
Type: Master's thesis
Year: 2011
Downloads: 57
Quote: 1
Read: Download Dissertation
Abstract
|
With the rapid development of Internet technology,information resources which are hidden in web databases have received extensive attention because of its large amount of data and structure integrity.These information resources are displayed to users in the form of HTML pages after users have submitted search queries on web query page.And researchers usually call these information resources Deep Web. To maximize the use of these Deep Web resources,those semi-structured and unstructured data on the web page need to be extracted through a variety of technical means. Meanwhile, in order to make the extracted data possess a higher use value, semantic annotations must be added to these data so that they can be understood by machines.This paper studies the technology of information extraction and data annotation in Deep Web for specific field. At first, lead type information of nodes into extraction of data records and then achieve the semantic annotation based on ontology. Finally, a prototype system is designed combined with my project experience. The main research work of this paper include:1) This paper gives a presentation about the development history、evaluation criteria and related technologies of Web information extraction in brief and analyses existed information extraction method in depth.2) Combined with characteristics of result page itself in Deep Web and using the features of vision and content on page layout, a page purification method is proposed which involves tag filter, visual feature filter and content rule filter. The experiments show that the approach can effectively improve the efficiency and precision of subsequent data extraction.3) This paper has proposed a new extraction method of data record based on node type. First, the HTML node is divided into four types: block type, style type, text type and image type. And then assign each type a weight value. Thirdly, calculate entropy value for each property node in data records of result page according to different node type. Finally determine the nodes which represent corresponding data records with the entropy value and achieve extraction of nodes in data records. Compared with other methods, this method has higher efficiency.4) View the domain ontology as the global schema followed by web databases and achieve mapping from ontology to schema through the methods such as kernel density and K-L divergence and so on. The experiment shows that the approach has some certain advantages.5) Design an information integrated platform oriented to the biomedical field based on above work.
|
Related Dissertations
- The Design and Implementation of DICOM Middle Software and Access Control Model in Formation Integration Platform,TP311.13
- Active faults based radar image information extraction method applied research and demonstration,P542.3
- Ontology-Based Hazard Information Extraction from Chinese Food Complaint Documents,TP391.1
- Tracking Events for Food Complaint Documents Based on Ontology,TP391.1
- The Research of Incremental Deep Web Acquisition Technology,TP393.09
- Analysis of Deep Web Page’s Structure and Its Rich-Content Extraction,TP393.092
- Research and Application of Semantic-Based Information Extraction,TP393.09
- Printers based on natural language HCI Research and implementation,TP11
- Research on Domain Ontology-Based Web Information Extraction,TP393.09
- Research of Data Source Selection with Similar Theme in Deep Web Integrated System,TP311.13
- Deep Web Data Cleaning Method Research and Application,TP393.09
- Research on Crawling Deep Web Information,TP393.09
- The Study on Deep Web Interface Integration and Search Strategy,TP393.09
- Deep Web Interface Discovery Based on Domain Knowledge,TP393.09
- Body in the integration of medical information and research,TP391.1
- Based on Maximum Entropy Model Research on extracting Chinese entity relation,TP391.1
- Research on the Technology of the Corporate Relations Mining,TP391.1
- Research of Enterprise Competitive Intelligence System Based on Data Processing Center,F272
- The Application of Beijing-1 Image Fusion and Its Information Extraction from Water Body,P
- Research on Management Information System Establishment and Development in Jinxi Petrochemical Branch,TP311.52
- Development of human technology - based web content extraction system,TP393.092
CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer software > Program design,software engineering > Programming > Database theory and systems
© 2012 www.DissertationTopic.Net Mobile
|