Dissertation > Excellent graduate degree dissertation topics show
Concept tree based Web Information Extraction Technology Research
Author: GuWen
Tutor: WangHongMei
School: Changchun University of
Course: Computer Software and Theory
Keywords: Information Extraction Concept tree Semantic templates Extraction rules
CLC: TP391.1
Type: Master's thesis
Year: 2010
Downloads: 77
Quote: 1
Read: Download Dissertation
Abstract
|
The rapid development of the Internet, people living in the world of data. Every day there are a large number of Web information posted on the Internet reproduced, every day there are a lot of information on the Internet upload and download. Information on the Internet is like the ocean general boundless. There is a growing need for a technology and tools to help them quickly find the information they need. It is hoped that this approach and tools can have high accuracy, efficiency and speed and artificial intelligence features. Web information extraction technology has become the focus of attention. At present, has developed a variety of Web information extraction technology and systems in the field of information extraction has also made important achievements, but they also showed some differences between and shortcomings. Extraction mechanism which is mainly the differences and differences in the final form of the drawing result. Pursuant to the adoption of the principle can be divided into six kinds of Web information extraction methods, such as wrapper-based information extraction language, HTML-based structure information extraction, natural language processing method based on information extraction and so on. In establishing extraction template, the part of the system can only generate single-slot extraction rules, this has led to results very simple, but the effect is not ideal; still others are multi-slot extraction rules, but requires manual preparation of professionals, so need to extract the content is very relevant familiar with the implementation is relatively complex; other systems to extract the content to very demanding, only one type of the text or a piece of content extraction, thus weakening the applicability of information extraction. In response to these problems, this paper uses the concept-based information extraction expanded way to create templates and concept-based extraction rules. Purpose of this paper is to study an effective learning algorithm to automatically generate extraction rules, so that even non-professionals can successfully guide the generation and extraction rules from a similar structure to extract pages to the information they need. In this paper, the concept of expansion based on semantic mechanisms through appropriate artificial settings, improve system stability and effectiveness. Semantic concept uses a concept tree expansion mechanism, this paper focuses concept tree establishment and coding concept tree research. Information extraction process takes advantage of the concept of extension mechanism, from the extracted template creation, to extract the text information filtering and information mapping mechanism to the final text database query mechanism embody the concept of expanded role. Of course, to extract the text information extraction process is also important, this study for the three main issues discussed are named entity recognition, refer to digestion process and the time information. In this paper, the concept of Web-based information extraction technology compared with the conventional technique, the formation of a multi-slot template, expanding the range of extraction results; extraction rules applied in the conceptual mechanisms, but also reflects the characteristics of artificial intelligence, more reflects the humane extract information.
|
Related Dissertations
- Study on Growth Monitoring Technique Based on Pixel Un-Mixing Method and HJ Remote Sensing Images in Paddy Rice,S511
- Based on high-resolution remote sensing data mining houses information extraction,TP751
- Reptiles theme for Education News Design and Implementation,TP391.3
- Topic search engine key technology research,TP391.3
- Study on Extraction of Coniferous Forest Information in Southern China,TP79
- Based on multi-feature classification comparison Urban Remote Sensing Dynamic Monitoring,TP873
- Classification of Hyperspectral Remote Sensing Image Based on the Cloud Model Theory,TP751
- Topic Detection and Ranking in Web Community,O226
- Hull section robotic welding path planning and offline programming,TP242
- Research and Implementation of the Deep Web information extraction system,TP311.52
- Multi-language support program comprehension understanding and information extraction technology research,TP311.52
- Based on semi- structured text transporter protein substrate information extraction system,Q811.4
- Digital Geomorphological Information Extraction and Spatial Pattern Analysisin Shanxi Province,P931
- The Research and Implementation of Directed Protein-protein Interaction Extraction,Q51
- The Design and Implementation of the Auto-Search and Comparison System for e-Products,TP311.52
- Design and Implementation for Image Search Engine Based on Mobile Phone,TP391.3
- Massive web collection system design,TP391.3
- Design and Implementation of Domain Specified Deep Web Information Integration System,TP311.52
- Auto-match between Text and Map Based on Chinese Word Segmentation,P208
- For multi- tasking, multi- channel parallel crawler technology research,TP391.3
CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer applications > Information processing (information processing) > Text Processing
© 2012 www.DissertationTopic.Net Mobile
|