Dissertation > Excellent graduate degree dissertation topics show
Research of a Suffix Tree Based Automatic Wrapper Generation Method
Author: ZhangYuLiang
Tutor: YuanSenZuo
School: Jilin University
Course: Computer Software and Theory
Keywords: Semi-structured data Automatically generated wrapper Suffix tree Information Extraction Web denoising
CLC: TP311.1
Type: Master's thesis
Year: 2005
Downloads: 148
Quote: 1
Read: Download Dissertation
Abstract
|
This paper studies the method automatically generates a wrapper for data-intensive pages . Firstly, the noise removal method of HTML pages . Controllable noise granularity of DOM - based web noise removal method . By setting a reasonable noise granularity , can effectively guarantee the accuracy of the web noise removal . Site Denoising can eliminate started noise generated wrapper , to improve the speed and accuracy of the wrapper generated . Next for denoised page , automatically generated based on the suffix tree wrapper . Connected to the first page parsing well-structured sequence of symbols , and the establishment of the suffix tree , the use of the suffix tree repeating substring discovery algorithm found in the sequence of symbols repeat mode ; generate the page repeat mode linked list through an iterative process , which found that the page nested mode. After that, the use page tree matching process page selection mode and the resulting page wrapper . The full advantage of the hierarchical nature of HTML , reducing the complexity of the algorithm in the generation process of the entire wrapper . This paper studies automatic wrapper generation method has low time complexity , there are some practical value .
|
Related Dissertations
- Research on Temporal Information Recognition and Normalization,TP391.1
- Study on Growth Monitoring Technique Based on Pixel Un-Mixing Method and HJ Remote Sensing Images in Paddy Rice,S511
- Based on high-resolution remote sensing data mining houses information extraction,TP751
- Object-Based Automatic Extraction of Change Information Based on High-Resolution Remote Sensing Image Research,P237
- Based on semi- structured text transporter protein substrate information extraction system,Q811.4
- The Research and Implementation of Protein Classification Algorithm on the Basic of String Kernel,TP301.6
- Research of Driver Fatigue Detection Based on Information Fusion,TP391.41
- Research and Application of Semantic-Based Information Extraction,TP393.09
- Segmentation of Social Media Web Page and Extraction of Topic Frequent Cluster,TP391.1
- Research of Chinese Meta Search Engine Based on Clustering,TP391.3
- On Research of Deep Search and Information Extraction for E-commerce Websites,TP393.092
- Mobile Web Site Development Based on DOM Tree Information Extraction,TP393.092
- The Research and Application of Web Information Extraction Technology,TP393.09
- Design and Implementation of Text Information Extracting Modules of Html Web Pages Based on DOM,TP393.092
- Research and Application of Information Extraction -based the multidimensional semantics Internet drugs,TP393.09
- The Relation and Analysis of Relation-network Analysis Technology Based on MongoDB,TP311.13
- The Images Interpretation and Information Extraction of the Burst and Potential Debris Flows in Remote Sensing Way,P642.23
- Study on Technologies of Remote Sensing Feature Analysis and Information Extaction of Earthquake Disaster,P237
- Study on Remote Sensing Geological Mineral Information Exrraction for Tibet Xiongcun Porphyry Coper-gold Deposit,P627
- Studying on Predict for Ore-hunting of Remote Sensing in Southern Area of Qinghai Zaduo,P627
- Research on Application of Detection and Extraction of Alteration Information in Different Natural Landscape Areas,P627
CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer software > Program design,software engineering > Programming
© 2012 www.DissertationTopic.Net Mobile
|