Dissertation > Excellent graduate degree dissertation topics show

Research on Automatic and Efficient Technologies for Web Information Extraction

Author: SongZuoYing
Tutor: HongXiaoWen
School: Harbin Institute of Technology
Course: Computer Science and Technology
Keywords: Web Information Extraction Wrapper Template DOM Tree Matching TagPaths
CLC: TP393.092
Type: PhD thesis
Year: 2013
Downloads: 1
Quote: 0
Read: Download Dissertation

Abstract


The World Wide Web has become an important resource of information due to itsexplosive growth and spread in the past two decades. The tremendous amount of web datahas opened a new era for data analysis and mining systems. More and more web applica-tions need to extract, mine, and integrate data from enormous data sources. However, dueto the semi-structure characteristic of web pages, web data exhibited on web pages is notdirectly consumable by machines. Web information extraction aims at extracting struc-tured data from web pages, which is a very challenging problem due to the large-scaleand highly-heterogeneous characteristic of web information.Aiming at handling the large-scale and highly-heterogeneous characteristics of webinformation, this dissertation studies automatic and efcient technologies for web infor-mation extraction, conducted on two levels of data records and data units respectively.The research content includes:1. Targeting the high-heterogeneity of web information, a novel automatic datarecord extraction method called MiBAT (Mining data records Based on Anchor Trees) isproposed. Existing similarity-based automatic approaches cannot extract web data record-s accurately when a large amount of unstructured content exists (e.g., user-generated con-tent). This paper presents the concept of pivots, which correspond to some key data units.For example, almost every data record created and posted by users (e.g., online forumposts, user reviews, etc.) contains the publication date as a key data unit, which is apivot derived by domain constraints. The proposed MiBAT method detects pivots basedon domain constraints, identifies anchor trees that are DOM (Document Object Model)sub-trees containing the pivots, and finally extracts data records around the anchor treesautomatically. Experimental results show that, compared to existing approaches, MiBATis able to overcome the irregularity of data records caused by unstructured content, result-ing in high accuracy.2. Targeting the large-scale of web information on the level of data records, a fastand efcient anchor tree finding algorithm is proposed. In web mining community, thetraditional mining approach is to enumerate sub-trees in a top-down manner; followingthis approach, the time complexity of MiBAT is O(n2), where n is the number of nodeson the DOM tree of the web page. In this paper, a novel anchor tree finding algorithm is presented based on aggregating tag paths in a bottom-up fashion, which enables MiBAT torun in O(n log n) time. Experimental results demonstrate that the new method significantlyimproves the efciency of MiBAT while remaining high accuracy.3. Targeting the cross-domain high-heterogeneity of web information, the conceptof generic pivots is proposed. The concept of pivots origins from domain constraints,corresponding to some domain-dependent key data units. In real applications, diferentdomain constraints are required to be identified for diferent domains, which limits theapplicability of MiBAT to some extent. To resolve the domain dependency, this paperexpands the concept of pivots and proposes generic pivots. Experimental results suggestthat, when using generic pivots, MiBAT is applicable to diferent domains achieving highaccuracy, without any pre-defined domain constraints.4. Targeting the large-scale of web information on the level of data units, a fastand efcient method for DOM tree matching is proposed for data unit alignment andextraction. The most widely used tree matching algorithm runs in O(n2) time, whichis not appropriate for web-scale processing. This paper proposes a novel tree matchingmethod based on the longest common subsequence (LCS) of the tag path sequences ofDOM trees. By exploring the inherent sparsity of the LCS problem, the proposed treematching method runs in O(r log n) time, where r is the number of pairs of nodes thathave identical tag paths from the two trees; when the matching is sparse, r≈O(n),and the algorithm runs in O(n log n) time approximately. Extensive experimental resultsdemonstrate that, compared to the existing method, the proposed approach significantlyimproves the running efciency and also achieves similar tree matching results as well asdata unit alignment results.In summary, this dissertation presents technologies for automatic and efcient webinformation extraction on both levels of data records and data units, which can well handlethe large-scale and highly-heterogeneous characteristics of web information with boththeoretical and application value.

Related Dissertations

  1. Research on Algorithms of 2D Face Template Protection,TP391.41
  2. Incomplete information on the completeness of the system and its knowledge acquisition,TP311.13
  3. On the television program template of intellectual property protection,G222
  4. User-Steered Development of Personalized Information Sevice,TP393.09
  5. Tracking Events for Food Complaint Documents Based on Ontology,TP391.1
  6. Web Page Attribute Extraction Method Research,TP391.1
  7. User Experience Design for Mobile-based Enterprise Management Software,TP311.52
  8. Research on Vision Detection Algorithm for Tracing Printing System,TP391.41
  9. Visual Servoing Approaches Based on Parallel Mechanism,TP391.41
  10. Web Service Composition Based on QoS Awareness,TP393.09
  11. Template- based virtual machine memory cloud computing services system,TP393.09
  12. Video Mean-Shift tracking algorithm applied research,TP391.41
  13. Research on Domain Ontology-Based Web Information Extraction,TP393.09
  14. Template independent web information extraction,TP393.092
  15. Research on Data Acquisition and Topic Analysis of Online Public Opinion,TP393.09
  16. Research on Crawling Deep Web Information,TP393.09
  17. Design and Implementation of Web Information Extraction Based on DOM,TP393.09
  18. An Approach to the Key Problems of Web Information Extraction Based on Prefix Expression,TP391.1
  19. A Unified Method of Heterogeneous Data Resources in Network Environment,TP311.13
  20. Swing Detection System of Straddle-type Monorail Track Beam Based on MV,TP274
  21. The Accuracy of Guided Template in Implant Placement of Computer-based Guded Surgery: a Pilot Study,R783.6

CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer applications > Computer network > General issues > The application of computer network > Web browser
© 2012 www.DissertationTopic.Net  Mobile