Dissertation > Excellent graduate degree dissertation topics show

The Design and Implementation of Web Information Extraction System Based on Dom

Author: ZhangLiYan
Tutor: ZhaoHongWei
School: Jilin University
Course: Computer System Architecture
Keywords: DOM information extraction partition STU-DOM degree of correlation
CLC: TP393.09
Type: Master's thesis
Year: 2011
Downloads: 65
Quote: 0
Read: Download Dissertation

Abstract


The rapid growth of Internet and its technologies has made Internet a vast resource of information. But these information is not always useful for users, some useful information in a web page is always mixed up with unimportant words, which makes it difficult for the users to acquire the theme information effectively and limits the availability of web. The effection of web information extraction system is to collect data from internet and integrate it into XML or relational database so as to provide information services such as data retrieval, data mining and OLAP etc. But data in HTML pages is unstructured, in HTML pages there are a great deal of labels, pictures and flash advertisements etc. which bring a lot of difficulties for the system. In order to solve this problem, related scholars has done mass researches, thereupon some techniques related to the network appear. By deleting redundant labels and unrelated pictures and flashes, we extract real topical information which can reduce the web size and improve the efficiency and serviceability of information, thus to improve its effeciency and accuracy, and to lay a foundation of subsequent data retrieval, data mining and OLAP services. Therefore, extraction of web information has great significance in both theory and application, and it has become a hot topic in the information field.After a great deal of research, we have found the advantages and disadvantages of the current extraction methods. This thesis puts forward a new method which based on STU-DOM model, offers a algorithm of page structrue colation and theme related pruning, based on which we designed and made the system come true. We designed and implemented the web extractor system on this algorithm.STU tree model and STU-DOM model are based on the partition theory. STU-DOM can effectively describe the structure, content and partition layout of web page, which improves the accuracy, reliability and expandability of the methodology.Based on the STU-DOM model, we provide HTML structure filter and partition algorithms, and prune algorithm based on topical correlativity. The algorithms can automatically extract topical information from heterogeneous web pages and are highly precise and universal. We also proposed and implemented some optimization policies, such as improving the partition degree, providing special lists for empty words and key words, and designing weighted topical correlativity. The optimizations provided significant improvements in efficiency and accuracy, and minimize the redundant information in web pages.Experimental results has proved that our approach can automatically and accurately extract topical information. It is independent of the content, structure and layout of web pages, so it ahs higher research significance and application value.

Related Dissertations

  1. Research on Domain Entity Attribute and Event Extraction Technology,TP391.1
  2. Research on Temporal Information Recognition and Normalization,TP391.1
  3. Task Partition of Network Simulation under Large-Scale Computing,TP393.01
  4. The Research of Chaos M-ary Modulation and Demodulation Method Based on Chaotic Oscillator,TN915.05
  5. Study on Growth Monitoring Technique Based on Pixel Un-Mixing Method and HJ Remote Sensing Images in Paddy Rice,S511
  6. Land Desertification in Qinghai Lake Landscape Pattern Change,X171
  7. Active faults based radar image information extraction method applied research and demonstration,P542.3
  8. Based on high-resolution remote sensing data mining houses information extraction,TP751
  9. Research on Solution of the Mortgage House in Divorce Case,D923.9
  10. Web Page Attribute Extraction Method Research,TP391.1
  11. High-performed Kernel Classification Methods Based on Multi-kernel Learning,TP391.41
  12. The General Methods of Studying the Spectra of Graph,O157.5
  13. Research on Image Segmentation and 3D Modeling Based on Medical Image Sequence Matching,TP391.41
  14. Research on Application of Visualization of Analysis and Calculation in Power System,TM769
  15. Research on Cluster-Based Coverage Scheme in Wireless Sensor Networks,TP212.9
  16. The Research for Named Entity Recognition and Relation Extraction in Text,TP391.1
  17. The MObile Widget Engine Researching and Implementing Based on the Webkit,TP391.3
  18. The key component vertical search engine technology research,TP391.3
  19. Reptiles theme for Education News Design and Implementation,TP391.3
  20. GPU-based image search Chinese Research on key technologies of the retrieval,TP391.1
  21. Home Academic Information Extraction System,TP393.092

CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer applications > Computer network > General issues > The application of computer network
© 2012 www.DissertationTopic.Net  Mobile