Dissertation > Excellent graduate degree dissertation topics show

Design and Implementation of Web Information Extraction Subsystem in the Public Opinion System

Author: LiZuo
Tutor: LiuDan
School: University of Electronic Science and Technology
Course: Computer technology
Keywords: information extraction automatic generate templates Web page clustering
CLC: TP393.092
Type: Master's thesis
Year: 2013
Downloads: 30
Quote: 0
Read: Download Dissertation

Abstract


Network public opinion can reflect the attitude of the general public on a varietyof events, it is an important channel for the relevant departments to understand thepublic opinion. Web information extraction is the input of public opinion in publicopinion analysis system, it directs the source and judgment of public opinioninformation. With the rapid development of the Internet, there is more and more Webforms. In order to obtain the information of public opinion quickly and accurately,there is increasingly high demand for Web information.This thesis studies the Web information extraction technology on the above issues,on the basis of in-depth analysis of the current Web page forms and the structure of thepage, combine with analysis requirements of public opinion, and proposes differentextraction methods for the four different sources of public opinion information page,such as news, blogs, forums and microblogging.The research says:1. This thesis research Web information extraction technology of news anf blogclass, use common page text extraction technology to extract mian body, and useregular expressions to extract other data items. The method does not depend on pagestructure, at the same time, it has high extraction speed, high accuracy and versatility.2. This thesis researches Web clustering algorithm, and proposes a clusteringmethod based HTML tag tree, the method is based on forum page structure, itcalculates the node’s value of HTML tag tree, it uses weighted cosine similarityformula to calculate the similarity of two Web tag tree. The method has good clusteringresults, and its time complexity is O(n).3. This thesis researches Web information extraction technology automaticallybased on similarity comparison of web-based structure, and propose a forum pageautomatic information extraction method. This method can automatically generate anextraction template for each forum website, the template uses entropy、 structuralsimilarity and others features to confirm specific information in forums, in the end, it isused to extract information of other pages on this web.4.This thesis researches Web information extraction of microblog class pages, we propose a new extraction method combined with tag attribute and regular expression,the method takes many kinds of characteristics of the data items into account, it usesthe label attributes and attribute values to locate, and uses regular expressions tocomplete the precise extraction.The experiments show that, this thesis proposes four website informationextraction method for news、blogs、forums and microblog pages, this methods canextract information from the massive network rapidly and accurately. They had highrecall and precision rates, and can transform extract data into structured data which arestored in the database. All these methods meet the requirements of public opinion dataanalysis system.

Related Dissertations

  1. Research on Domain Entity Attribute and Event Extraction Technology,TP391.1
  2. Research on Temporal Information Recognition and Normalization,TP391.1
  3. Study on Growth Monitoring Technique Based on Pixel Un-Mixing Method and HJ Remote Sensing Images in Paddy Rice,S511
  4. Land Desertification in Qinghai Lake Landscape Pattern Change,X171
  5. Active faults based radar image information extraction method applied research and demonstration,P542.3
  6. Based on high-resolution remote sensing data mining houses information extraction,TP751
  7. Scholar Resume Automatic Generation Based on Text Mining,TP391.1
  8. Research and Application of Information Extraction Based on Query Expansion,TP391.1
  9. Ontology-Based Hazard Information Extraction from Chinese Food Complaint Documents,TP391.1
  10. Tracking Events for Food Complaint Documents Based on Ontology,TP391.1
  11. Study on Data Extraction and Semantic Annotation for Specific Field Deep Web,TP311.13
  12. Object-Based Automatic Extraction of Change Information Based on High-Resolution Remote Sensing Image Research,P237
  13. Applications of Bibliometrics and Text Mining in the Life Science,TP391.1
  14. Research on Hyperlinks Extraction Based on Hotspot Website Content Analysis,TP393.092
  15. Comparative Experimentation Study on Information Extraction of Object-Oriented Based on Quick Bird Image,P237
  16. Web Page Attribute Extraction Method Research,TP391.1
  17. Rule-based Web text information extraction technology research,TP391.1
  18. The Research for Named Entity Recognition and Relation Extraction in Text,TP391.1
  19. Research of the Algorithm of Moving Vehicle Detection &tracking Based on Surveillance Video and Its Implemen Tation,TP391.41
  20. Product Review Mining Based on Opinion Words,TP311.13
  21. Research and Application of Semantic-Based Information Extraction,TP393.09

CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer applications > Computer network > General issues > The application of computer network > Web browser
© 2012 www.DissertationTopic.Net  Mobile