Dissertation > Excellent graduate degree dissertation topics show

Web information extraction based on the key to improving the HITS algorithm and location information

Author: ChenHanSheng
Tutor: ZhangShiYong
School: Fudan University
Course: Applied Computer Technology
Keywords: Link Analysis web information extraction SVM
CLC: TP301.6
Type: Master's thesis
Year: 2009
Downloads: 161
Quote: 0
Read: Download Dissertation

Abstract


With the enhancement of informatization,more and more data is shared by the Internet.In a large database like the Internet,how to obtain the necessary information is a difficult issue.This involves of two core problems:how to acquire the important web pages and how to extract the structured information of web pages.Search engines are one of the Internet information retrieval tools.Because of their general-purpose uses,they must treat each web page fairly.And therefore,they are not suitable to handle the problem of information retrieval of specific domains.There is a considerable portion of documents which are unstructured or semi-structured.The traditional information extraction methods are generally based on structured data from text documents.Therefore,how to extract information from web pages has become one of the research hotspot in recent years.This also led to a new research sub-area,namely,web information extraction.This paper focuses on the research of a method of acquiring important web pages and extracting structured information from them.First of all,by the analysis of the advantages and disadvantages of two link analysis methods,namely HITS and PageRank,this paper chooses HITS algorithm as the basic method.In the corresponding experiment,it shows that the traditional HITS method has two flaws:ignorant of new web pages and vulnerable to "spam links",so it is not suitable for applications which deals with up-to-date information such as news.Based on the previous research which solves the "spam links" problem by adding a filter,this paper presents a new algorithm called TimeWeightedHits,which adds a time factor to improve the HITS method further.The experiments proved that. it can filter the unwanted web pages effectively and acquire the appropriate up-to-date important web pages.Secondly,to extract information from these web pages,the paper presents a position-based information extraction method.By simulating the rendering process of web browsers,it gets the exact position information on the browser screen of each tag of HTML documents to extract the position features.Then it uses parts of the pages set retrieved by TimeWeightedHits method as the training set to let SVM generate a classification model,which is used to predict the test set.Compared with another method which based on manually defining heuristic rules,it improves accuracy greatly.Moreover,it reduces the work during training process.

Related Dissertations

  1. Soft Sensor of Naphtha Dry Point on Support Vector Machines Regression,TE622.1
  2. The Research of the Fault Diagnoses Algorithm for the Liquid Rocket Engine Testing Bed Based on PCA-SVM,V433.9
  3. ISAR Imaging Simulation of Space Targets and Target Recognition Based on ISAR Images,TN957.52
  4. Research on Autamatic Music Structrue Analysis,TN912.3
  5. Research on Feature Extraction and Classification of Pulse Waveform for Cholecystitis and Nephrotic Syndrome Diagnosis,TP391.41
  6. Research on Classification Method of Tongue Substance Color and Tongue Coating Color Based on SVM,TP391.41
  7. The Research on Paper Currency Classification Method Based on Harr-Like Feature and Minimal Ball Including Samples,TP391.41
  8. Research on Focused Crawler Based on SVM Classification Algorithm,TP391.3
  9. Research on Predicting Intrinsic Disorder Protein Structure Based on Supervision Manifold Learning Algorithm,Q51
  10. Study on the Road Condition Monitoring Based on Vehicular 3D Acceleration Sensor,TP274
  11. Research of Orange Quality Classification Technology Based on Computer Vision,TP391.41
  12. Analysis on Synoptic Climatology Characteristics and Forecast Methods of Fog in Hainan,P457
  13. Research of License Plate Recognition Based on Rough Sets and Fuzzy SVM,TP391.41
  14. Study on Visual Target Detection Based on SVM,TP391.41
  15. Web Page Attribute Extraction Method Research,TP391.1
  16. The Research on Electrode 3D Model Classification and Retrieval Based on SVM and Shape Features,TP391.41
  17. Multi-step-ahead Stock Price Index Forecasting Based on Hybrid Models,F224
  18. Research on ECG Feature Extraction and Classification Method,TN911.7
  19. Multi-feature fusion of visual tracking algorithm,TP391.41
  20. Or figure based on license plate detection and recognition,TP391.41
  21. EEG -based image retrieval of emotional,TP391.41

CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > General issues > Theories, methods > Algorithm Theory
© 2012 www.DissertationTopic.Net  Mobile