Dissertation > Excellent graduate degree dissertation topics show

On Similarity Search Among Web Pages

Author: JinZuoLu
Tutor: ZhangYueQin
School: Taiyuan University of Technology
Course: Software Engineering
Keywords: web page network similarity search SimRank static pruning
CLC: TP393.09
Type: Master's thesis
Year: 2013
Downloads: 7
Quote: 0
Read: Download Dissertation

Abstract


With the popularity and rapid development of Web technology, all kinds of information that people can easily access through a browser. But the users not proceed that the rapid growth of web data is often faced with a lot of information.requires an effective Web page similarity search methods to obtain useful information from the Web. It requires an effective similarity search methods of web page to obtain useful information from the web. The problem on similarity search among web page to be solved is:according to a given web page, you can return a number of other similar pages. Furthermore, hyperlinks between pages can reflect the possible direction of information transfer between each page. It can provide an important basis for the discovery of similar pages. To solve this problem is to find an efficient and accurate similarity algorithm, and be able to quickly respond to user requests.Web page network usually presents large-scale and rapid growth, it makes the similarity calculation seems more challenging. But traditional algorithm of content-based similarity may be related to some privacy in the contents of the user, and is not accurate enough. In contrast, based on the structure of the similarity algorithm on the accuracy of calculation will be much higher. SimRank is a classical algorithm based on structural similarity, and has the characteristic "rapid convergence", however, it can not be applied to large networks well since its space and time cost is very high. To solve the above problem, we proposed an efficient similarity search method among Web pages (WSR) based on a SimRank. It can significantly reduce the space cost and pre-computation time cost. Using static pruning technique to optimize the web network and proposed WSR-pruning algorithm, and improve the efficiency of pre-computed efficiency and online query processing.The work includes mainly the following sections:1. We explained the background and significance of the research, and the development and status on similarity search among web pages.2. Introduced the classic link analysis algorithm, as well as several similarity calculation method.3. Proposed on similarity search among web pages based on SimRank. Elaborated the basic idea and recursive iterative process on SimRank, and analyzed its strengths and weaknesses.4. To the problem of calculation on the time and space overhead huge for SimRank, we proposed an efficient similarity search method among Web pages (WSR) based on a SimRank and Network of relationships to Web pages. It only compute the2-hop similarities among Web pages, i.e., for a given query, we computing the2-hop similarity between query and each page in the network based on the pre-computed1-hop similarity matrix. Given a algorithms of online query processing, and analysis online query processing algorithm time complexity and error precision Theoretical. There are some of the Web page in the network is not important page link relationship, so similarity calculation is not of great value, and the page can not be deleted, then consider to reduce these unimportant link relationship in order to improve the calculation efficiency. Using static pruning technique to optimize the web network and proposed WSR-pruning algorithm.5. Comparative analysis by experiment. The comparative analysis of results show, compared with traditional SimRank, WSR and WSR-pruning reduce the storage overhead and computational time overhead. They have a higher accuracy and rapid query response time.

Related Dissertations

  1. Graph-based similarity search scientific literature key technology research,TP391.3
  2. Research and Application of Similarity-based Mining of Financial Data Analysis System,TP311.13
  3. Web -based social tagging Retrieval,TP393.092
  4. Graphical Mining Based on Time-series,TP311.13
  5. The Design and Implement of Graduate Occupation Recommending System,TP311.52
  6. Semantic Web and Its Application in Search Engines,TP391.3
  7. Large-Scale Image Copy Detection Algorithm Research Based on Local Media-Fingerprints,TP391.41
  8. Research and Implementation of Content Forecasting Mechanism in CDN-P2P System Based on Hadoop,TP393.02
  9. Research of Deep-sea Floating Platform Monitoring and Data Similarity,TE953
  10. The Research on a Few Key Issues in Querying Algorithms over Data Streams,TP301.6
  11. Research on Mining and Its Application in Time Series Database,TP301.6
  12. The Research on Time Series and Cluster Analysis,TP311.13
  13. Similarity Search in Peer-to-Peer Systems,TP301
  14. Research on Vector Approximation Method in High-dimensional Index Technology,TP391.4
  15. Similarity Search and Outlier Detection in Time Series,TP311.11
  16. Research on Several Key Issues in Financial Time Series Mining Based on Feature Analysis,TP311.13
  17. Research on Mining and Similarity Searching in Time Series Database,TP311.13
  18. Several high-dimensional data mining research on key issues,TP311.13
  19. The Study of Some Security Problems for the Information Integrated Platform in Manufacturing Enterprises,TP309
  20. Research and Application of Financial Multivariate Time Series Mining Methods,F830

CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer applications > Computer network > General issues > The application of computer network
© 2012 www.DissertationTopic.Net  Mobile