Dissertation > Excellent graduate degree dissertation topics show

Research of Deleting Duplicate Web Pages on Campus Search Engine

Author: ZhangFang
Tutor: GaoYongBing
School: Inner Mongolia University of Science and Technology
Course: Applied Computer Technology
Keywords: Duplicate Detection Algorithm Fingerprint Signature Fourier Transform Campus Network Search Engine Nutch
CLC: TP391.3
Type: Master's thesis
Year: 2012
Downloads: 40
Quote: 0
Read: Download Dissertation

Abstract


With the rapid development of the campus network construction, campus network informationresources increase quickly, which makes difficult for teachers and students to locate valuableinformation quickly, and it is also a waste of time and efficiency. Based on the characteristics of thecampus network itself, mature general search engines cannot be applicable to the campus networkcompletely, and retrieval results have too many duplicate web pages due to the copy messages fromother websites. By analysis of the characteristics of the pages of the campus network and theexisting duplicate detection technology, the problem that the retrieval results have to many duplicatepages is solved. For different types of duplicate pages, the index and real-time search duplicatedetection strategies are used to build a campus network search engine. The work is finished asfollows:Firstly, web page duplicate detection algorithms are researched and analyzed. Foremost thecauses of web page noise, the definition and types of noise are analyzed, and correlation contentmerger technology is used to get rid of noise and extract content in order to obtain the content of theweb pages. Then, after the study and analysis of Chinese word segmentation technologies, Paodingword segmentation analysis is applied to Nutch secondary development--Nutch source is modifiedand Chinese word segmentation is realized.Secondly, duplication detection algorithm is researched and improved in the index phase. Afterthe analysis of existing algorithms, the duplicate detection algorithm based on the longest paragraphand fingerprint is proposed for the complete or part duplication of web pages. Firstly, duplicatepages are removed from the entire documents. Secondly, the document from which duplicate pageshave been removed and filleted is segmented. Paragraphs sorted are sorted, and then the first Nparagraphs are taken for fingerprint signatures which are used for the characteristics of the document.When the number of paragraphs which are the same in the two documents exceeds a given threshold,which determines whether these two documents are duplicate documents. The first N segments havebeen extracted and the fingerprint is sorted, which greatly reduces the computational complexity.Experiments show that this method has higher duplicate detection accuracy.Thirdly, duplicate pages modified minimum from the original pages use the optimized Fouriertransform duplicate detection algorithm. The algorithm maps each word of each document into anumerical fingerprint, and each document can be expressed as a discrete sequence of numbers. TheFourier coefficients are obtained by use of the Fourier transform to process the numerical sequence,and the similarity of the two series can be obtained by comparing the first several coefficients. The experiments show that recall rate and duplicate detection rate both can be taken into account basedon the optimized Fourier transform algorithm for the web pages which have been modified.Nutch is used as a system development tool and through the modification of the Nutch sourcecode,duplicate detection algorithm is achieved when the web pages are indexed. Web pageduplicate detection algorithm is achieved by the form of plug-in when the pages are being retrieved.Campus network search engine is designed and implemented on the basis of the Nutch. The campusnetwork search engine system development process and methods are described in detail. Finally,using Nutch to crawl the pages of the campus network as the experimental data sets, duplicatedetection strategy proposed is tested experimentally. The results show that the accuracy rate ofsearch results and the duplicate detection are improved by the combination of the two algorithms,and the campus network search engine system built can be runned effectively and steady.

Related Dissertations

  1. Web search engine related technology research,G354
  2. The Design and Implementation of Lucene-Based Network Literature Vertical Search Engine,TP391.3
  3. Study on Web-based Translation Technology for Out-of-Vocabulary,TP391.2
  4. The Research of Link Structure in Tibetan Web Base on Social Network Analysis,TP393.09
  5. Research of Intelligent Search Engine Based on Semantic Web,TP391.3
  6. Meta Search Engine Based on BP Network,TP391.3
  7. Search Engine Provider Copyright Infringement Liability Standard Discussion,D923.41
  8. Motivations of Young Spectators of Human-powered Search Engine,G206
  9. The Method Research of the Network Dynamic Vocabulary System Construction,G354
  10. The Research on Deep Web Interfaces Integration and Query Results Ranking,TP274
  11. Study on Safety Education Management Information System Based on Struts2,TP311.52
  12. Study on the Search Engine Marketing of Enterprises,TP391.3
  13. The Study of Topic-Oriented IT News with Search Enging and Web Page Analysing,TP393.092
  14. Search engines trademark infringement and Responsibility,D923.43
  15. Research and Implementation of Vertical Search Engine Key Technology,TP311.52
  16. Based on analysis of public opinion judged the emergency warning platform for the design and implementation of,TP311.52
  17. The Research on Key Technologies for Web Information Personalization Collection and Management,TP393.09
  18. The Research and Design of Monitoring System for BBS,G206
  19. The Long Tail Theory Applied in Baidu’s Online Advertising,F49
  20. Mobile Internet environment Wireless Vertical Search Engine Research and Implementation,TP391.3
  21. Fire domain search engine system design and implementation,TP391.3

CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer applications > Information processing (information processing) > Retrieval machine
© 2012 www.DissertationTopic.Net  Mobile