Dissertation > Excellent graduate degree dissertation topics show

Website Crawler and Retrieval System Based on Lucene

Author: XiQi
Tutor: LvZhenLiao
School: Northeastern University
Course: Computer Software and Theory
Keywords: Common subsequence Web Spiders Lucene Inverted index Full-text search
CLC: TP393.092
Type: Master's thesis
Year: 2008
Downloads: 168
Quote: 0
Read: Download Dissertation

Abstract


With the development of network and Web resources rich Web text information retrieval system to obtain the required information has become an important part of people's daily lives, and users are increasingly concerned about how to more accurately and efficiently find information. Theory and technology of Web information retrieval system are introduced, and a more in-depth practice of information retrieval in the Web text information retrieval applications. In the second chapter, theoretical papers, such as the type of the search engine, the Chinese word segmentation method, the theory of inverted index used in this thesis Lucene theory and use in the third chapter, according to the the site features two page templates analysis algorithms, the first algorithm is based on the longest common subsequence model using dynamic programming method to find the optimal solution, optimized and extended algorithm of the original literature, in order to achieve find the web template string and the purpose of the string is inserted, the second algorithm uses the statistical theory and principle, the page template abstracted into a mathematical model, to extract the common identity of the web page to start position and end position , different variance different calculated according to the length of the body to determine the position of the body in the into the original, which can be extracted from a web page into the contents of the body, in order to achieve the purpose of saving space and reducing the time of the indexing and search index, and finally compare the advantages and disadvantages of the two algorithms. In the fourth chapter, a web spider Java developers, including the processing of heterogeneous data, such as word, pdf, rtf file content extraction, addition of an HTML file parsing method and multi-threaded use. In the fifth chapter, to achieve a web crawling system, the information specified by the user from the Internet to download pages, including content Next, in order to improve the efficiency of information retrieval using Lucene package, the establishment of full-text search to speed up the search speed and improve the accuracy and timeliness of the information, and save a lot of storage space. From a theoretical and practical point of view, both the design and analysis of algorithms, but also specific program implementation, using Oracle, Tomcat, Jsp, Java, Eclipse, Lucene software and language original HTML parsing method for the Users save time, improve work efficiency.

Related Dissertations

  1. Web Service Platform Research and Implementation Based on MVC Design Pattern,TP311.52
  2. Network Intelligent Answering System Research and Implementation,TP393.09
  3. Research on Information Retrieval Technology Based on Semantic Web,TP391.3
  4. Research and Implementation of Website Search Technology Based on Ajax/Lucene,TP393.092
  5. Vertical Search Engine Technology in Network Control in the Patrol Public Opinion Research and Application,TP391.3
  6. A Study on Compression Algorithm Performance Based Inverted Index,TP391.3
  7. Study on an Improved Chinese Segmentation Algorithm and Its Application in Lucene,TP391.1
  8. Distributed Image Search Engine Design and Implementation,TP391.41
  9. Research and Implementation of full - text retrieval system based on lucene Chinese,G354
  10. Java-based Zhejiang Textile \u0026 Fashion College campus network search engine,TP393.18
  11. A Multi-semantic Search Engine for Traditional Chinese Medicine,TP391.3
  12. Research and Implementation of Search Engine Prototype Based on Deep Web Crawler,TP391.3
  13. The Design and Implementation of Second-Full-Text Retrieval System Based on Lucene,TP391.3
  14. The Research of Subject-Oriented Search Engine Based on Lucene,TP391.3
  15. Research and Design of Military Document Retrieval System Base on LUCENE,TP391.3
  16. Design and Implementation of a Lucene Based Intra-site Information Retrieval System for a Journal Site,TP391.3
  17. J2EE frameworks and full-text information management system,TP311.52
  18. Research and Implementation of Vertical Search Engine Techniques,TP391.3
  19. Figure Nutch -based search engines love blog design and implementation,G250.73
  20. Research and Implementation on Trading Information Oriented Vertical Search Engine,TP391.3

CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer applications > Computer network > General issues > The application of computer network > Web browser
© 2012 www.DissertationTopic.Net  Mobile