Dissertation > Excellent graduate degree dissertation topics show
An Application Research of Information Extraction on Topic Search Engine
Author: XuCheng
Tutor: HeYouQuan
School: Chongqing Jiaotong University
Course: Management Science and Engineering
Keywords: Topic search engine Page denoising Feature Extraction Text Classification
CLC: TP391.3
Type: Master's thesis
Year: 2011
Downloads: 76
Quote: 1
Read: Download Dissertation
Abstract
|
The rapid development of the network and information technology makes the Internet to become an important resource for the dissemination and sharing of global information platform. The geometrically growing amount of information in the Internet, makes difficult to obtain useful information from the Web has become increasingly large. Information overload has become a pressing problem. There is an urgent need for a like query the database query the exact query methods of data on the Web. Therefore, how to effectively organize and analyze vast amounts of Web information resources and extract useful information to many research hope to solve the problem. Search engine in recent years had a theme, it is professional or theme-oriented search engine, only collecting information related to the user's needs, access to knowledge and information they need quickly, which need to be extracted from the Web repository a specific topic related information. At present, the key technologies involved in the topic search engine, structured information extraction technology is different from the key points of the general search engines. In information extraction technology, multi entire web page to extract structured information and generate a wrapper, thereby generating wrappers information extraction accuracy rate is not high. Interference due to \Improve the efficiency and accuracy of information extraction can greatly improve the adaptability topic search engine, and provides users with efficient and accurate search results. This article is for pages structured information extraction technology research, an effective and reliable method to improve and enhance the performance and efficiency of the information extraction links, and to automate the information extraction process as much as possible, in order to avoid artificial excessive interference; strive able full and effective use of resources, reasonable construction of the access to information system architecture. This thesis mainly include the following aspects: 1. Bulk acquisition research Web page. Professional field of information needed in order to get the theme search engine, you must grab the relevant page from the internet as the original. Web pages crawled need technology and crawl submodule structure, and a simple Web page batch fetching. 2 Web page denoising. Due to the presence of noise, serious impact on the identification of the content of the page theme, which affects the quality and efficiency of the final search results. Clear the noise content pages become an important guarantee to improve the accuracy of the search engines and the premise. In this paper, the original all kinds of pages block the advantages and disadvantages of denoising model analysis, statistical methods into the DOM-based web pages, in order to enhance the effectiveness and efficiency of denoising block model. 3 study Chinese text vector model and feature vector extraction method. The text feature indicates the merits of a direct impact on the extraction of the characteristic properties, thus affecting the classification of the latter part of the text. Look at the results from the previous studies, the vector space model of text feature indicates how much of the feature words the number of model classification algorithm efficiency is closely related to this study the extraction method of feature words, analyzed based on several popular features of the vector space model extraction algorithm, and in-depth study of the CHI algorithm and its improvements, to improve the performance and effectiveness of the algorithm, better late text classification. 4. Study the text of the Web page classification. Due to the huge number of Web pages, the lack of reliable label information characteristics, its use them effectively, you must first automatic classification to obtain more reliable category label information to categorize the subsequent processing. Classic Naive Bayes algorithm based on the algorithm to improve, and the coefficient of adjustment, and taking into account the unique semi-structured characteristics of web text data based on the extraction and use of the structural information of Web text classification process to to strengthen the classification results, further reducing the classification error. The final text of Web pages classification comparative experiments, the experimental results show that the proposed algorithm can effectively improve the classification accuracy and reduce the computation time, and apply to the topic of search engine information extraction. Finally design theme search engine information extraction system overall architecture and sub-module flowchart and part specific implementation code for future researchers reference. The innovation of this paper is that the demand for information extraction for topic search engine denoising improved website page text feature extraction and classification algorithm, after experimental verification compared to similar algorithm can effectively improve the accuracy and reduce the time complexity. Well adapted topic search engine web text information extraction needs.
|
Related Dissertations
- Research on Automatic Detection Algorithm for Substructure Distress of Highway Pavement Based on SVM,U418.6
- ISAR Imaging Simulation of Space Targets and Target Recognition Based on ISAR Images,TN957.52
- Research on Feature Extraction and Classification of Pulse Waveform for Cholecystitis and Nephrotic Syndrome Diagnosis,TP391.41
- Application of Q-Learning in the Content-Based Image Retrieval Technology,TP391.41
- Research on Transductive Support Vector Machine and Its Application in Image Retrieval,TP391.41
- Research on Feature Extraction and Classification of Tongue Shape and Tooth-Marked Tongue in TCM Tongue Diagnosis,TP391.41
- Research on Text Classification Based on Biomimetic Pattern Recongnition,TP391.1
- Tourism Comments on the Internet’s Semantic Analysis and Usefulness Research,TP391.1
- Research on Visual Measurement for Spacecraft Rendezvous and Approach,TP391.41
- Research on the Image Real-Time Acquisition, Storage and Image Processing System,TP391.41
- Feature Extraction, Selection and Combination in Lipreading,TP391.41
- Multi-currency Notes Technology Research and Implementation,TP391.41
- The Research on Paper Currency Classification Method Based on Harr-Like Feature and Minimal Ball Including Samples,TP391.41
- Pavement Distress Recognition Based on Image,TP391.41
- Research on Visual Detection and Tracking of Mobile Robots,TP242.62
- Research on Fusion Algorithm of Hyper Spectral and High Spatial Resolution Remote Sensing Image,TP751
- An Approach for Identifying a Plant Resistance Gene Based on the Random Forest,Q943
- Tobacco Diseases Auto-Recognition Research Based on Image Processing Technology,S435.72
- Research on Nondestructive Detection Technology for External Qualities of Papayas Based-on Vision,S667.9
- Research on Identification System of Cashmere and Wool Fiber,TS101.921
- Research for Infrared Image Target Identification and Tracking Technology,TP391.41
CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer applications > Information processing (information processing) > Retrieval machine
© 2012 www.DissertationTopic.Net Mobile
|