Dissertation > Excellent graduate degree dissertation topics show
Research on Hyperlinks Extraction Based on Hotspot Website Content Analysis
Author: LangFengJu
Tutor: ZhangWei
School: Ocean University of China
Course: Applied Computer Technology
Keywords: HTML Parser Information extraction Web analysis Double feedback
CLC: TP393.092
Type: Master's thesis
Year: 2010
Downloads: 57
Quote: 0
Read: Download Dissertation
Abstract
|
Information on the Internet is very extensive, and that there are many hotspot information people are concerned about, these content of hotspot website is on behalf of the most interested part of the internet information, the purpose of this paper is on NBA hotspot website content analysis, parser hyperlinks and corresponding text information, and then through the URL and text feedback the heat degree of website.Firstly,This paper introduces the overview of web information extraction,its developing history and current status,analyses technique classification of web information extraction and its common algorithm and elaborates the technique and weighing measure of hotspot web information extraction. Secondly, this paper analyses the feature of the hotspot website page content analysis, including the characteristics of the hotspot website and hotspot website sports NBA, the characteristics and comparison of Sohu and NetEase NBA page, analysis hotspot website, hotspot content analysis and internal relationship from Web language features. By contrast,summarizes the features of the NBA class hotspot website content, and these characteristics are ideal for some HTML Parser to parse. Through the contrast of the characteristics of mainstream HTML Parser, elaborate the superiority of HTML Parser on analyzing hotspot web page, through the realization of hotspot web page gathering, has further confirmed web page internal composition structure and constitution characteristic and so on, puts forward a kind of hotspot double feedback URL and text extraction strategy based on HTML Parser:First through HTML Parser extracts web page’s URL, then extracts text from URL, feedbacks URL heat degree through the extraction text, feedbacks the heat degree of entire web page through URL again.Finally, this paper realizes website hyperlink information extraction based on HTML Parser. This paper mainly introduces the realization of hotspot website hyperlink and text information extraction using two algorithms. System operation results and effectiveness evaluation mainly through querying Sohu sports NBA and NetEase sports NBA extraction results, the two performance indicators of test precision and recall rate to compare two hotspot website links extraction effect, and through URL and text feedback the heat degree of website.The hyperlinks extraction based on hotspot website content analysis this paper is studying has already realized the simple web page analysis, can filter some garbage information and clean noise, but whether it can satisfy the user request truly,enhance useful information feasibility which is analyzed, also needs to study further.
|
Related Dissertations
- Research on Domain Entity Attribute and Event Extraction Technology,TP391.1
- Research on Temporal Information Recognition and Normalization,TP391.1
- Based on high-resolution remote sensing data mining houses information extraction,TP751
- Object-Based Automatic Extraction of Change Information Based on High-Resolution Remote Sensing Image Research,P237
- Comparative Experimentation Study on Information Extraction of Object-Oriented Based on Quick Bird Image,P237
- Object-oriented Information Extraction of woodland,P237
- Research on Object-oriented Remote Sensing Image Information Extraction Technology,P237
- Study on Technologies of Remote Sensing Feature Analysis and Information Extaction of Earthquake Disaster,P237
- City names addresses coding Matching,P208
- Criminal cases text information extraction,G353.1
- Research on Technology of Deep Web Oriented Data Extraction and Semantic Annotation,TP311.13
- Crawler and Incremental Update Strategy Research in Deep Web,TP391.3
- Statistic-based Automatic Keypharse Extraction and Summarization from Multi-document,TP391.1
- Related Studied on Information Extraction and Information Recommendation Based on Web Data Mining,TP393.09
- Research on Spectral Analysis and Its Application to Mineralization Information Extraction,P237
- The Research of Man-made Objects Extraction Methods Based on QuickBird Images,P237
- Research and Implementation of Topic-Based Event Fusion,TP393.09
- Chinese BBS Information Extraction and Classification,TP393.094
- The Research of Paper in Quality Evaluation Based on Sciencepaper Online,TP393.092
- Personal Position and Title Information Extraction in Large-scale Real Texts,H08
- Research on Web Information Extraction Tool,TP393.092
CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer applications > Computer network > General issues > The application of computer network > Web browser
© 2012 www.DissertationTopic.Net Mobile
|