Dissertation > Excellent graduate degree dissertation topics show
Block Based Web Community Identification
Author: GaoLiang
Tutor: ZhangXianChao
School: Dalian University of Technology
Course: Applied Computer Technology
Keywords: Community Identification Web Page Segmentation Link Analysis Max Flow Algorithm
CLC: TP393.092
Type: Master's thesis
Year: 2011
Downloads: 14
Quote: 0
Read: Download Dissertation
Abstract
|
With the rapid development of the Internet, Web has been the most important platform for modern people to share information and resources. Mining the characteristics of the Web is essential to obtain and understand the information on the Web. Because of the self-organization of the Web, there are many communities in the Web, and how to extract and make use of these communities has become one of the most important research topics in Web Data Mining.A Web community is defined as a set of Web pages created by people with the similar topic. Community is of great value in Web research, it reflects the social action of Web users, evaluation history and inter-relation of Web. It provides the most creditable resources on certain topic. Moreover, it provides an effective way to improve the efficiency of searching result.Currently, in community identification research, a Web page is usually considered to concern on only one interest, so one node in Web graph is correspond to one Web page. However, in fact, there are often multiple interests in one Web page, and it may generate different kind of links for different interests. If we don’t do segmentation for this kind of Web pages, the community from link analysis result will contain many irrelevant pages. To solve this problem, we propose a block based Web community identification algorithm. In this algorithm, we begin with do segmentation for every Web page, and then, construct the Web graph with the blocks, finally, indentify the communities in the block based Web graph. The experimental results indicate that our approach has better performance than page based algorithm in avoiding topic drifting and improving the accuracy of the extracted community.
|
Related Dissertations
- Study on the Improvement of Hits Algorithm of Web Structure Mining,TP311.13
- The Link Analysis of Chinese Provincial Government Websites,D625
- Spider Crawling on Mobile Search Research and Implementation Strategy,TP391.3
- A Study on Methods of Web Page Topical Information Extraction,TP393.092
- Link analysis based evaluation system of digital library,G250.76
- Research of Small World Effect in World Wide Web,O157.5
- Research on Financial Blog Crawling and Ranking Algorithm,TP393.092
- Web Link and Web Traffic Analysis for Websites of 211 University Libraries Based on Webmetrics,G250.7
- Web Page Sorting Algorithms Based on the Analysis of the Linking Structure,TP393.092
- Research on Extraction and Analysis of Information Networks and Rank Related Problems in INTERNET,TP391.1
- Knowledge Discovery in Time Series,TP311.13
- Research on Entity-level Search Crawler and Information Extraction,TP391.3
- The link similarity web - based Sorting Algorithm,TP391.3
- Research of Blog Ranking Algorithm Based on Link Analysis and Blog Characteristic Features,TP391.3
- Research and Improvement on Link Analysis Based on HITS Algorithm,TP393.01
- Approach for Topic Tracking Based on Semantic and Hyperlink,TP391.1
- Study of Web Information Retrieval Based on Structure and Subject,TP391.3
- The Study of Retrieval Method and Searcher of Internet Search Engine,TP391.3
- Research on the Hyperlink-based Ranking Algorithms in Web Information Retrieval,TP393.09
- Webometrics : Theories , Methods and Case,G350
CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer applications > Computer network > General issues > The application of computer network > Web browser
© 2012 www.DissertationTopic.Net Mobile
|