Dissertation > Excellent graduate degree dissertation topics show

Web Document Clustering Based on Knowledge Granularity

Author: HuangFaLiang
Tutor: ZhangShiChao;YanXiaoWei
School: Guangxi Normal University
Course: Computer Software and Theory
Keywords: Data Mining Web Mining Web Document Clustering Rough Set Knowledge Granularity
CLC: TP393.092
Type: Master's thesis
Year: 2005
Downloads: 367
Quote: 1
Read: Download Dissertation

Abstract


The rapid development of the Internet (WWW) has greatly changed people's lives, it has become the mainstream channels to exchange ideas and access to information. Hidden in the vast expanses of the network data marine wealth of valuable knowledge quickly and efficiently obtain useful knowledge from the this mass data source in all users, including businesses, individuals and other must face and resolve the problem. Thus, the use of data mining (Data Mining) technology for Web Data Mining (Web Data Mining) become an important research focus in the field of data analysis, a wide range of experts and scholars. After nearly a decade of growth, Web data mining and achieved fruitful results, many related technologies mature and stable and good application in the actual production and life, such as search engine to hunt people a tremendous the convenience of e-commerce for industrial business community provides a whole new way of doing business. Compared with the traditional data, network data has a complex structure, forms and wide, and the functional requirements of the users of the Web data mining is a wide variety of which presented a greater challenge to the field of data analysis. Web data mining can be roughly divided into three parts: content mining, usage mining, and structure mining. They are the main technical: correlation analysis, time series analysis, cluster analysis. Web data clustering analysis is the basis of a core Web data mining research topic. Cluster analysis multifaceted role in the search space, with compression to speed up the retrieval speed. Knowledge workers can help to efficiently and accurately find the most similar document with a document; improve information retrieval system, the rate of return (Recall) and precision (Precision); improve the degree of personalization of the search engine. The network on the most common and most important form of data is a markup language that Web documents. Therefore, the cluster analysis of Web documents is a very important and valuable work. In this paper, in-depth understanding of the existing Web data mining technology, the basis of Web Document Clustering analysis technology to analyze the model of traditional text and text clustering algorithm, analysis of the existing representation model with existing clustering algorithm advantages and defects. In order to overcome the deficiencies of the existing clustering algorithm, this article will introduce the theory of knowledge granularity, Web document clustering method based on knowledge granularity. This thesis work mainly include the following aspects: (1) traditional Web clustering method is mainly based on two knowledge granularity of the document - feature words, this will lead to a \based on multi-level granularity Web document representation mechanism and theory, and gives a concrete model Web documents based on multi-level granularity: document - paragraph - characteristic words \model); (2) In this model, we noted that the VSM-based similarity measure commonly used \generated based on these problems, we introduce tolerance rough set theory, said the extended model based on rough set text: EVSM; (3) in the selection process of the clustering algorithm, both take into account the traditional K-means clustering method is suitable for processing of the massive set of documents, and consider it more sensitive to outlier data (which is not ideal for non-spherical data clustering effect) on the basis of the traditional K-means, therefore, we propose an improved K-means clustering algorithm: NK-means. (4) Finally, we proposed and a platform for Web data analysis WebAnalyser further on this platform for Web document clustering analysis of WCBGK of algorithms.

Related Dissertations

  1. Research and Implementation of Mining Implicit User Interest,TP311.13
  2. Fault Diagnosis Method Based on Support Vector Machine,TP18
  3. A Study on Healthcare Product Marketing Based on Data Mining Technology,F426.72
  4. Gao Zhong-ying academic thought and experience and use of Bufei Decoction treatment of common diseases of the respiratory system drug law,R249.2
  5. Bing- thick academic thought and clinical experience and empirical studies apply to turtle soups treatment of chronic kidney disease,R249.2
  6. The Design and Implementation of Bicluster Data Analyzing Software,TP311.52
  7. Research on Clustering Algorithm Based on Mutation Particle Swarm Optimization,TP18
  8. Research on Fuzzy C-Mean Clustering Algorithm Based on Particle Swarm Optimization and Shuffled Frog Leaping Algorithm,TP18
  9. Research on Clustering Algorithm Based on Genetic Algorithm and Rough Set Theory,TP18
  10. Based on Rough Set of Urban Areas When Traffic Green Control System Research,TP18
  11. Based on data mining research tax audit case selection,F812.42
  12. Community-oriented education, personalized learning system and its implementation,TP391.6
  13. Association rule mining based Intrusion Detection System Research and Implementation,TP393.08
  14. Data warehouse technology in the banking customer management systems research and implementation,TP315
  15. Design to E-learning System in Senior Vocational School Base on Moodle,TP311.52
  16. Design and Development of Teaching Quality Assessment System Based on Data Mining,TP311.13
  17. Incremental rough set attribute reduction,TP18
  18. Calculation of Knowledge Granulation and Study of Its Application in Attribute Reduction,TP18
  19. The Application of Association Rules Algorithm in Higher Vocational Colleges’ Endorsement of Impoverished Students,G717
  20. Based on Data Mining Technologies in Urban Water Supply Analysis and Decision,F299.24;F224
  21. Research on Application of Data Mining Technology in Degree of Satisfaction Analysis of Television Customers,TP311.13

CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer applications > Computer network > General issues > The application of computer network > Web browser
© 2012 www.DissertationTopic.Net  Mobile