Dissertation > Excellent graduate degree dissertation topics show
Study on Key Technologies of Frequent Items Mining and Clustering on Data Streams
Author: TuLi
Tutor: Chen
School: Nanjing University of Aeronautics and Astronautics
Course: Applied Computer Technology
Keywords: Stream data Data Mining Frequent items Sliding window Time decay model Data stream clustering Multiple Data Stream Clustering
CLC: TP311.13
Type: PhD thesis
Year: 2009
Downloads: 351
Quote: 2
Read: Download Dissertation
Abstract
|
With the rapid development of information technology, streaming data in different ways in many areas of the application. Including network traffic monitoring, financial applications, communications, data management, network security monitoring, sensor networks and so on. In these applications, the novel forms of streaming data as well as the related art, the research is very important. Therefore, the data on the data stream mining become a hot research field of data mining. Mainly current data stream mining: frequent item (set) mining, cluster analysis, classification, anomaly analysis. This paper analyzes the current stream data mining algorithms, cluster analysis on single and multiple data streams for mining frequent item on the data stream, and a more efficient algorithm. The main contribution of this paper is as follows: (1) most of the existing data stream mining frequent item algorithms and there is not enough emphasis on the importance of the current data. A sliding window of the most recent period of data mining technology. Therefore, we propose a sliding window of stream data frequent items mining algorithm. The algorithm uses a linked list queue strategy can be greatly simplified algorithm to improve the efficiency of the mining. Threshold value S, the error ε, and the window length for a given n, the algorithm can be in εn error detection window frequency data stream over the Sn frequent, and their complexity is O (ε -1 sup >), the processing and query time of each data item are O (1). On this basis, we will also have the algorithm extends the to frequent different stream data mining algorithms, parameter changes, making the algorithm can be adjusted between the time and space complexity. By a large number of experiments show that the proposed algorithm has better accuracy and time and space efficiency than other similar algorithms. (2) weakened the time decay of the importance of the \FC2 space algorithm complexity is O (ε -1 sup>), each data item, the processing time is O (1). By a large number of experiments prove, FC2 has higher accuracy than other similar methods, faster processing speed and less memory requirements. Then, we propose a simple and fast mining data streams frequent the λ-Count algorithm. The algorithm can be in the spatial complexity of O (logλε) detecting ε-approximate frequent items, the processing time for each data item is O (1). By a large number of experiments show that, λ-Count correct rate, memory requirements and processing speed than other similar methods. (3) most existing real-time streaming data clustering algorithm, as CluStream, are based on the k-means algorithm. These algorithms mining clusters of arbitrary shape, and processing of isolated point, there is a difficulty, but these algorithms require a priori knowledge to determine the number of clusters k and the length of the time window of the user-defined. To solve these problems, we propose a D-Stream framework of stream data clustering algorithm based on density and have proposed a strategy of the algorithm based on this framework, the algorithm DS0 and introduced to attract degrees DS1. Attract potential link between the degree of data density and cluster structure by exploring the attenuation coefficient, the algorithm can effectively generate clustering and real-time adjustment, to detect and remove those singular cells isolated point mapping system to dynamically increase the space and time efficiency. Experimental results show that the algorithm has high quality and efficiency, can accurately reflect the process of the evolution of the real-time data streams. (4) the clustering of multiple data streams are usually Euclidean distance to measure the similarity between the data stream. Euclidean distance has a lot of limitations, it ignores the trend of the data stream and sequence mode. For the user, this information is often more important. To this end, we propose multiple data stream clustering algorithm based on Kendall correlation coefficient. The algorithm uses the AU statistic multiple data streams of raw data quickly compressed into a statistical summary. According to these statistics Summary incremental calculation Kendall correlation coefficient to measure the similarity between the data. We also propose the k-means algorithm to generate a dynamic clustering results. Dynamic k-means algorithm dynamically, in real-time adjustment of the number of clusters, the timely detection of the development and changes of the data stream. The algorithm is applied to the clustering problem (COD) in accordance with user requirements, so that the user can query clustering results in any time interval. Fragments divided by a reasonable time, allows the user to specify an arbitrary time interval can be a combination of these time pieces. Experimental results show that the algorithm is better than other similar method clustering quality, speed and stability, in real-time to reflect the changes in the data stream.
|
Related Dissertations
- A Study on Healthcare Product Marketing Based on Data Mining Technology,F426.72
- Gao Zhong-ying academic thought and experience and use of Bufei Decoction treatment of common diseases of the respiratory system drug law,R249.2
- Bing- thick academic thought and clinical experience and empirical studies apply to turtle soups treatment of chronic kidney disease,R249.2
- The Design and Implementation of Bicluster Data Analyzing Software,TP311.52
- Research on Clustering Algorithm Based on Mutation Particle Swarm Optimization,TP18
- Research on Fuzzy C-Mean Clustering Algorithm Based on Particle Swarm Optimization and Shuffled Frog Leaping Algorithm,TP18
- Research on Clustering Algorithm Based on Genetic Algorithm and Rough Set Theory,TP18
- Based on data mining research tax audit case selection,F812.42
- Community-oriented education, personalized learning system and its implementation,TP391.6
- Association rule mining based Intrusion Detection System Research and Implementation,TP393.08
- Data warehouse technology in the banking customer management systems research and implementation,TP315
- Design to E-learning System in Senior Vocational School Base on Moodle,TP311.52
- Design and Development of Teaching Quality Assessment System Based on Data Mining,TP311.13
- The Application of Association Rules Algorithm in Higher Vocational Colleges’ Endorsement of Impoverished Students,G717
- Based on Data Mining Technologies in Urban Water Supply Analysis and Decision,F299.24;F224
- Research on Application of Data Mining Technology in Degree of Satisfaction Analysis of Television Customers,TP311.13
- Web Usage Mining and the Research of Personalized Recommendation,TP311.13
- Data Mining of Application in the School Management and Training Students,TP311.13
- Research on Employment Monitoring System of University Graduate,G647.38
- Design and Implementation for Decision Support System of Drug Administration Based on Data Warehouse,TP311.13
- A Research on the Credit Card Client Activating and Response Extent Based on Data Mining,F832.2
CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer software > Program design,software engineering > Programming > Database theory and systems
© 2012 www.DissertationTopic.Net Mobile
|