Dissertation > Excellent graduate degree dissertation topics show

Research on Algorithm of Mining Frequent Itemsets from Uncertaln Data

Author: WangJinMiao
Tutor: ZhangLongBo
School: Shandong University of Technology
Course: Applied Computer Technology
Keywords: uncertain data frequent itemsets UF-growth algorithm maximal frequentitemsets
CLC: TP311.13
Type: Master's thesis
Year: 2012
Downloads: 118
Quote: 1
Read: Download Dissertation

Abstract


Uncertain data is widespread in some application fields such as sensor network, RFID (radio frequent identification), Web applications and so on. The research on uncertain data mining has been become a new hotspot in the area of data mining recently. Uncertain data mining includes clustering, classification, frequent itemsets mining and outlier detection, etc, among which the frequent itemsets mining is one of the focus issues.This paper elaborates the causes of uncertainty and the data types of uncertain data, after that the research status of uncertain data mining and the classical algorithms of mining frequent itemsets from certain data (traditional data) is summarized. And then some new algorithms used to mine frequent itemsets from uncertain data (such as U-Apriori and UF-growth) and uncertain data stream (such as UF-streaming and SUF-growth) is mainly discussed. The U-Apriori and UF-growth are the expansion and improvement of the classical algorithm Apriori and FP-growth respectively, while the UF-streaming and SUF-growth are all based on the tree structure. All of them are efficient algorithms in the area of mining frequent itemsets from uncertain data. But through several analysis we found that the research on mining frequent itemsets from uncertain data mostly concentrated on the complete frequent itemsets, and there is few algorithm used to mine maximal or closet ones.A new algorithm UMF-growth used to mine maximal frequent itemsets from uncertain data is proposed in the paper. The main idea of the algorithm is detailed illustrated with an example. The UMF-growth is based on the UF-growth algorithm and only need to scan the original database twice to complete the mining of maximal frequent itemsets. Different with the UF-growth algorithm, the mining process of the UMF-growth is divided into two steps:The first step is to find out all of the local maximal frequent itemsets with the frequent1-item as suffixes, respectively. The second step is to insert all of the local maximal frequent itemsets into the UMF-Tree in a similar fashion as the construction of FP-Tree. And then we will capture all of the maximal frequent itemsets of the original database. In order to further improve the efficiency of our algorithm, an improvement strategy, namely to discretize the expected support of the UF-Tree node up to n decimal places is proposed. The experimental results show that the performance of UMF-growth is very good and especially suitable for the dense database, and the improvement strategy could effectively improve the efficiency.

Related Dissertations

  1. Study on Skyline Query Processing Techniques on Uncertain Data,TP311.13
  2. Research of Closed Frequent Itemsets Mining Algorithm in Data Steams,TP311.13
  3. Study on Discovering the Relationships among Data Resources in DataSpace,TP311.13
  4. Modeling and Storing the Uncertain Data in the k-anonymity Privacy Protection Model,TP309
  5. Research on Uncertain Time Series Similarity Matching,TP311.13
  6. Skyline query uncertain data probability algorithm,TP311.13
  7. Association rules algorithm and its pharmacy system in the intelligent application of,TP311.13
  8. Research on Urology Clinical Decision Support System Based on Data Mining,TP311.13
  9. Research on Continuous Probabilistic Skyline Queries for Moving Objects,TP311.13
  10. Weighted Association Rules Mining Algorithm Based on Matrix,TP311.13
  11. Efficient Frequent Item Set Discovery Methods and Improved Apriori,TP311.13
  12. Mining Probability Frequent Patterns to Recover Uncertain RFID Data Stream,TP391.44
  13. Application and research association rule mining in the tax system,TP311.13
  14. Association rules based on coronary heart disease database data mining system design and implementation,TP311.13
  15. The Appliction of Data Mining in Colliery Safety Monitoring,TP311.13
  16. Research of Skyline Query Processing Technology on Uncertain Data Stream,TP311.13
  17. Research on Frequent Itemsets Algorithm Based on Projection Array and FP-tree,TP311.13
  18. Improved Linear Regression Forecast Algorithm Based on Association Rules,TP311.13
  19. Multi-label Text Classification Algorithm Research,TP391.1
  20. Study on Processing Uncertain Data in Deep Web,TP311.13
  21. Frequent Itemsets Mining Algorithm,TP311.13

CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer software > Program design,software engineering > Programming > Database theory and systems
© 2012 www.DissertationTopic.Net  Mobile