Dissertation > Excellent graduate degree dissertation topics show

Top-k Query Technology of Massive Uncertain Data in Cloud Environments

Author: LuZuo
Tutor: ChenHuaHui
School: Ningbo University
Course: Applied Computer Technology
Keywords: Topk Uncertain Data Cloud Environment MapReduce
CLC: TP311.13
Type: Master's thesis
Year: 2013
Downloads: 2
Quote: 0
Read: Download Dissertation

Abstract


With the development of information technology, the amount of data which could be obtained from the network explosively increased. What people faced with the challenge is not the lack of big enough information, but is how to find the valuable information which we need. To solve the problem, the Top-k query showed a great vitality. The Top-k query is a very important technology in the application of the large data interaction. According to the sorting with the user’s query conditions, the Top-k query result is the set of tuples ranked in the top k. At the same time, the data offen has a lot of noise, missing values, inconsistent factors, etc; the uncertainty is prevalent among the mass data. The Top-k query on uncertain data will be more complex than the traditional Top-k query on certain data, no matter from the query semantics or the query algorithm. The Top-k query on uncertain data gradually attracted the scholars’ attention.Since the concept of cloud computing introduced by Google, it has been strongly supported and developed by the academic and business communities. The design concept of cloud computing is allowing dynamic allocation of computing power, network resources, storage resources, on-demand services. Able to provide powerful computing and storage services, cloud computing can deal with the massive information at a relatively low cost, and thus get the favor of many IT companies.As cloud computing has a strong capabilities of processing the mass data, the Top-k query techniques will significantly improve the efficiency by using some technologies in cloud computing. The main work is as follows:1. To deal with the datasets which tuple is "tuple-level" uncertainty, we analyzed the Top-k query semantics base on the parameterized ranking functions and designed an algorithm to compute the upper bound of the tuple’s parameterized ranking function, which tuple has not been retrieved. In that way, we could avoid computing all the tuples’ value of ranking function in the dataset, and solve the problem of pruning in the Top-k query. As the experiments show, our algorithm is more effective to deal with the Top-k queries for the massive uncertain data on running time.2. In view of the uncertain dataset, we proposed a query semantic of Top-k frequent items and presented a query algorithm based on the generating function. At the same time, three pruning rules were proposed to filter out the items which can’t be the Top-k frequent items. 3. We built a cloud computing environment. In this cloud environment, we designed two algorithms based on the MapReduce programming model to achieve the distributed parallel computing of Top-k queries. As the experiments show, our algorithms are more effective to deal with the Top-k queries for the massive uncertain data on running time.

Related Dissertations

  1. 3D Mannequins Generating Engine Based on eMTM with MapReduce,TP391.41
  2. An Intrusion Detection System for High-Speed Networks,TP393.08
  3. The Research of Text Classification Based on Hadoop,TP391.1
  4. Resarch of Task-level Data Processing Based on Multicore CPU and Test of Its Performance on Cluster Platform,TP274
  5. Fault Tolerance for MapReduce in the Cloud Environment,TP302.8
  6. The Optimization of High Performance MapReduce FairScheduler and the Implementation on Simulator of Huge Scale Cluster,TP311.13
  7. Research and Application of column storage management technology based on RFID data,TP315
  8. Analysis and Optimization of Massive Data Processing on High Performance Computing Architecture,TP338
  9. Information Flow Control Model in Distributed Systems,TP316.4
  10. IaaS cloud computing - based Web application technology research,TP393.09
  11. Research and Implement of Job Scheduling Method for Multi_User MapReduce Clusters,TP311.13
  12. A Study on Word Document Decryption Using Time-memory Trade-Off Algorithm,TP391.12
  13. Research of Adaptive Local Hyperplane K-nearest Neighbor Classification Algorithm Based on Hadoop Platform,TP301.6
  14. Associative Classifier for Uncertain Data,TP311.13
  15. Research on Optimal Scheduling of Jobs Approach in Cloud,TP3
  16. Reseach of Hadoop Cluster Based on Eucalyptus Cloud Platform,TP338.8
  17. Research on the Processing of Uncertain Data on Wireless Sensor Network,TN929.5
  18. The Performance Optimization and Improvement of MapReduce in Hadoop,TP338.8
  19. The Research on Task Scheduling Model in Data-intensive Super Computing,TP311.52
  20. Research and Implementation of Distributed Web Crawl Based on Hadoop Architecture,TP391.3
  21. Hadoop MapReduce model to achieve the optimal performance analysis and improvement,TP338.6

CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer software > Program design,software engineering > Programming > Database theory and systems
© 2012 www.DissertationTopic.Net  Mobile