Dissertation > Excellent graduate degree dissertation topics show

The Research and Application of Cloud Computing Model Based on Hadoop

Author: CaoFengBing
Tutor: WuKaiGui
School: Chongqing University
Course: Computer System Architecture
Keywords: Cloud Computing Massive Data Hadoop MapReduce
CLC: TP3
Type: Master's thesis
Year: 2011
Downloads: 1042
Quote: 2
Read: Download Dissertation

Abstract


With the development of Internet technology and society, the amount and scales of network data are rapidly growing. No matter who enterprises or individuals face the problem, that is, how to store and analyze these massive data. With the growing amount of data, the traditional methods for processing massive data have many shortcomings, such as, higher cost of data storage, more difficult data management, less reliable, lower efficiency, more difficult parallel programming, and so on. A novel technology is needed to solve these problems, so the cloud computing has been brought naturally.Cloud computing integrates the free resource, provides strong storage and computes power through Internet. Hadoop is an open source distributed data processing framework. It is designed to efficiently handle massive data. Because of the features of scalability, high reliability, low cost and high efficiency, the Hadoop has become a popular development platform for cloud computing. But Hadoop has some defects, for example, the single NameNode manages the entire file system metadata and performs all the important file system operations in the Hadoop cluster, which limits the scalability and availability of the Hadoop cluster. In addition, in the traditional Hadoop, the computing and storage are tightly coupled, which also limits the performance of the cluster.The topic of this paper is the research and application of cloud computing model based on Hadoop. There are two main purposes: One is to analyze the performance bottlenecks of the tradition Hadoop framework, and to devise a new Hadoop framework; The other is, on the new Hadoop framework, to design a cloud computing model based on improved Hadoop, and to implement appropriate cloud computing platform, which focuses on solving the problem of the massive data storage and its analysis.The main works and achievements of this paper are as follows:①The basic theories of cloud computing are introduced, including the concept, characteristics and the key technology of cloud computing and some typical applications of the cloud computing platforms.②The Hadoop infrastructure is described, including the Hadoop distributed file system (HDFS) , the theory and implementation of the MapReduce parallel programming model, the optimization measures of the Hadoop framework and so on. Furthermore, the mechanisms of data management, performance guarantee and reading and writing files in the HDFS are also detailed and analyzed.③The performance bottlenecks of the Hadoop is discussed, and a new type Hadoop framework is designed to solve the two issues of the tradition Hadoop which are the bottleneck of single NameNode and the tight coupling of computing and storage. The theoretical analysis and experimental comparison demonstrate the efficiency of the new framework. The new framework has high availability and good scalability.④A cloud computing model based on improved Hadoop is designed and the main functional modules in this model are implemented by WEB technology.The newly proposed cloud computing model has many advantages, such as high efficiency, high reliability, low cost, and more convenient for users. This system model can be applied to enterprises for data storage and computation, scientific research and other fields. It provides a high reliability and low cost solution for the massive data storage and computing, and offers a novel idea for the development of cloud computing technology and the application of Hadoop.

Related Dissertations

  1. Research on Compression, Operation and Query Processing Methods of Massive Datasets,TP311.13
  2. The Research of Dynamic Trust Model on Cloud Computing Platform,TP309
  3. Cloud-based digital library service model,G250.76
  4. Research and Application of Map/Reduce Based Distributed Log Analyzer,TP311.52
  5. Design and Implementation of Online Shopping Prototype System Based on Hadoop,TP311.52
  6. Design and Implementation based the Google platform promotional modules commodity module,TP311.52
  7. The Operation and Maintenance of ITIL Based on Cloud Computing,TP311.52
  8. The Research of Software Service Platform Based on Cloud Computing,TP311.52
  9. An Intrusion Detection System for High-Speed Networks,TP393.08
  10. Research of Resource Provisioning Technique in Telecom Business Supporting Based on Cloud Computing Platform,TP3
  11. Hadoop data center deployment and tracking systems research,TP308
  12. Linux-based platform LAN Cloud Monitoring System Analysis and Implementation,TP311.52
  13. Logically oriented virtual domain multi -level access control system,TP309
  14. Fault Tolerance for MapReduce in the Cloud Environment,TP302.8
  15. Simulation Resource cloud storage technology research and implementation,TP333
  16. Cloud data storage technology research homeostasis,TP333
  17. Cloud-based computing and immune reconfigurable embedded systems research network,TP368.1
  18. Unbalanced data set classification method and its application in the telecommunications industry,TP311.13
  19. Cloud-based distributed real-time database performance data storage and retrieval mechanism,TP311.13
  20. Robot cloud platform Research and Implementation,TP242
  21. A scalable prototype design and implementation of MapReduce,TP311.52

CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology
© 2012 www.DissertationTopic.Net  Mobile