Dissertation > Excellent graduate degree dissertation topics show
The Research of Performance Optimization of Hadoop in Big Data
Author: CaoYing
Tutor: LiYanZuo
School: Dalian Maritime University
Course: Computer Science and Technology
Keywords: Hadoop MapReduce Hadoop Speculate Execution DBInputFormat
CLC: TP311.13
Type: Master's thesis
Year: 2013
Downloads: 1030
Quote: 0
Read: Download Dissertation
Abstract
|
With the rapid development of the technology of Internet and Web, audio, video, Web logs, Internet search index and the text files of Internet etc,which have been widely used have brought the sharply increase in amount of data, it indicates that the era of big data is coming. In the era of big data the data has characteristics of the data amount increasing sharply and the more complicated data structure, the situation like that can lead to the more difficulty in data storage and processing. While the emergence of the Hadoop greatly simplifies the problem of data storage and processing in the era of big data, so the paper researching on Hadoop and its’ optimization has important practical significance.The main research in this this paper is following. First of all, the theory of HDFS and MapReduce that are core technology of Hadoop which is studied and analyzed. And its’ study in detail considers the following several aspects:NameNode, DataNode, interface, class, call relationship, and analyzes the working mechanism of HDFS and MapReduce. Meanwhile, this paper aims at two performance problems which the Hadoop exists so far and then it puts forward a improved program preliminarily based on the in-depth study on source code. Secondly, this paper studies and analyses the poor performance of the Hadoop speculate execution algorithm in heterogeneous environme-nts which is the first performance problem. A new improved algorithm for the problem is put forward. The new algorithm can adjust the execution of backup task automatically according to system load condition to make it balanced, and getting more precise stragglers using the way of putting the task that residual time value that is based on historical average completion time putting forward by Zaharia is greater than0.2in queue. The new algorithm to a certain extent improves the performance of speculate execution in the heterogeneous environment. Finally, one of the second performance problem of Hadoop is that it brings about the defects of the performance when DBInputFormat processes huge amounts of data in a relational database. To solve the problem, improving DBInputFormat interface, putting forward a new sharding strategy and building the improved interface are done. It improves the efficiency and performan- ce of Hadoop processing the relational database.Building experiment platform and experimenting the the proposed algorithm and the improved interface, they are verified that they can, to some extent, improve the performance of Hadoop.
|
Related Dissertations
- Research and Application of Map/Reduce Based Distributed Log Analyzer,TP311.52
- Design and Implementation of Online Shopping Prototype System Based on Hadoop,TP311.52
- Design of the Mobile Learning System Based on Hadoop,G434
- HADOOP architecture based on the social security project web log analysis system,TP311.52
- 3D Mannequins Generating Engine Based on eMTM with MapReduce,TP391.41
- The Research of Software Service Platform Based on Cloud Computing,TP311.52
- Research on Fast Queryalgorithm of Massive Data,TP311.13
- An Intrusion Detection System for High-Speed Networks,TP393.08
- Incremental Learning Method Based on Cloud Computing,TP311.13
- Cloudqueue: An Internet-Scale Messaging Infrastructure Based on Hadoop,TP311.52
- The Research of Text Classification Based on Hadoop,TP391.1
- Research and Implementation on a Distributed Service Registry Based on HADOOP Platform,TP393.09
- Resarch of Task-level Data Processing Based on Multicore CPU and Test of Its Performance on Cluster Platform,TP274
- Hadoop-based video transcoding system design and implementation,TN919.81
- Hadoop data center deployment and tracking systems research,TP308
- Fault Tolerance for MapReduce in the Cloud Environment,TP302.8
- Cloud-based mobile data storage backup system,TP309.3
- A scalable prototype design and implementation of MapReduce,TP311.52
- The Research of Distributed Text-based Data Filtering Technology and System Implementation Based on MapReduce,TP391.1
- Design and Implementation of Massive Web Log Analysis System Based on Hadoop/Hive,TP393.09
- Large-scale approximation paragraph fingerprint - based page detection algorithm research,TP393.092
CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer software > Program design,software engineering > Programming > Database theory and systems
© 2012 www.DissertationTopic.Net Mobile
|