Dissertation > Excellent graduate degree dissertation topics show
Research and Implement of Job Scheduling Method for Multi_User MapReduce Clusters
Author: WangKai
Tutor: WuQuanYuan;YangShuQiang
School: National University of Defense Science and Technology
Course: Computer Science and Technology
Keywords: Distributed Computing MapReduce Hadoop Job scheduling Waiting to be scheduled Multi-user sharing
CLC: TP311.13
Type: Master's thesis
Year: 2010
Downloads: 350
Quote: 2
Read: Download Dissertation
Abstract
|
PB -level data sets need to deal with the current data -intensive computing and GB - level data stream , facing large-scale data management , and management of complex computing environments , scalable computing platform and other aspects of the problem . Hadoop is an easily scalable distributed computing architecture that the cheap PC node can be combined to provide computing services , the MapReduce framework to provide users with easy-to-use programming model for large-scale data parallel processing . In this paper, on the basis of analysis of the existing Hadoop cluster job scheduling method , the problem of poor data locality of the existing MapReduce cluster multi-user job scheduling method conducted in-depth research . And local issues not support tasks for the Hadoop existing scheduling algorithm data , proposed a scheduling method based on the time of waiting for the priority task scheduler to run on the node where required data in order to achieve the purpose of data locality , to effectively reduce the overhead of the calculation process of the IO , improve system throughput and to reduce the average response time for a single job . To verify the effectiveness of the method we proposed job scheduling gives the design and implementation, and experimental verification . The results show that , waiting to be scheduled based on the time the guarantee based on fair multi-user shared cluster node data locality has been greatly improved , effectively increase the throughput of the cluster system , effectively reduce the average response time of a single job .
|
Related Dissertations
- Research and Application of Map/Reduce Based Distributed Log Analyzer,TP311.52
- Design and Implementation of Online Shopping Prototype System Based on Hadoop,TP311.52
- Research and Implementation of job scheduling algorithm in the Linux cluster environment,TP301.6
- Research and Implementation of Key Techniques of Building Distributed System,TP338.8
- An Intrusion Detection System for High-Speed Networks,TP393.08
- Hadoop data center deployment and tracking systems research,TP308
- Linux-based platform LAN Cloud Monitoring System Analysis and Implementation,TP311.52
- Fault Tolerance for MapReduce in the Cloud Environment,TP302.8
- New Container Terminal mixed loading optimization scheduling strategy,U691.3
- A scalable prototype design and implementation of MapReduce,TP311.52
- The Research of Distributed Text-based Data Filtering Technology and System Implementation Based on MapReduce,TP391.1
- Large-scale approximation paragraph fingerprint - based page detection algorithm research,TP393.092
- IaaS cloud computing - based Web application technology research,TP393.09
- Study on Hadoop-based Inverted Index,TP391.3
- Performance Optimization and Applications of MapReduce in Cloud Computing,TP3
- Research and Application of Remote Sensing Image Subblock and Classification,TP751
- Research of Semantic Web Services’ Search and Combination Model Based on P2P Technology,TP393.09
- Research on the Key Technology in Distributed Computing System,TP393.09
- Research on Unified Access Plantform for Unstructured Data and Index Technology,TP311.52
- Research on Parallelization of Clustering Algorithm Based on MapReduce,TP311.13
CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer software > Program design,software engineering > Programming > Database theory and systems
© 2012 www.DissertationTopic.Net Mobile
|