A delay scheduling algorithm based on locality resource forecasting(LRFD)is proposed to address the unreasonable waiting problem generalized by the static time-wait threshold in delay scheduling algorithm of YARN platform for short jobs. The algorithm takes both the characteristics of short jobs and dynamic resource availability into consideration to assign tasks. It estimates local resources to make reasonable waiting according to both the task progress on the nodes and the unhandled splits distribution in the cluster of the job. Experimental results and comparison with the traditional delay scheduling algorithm show that LRFD gets a better stability and improves the performance about 10% for short jobs on average and achieves a maximum speedup up to three times.
关键词
Keywords
references
DEAN J, GHEMAWAT S. MapReduce: simplified data processing on large clusters [J]. Communications of the ACM, 2008, 51(1): 107-113.
CHEN Y, ALSPAUGH S, KATZ R. Interactive analytical processing in big data systems: a cross-industry study of MapReduce workloads [J]. Proceedings of the Very Large Data Base Endowmat, 2012, 5(12): 1802-1813.
ISARD M, PRABHAKARAN V, CURREY J, et al. Quincy: fair scheduling for distributed computing clusters [C] ∥Proceedings of the ACM SIGOPS 22nd Symposium on Operating Systems Principles. New York, NY, USA: ACM, 2009: 261-276.
REN Z, XU X, WAN J, et al. Workload analysis, implications, and optimization on a production hadoop cluster: a case study on Taobao [J]. IEEE Transactions on Services Computing, 2014, 7(2): 307-321.
ZAHARIA M, CHOWDHURY M, FRANKLIN M J, et al. Spark: cluster computing with working sets [C]∥Proceedings of the 2nd USENIX Conference on Hot Topics in Cloud Computing. Berkeley, CA, USA: USENIX Association, 2010: 10.
HENDERSON R L. Job scheduling under the portable batch system [C]∥Proceedings of the Workshop on Job Scheduling Strategies for Parallel Processing. Berlin, Germany,: Springer-Verlag, 1995: 279-294.
FREY J, TANNENBAUM T, LIVNY M, et al. Condor-G: a computation management agent for multi-institutional grids [J]. Cluster Computing, 2002, 5(3): 237-246.
HINDMAN B, KONWINSKI A, ZAHARIA M, et al. Mesos: a platform for fine-grained resource sharing in the data center [C]∥Proceedings of the 8th USENIX Conference on Networked Systems Design and Implementation. Berkeley, CA, USA: USENIX Association, 2011: 295-308.
VAVILAPALLI V K, MURTHY A C, DOUGLAS C, et al. Apache hadoop yarn: yet another resource negotiator [C]∥Proceedings of the 4th Annual Symposium on Cloud Computing. New York, NY, USA: ACM, 2013: 1-16.
ZAHARIA M, BORTHAKUR D, SARMA J S, et al. Job scheduling for multi-user Mapreduce clusters[R]. Berkeley, California, USA: EECS Department, University of California, Berkeley, 2009: 1-16.
ZAHARIA M, BORTHAKUR D, SEN SARMA J, et al. Delay scheduling: a simple technique for achieving locality and fairness in cluster scheduling [C]∥Proceedings of the 5th European Conference on Computer Systems, New York, NY, USA: ACM, 2010: 265-278.
ELMELEEGY K. Piranha: optimizing short jobs in hadoop [J]. Proceedings of the Very Large Data Base Endowment, 2013, 6(11): 985-996.
ZAHARIA M, KONWINSKI A, JOSEPH A D, et al. Improving MapReduce performance in heterogeneous environments [C]∥Proceedings of the 8th USENIX Symposium on Operating Systems Design and Implementation. Berkeley, CA, USA: USENIX Association, 2008: 29-42.