Based on the restless multi-arm bandit(RMAB)approach
a multiplexing strategy of backup virtual machines(VMs)is proposed to resolve the problem of low utilization of backup VMs in the cloud environment
and the optimal condition is given. This strategy regards an individual backup VM as a Markov process with two states
namely “idle”(1)and “backup”(0)
and models the scheduling of multiple backup VMs as a Markov decision problem(MDP)consisting of multiple Markov processes. The goal of this strategy is to maximize the utilization of backup VMs without obvious reduction in the system availability under the constraint of limited backup VMs. However
this problem is computationally intractable with traditional dynamic programming methods due to the curse of dimensionality. Therefore
this paper transforms the original MDP problem to a RMAB problem and adopts a simple single-step heuristic policy to resolve it. By calculating the single-step optimal solution
the long-term optimal solution can be obtained. Under specific conditions
the optimal solution of this strategy is guaranteed. The results of simulation experiments show that the proposed policy can achieve the goal of extending the backup ratio between backup VMs to service VMs from 1:1 to 1:M(M 1)while the failed VM assurance rate is no lower than 96%. Correspondingly
the utilization of backup resources is significantly enhanced. When the failure rate of service VM is low
the utilization of backup resources can be raised 89% compared with the 1:1 backup. The building and operation costs of a cloud platform can be reduced with the help of this backup VM scheduling strategy.
关键词
Keywords
references
JAYASINGHE D, PU C. Improving performance and availability of services hosted on IaaS clouds with structural constraint-aware virtual machine placement [C]∥Proceedings of the 8th IEEE International Conference on Services Computing. Piscataway, NJ, USA: IEEE, 2011: 72-79.
VERMA A, DASGUPTA G, TAPAN K, et al. Server workload analysis for power minimization using consolidation [C]∥Proceedings of the 2009 Conference on USENIX Annual Technical Conference. Berkeley, CA, USA: USENIX, 2009: 28.
BIN E, BIRAN O. Guaranteeing high availability goals for virtual machine placement [C]∥Proceedings of the 31st International Conference on Distributed Computing Systems. Piscataway, NJ, USA: IEEE, 2011: 700-709.
WHITTLE P. Restless bandits: activity allocation in a changing world [J]. Journal of Applied Probability, 1988, 25(2): 287-298.
SCHROEDER B, GIBSON G A. A large-scale study of failures in high-performance computing systems [J]. IEEE Transactions on Dependable and Secure Computing, 2010, 7(4): 337-350.
OKAMURA H, TADASHI D, SHUNJI O. Software reliability growth models with normal failure time distributions [J]. Reliability Engineering System Safety, 2013, 11(6): 135-141.
CULLY B, LEFEBVRE G. Remus: high availability via asynchronous virtual machine replication [C]∥Proceedings of the 5th USENIX Symposium on Networked Systems Design and Implementation. Berkeley, CA, USA: USENIX, 2008: 161-174.
SINGH R, IRWIN D, SHENOY P, et al. Yank: enabling green data centers to pull the plug [C]∥Proceedings of the 10th USENIX Symposium on Networked Systems Design and Implementation. Berkeley, CA, USA: USENIX, 2013: 143-155.
WANG D, GOVINDAN S, ANAND S, et al. Under provisioning backup power infrastructure for datacenters [C]∥Proceedings of the 19th International Conference on Architectural Support for Programming Languages and Operating Systems. New York, USA: ACM, 2013: 177-192.
AHMAD S, HAJI A, LIU M. Multi-channel opportunistic access: a case of restless bandits with multiple plays [C]∥Proceedings of the 47th Annual Allerton Conference on Communication, Control and Computing. Piscataway, NJ, USA: IEEE, 2009: 1361-1368.