Improved Job Scheduling Algorithm with Fault Tolerance in Grid Computing
Abstract
Job scheduling in grid computing remains challenging due to heterogeneous resources, dynamic workloads, and frequent failures. Traditional algorithms such as First-Come-First-Served and Round Robin lack adaptive mechanisms for reliability in failure-prone environments. This study develops an improved job scheduling algorithm that incorporates fault tolerance to enhance reliability, efficiency, and task completion. Objectives were to design a scheduler integrating checkpointing, selective replication, and adaptive rescheduling; implement it using Python (Flask); and evaluate its performance against conventional methods. A Design Science Research methodology was adopted, involving iterative development, simulation, and comparative testing. Results show the improved algorithm achieved higher completion rates, reduced recovery times, and better scalability than baseline schedulers, confirming that the proposed solution addresses the limitations of existing approaches and offers a practical framework for resilient distributed computing.