Fault-Tolerant Topology Reconfiguration Framework for NoC-Based Multiprocessor Arrays via Bipartite Matching and Reinforcement Learning
Abstract
As Network-on-Chip (NoC)-based multiprocessor arrays continue to scale, permanent faults in processing elements (PEs) increasingly jeopardize system reliability and performance. Topology reconfiguration has become a critical mechanism to tolerate such faults by restoring logical interconnects. However, existing heuristic-based methods often struggle to balance communication efficiency, load distribution, and reconfiguration overhead, especially under high fault densities. To address these limitations, this paper proposes KM-RL, a two-stage fault-tolerant reconfiguration framework that integrates bipartite matching and reinforcement learning to optimize PE allocation and migration. In the first stage, the allocation of redundant PEs to faulty ones is formulated as a weighted bipartite matching problem and solved optimally using the Kuhn-Munkres algorithm, followed by a layout refinement strategy to improve load balance. The second stage employs Q-learning to guide migration planning, reducing path lengths and alleviating network congestion. Extensive experiments demonstrate that KM-RL consistently outperforms state-of-the-art methods across a wide range of array sizes and fault densities. Static evaluations indicate that KM-RL yields topologies with reduced congestion and shorter path lengths, while dynamic simulations reveal superior communication performance, combined with a 99.8% reduction in reconfiguration time. This framework provides both theoretical and practical support for fault-tolerant communication in NoC-based multiprocessor arrays.