An adaptive CNN–transformer feature fusion framework for multi-class PCB surface defect detection
Abstract
PCBs are required in the final products of electronic manufacturing, and quality inspection must be carried out. Uncorrected surface defects may persist through subsequent production stages and degrade device performance. Convolutional Neural Networks (CNNs) have been widely applied to PCB defect detection because of their excellent local-texture- capturing capabilities, but they are still relatively poor at modelling long-range spatial dependencies across the entire board. Transformer-based models address this problem by learning global context representations, but they are generally much more computationally expensive. To use the combined strengths of both architectures, a hybrid feature fusion framework for multi-class PCB defect detection is proposed in this paper. A lightweight CNN branch extracts fine-grained local features of the defect region, and a Transformer-style encoder obtains global context from the shared feature backbone. The above features are integrated via an attention-based fusion module with adaptive weighting and reduce redundant computations through a shared extraction path. Then, the fused representation is used by the classification and localization networks to jointly predict defect categories and locations. The mean Average Precision (mAP) of the proposed framework on the public PCB defect benchmark dataset reaches 96.847%. Experimental results show that the proposed framework outperforms Faster R-CNN, YOLOv5, RetinaNet, EfficientDet and the original Swin Transformer. It has almost the same accuracy as a pure Transformer-based system but at a lower computational cost; thus, it is more feasible in practice.