Multimodal Biometric Authentication Using Fingerprint and Finger Vein Recognition: A Frozen Pre-Trained CNN Backbone Framework with Lightweight Classifiers
Abstract
Security systems that depend on a single biometric modality have well-known weaknesses, including vulnerability to presentation attacks and sensitivity to acquisition conditions. This work proposes a multimodal biometric framework that combines fingerprint and finger vein recognition for identity authentication. The distinguishing feature of the proposed approach is that no convolutional backbone is trained or fine-tuned at any stage: three ImageNet pre-trained architectures- VGG16, ResNet50 and EfficientNet-B0 -are used strictly as frozen feature extractors, and the only fitted components are two lightweight, low-capacity classifiers and two scalar attention weights. After extraction, Global Average Pooling (GAP) produces compact per-backbone descriptors that are L2-normalised and concatenated into a single 3,840-dimensional joint descriptor; Principal Component Analysis (PCA) then reduces this descriptor to 206 components for fingerprints and 283 components for finger veins while retaining 95 % of the variance in each case. An attention module with channel, spatial and dynamic components is integrated to emphasise the most discriminative regions of the biometric image. For matching, fingerprint recognition uses a Random Forest classifier and finger vein recognition uses a K-Nearest Neighbour classifier; in both cases the accept/reject decision is taken on a calibrated cosine similarity score. The two modalities are combined through a conservative decision-level gate in which both must independently exceed a common calibrated threshold of 0.74 before access is granted, while a weighted score (0.6 fingerprint, 0.4 finger vein) is retained only as a logged confidence value. The system was evaluated on two public datasets - SOCOFing for fingerprints and SDUMLA-HMT for finger veins - linked by an explicitly documented virtual pairing protocol. Results show 98.90 % accuracy for fingerprint recognition and 98.30 % for finger vein recognition. After fusion the overall accuracy reaches 98.88 % with an Equal Error Rate (EER) of 0.03 %. Ten-fold subject-disjoint cross-validation gives 99.64 % ± 0.43 % mean accuracy (95 % confidence interval 99.33 %–99.95 %). The system requires 156.2 ms per authentication request and about 24 % of CPU on average on a machine with no GPU, which confirms that it can operate in environments with limited computational resources.