Skip to content

Cascaded Multi-Head Attention Transformer Framework for Direction of Arrival Estimation

2026 · IEEE Transactions on Cognitive Communications and Networking · Vol 12, pp. 10450-10465 · 0 citations · 49 references
Computer Science

Abstract

Deep neural networks have demonstrated significant potential in direction of arrival (DOA) estimation. However, some existing architectures, especially convolution-based ones, mainly emphasize local feature extraction and may not sufficiently capture long-range dependencies in array observations. To better model such nonlocal correlations, this paper presents a Cascaded Multi-Head Attention Transformer (CMA-Former) for grid-based DOA estimation. The sample covariance matrix is first converted into a compact token sequence using its upper-triangular off-diagonal entries. For each entry, the real part, the imaginary part, and the sine and cosine of its phase are stacked as input features, providing a periodic phase encoding that avoids the discontinuity inherent in raw phase values. A stack of customized Transformer encoders, each equipped with cascaded multi-head attention modules whose head count increases progressively, is then employed to capture sensor-pair correlations across multiple representation subspaces and scales. Finally, a classification token together with a classification head produces confidence scores over a predefined angular grid. Simulation results show that CMA-Former achieves an RMSE lower than or comparable to that of the deep-learning baselines considered. Moreover, it attains a higher estimation success rate for closely spaced sources, indicating an improved capability to resolve adjacent targets. At high SNR, the performance of all grid-based methods is bounded by the off-grid error floor imposed by the fixed angular grid. In addition, a hardware experiment using a cascaded mmWave radar platform further demonstrates the feasibility of applying CMA-Former to real radar measurements without retraining. The source code is publicly available at https://github.com/Syyyt/CMA-Former-official

View source

Similar papers

Open access Aug 2026

Two-Dimensional DOA Estimation Based on Dual-Branch CNN

A dual-branch convolutional neural network (CNN) for 2-D DOA estimation based on uniform rectangular arrays that obtains smaller root mean square errors than methods with multiple signal classification (MUSIC), estimation of signal parameters via rotational invariance techniques and ordinary CNN methods, with milliseco...

Fangyu Liu, Gui-Mei Zheng, Yu-Wei Song et al. · 0 citations
Open access Aug 2026

CoDAT: Collaborative Dual-Attention Transformer with Low-Cost Temporal Modeling for Efficient Edge Action Recognition

CoDAT is proposed, a Collaborative Dual-Attention Transformer that replaces conventional multi-head attention with a lightweight dual-branch module: Spatial Convolutional Attention (SCA) for local aggregation and Strided Single-Head Attention (SSHA) for global context.

Novendra Setyawan, Chi-Chia Sun, Mao-Hsiu Hsu et al. · 0 citations
Preprint Aug 2026

Neural Array-Generic Direction-of-Arrival Estimation Exploiting Array Transfer Functions

Direction-of-arrival (DoA) estimation is a key component of multichannel audio processing, yet many deep learning approaches remain tied to the microphone arrays used during training and generalize poorly to unseen devices. This paper proposes an array-generic neural DoA estimation framework using measured or simulated...

Mikko Heikkinen, A. Politis, K. Drossos et al. · 1 citation
Open access Aug 2026

DOA Estimation using Multi-head Self-Attention with Relative Positional Encoding and Hybrid Multi-Objective Grey Wolf Optimization

The performance of the suggested method for DOA estimation was superior to that of conventional existing techniques of MUSIC, SVR, and ESPRIT, establishing the efficacy of the suggested method for DOA estimation in realistic scenarios.

Zainab H. Mohammad, S. Hamad, Safaa K. Hussaine et al. · 1 citation
Open access Aug 2026

Deepfake Detection via Frequency-Aware Vision Transformer and Bidirectional Cross-Attention Fusion with Post-Processing Robustness

FAViT (Frequency-Aware Vision Transformer), a hybrid architecture capable of jointly utilizing spatial- and frequency-domain forensic information by the means of a bidirectional cross-attention fusion scheme, is presented.

Wasin Alkishri, Shahid Kamal, Jabar H. Yousif · 0 citations
Open access 2026

Vision Mamba With Joint Spatiotemporal Features for Efficient Video Representation Learning in Self-Supervised Scheme

This study introduces Vision Mamba (ViM), leveraging a Selective State Space Model to capture long-range temporal dependencies with linear computational complexity, validating the ViM as a highly efficient solution that balances computational feasibility with good performance in detecting complex criminal activities at...

Rahman Indra Kesuma, M. L. Khodra, B. R. Trilaksono · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.