A unified backbone-expert framework with a common convolutional state-space backbone and two specialized interfaces for automatic modulation recognition, confirming the benefit of expert-interface decoupling over one-size-fits-all architectures.
Abstract
Automatic modulation recognition (AMR) faces distinct representation bottlenecks under varying observation lengths, where a single model architecture often fails to excel. To address this, we propose a unified backbone-expert framework with a common convolutional state-space backbone and two specialized interfaces. For short sequences, we inject explicit lag-aware complex-plane descriptors as relation tokens before encoding to compensate for information loss. For long sequences, we design a gated multi-scale residual refinement module to correct the feature map, combined with a fixed-averaging classifier collaboration to harness complementary evidence. Our framework achieves overall average accuracies of 67.28 \pm 0.14% on RML2016.10b and 87.19 \pm 0.77% on HisarMod2019 (mean \pm sample standard deviation over three runs), respectively. The framework's efficacy is further validated through three-seed ablations, native-length cross-configuration tests, and controlled window studies, confirming the benefit of expert-interface decoupling over one-size-fits-all architectures.
ARMDIL is an ensemble that uses a multimodal large language model (MLLM) agent to dynamically route each image to the most suitable vision backbone, drastically improves adaptability by allowing new information to be integrated via simple prompt modifications, while enhancing interpretability through natural language r...
Daniel A. Perkins, J. Squires, Janou Milligan et al.· 0 citations
This work proposes a hierarchical Mixture-of-Experts (HMoE) Transformer that processes INR weights using conditional computation aligned with the structure of the underlying implicit network and develops weight-space attribution and pruning methods that identify parameters most relevant for classification.
Stanislaw Janik, Michal Byra· arXiv.org· 0 citations
Frequency-Domain Latent-attention Gated Pooling (FLaG), a plug-in aggregation module that re-expresses encoder outputs in the Fourier domain before final pooling, provides a transferable frequency-domain aggregation bias across protein, visual, and textual representations, with benefits that depend on the backbone and...
Ke-Wei Li, Rong Zhang, Xuelin Wang et al.· 0 citations
This study introduces ClustRS, a two-part, training-free algorithm for robust token pruning, demonstrating a simple yet powerful alternative to both score-only and diversity-only pruning rules, paving the way for compute-efficient and noise-resilient VLM deployment.
Baptiste Rossigneux, Inna Kucher, Vincent Lorrain et al.· 0 citations
Recent vision backbones increasingly rely on sophisticated modules, whereas long-used operators such as residual connections, activations, and normalization layers remain less explored for lightweight redesign. This paper revisits these operators and proposes three operator-level redesigns: Subtractive Residual Connect...
MLLM-based segmentation faces a core segmentation trilemma: high segmentation performance, preserved dialogue ability, and fast inference. Embedding-prediction methods may disrupt language modeling through pixel-level objectives, whereas next-token generation is inefficient for dense masks. We propose All-Mask Predicti...
Jiazhen Liu, Mingkuan Feng, Long Chen· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.