Skip to content
Preprint

A Unified Backbone--Expert Framework with Relation-Token and Residual--Classifier Interfaces for Automatic Modulation Recognition

Aug 2026 · 0 citations · 28 references
Computer Science

TL;DR

A unified backbone-expert framework with a common convolutional state-space backbone and two specialized interfaces for automatic modulation recognition, confirming the benefit of expert-interface decoupling over one-size-fits-all architectures.

Abstract

Automatic modulation recognition (AMR) faces distinct representation bottlenecks under varying observation lengths, where a single model architecture often fails to excel. To address this, we propose a unified backbone-expert framework with a common convolutional state-space backbone and two specialized interfaces. For short sequences, we inject explicit lag-aware complex-plane descriptors as relation tokens before encoding to compensate for information loss. For long sequences, we design a gated multi-scale residual refinement module to correct the feature map, combined with a fixed-averaging classifier collaboration to harness complementary evidence. Our framework achieves overall average accuracies of 67.28 \pm 0.14% on RML2016.10b and 87.19 \pm 0.77% on HisarMod2019 (mean \pm sample standard deviation over three runs), respectively. The framework's efficacy is further validated through three-seed ablations, native-length cross-configuration tests, and controlled window studies, confirming the benefit of expert-interface decoupling over one-size-fits-all architectures.

View source

Similar papers

Preprint Aug 2026

MLLM-Routed Heterogeneous Ensembles for Robust Cross-Dataset Image Classification

ARMDIL is an ensemble that uses a multimodal large language model (MLLM) agent to dynamically route each image to the most suitable vision backbone, drastically improves adaptability by allowing new information to be integrated via simple prompt modifications, while enhancing interpretability through natural language r...

Daniel A. Perkins, J. Squires, Janou Milligan et al. · 0 citations
Jul 2026

Weight-Space Mixture-of-Experts for Implicit Neural Representation Classification

This work proposes a hierarchical Mixture-of-Experts (HMoE) Transformer that processes INR weights using conditional computation aligned with the structure of the underlying implicit network and develops weight-space attribution and pruning methods that identify parameters most relevant for classification.

Stanislaw Janik, Michal Byra · 0 citations
#artificial intelligence Preprint Sep 2026

FLaG: Frequency-Domain Latent-attention Gated Pooling for Token Aggregation

Frequency-Domain Latent-attention Gated Pooling (FLaG), a plug-in aggregation module that re-expresses encoder outputs in the Fourier domain before final pooling, provides a transferable frequency-domain aggregation bias across protein, visual, and textual representations, with benefits that depend on the backbone and...

Ke-Wei Li, Rong Zhang, Xuelin Wang et al. · 0 citations
Preprint Aug 2026

Clustering and Token Denoising for Faster and More Robust VLMs

This study introduces ClustRS, a two-part, training-free algorithm for robust token pruning, demonstrating a simple yet powerful alternative to both score-only and diversity-only pruning rules, paving the way for compute-efficient and noise-resilient VLM deployment.

Baptiste Rossigneux, Inna Kucher, Vincent Lorrain et al. · 0 citations
Open access Aug 2026

Lightweight Redesign of Long-Used Operators in Vision Backbones for Efficient Visual Recognition

Recent vision backbones increasingly rely on sophisticated modules, whereas long-used operators such as residual connections, activations, and normalization layers remain less explored for lightweight redesign. This paper revisits these operators and proposes three operator-level redesigns: Subtractive Residual Connect...

Zhan-Yi Lian, Kepeng Luo, Yunfeng Wang · 0 citations
Preprint Aug 2026

Better, Stronger, Faster, and Broader: Structured All-Mask Prediction for MLLM-Based Segmentation

MLLM-based segmentation faces a core segmentation trilemma: high segmentation performance, preserved dialogue ability, and fast inference. Embedding-prediction methods may disrupt language modeling through pixel-level objectives, whereas next-token generation is inefficient for dense masks. We propose All-Mask Predicti...

Jiazhen Liu, Mingkuan Feng, Long Chen · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.