H-TATC: Hierarchical Task-Aware Token Compression for Multimodal Semantic Communication
Abstract
Multimodal semantic communication reduces uplink overhead by transmitting compact task-relevant representations. However, existing token compression methods either aggregate tokens using fixed structures or select tokens under fixed modality-wise budgets, without jointly adapting modality relevance and token utility to the task query. We propose hierarchical task-aware token compression (H-TATC), which allocates a shared non-text token budget across modalities and subsequently selects query-relevant tokens within each modality. On MUSIC-AVQA, H-TATC improves accuracy by 2.4 percentage points over TokCom-MLMs at B=64 and SNR =12 dB, and reaches 72.1% at B=128, only 0.7 points below the full-token reference, with negligible additional end-to-end latency.