Skip to content

Representation Editing for Multimodal Test-Time Adaptation

Sep 2026 · 0 citations
Computer Science

TL;DR

This work proposes FourIer Representation Editor (FIRE), a novel multimodal TTA approach that directly edits semantically rich intermediate representations, and demonstrates the superiority of FIRE over existing multimodal TTA methods.

Abstract

Multimodal test-time adaptation (TTA) aims to adapt a pretrained multimodal model online to distribution shift across modalities using unlabeled test data, showing broad potential in real-world applications. However, existing methods primarily focus on adjusting fused features to bridge the source-target gap, lacking explicit control over intermediate representation misalignment, which is a key driver of performance drop under distribution shift. In this work, we tackle this challenge from the perspective of representation engineering. Unlike previous TTA methods that update fusion weights in place, we propose FourIer Representation Editor (FIRE), a novel multimodal TTA approach that directly edits semantically rich intermediate representations. Specifically, we first adopt representation editors into each intermediate layer of the unimodal encoders, enabling layer-wise calibration of unimodal representations. To further enhance the diversity and stability of the low-rank editing subspaces, each representation editor performs frequency domain mixing via the fast Fourier transform to construct structured bases. Moreover, we introduce multi-level adaptation objectives to optimize these editors, jointly promoting cross-modal semantic alignment, source-target statistical alignment, and asymmetric prediction consistency. In this way, FIRE yields aligned unimodal representations for fusion and further improves prediction reliability. Extensive experiments on two widely used multimodal benchmarks under various corruption types demonstrate the superiority of FIRE over existing multimodal TTA methods.

View source

Similar papers

Oct 2026

Exploring Multimodal Adapters with Multi-Task Decoders for Video Action Recognition.

Large-scale vision-language pretrained models like CLIP, paired with parameter-efficient fine-tuning techniques, have emerged as promising solutions for image-to-video transfer in video action recognition. However, existing methods often prioritize strong supervised performance at the expense of transferability and gen...

Meng-Meng Wang, Ze-Yi Huang, Bo-Yuan Jiang et al. · 0 citations
#machine learning Preprint Sep 2026

SPeaR: Test-Time Adaptation with Steering Primitives for Realigning Representations

Test-time adaptation (TTA) addresses distribution shift using only unlabeled test data. Existing methods typically adapt pretrained models by updating their parameters, limiting both what is adapted and where adaptation can occur within the network. We instead keep the pretrained network frozen and steer its intermedia...

Muhammad Sudipto Siam Dip, Ali Etemad · 0 citations
Open access Aug 2026

TCFNet: an end-to-end framework for multimodal action quality assessment via temporal enhancement and contrastive fusion

Existing Action Quality Assessment (AQA) methods have limitations, such as over-reliance on unimodal, insufficient long-term temporal modeling, and modality alignment biases in multimodal models. To address these issues, we propose TCFNet, an AQA approach from the perspective of multimodal fusion. Compared to previous...

Zhenxian Lin, Ming-Hui Zhang, Cheng-Mao Wu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

PACE: Progressive Angular-to-Norm Contrastive Embedding

Multimodal embedding models encode heterogeneous inputs into a shared embedding space, enabling efficient similarity computation across modalities and tasks. Most existing methods optimize cosine-based contrastive objectives, which promote stable training but restrict semantic compatibility to angular geometry, preclud...

Yan-Ping Li, Wei Zhou, Ya-Wen Liu et al. · 1 citation
Preprint Aug 2026

Towards Purified Multi-Label Test-Time Adaptation of Vision-Language Models

PuRF is introduced, a novel PuRiFication-driven cache-based method for multi-label test-time adaptation of vision-language models that consistently outperforms state-of-the-art methods on ViT-B/32 across five datasets.

Yiwen Liang, Hui Chen, Yizhe Xiong et al. · 0 citations
Preprint Aug 2026

Exploring the Design Space of Representation Learning for Audio Transformations

This framework produces both a transformation embedding and a processed-audio embedding, and it finds that the two play complementary roles: distance-based tasks favor the former, while probe-based tasks favor the latter.

Sungho Lee, Marco A. Martínez-Ramírez, Junghyun Koo et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.