Skip to content
Preprint

Task-State Adaptation with Prototype Memory for Multi-Task Dense Prediction

Aug 2026 · 0 citations · 43 references
Computer Science

TL;DR

MemMTL, a multi-task dense prediction framework that estimates a compact task state from global visual context and refines it through a learnable task-state prototype memory, is proposed.

Abstract

Vision foundation backbones provide strong representations for dense prediction, yet a single shared feature still needs to support tasks with different, image-dependent adaptation requirements. We propose MemMTL, a multi-task dense prediction framework that estimates a compact task state from global visual context and refines it through a learnable task-state prototype memory. The refined state is converted into task-conditioned expert logits and combined with token-level logits before sparse top-$k$ selection over a local expert bank shared by all tasks. A separate task-agnostic residual bank provides a common adaptation path, and both paths are added once to the backbone feature before task-specific prediction. We specify a matched evaluation protocol on NYUD-v2 and PASCAL-Context with SAM 3 and ViT-L backbones to measure predictive quality, computational cost, and the contributions of task-state conditioning, prototype retrieval, and sparse routing. The numerical record in the present working draft predates this canonical implementation and must be regenerated before it can support empirical claims.

View source

Similar papers

Preprint Aug 2026

MoTE: Mixture of Task Experts for Multi-Task Video Understanding

Procedural video-language models must solve heterogeneous tasks from the same visual evidence, including action recognition, forecasting, and procedure prediction. Dense transformer decoders share the same feed-forward networks across tasks, which can entangle task behavior and make controlled capability expansion diff...

Muhammad Asad Ali, U. Khan, Nadia Robertini et al. · 0 citations
Preprint Aug 2026

Task-Anchored Representation Shaping for Pre-Trained Model-Based Continual Learning

TAILS resolves cross-task ambiguity at the representation level, while leaving the original PTM, method-specific modules, and classifier unchanged, and can improve classification and task-inference performance with modest parameter overhead and negligible inference cost.

Zhiming Xu, Huiyu Yi, Zheng-He Xie et al. · 0 citations
Open access Sep 2026

TAST: Task-Aware Sparse Topology Learning for Multi-Task Recommendation

Multi-task recommendation improves predictive performance by sharing knowledge across related tasks, but existing dense architectures and predefined expert-routing mechanisms do not explicitly adapt fine-grained parameter connectivity to individual tasks, potentially leading to parameter competition, negative transfer,...

Ming-Zhu Zhang, Zhen Yang, Rui-Ping Yin · 0 citations
Preprint Aug 2026

Hierarchical Skill Retrieval for Data-Efficient Adaptation of Vision-Language-Action Models

While Vision-Language-Action (VLA) models pretrained on large-scale robot datasets provide a strong foundation for robot manipulation, their performance can degrade when adapted to new tasks with limited task-specific demonstrations. Retrieval offers a practical way to reuse existing demonstrations for data-efficient a...

Haoran Hao, Shahram Najam Syed, Jeff G. Schneider et al. · 0 citations
#machine learning Preprint Sep 2026

StateMem: Single-State Residual Memory with Adaptive Inference for Vision-Language-Action Policies

Memory-dependent robotic manipulation often requires later actions to use information from earlier interactions. Existing vision-language-action (VLA) policies primarily rely on current observations, limiting historical information retention. Memory-augmented VLAs, such as MemoryVLA, address this limitation with extern...

Wen-Zhuo Li, Q. Shi, Yi Zhou · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.