Skip to content

DAGR: State-Conditioned Goal Representations via Difference-Aware Goal Cross-Attention

Jul 2026 · arXiv.org · Vol abs/2607.13731 · 0 citations
Computer Science Mathematics

TL;DR

DAGR is proposed, which refines the static embedding of any late-fusion encoder into a state-conditioned one through multi-scale gated cross-attention, and traces the gain to the gated residual rather than to the difference bias that names the method.

Abstract

Goal-conditioned reinforcement learning hinges on how the goal is encoded. Contrastive, metric, temporal-distance and information-theoretic encoders disagree on the objective. They agree on one thing. None of them sees the current state, so the embedding cannot mark which part of the goal still needs action, and the policy must recover that cue by inverting both encoders. We propose DAGR, which refines the static embedding of any late-fusion encoder into a state-conditioned one through multi-scale gated cross-attention. A gated residual holds the refinement near the base, and a difference-aware attention rule biases the scores by a per-token state-goal mismatch. A single condition decides what such a refinement can guarantee, namely whether the block returns its input at closed gates. We prove that the usual post-norm placement violates it, measure the consequence on frozen checkpoints, and recover part of the resulting loss by restoring the condition. On OGBench DAGR improves navigation and matches or trails the base elsewhere. Our ablations trace the gain to the gated residual rather than to the difference bias that names the method. Code is available at https://github.com/leixingxing1/DAGR

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.