It is suggested that correlated prompts alter not only the effective sample size of in-context learning, but also which attention architectures are best matched to the task.
Abstract
Modern sequence models have a striking capacity for in-context learning (ICL); they can perform new tasks based only on examples given in the prompt. Understanding how this ability emerges requires theory that captures important properties of natural data. Linear regression has served as a useful sandbox for ICL theory, but existing work has largely focused on prompts with independent examples. In this work, we extend this setting to sequentially correlated data, a basic feature of real sequences. We present a solvable model based on linear attention and test our predictions on realistic transformer architectures. We identify two distinct effects: First, when the query token is independent of the context, within-context correlations induce an effective context length: correlated prompts behave like shorter i.i.d. prompts. Second, when the query is also correlated with its context, test error is reduced, particularly for softmax attention when compared to linear attention. These results suggest that correlated prompts alter not only the effective sample size of in-context learning, but also which attention architectures are best matched to the task.
Large language models can process increasingly long prompts, yet their ability to locate and use decisive evidence may degrade as irrelevant or confusable context is added. We formulate this phenomenon, which we call context poisoning, as extreme-value interference in attention: the decisive-evidence score is upper-bou...
Meysam Ghaffari, Nina Fatehi, Bhaskar Sen et al.· 0 citations
Long-context understanding requires large language models (LLMs) to reason over lengthy documents, conversations, and code, yet task-relevant evidence is often sparse and scattered amid substantial irrelevant and redundant content. We propose Highlight-Then-Summarize (H2S), a compress-then-reason paradigm that first id...
The Information Abundance Paradox is supported and it is suggested that scaling toward near-infinite context is not simply a matter of supplying more data, even when high-quality long-context data is abundant.
Arda Uzunouglu, Benjamin Van Durme, Daniel Khashabi· 2 citations
This survey argues that context injection strategy, rather than context capacity, is the defining research challenge for long-context LLM deployment, and proposes a three-axis analytical framework revealing that injection performance is jointly governed by selection, representation, and scheduling.
Aicha Dakir, M. El Hajji, Tarek Ait Baha et al.· EPJ Web of Conferences· 0 citations
TopoCompress is introduced, a training-free and model-agnostic framework that compresses long contexts by selecting coherent semantic spans by selecting coherent semantic spans and achieves performance comparable to the strongest baseline while using a 4x smaller compression budget.
Implicit multimodal in-context learning compresses demonstrations into internal interventions, ranging from static task vectors to query-conditioned transformations and attention routing. Despite their common goal, these methods differ substantially in how the intervention depends on the query and where it modifies the...
Jiaqian Li· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.