Skip to content
Preprint

ARASH: Adaptive Retrieval And Shot Selection for Tabular Prediction

Aug 2026 · 0 citations · 35 references
Computer Science

TL;DR

ARASH (Adaptive, query-specific Retrieval And Shot selection), a method that improves TFM efficiency by selecting optimal shots based on local neighborhood analysis within the training set by reducing the prompt length and memory usage of TabPFN.

Abstract

Tabular prediction is a critical task across numerous applications. The recent success of large language models has sparked various approaches for adapting them to the tabular domain. A prevalent strategy involves training or fine-tuning specialized Tabular Foundation Models (TFMs) such as TabPFN. However, TFMs require substantial computational resources, and frequent model retraining is often impractical. In-context learning (ICL), specifically, few-shot prompting, offers a resource-efficient alternative to enhance performance. Yet, identifying the most relevant rows to serve as shots remains a challenge for tabular data. This paper introduces ARASH (Adaptive, query-specific Retrieval And Shot selection), a method that improves TFM efficiency by selecting optimal shots based on local neighborhood analysis within the training set. Our results demonstrate that ARASH reduces the prompt length and memory usage of TabPFN by 1261.5$\times$ and 2.56$\times$, respectively, while providing comparable accuracy.

View source

Similar papers

Preprint Aug 2026

TabDPT-Turbo: Efficient In-Context Learning for Tabular Prediction

This work adopts an alternate approach, sticking with row-based attention while incorporating long context pre-training to eliminate the need for retrieval in TabDPT-Turbo, a model that provides comparable default performance to TabDPT v1.1 on TabArena-Lite, CC18, and CTR23, at orders of magnitude faster.

Rasa Hosseinzadeh, Alex Labach, Zexin Xue et al. · 4 citations
Preprint Aug 2026

Localized TabICLv2: Scaling Tabular In-Context Learning through k-NN

Localized TabICLv2 introduces a method that reduces the inference cost of TabICLv2 by retrieving only the k nearest training neighbours for each test point, measured by similarity in the model's Stage 2 row-representation space, rather than using the full training context.

Beimnet Bekele Guta · 0 citations
Preprint Aug 2026

Loss-Based Active Learning for Neural Abstractive Summarization

This work proposes LOBSTER, a novel active learning framework designed specifically for abstractive summarization that improves performance by prioritizing unlabeled instances semantically similar to the model's current high-loss training examples, enabling the model to explicitly correct its specific weaknesses.

M. Ioannou, Tatiana Passali, George Michalopoulos et al. · 0 citations
Review Sep 2026

One-Step Retrieval Framework for Real-Time Sponsored Search Ads Using Hierarchical Text Representations

Traditional retrieval systems typically use multi-stage cascading architectures (MCA), where each module is optimized independently, leading to inconsistent objectives and the premature elimination of high-potential candidates. Recent LLM-based generation methods offer end-to-end solutions but use discrete semantic ide...

Tong-Tong Liu, Ren-Yu Zhang, Jia-Yu Ding et al. · 0 citations
Jul 2026

The Case Against Generation for Retrieval: Discriminative Language Models as Effective Retrievers

Large Language Models (LLMs) have emerged as powerful assets for recommender systems. However, deploying them as generative recommenders or zero-shot rankers at web-scale remains bottlenecked by prohibitive computational overhead and grounding challenges. In this paper, we revitalize the classic, highly efficient two-t...

Zhe Xu, Prachi Agrawal, Kavosh Asadi et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.