Skip to content
Preprint

When Do Corrective Features Help? An Agent for Corrective Feature Discovery on Black-Box Forecasters

Aug 2026 · 0 citations · 37 references
Computer Science

TL;DR

This work presents CRAFTER (Corrective Residual Agent with Feature-based Temporal Exploration and Reasoning), which keeps the backbone frozen and mines its residual with two complementary generators: a compositional search over the raw input channels, and a large language model that proposes named feature combinations, binary flags, and short executable code.

Abstract

Frozen pretrained forecasters often fail in structured, recurring ways that are costly to repair through fine-tuning. We study corrective feature discovery: mining interpretable features of a frozen forecaster's residual to drive a lightweight post-hoc corrector. Prior automated feature engineering models the data-generating process; corrective features instead model the model-failure process. We present CRAFTER (Corrective Residual Agent with Feature-based Temporal Exploration and Reasoning), which keeps the backbone frozen and mines its residual with two complementary generators: a compositional search over the raw input channels, and a large language model (LLM) that proposes named feature combinations, binary flags, and short executable code. A single validation-grounded gate accepts or rejects every candidate regardless of its origin, and a validation-selected corrector applies the accepted features or leaves the forecast unchanged. This source-agnostic pipeline also allows prior feature-engineering systems to be evaluated under identical conditions, making CRAFTER an instrument for attributing forecast improvements to the feature source alone. Across six public datasets and six frozen backbones, CRAFTER surpasses every dedicated feature-engineering system at every feature budget, roughly doubling the improvement achieved by the corrector alone and reducing the error of the weakest backbones by up to 27%. These gains are robust across different LLM backends and persist even when applied on top of fine-tuned backbones.

View source

Similar papers

Open access Sep 2026

MARSAD: Multi-Agent Requirements System for Anomaly Detection in Software Engineering

Software requirements anomalies significantly impact system quality and development costs, yet existing detection approaches often focus on single anomaly types or lack automated correction capabilities. This paper presents MARSAD, a novel multi-agent requirements system that leverages Large Language Models (LLMs) to s...

Khaled Hamdan, Mohammed Mediani, Abdelhadi Hireche · 0 citations
Preprint Aug 2026

Where World Models Break: Natural-Input Failure Discovery

World models predict action-conditioned futures and serve as critical internal simulators for downstream planning and control. However, catastrophic prediction failures of world models could dangerously propagate through the control pipeline, as subsequent agent or model training and decision-making depend heavily on t...

Zhanpeng Shi, Zi Liang, Rong Feng et al. · 0 citations
Review Aug 2026

TRACE: TRajectory Attribution for Automated Context Engineering

This work presents TRACE (TRajectory Attribution for Automated Context Engineering), an automated feedback loop that mines historical agent trajectories to diagnose and remediate context failures, showing that over 80% of context-layer failures can be automatically diagnosed and remediated by mining historical trajecto...

Yikai Zhao, Pradeep Kumar Misra, Saurabh Pandey · 0 citations
Preprint Aug 2026

PILOT Technical Report

PILOT (Proactive Insight Learner for Online Tree-Experiments), an LLM-agent framework that organizes three roles within a constrained control loop where deterministic services enforce all safety, statistical, and permission boundaries, is presented.

Jiuning Lin, Ruiquan Lan, Xiaodong Zhu et al. · 0 citations
Preprint Sep 2026

KnowFeat: Knowledge-Guided Feature Engineering via LLM Agents

Automated feature engineering with large language models (LLMs) can produce semantically meaningful features for tabular data, yet existing methods lack structured domain knowledge, rigorous verification, and explainable provenance. We propose KnowFeat, a knowledge-guided feature engineering framework that organizes do...

Chengsong You, Wangyue Li, Wei-Qiao Que et al. · 0 citations
#artificial intelligence Review Aug 2026

RiskBlend: A Multi-Signal Framework for Test Input Prioritization in Machine Learning Regression Testing

RiskBlend is proposed, a classifier-agnostic prioritization framework that combines four complementary risk signals: historical failure patterns, prediction shift, decision-boundary shift, and neighborhood change that achieves the highest average APFD in all 80 dataset-classifier-scenario combinations.

Madhusudan Srinivasan, Namith Nishal Raphae · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.