Skip to content
Preprint

Population-Robust Feature Selection via Generalized Welfare Optimization

Aug 2026 · 0 citations · 24 references
Computer Science

TL;DR

PopFS is introduced, a method for learning one shared, deployable feature set that is robust to population differences while letting each pop- ulation train its own model.

Abstract

Choosing which features to collect is a deployment decision: the same limited questionnaire, test panel, or sensor set may need to serve several heterogeneous populations. Standard feature-selection methods typically optimize for one large population, while existing robust approaches tend to learn one shared model for every population. We introduce PopFS, a method for learning one shared, deployable feature set that is robust to population differences while letting each pop- ulation train its own model. PopFS uses a tunable welfare objective that lets practitioners balance overall predictive ben- efit against stronger protection of the populations that benefit least. To make this objective practical at scale, PopFS first uses multitask sparse learning to reduce the candidate pool, then searches directly over hard feature sets by ranking promising additions and swaps and fully refitting only a shortlist. Across eight population splits from six prediction tasks drawn from five tabular and public-health datasets, PopFS consistently achieves strong average and worst-population performance while scaling to thousands of candidate features. A 43-state COVID-19 nowcasting study further shows that changing the welfare objective can improve the least-served states with lit- tle change in average performance and yields an interpretable change in the selected symptom signals. Our code is available at https://github.com/Rachel-Lyu/PopFS.

View source

Similar papers

Aug 2026

Supervised Feature Selection via Collective First-Order Neural Dynamics.

These findings validate that the synergy between FND and the collective mechanism effectively balances local exploitation and global exploration, leading to robust feature selection performance across diverse datasets and parameter settings.

Bolin Liao, Yufei Wang, Shuai Li et al. · 0 citations
#machine learning Preprint Sep 2026

Scaling Optimal Classification Trees via Adaptive Feature and Sample Reduction

Dynamic programming for optimal classification trees becomes computationally expensive as the numbers of features and training samples increase. We develop a joint feature- and sample-space reduction framework based on STreeD. Weighted STreeD merges duplicate records created after projection onto a fixed candidate set...

Jian-Cheng Tu, Wen-Qi Fan · 0 citations
Book Open access Sep 2026

Which LLM to Fine-Tune? Agent-Driven Model Selection at Scale

Open-source model hubs now host over two million public AI models, yet teams building customer-facing AI systems must still determine which model to fine-tune for production deployment—a decision that shapes the quality, latency, and cost experienced by hundreds of millions of users. At Amazon, we spent over years of i...

Chen Luo, Yu-Lin Liu, Yi Liu et al. · 0 citations
Open access Aug 2026

Elastic net–guided NSGA-II for scalable multi-objective feature selection in high-dimensional data analytics

Redundant, irrelevant, and noisy features make it very hard to analyse high-dimensional data, especially when the number of features is much larger than the number of samples. Conventional feature selection methods, such as filter, wrapper, and embedded methods, are unable to balance predictive accuracy, feature subset...

S. Subhani, Gudipati Murali, Chandra Sekhar Sanaboina · 0 citations
Open access Apr 2025

Comprehend, Divide, and Conquer: Feature Subspace Exploration via Multi-Agent Hierarchical Reinforcement Learning

Feature selection aims to preprocess the target dataset, find an optimal and most streamlined feature subset, and enhance the downstream machine learning task. Among filter, wrapper, and embedded-based approaches, the reinforcement learning (RL)-based subspace exploration strategy provides a novel objective optimizatio...

Wei-Liang Zhang, Xiaohan Huang, Yi Du et al. · 1 citation
Open access Aug 2026

Simultaneous Multi-Objective Evolutionary Optimization of Heterogeneous Ensembles, Learner-Specific Feature Subsets, and Aggregation Weights

This paper introduces an integrated multi-objective evolutionary framework for synthesis of heterogeneous regression ensembles featuring localized, learner-specific feature selection. Rather than enforcing global feature spaces, the proposed paradigm simultaneously optimizes base estimator activation patterns, customiz...

J. Galván, G. Sánchez, Fernando Jiménez · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.