Skip to content
Preprint

Adaptive Finite-Budget Training for CVaR Risk-Aware Q-Learning

Aug 2026 · 0 citations · 36 references
Computer Science Economics

TL;DR

Results demonstrate that adaptive finite-budget training design, applied solely to the training procedure without altering the risk objective, can materially improve the reliability and risk-adjusted performance of risk-aware Q-learning in financial applications.

Abstract

Risk-aware Q-learning (RaQL) provides a model-free, two-timescale estimator for dynamic risk objectives, but its finite-budget behavior remains fragile: fixed inner-loop hyperparameters can produce unstable value estimates, persistent Bellman residuals, and inefficient sample reuse. This paper proposes an adaptive training controller for Conditional Value-at-Risk (CVaR) RaQL and evaluates it on a daily Bitcoin trading task. The controller preserves the original CVaR estimator and Bellman fixed point; instead, it redesigns the training procedure through six coordinated mechanisms: per-cell inner-step sizing, outer-rate-matched decay synchronization, a short early correction for the VaR-like inner variable, a coverage-first-then-greedy sample allocation rule, progressive suffix aggregation of mature inner estimates, and data-driven calibration of key scales from online-observable quantities. Across 20 random seeds and 856,000 inner-transition samples, the controller reduces the mean empirical CVaR Bellman residual by approximately 85% relative to the fixed-parameter baseline (MeanBEQ: 1.2202 to 0.1854; MeanBEV: 1.1624 to 0.0535) and maintains stability across CVaR levels, discount factors, and training budgets. On the chronological out-of-sample test set, the learned policy attains a Sharpe ratio of 0.9281 with a maximum drawdown of 6.46% after transaction costs. Although buy-and-hold yields a higher cumulative return (35.43% vs. 23.61%), the adaptive policy achieves far lower volatility (9.57% vs. 47.93%), drawdown, and CVaR loss. These results demonstrate that adaptive finite-budget training design, applied solely to the training procedure without altering the risk objective, can materially improve the reliability and risk-adjusted performance of risk-aware Q-learning in financial applications.

View source

Similar papers

Preprint Sep 2026

Feedback-Aware Tuning of Recursive Q-Learning

Model choice in backward Q-learning is recursive because a later-stage choice changes the response supplied to an earlier regression and can alter its model-comparison statistic. Separate stagewise criteria do not directly assess the target-stage prediction risk of a completed Q-learning fit. We address this mismatch b...

M. Kojima · 0 citations
Preprint Aug 2026

Revisiting TD Target Aggregation under Uncertainty in Q-Learning

The proposed SADQ is a simple modification to Q-learning that regularizes how the TD target is formed, and consistently improves training stability across classical control tasks, real-world vector-based environments, and Atari benchmarks when compared to strong DQN variants.

Li-Peng Zu, Xiaonan Zhang · 0 citations
Preprint Aug 2026

BRiG-AFA: Bellman Risk-to-Go Learning for Non-Myopic Active Feature Acquisition

A reproducible mechanism-level case for direct Bellman risk regression is established and the experiments still needed for state-of-the-art comparison are delimited, establishing a reproducible mechanism-level case for direct Bellman risk regression.

Jiaorong Feng, Qian Li, Ying Li · 0 citations
#machine learning Preprint Sep 2026

Model-based Bootstrap for Offline Policy Evaluation in Tabular Reinforcement Learning

Offline policy evaluation (OPE) is crucial in high-stakes reinforcement learning applications, where new policies must be assessed reliably before deployment. In such settings, point estimates alone are insufficient; principled uncertainty quantification, such as confidence intervals and variance estimates, is essentia...

Wei-Wei Wang, Yu-Qiang Li, Xian-Yi Wu et al. · 0 citations
Preprint Aug 2026

Risk-Sensitive Reinforcement Learning with Smoothed Quantile Objectives

UCB-BQRL is developed, a model-based optimistic learning algorithm that maintains confidence sets for the transition kernel and plans using a lower-buffered quantile criterion, and an information-theoretic lower bound of $\Omega(H/\rho_\tau\sqrt{AT})$ for the regret of any algorithm dealing with a quantile objective fu...

Mohammad Alipour-Vaezi, Huai-Yang Zhong, S. Khodadadian · 1 citation
#machine learning Preprint Sep 2026

Online Self-Weighted Fine-Tuning

Online Self-Weighted Fine-Tuning is proposed, a simple method that augments SFT with online, trajectory-level weighting and offers a favorable compute-performance trade-off as a practical approach for fine-tuning small-to-medium LLMs on binary-verifiable reasoning tasks with only 2 online rollouts.

Hai-Quan Wen, Yiwei He, Bei Peng et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.