Skip to content

AlphaZeroBeta: Deep Reinforcement Learning for Market-Neutral Portfolios

Jul 2026 · Financial Innovation · Vol 12 · 0 citations
Economics

TL;DR

Backtests covering 2014-2024 across seven equity indices show that the model achieves higher Sharpe ratios than the baselines while maintaining near-zero benchmark correlations and competitive drawdowns.

Abstract

Market-neutral portfolios aim to generate consistent returns while offsetting systematic market risk. Traditional approaches based on factor models or convex optimization often underperform during market regime shifts or when structural assumptions break down. We propose AlphaZeroBeta, a deep reinforcement learning framework designed to deliver benchmark-relative alpha (excess returns) with near-zero beta (market neutrality). AlphaZeroBeta combines a composite reward function that balances risk-adjusted excess return, benchmark correlation, and transaction costs with a CNN-GRU policy trained end-to-end via Recurrent PPO and evaluated through a rolling walk-forward protocol. Backtests covering 2014-2024 across seven equity indices show that the model achieves higher Sharpe ratios than the baselines while maintaining near-zero benchmark correlations and competitive drawdowns.

View source

Similar papers

Book Open access Aug 2026

Reinforcement Learning with Scenario-Context Rollout in Portfolio Management

When economic structures and market dynamics shift, classic portfolio rebalancing algorithms often suffer from unstable and degraded performance. To improve the return and robustness of portfolio management, we explore reinforcement learning (RL) and propose Scenario-Context Rollout (SCR), a macroeconomics-guided feedb...

Vanya Priscillia Bendatu, Yao Lu · 0 citations
Preprint Aug 2026

Tabular Deep Learning for Algorithmic Trading: Cross-Regime Bayesian Optimisation for Equity Signal Generation

No individual tabular deep learning architecture outperforms gradient-boosted trees, but combining XGBoost and TabNet using rank aggregation produces a Hybrid ensemble with an annualised return of 51.26%, a Sharpe ratio of 2.44, and a statistically significant CAPM alpha of 0.423.

Joshua Le Grice · 0 citations
Open access Aug 2026

Benchmarking deep reinforcement learning and classical models for portfolio optimization across market efficiency regimes

The results show that deep learning models perform best in highly efficient markets where signals are weak but consistent, and in moderately and least efficient markets, traditional strategies often achieve similar or better returns.

H. Sahu, Avishek Bhandari · 0 citations
Conference Aug 2026

Advances in Portfolio Optimization from Mean-Variance to Reinforcement Learning

Portfolio optimization is a fundamental problem in finance which has normally been addressed by mean-variance frameworks and their extensions. However, these methods rely on assumptions such as normally distributed returns and covariance estimates which often fail to capture the dynamics of real markets. Advances in ma...

Steven Itti Leon, Rishi V. N., Venkatakrishnan K. V. et al. · 0 citations
Conference Aug 2026

Uncertainty-Aware Forecast-Conditioned Reinforcement Learning for Multi-Asset Algorithmic Trading

Financial markets are challenging to navigate due to changing regimes, high volatility, and unpredictable investor behavior, often leading to model misspecification in classical stationary frameworks like Moving Average (MA) and Autoregressive (AR) models. To address this, we propose an uncertainty-aware framework that...

A. Verma, Arti M. K., Surjeet Kumar · 0 citations

Related blog posts

MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.