Skip to content

CLaC@FinMMEval 2026 Task 3: Sentiment-Augmented Deep Reinforcement Learning for Active Trading - An Alpha-Reward Approach

Jul 2026 · arXiv.org · Vol abs/2607.16028 · 0 citations · 27 references
Computer Science

TL;DR

This paper presents the system for Task 3 of the CLEF 2026 FinMMEval Lab, which requires daily long, flat, or short trading decisions for Bitcoin and Tesla using news and historical market data, and achieves the strongest overall performance.

Abstract

This paper presents our system for Task 3 of the CLEF 2026 FinMMEval Lab, which requires daily long, flat, or short trading decisions for Bitcoin (BTC) and Tesla (TSLA) using news and historical market data. We formulate the problem as a discrete-action Markov Decision Process and compare four deep reinforcement learning algorithms: Policy Gradient (PG), Proximal Policy Optimization (PPO), Deep Q-Learning (DQL), and Deep Deterministic Policy Gradient (DDPG). The agents use technical indicators, cyclical calendar encodings, and daily news sentiment scores produced by LLaMA 3.2 1B. To reduce overfitting and align training with the objective of outperforming buy-and-hold, we introduce an alpha reward based on excess market return and randomize episode start dates. Hyperparameters are optimized with Ray Tune over 180 trials per algorithm-asset pair, with early stopping and model selection based on validation Sharpe ratio. On the CLEF Task 3 test set, DDPG achieves the strongest overall performance. DQL was selected a priori for the live endpoint because it obtained the highest validation Sharpe ratio, with selection performed without access to the test period. For TSLA, DDPG and DQL achieve cumulative returns of 54.96% and 52.62%, respectively, compared with 16.45% for buy-and-hold. For BTC, DDPG achieves a positive return of 1.58% while buy-and-hold declines by -34.27%. The results also reveal a substantial validation-to-test generalization gap, highlighting the difficulty of transferring policies selected in bull-market conditions to a bear-market regime.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Deep Reinforcement Learning for Equity Trading: Benchmarking Actor-Critic Methods with Forward Retraining

Consistently profitable trading is difficult because equity markets are noisy, non-stationary, and only partially predictable from historical data. We benchmark five deep reinforcement learning (DRL) actor-critic methods: A2C, PPO, DDPG, TD3, and SAC, that learn trading actions end-to-end from market states, and compar...

Bi-Cheng Wang, Xin-Yi Zhang · 0 citations
#reinforcement learning Open access Sep 2026

Reinforcement Learning-Driven Dynamic Trading Strategies for Financial Markets

This paper presents the RL-DynTrade framework by using a cutting-edge deep reinforcement learning method, Proximal Policy Optimization (PPO), with a Deep Q Network (DQN) agent to dynamically adapt to changing risk-reward dynamics. PPO enables real-time, fine-grained, risk-reward adaptation via an actor-critic design wi...

Xi-Jing Ou, Jie Huang · 0 citations
#machine learning Preprint Oct 2026

Do Your Own Research: Learning to Forecast by Learning to Search

Outcome-based reinforcement learning can train language models to forecast real-world events, but prior forecasting work either freezes research context before training or deploys agentic research only at test time, so the skill of gathering evidence is never shaped by the reward. We introduce an agentic forecasting en...

Yusuf Afifi, Artur Kiulian, Anton Polishko et al. · 0 citations
Preprint Aug 2026

SAPO: Single-Rollout Autoregressive Policy Optimization for Agentic Reinforcement Learning

This work proposes Single-rollout Autoregressive Policy Optimization (SAPO), a low-memory and compute-efficient framework in which the policy and value functions share a single autoregressive backbone, and introduces a trajectory-level generalized advantage estimator that combines lambda-returns with batch normalizatio...

D. Liang, Lang Feng, Bo An et al. · 2 citations
Review Open access Aug 2026

A survey on LLM-enhanced reinforcement learning in financial markets

A three-paradigm taxonomy (feature-based, auxiliary-based, and policy-based) based on the functional role of LLMs within the RL pipeline is proposed, which provides superior scalability and stability, though often at the expense of representational depth.

Ghusoon Hadi al-Aldaffaie, Alireza Taheri, Amirfarhad Farhadi et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.