Skip to content

Category

reinforcement learning

180 papers

#reinforcement learning Open access Aug 2026

Adaptive Bayesian Optimization for Reinforcement Learning Reward Shaping

Reward shaping is a critical technique in reinforcement learning (RL) that aims to accelerate learning by providing the agent with informative rewards. However, designing effective reward shaping functions can be a challenging and often tedious process, requiring domain expertise and extensive manual tuning. This paper proposes an adaptive Bayesian optimization approach to automate the reward shaping process. The system iteratively explores the space of potential reward functions, leveraging the agent's performance as feedback to refine the search strategy. We demonstrate that this approach can learn optimal reward shaping functions, leading to significant improvements in learning speed and agent performance compared to traditional reward shaping methods. The core claim is that an adaptive Bayesian optimization framework can effectively automate the reward shaping process, offering a more robust and efficient solution for complex RL problems. The system utilizes a Gaussian Process (GP) surrogate model to approximate the reward function landscape and employs an acquisition function, such as Expected Improvement, to guide the exploration process. The approach is evaluated on a suite of benchmark RL environments, showcasing its effectiveness across diverse scenarios. This work contributes to the broader field of RL by providing a practical and automated method for reward shaping, potentially unlocking new possibilities for tackling challenging RL problems.

Jincheng Zhang · 0 citations
#reinforcement learning Open access Aug 2026

Based on Multi-modal Data Fusion for Subconscious Decision Simulation

This paper presents a novel approach to simulating subconscious decision-making processes by leveraging multi-modal data fusion. The core idea is to construct a computational model capable of mirroring the complexities of human subconscious decision-making, moving beyond traditional behavioral analysis. We employ a graph neural network (GNN) architecture for robust multi-modal data integration, transforming diverse data streams – including visual, auditory, and tactile information – into a unified representation. This representation is then utilized within a reinforcement learning framework to simulate the subconscious decision-making process, explicitly modeling the interactive effects between different modalities. The resulting model provides a deeper understanding of how individuals make decisions without conscious awareness, offering potential applications in fields such as robotics, human-computer interaction, and cognitive modeling. The key innovation lies in the comprehensive incorporation of multi-modal interactions, providing a more accurate representation of the human subconscious than existing approaches. We define the following key equations to represent the core processes within the model: Let *xi* represent the input vector for modality *i*, where *i* ∈ {V, A, T}, representing Visual, Auditory, and Tactile modalities, respectively. The dimensionality of each *xi* is denoted as *di*. The multi-modal fusion process can be expressed as: * *xfused* = FusionNetwork(*xV*, *xA*, *xT*) Where *xfused* is the fused representation and FusionNetwork is the graph neural network. The reinforcement learning agent's decision-making process is governed by the following equation: * *ai* = argmaxj [Q( *xfused*, *aj* ) + β * R( *xfused*, *aj*)] Where *ai* is the action taken, *Q* is the Q-function estimating the expected reward, *R* is the reward function, and β is a weighting factor. The model's training objective can be formalized as: Minimize Eτ [ Σt=0T γt *R( *xfused*, *at* )] Where τ is a trajectory, *R* is the reward function, γ is the discount factor, and T is the time horizon.

Jincheng Zhang · 0 citations
#reinforcement learning Open access Aug 2026

Hierarchical Reinforcement Learning with Intrinsic Motivation and Meta-Learning

This paper proposes a novel hierarchical reinforcement learning (HRL) framework that leverages intrinsic motivation, meta-learning, and a hierarchical architecture to address the limitations of traditional HRL methods concerning exploration and generalization. The core idea is to integrate these three components to create a more adaptive and efficient learning system. At each level of the hierarchy, intrinsic motivation, specifically novelty seeking, encourages exploration. Simultaneously, meta-learning dynamically adjusts the learning rate and policy updates, enabling rapid adaptation to diverse tasks. We demonstrate the effectiveness of this approach through a theoretical analysis and a conceptual framework, outlining the key components and their interactions. The resulting system offers improved learning speed and robustness compared to standard HRL techniques, particularly when facing complex and varied environments. This work lays the groundwork for future research in developing truly adaptable and intelligent hierarchical agents.

Jincheng Zhang · 0 citations
#reinforcement learning Open access Aug 2026

Adaptive Resource Allocation in Cloud Computing (ARAC)

This paper presents Adaptive Resource Allocation in Cloud Computing (ARAC), a novel approach to cloud management that leverages reinforcement learning for dynamic resource allocation. Traditional cloud platforms often rely on manual configuration and pre-defined rules, leading to suboptimal resource utilization and potentially degraded user experience. ARAC addresses this limitation by employing a reinforcement learning-based resource scheduling algorithm. This algorithm continuously learns and adapts to changing conditions, optimizing the allocation of virtual machines, storage, and bandwidth based on user requests, resource utilization rates, and system load. The core claim of ARAC is to design a cloud platform capable of automatically adjusting resource allocations in response to evolving demands. The system's mechanism involves a dynamic adjustment of resources, aiming for optimal utilization and a superior user experience. This paper outlines the architecture, the reinforcement learning framework, and the key components of ARAC, demonstrating its potential to significantly improve cloud computing efficiency and responsiveness. ---

Jincheng Zhang · 0 citations
#reinforcement learning Open access Aug 2026

Adaptive Meta-Learning via Simulated Environment Dynamics

Meta-learning, the learning to learn, has shown significant promise in tackling complex tasks. However, a prevalent limitation lies in the reliance on static reward functions and environment dynamics, often simplifying the learning process and potentially hindering generalization to real-world scenarios. This paper introduces a novel adaptive meta-learning framework that addresses this limitation by dynamically adjusting the simulated environment's dynamics during the meta-training phase. The core idea is to utilize a learned Markov model to govern the environment's behavior, and to adapt the model parameters based on the agent's performance. This creates a continually evolving training environment, mirroring the inherent dynamism and uncertainty of real-world systems. We demonstrate that this approach leads to improved meta-learning performance compared to traditional static environment meta-learning methods. The algorithm incorporates key elements of reinforcement learning and Bayesian modeling to achieve adaptability and robustness. The primary formula representing the updated Markov model is: (Qt+1 | Qt, At) = f(Qt, At), where Qt+1 represents the state of the Markov model at time t+1, Qt is the state at time t, and At is the agent's action at time t. The function f is a parameterized function that is updated during the meta-training process. We explore the theoretical implications of this dynamic adaptation and discuss potential avenues for future research.

Jincheng Zhang · 0 citations
#reinforcement learning Open access Aug 2026

Massed and Distributed Training in a Canine Match-to-Sample Task: A Single-Subject Exploratory Study

The scheduling of practice may influence learning performance, behavioral engagement, and response speed across species. This exploratory study compared massed and distributed training in a visual match-to-sample task involving a three-year-old male mixed-breed domestic dog. A single-subject within-subject design was employed in which the dog underwent both training conditions. The massed condition consisted of 15 planned trials conducted in one continuous 12-minute session, whereas the distributed condition comprised the same number of planned trials divided across three sessions separated by four- to six-hour intervals. Performance was assessed descriptively through task completion, accuracy among completed trials, response latency, behavioral engagement, and positional choice. The dog completed 11 trials during massed training and 12 during distributed training. Accuracy remained broadly comparable, increasing from 64% under the massed condition to 67% under the distributed condition. Average response latency decreased from 3.0 seconds to 1.5 seconds, while fewer nonresponses and greater behavioral engagement were observed during the distributed condition. However, a pronounced left-side preference emerged during distributed training, indicating that reinforcement placement or other unintended environmental cues may have influenced performance. The findings are descriptively consistent with the potential value of spaced practice for maintaining response speed and engagement, but they do not establish superior match-to-sample learning. The study underscores the importance of counterbalancing, neutral reinforcement delivery, repeated observations, and rigorous control of positional cues in future canine-cognition experiments.

Jena Lorraine F. Foncardas, Felisse Marianne Z. San Juan · 0 citations
#reinforcement learning Open access Aug 2026

Quantum-Cognitive Reinforcement Learning via Penrose Objective Reduction

Classical reinforcement learning (RL) and decision theory rely on Kolmogorovian probability spaces and independent utility metrics. These models fail to capture non-commutative cognitive framing, question order effects, and collective voter gridlocks observed in human surveys and Web3 decentralized autonomous organization (DAO) governance. Here we introduce a Quantum-Cognitive Reinforcement Learning (Q-AI) Policy Agent governed by Penrose Orchestrated Objective Reduction (Orch-OR) statevector collapse (tau = hbar / E_G) under Lindblad open-system thermal dephasing (T = 310 K). We validate our architecture against two empirical datasets:1. Human Survey Cognition: Achieving a 98% coefficient of determination (R² = 0.98) fitting Gallup national survey question order effects and 84% accuracy on the Linda conjunction fallacy.2. Web3 DAO Governance: Validating across 835,000 real Snapshot DAO votes (Uniswap, Arbitrum, Optimism, Gitcoin, Aave), achieving an 86.7% Mean Absolute Error reduction (1.3% MAE vs 9.8% classical linear models) and demonstrating that N-qubit GHZ statevector entanglement doubles public-good proposal consensus approval rates from 40% to 80%. Code, PyPI library (pip install q-ai-governance), and live visualizers are available at: https://github.com/JonathanReiser/quantum-orch-or

Jonathan Reiser · 0 citations
#reinforcement learning Open access Aug 2026

Sip and Strum: Matcha Intake and a Practice-Reduction Reinforcement Contingency in Beginner Ukulele Skill Acquisition

This study examined differences in beginner ukulele skill acquisition under three combinations of matcha intake and a practice-reduction reinforcement contingency. A quasi-experimental three-group design was employed involving 15 students aged 18–25 years with no prior musical-instrument experience. Participants were assigned to one of three conditions: no matcha with reinforcement, matcha with reinforcement, or matcha without reinforcement. The reinforcement contingency consisted of shortening the remaining practice period when a participant attained at least 85% performance accuracy at the prescribed tempo twice consecutively. Participants completed two structured practice sessions, followed by retention assessments conducted 24–48 hours after learning. Musical performance was evaluated using accuracy, tempo maintenance, criterion attainment, correct chord progressions, and retention scores. Descriptive results showed that the matcha-with-reinforcement group exhibited the strongest overall performance, including more frequent tempo maintenance, higher criterion attainment, and more correct chord progressions. This group also obtained the highest mean retention score. Welch’s one-way analysis of variance indicated a statistically significant difference in retention across the three conditions, with the matcha-with-reinforcement group outperforming the matcha-only group. However, the small sample and absence of a no-matcha-without-reinforcement condition preclude separate estimates of the two interventions and any conclusion regarding their interaction. The findings provide preliminary evidence that combining physiological stimulation with a structured practice-reduction contingency may support beginner musical learning, although confirmation through a larger, complete factorial experiment is required.

Joycelyn Ann Faith S. Culajara, Felisse Marianne Z. San Juan · 0 citations
#reinforcement learning Open access Aug 2026

Quantum-Cognitive Reinforcement Learning via Penrose Objective Reduction

Classical reinforcement learning (RL) and decision theory rely on Kolmogorovian probability spaces and independent utility metrics. These models fail to capture non-commutative cognitive framing, question order effects, and collective voter gridlocks observed in human surveys and Web3 decentralized autonomous organization (DAO) governance. Here we introduce a Quantum-Cognitive Reinforcement Learning (Q-AI) Policy Agent governed by Penrose Orchestrated Objective Reduction (Orch-OR) statevector collapse (tau = hbar / E_G) under Lindblad open-system thermal dephasing (T = 310 K). We validate our architecture against two empirical datasets:1. Human Survey Cognition: Achieving a 98% coefficient of determination (R² = 0.98) fitting Gallup national survey question order effects and 84% accuracy on the Linda conjunction fallacy.2. Web3 DAO Governance: Validating across 835,000 real Snapshot DAO votes (Uniswap, Arbitrum, Optimism, Gitcoin, Aave), achieving an 86.7% Mean Absolute Error reduction (1.3% MAE vs 9.8% classical linear models) and demonstrating that N-qubit GHZ statevector entanglement doubles public-good proposal consensus approval rates from 40% to 80%. Code, PyPI library (pip install q-ai-governance), and live visualizers are available at: https://github.com/JonathanReiser/quantum-orch-or

Jonathan Reiser · 0 citations
#reinforcement learning Open access Aug 2026

Cognitive UAV-driven agro-surveillance framework for predicting crop stress–induced yield loss using spatio-temporal learning and adaptive irrigation control

Precision agriculture is becoming more and more of a challenge that requires the use of intelligent systems that are able to predict stress and prevent yield loss before it is too late. Traditional methods of agricultural surveillance are predominantly reactive with irrigation demands being based on thresholds or individual yield forecasts models that do not represent the intricate spatio-temporal interactions that exist between crop physiology, soil status, and environmental stresses. Besides, the majority of the current practices do not have an autonomous decision-making approach to preventive intervention which leads to inefficient use of water and slows down the response to stress. This paper suggests a cognitive UAV-assisted agro-surveillance system to predict yield vulnerability caused by crop stress and optimize adaptive irrigation with the help of spatio-temporal deep and reinforcement learning. The framework combines UAV-obtained RGB and multispectral and thermal imagery with measurements of soil sensors and meteorological data obtained with the Crop Health and Environmental Stress Dataset. A new GeoSpatio-TRiNet model is used to acquire long-range spatial relationship, time stress development, and diffusion of stresses across agricultural regions. The model predicts the vulnerability trajectories of the stress instead of the direct yield regression, and this allows early detection of yield risk. Such predictions serve to generate a cognitive environmental state of a Soft ActorCritic (SAC) reinforcement learning agent that autonomously computes zone-based irrigation behaviors to reduce the recurrence of stress at the minimum water usage cost. As shown by the results of the experiment, the proposed framework has a stress forecasting accuracy of 96.3% and performs much better than the traditional machine learning, CNN-based, and transformer-based baselines. The system also decreases the predicted yield vulnerability by 46.6 and enhances water-use efficiency by 41.1 as compared to irrigation strategies based on rules. The results confirm the usefulness of spatio-temporal intelligence with predictive control in terms of effectiveness, and the proposed framework is a scalable and sustainable solution to precision agriculture of the next generation.

S. Selvakumar, D. Venugopal · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.