Skip to content

Category

reinforcement learning

179 papers

Hybrid Reinforcement Learning “HRL” Approach with Energy Harvesting for Energy-Efficient Routing Protocols “EERP” “HRL-EERP” in Wireless Body Area Networks “WBANs” (Preprint)

BACKGROUND Wireless Body Area Networks (WBANs) revolutionize healthcare by utilizing miniature, low-power sensors, implanted or worn externally, to provide continuous, real-time tracking of vital signs like heart rate, blood pressure, temperature, and glucose levels. However, their potential is hindered by significant routing challenges, particularly energy efficiency, due to the limited battery life of sensor nodes and the dynamic, reliability-critical nature of medical applications. OBJECTIVE This paper introduces HRL-EERP, a Hybrid Reinforcement Learning-based Energy-Efficient Routing Protocol for WBANs, designed to address these issues. METHODS HRL-EERP integrates Deep Q-Network (DQN) to optimize multi-hop paths for low latency, Q-Learning to ensure efficient local routing decisions, Restricted Data Transmission (RDT) to minimize energy use by suppressing redundant packets, and Energy Harvesting (EH) to sustain node operation. RESULTS Evaluated using NS-3 with PhysioNet MIT-BIH Arrhythmia Database data (60-80 bpm baseline, 150 bpm anomalies, 86,400 samples/day), HRL-EERP achieved a 34-hour network lifetime, 0.019 mJ/s energy efficiency, 96.8% PDR, and 88 ms latency across 56,160 packets/day (post-35% RDT suppression), meeting ECG <100 ms requirement. RDT saved 1,512 mJ/day, while EH provided 120-720 mJ/day/node (0.05-0.3 mJ/s across head, chest, wrist). CONCLUSIONS HRL-EERP delivers a robust, energy-conscious solution for time-sensitive WBAN applications. CLINICALTRIAL None applicable

Sarah Allabun 6th · 0 citations
#reinforcement learning Dataset Open access Aug 2026

Data for "Auditing Single-Agent Reinforcement Learning for EV Charging Assignment: A Protocol-Amended Comparison of Trained, Untrained, and Heuristic Policies"

Data for "Auditing Single-Agent Reinforcement Learning for EV Charging Assignment: A Protocol-Amended Comparison of Trained, Untrained, and Heuristic Policies" Raw seed-level data and campaign manifests for a benchmark of five agents (Random, Adaptive Heuristic, Q-Learning, DQN, Double DQN) on EV charging-station assignment, simulated on real Rabat and Tangier (Morocco) road networks in SUMO. Includes: campaign manifests with SHA-256 provenance, raw per-seed CSVs for the confirmatory trained/untrained diagnostic (three scenarios) and the legacy 210-run benchmark, and the JSON summaries behind the manuscript's result tables. Integrity verifiable via the included SHA-256 manifest. Preliminary, data-only deposit. Simulation event logs and trained model weights are not included in this version; available from the corresponding author on request.

Nour-Eddine Moumni, Rachid Alaoui, Driss Kiouach · 0 citations

Asynchronous Federated Reinforcement Learning for Adaptive Resource Slicing and Low-Latency Task Offloading in Heterogeneous 6G Edge Computing Networks

The emerging paradigm of 6G wireless communication networks envisions ultra-reliable low-latency communication (URLLC), massive machine-type communications (mMTC), and pervasive edge computing intelligence. In heterogeneous mobile edge computing (MEC) networks, dynamically offloading compute-intensive tasks (e.g., augmented reality rendering, connected vehicular telemetry, autonomous robotic control) while orchestrating multi-tenant network slicing under time-varying channel conditions is an NP-hard stochastic optimization problem. Centralized reinforcement learning algorithms suffer from extreme communication overhead, severe backhaul congestion, and severe privacy vulnerabilities. Conversely, standard synchronous Federated Learning (FL) methods encounter severe 'straggler effects' caused by heterogeneous edge device processing capabilities. In this paper, we propose AF-EdgeRL, a novel Byzantine-resilient Asynchronous Federated Reinforcement Learning framework tailored for distributed resource allocation and dynamic task offloading. AF-EdgeRL deploys a distributed Proximal Policy Optimization (PPO) agent across edge servers and end-user devices, combined with a Staleness-Aware Adaptive Weight Aggregator (SAWA) that dynamically adjusts model update gradients based on hardware compute latency and channel state information (CSI). Furthermore, we establish theoretical convergence guarantees under non-convex reinforcement learning objectives. Evaluated on a high-fidelity 6G MEC simulator with real-world mobile mobility traces (Telecom Italia Milano dataset), AF-EdgeRL reduces end-to-end task execution latency by 41.2%, achieves 99.999% URLLC deadline compliance, and decreases edge energy consumption by 32.6% compared to state-of-the-art synchronous FedRL and centralized DRL baselines.

Daniel Merrow, Tember L. Nair, Lucas Farnandez · 0 citations
#reinforcement learning Open access Aug 2026

Antifragile Intelligence: A Triadic Framework for AI Governance, Digital Forensics, and Sovereignty in Emerging Economies

In an era defined by extreme Volatility, Uncertainty, Complexity, and Ambiguity (VUCA), artificial intelligence (AI) governance must transcend passive compliance checklists to become an embedded, adaptive socio-technical architecture. This paper proposes a triadic synthesis of Reinforcement Learning (RL), Generative AI (GenAI), and Cybersecurity, organized within a Seven-Layer Integrated Architecture spanning perception, cognition, adaptation, generation, protection, embodiment, and governance. Central to the framework is a formal isomorphism between Predictive Processing (PP) and Reinforcement Learning, in which both systems minimize prediction error through Bayesian updating (Friston, 2010; Friston et al., 2009). This isomorphism is operationalized through a safety-constrained objective function that treats variational free energy as a regularizer, mitigating the class of failures known as “reward hacking” (Laidlaw et al., 2025; Shihab et al., 2025; Skalse et al., 2022). Illustrative comparison of the Asynchronous Advantage Actor-Critic (A3C) algorithm against legacy Q-Learning suggests materially faster and more stable policy convergence under the resource-constrained, high-packet-loss conditions typical of emerging economies. By integrating the sub-Saharan African relational philosophy of Ubuntu/Unhu with global AI4People principles (Floridi et al., 2018; Van Norren, 2023; Yilma, 2025), the framework embeds explicit digital forensics workflows and blockchain-anchored chain-of-custody protocols (Atlam et al., 2024; Patil et al., 2024). The framework is further extended and empirically grounded through a twentyproject, four-cluster Edge-AI case portfolio spanning domestic safety, environmental intelligence, sustainable energy and agriculture, and healthcare accessibility in the Indian context, demonstrating the triadic architecture’s applicability from enterprise-scale governance to grassroots micro, small, and medium enterprise (MSME) innovation. This synthesis serves as a blueprint for organizations in the Southern African Development Community (SADC) and India to assert digital sovereignty, ensuring that autonomous systems are antifragile, context-sensitive, and designed for communal flourishing rather than extractive optimization.

Gabriel Kabanda · 0 citations
#reinforcement learning Open access Aug 2026

Primary Field Cybernetics

Primary Field Cybernetics: From Spectral-Phase Flow to a Symformic Theory of Control This article develops Primary Field Cybernetics (PFC) as a cybernetic theory derived from the ontology of Symformism and the field architecture of Dynamical Informational Field Theory (DIFT). Symformism treats an enduring form not as a static object but as a relational organization that preserves its identity through continuous change, exchange, perturbation, and reorganization. DIFT provides a physical research architecture for this intuition through the complex spectral-phase Primary Organizational Field, organized phase current, structural memory, adaptive access geometry, mobility, impedance, and dynamostatic persistence. From these relations, the article derives a domain-neutral cybernetic architecture: Primary Field → organized flow → retained history → differential impedance → differential accessibility → viable future action → control. The central proposition is that control cannot be reduced to selecting an action from a fixed repertoire. A system’s history can alter the practical accessibility of its future responses. Consequently, an enduring actant participates in reorganizing the conditions under which its own later regulation remains possible. On this basis, PFC distinguishes state control, access control, reflexive control, and relational metacontrol, and introduces the concept of accessible variety: regulatory variety understood not merely as nominally available responses, but as responses that remain practically reachable within relevant constraints of cost, delay, compatibility, and viability. The paper situates this proposal in relation to classical and second-order cybernetics, including Maxwell, Wiener, Ashby, Beer, von Foerster, Maturana and Varela, and gives particular attention to Marian Mazur’s theory of autonomous systems. It also distinguishes the proposed architecture from active inference, allostasis, reinforcement learning, eligibility traces, synaptic plasticity, and meta-learning. A deliberately limited numerical model is included as an illustration of one consequence of the theory rather than as validation of PFC or DIFT. It examines whether different histories can produce different future response accessibility under an otherwise matched challenge. Causal ablation and a single-timescale trace equivalence control are used to delimit what this example does and does not establish. The article concludes with a falsification programme based on matched-state, divergent-history experiments, in which present observables are matched, accessibility is measured before a decisive response, identical perturbations are applied, and the proposed access mechanism is selectively manipulated. Candidate applications include physical, biological, neural, artificial, and institutional systems. The resulting formulation shifts the fundamental cybernetic question from “How does a system correct its present state?” toward “How does an enduring organization preserve and reorganize the conditions under which viable future action remains accessible?” Keywords: Primary Field Cybernetics; Symformism; DIFT; Primary Organizational Field; spectral-phase flow; phase current; dynamostasis; impedance; access geometry; accessible variety; cybernetics; control; memory; resilience.

Sławomir Krakowski · 0 citations
#reinforcement learning Open access Aug 2026

The Cognitive Familiarity Supremacy Theory (CFST)

The Cognitive Familiarity Supremacy Theory (CFST) proposes that a substantial portion of human certainty, ideological attachment, collective identity formation, and perceived superiority emerges not primarily from objective rational evaluation, but from repeated familiarity encoding mechanisms operating within subconscious cognitive architectures.This framework argues that repeated environmental exposure, social conditioning, emotional reinforcement, identity fusion, symbolic repetition, and institutional amplification collectively construct familiarity-driven epistemic structures that are frequently mistaken for objective truth, rational certainty, or universal superiority. The theory synthesizes and mathematically formalizes principles from cognitive neuroscience, psychology, sociology, political theory, philosophy of mind, epistemology, systems theory, information theory, complexity science, behavioral economics, evolutionary biology, communication studies, artificial intelligence, anthropology, cybernetics, and cultural theory into a unified explanatory framework.CFST introduces a comprehensive causal chain model:Repeated Exposure \rightarrow Subconscious Encoding \rightarrow Identity Fusion \rightarrow Emotional Reinforcement \rightarrow Bias Formation \rightarrow Perceived Superiority.The theory proposes that human cognition operates through familiarity-weighted interpretive systems, where the subjective sensation of certainty often emerges from accumulated familiarity intensity rather than objective verification.The framework further integrates: Bayesian epistemology, predictive processing, Hebbian learning, social identity theory, information entropy, network propagation, algorithmic amplification, memetic evolution, cultural conditioning, political hegemony, and neurocognitive attractor-state dynamics.The theory also develops: formal mathematical models, belief topology equations, dynamic systems formulations, stochastic familiarity propagation systems, agent-based ideological simulations, network-theoretical belief diffusion structures, and computational cognitive equilibrium equations.At the civilizational level, CFST proposes that societies are partially constructed upon collectively reinforced familiarity architectures rather than purely objective truth systems. At the individual level, it explains ideological rigidity, nationalism, fanaticism, cultural supremacy perception, identity-protective cognition, and epistemic polarization.Finally, the theory proposes that genuine epistemic liberation requires conscious disruption of subconscious familiarity monopolies through critical reasoning, diversity exposure, meta-cognitive awareness, and reflective epistemological reconstruction.

Shamiul Hoque Shan · 0 citations
#reinforcement learning Open access Aug 2026

DIArc Foundational Note v0.1 — Minimum Claim Edition

Abstract The rapid development of artificial intelligence has significantly increased the availability of information, analytical capability, and machine-assisted reasoning. However, greater access to information does not necessarily produce better decisions. In many organizational contexts, the emerging bottleneck is no longer information acquisition, but the human and organizational capacity to determine what information is sufficient, when analysis should stop, when a decision should be made, and how outcomes should improve future judgment. This Foundational Note introduces Decision Intelligence Architecture (DIArc) as an architectural framework for Human–AI collaborative decision systems. DIArc is based on a central proposition: in the AI era, competitive advantage increasingly depends not on maximizing information, but on maximizing the rate at which high-quality decisions generate learning and improve judgment, under explicit constraints on information consumption and decision cycles. The architecture is organized into four theoretical layers. First, the Capability Inversion Hypothesis describes a structural shift in which information, knowledge, and analysis become increasingly abundant while judgment, commitment, execution, and learning become comparatively scarce capabilities. Second, Identity-driven Information Consumption (IDIC) describes a decision failure mechanism in which continued information consumption may serve identity reinforcement rather than decision improvement. Third, the Decision Constraint Architecture, comprising Decision Information Budget (DIB) and Decision Cycle Budget (DCB), introduces explicit constraints on information consumption and analytical iteration. Fourth, High-quality Decision Velocity (HQDV) describes the performance objective of accelerating completed high-quality decision loops, while Judgment Evolution Rate (JER) represents the longer-term evolutionary objective of improving judgment through outcome-based learning. This note constitutes the initial public disclosure of the DIArc architecture and establishes its theoretical baseline for subsequent research and branch concepts.

Lucas Xiaochun Xu · 0 citations
#reinforcement learning Review Open access Aug 2026

Learning by Consequence: A Narrative Review of Reinforcement Learning from Thorndike's Law of Effect to Deep Q-Networks and AlphaGo

Reinforcement learning---learning what to do from reward and punishment rather than from instruction---unifies animal psychology, optimal control, and machine learning into one computational program, and its deep-learning era delivered the field's most visible artificial intelligence achievements. This article presents a narrative review of the canonical line: Thorndike's 1911 law of effect, Bellman's 1957 dynamic programming, Samuel's 1959 checkers player, Sutton's 1988 temporal-difference learning, Watkins and Dayan's 1992 Q-learning, Tesauro's 1995 TD-Gammon, Sutton and Barto's 1998 synthesis, Mnih and colleagues' 2015 Deep Q-Network, Silver and colleagues' 2016 AlphaGo and 2017 AlphaGo Zero, Lillicrap and colleagues' continuous control with DDPG, and Schulman and colleagues' 2017 proximal policy optimization. The synthesis is organized around three themes: foundations, in which the credit-assignment problem received formal solutions in value functions and temporal difference; scaling, in which function approximation, experience replay, and self-play converted tabular theory into high-dimensional control; and algorithmic consolidation, in which actor-critic methods and policy gradients stabilized practice. It is concluded that reinforcement learning's contribution is a general grammar of goal-directed learning---and that its open problems, sample efficiency and reward specification, define the frontier between artificial and natural intelligence.

Zen Revista, 10 IA · 0 citations
#reinforcement learning Open access Aug 2026

Student Behavior Recognition and Intervention Methods in Intelligent Classrooms Based on Deep Reinforcement Learning

In response to the problem that traditional classroom student behavior analysis relies on manual observation and is difficult to achieve real-time, precise and personalized intervention. This paper constructs an end-to-end intelligent classroom intervention system based on deep reinforcement learning. This system adopts the Multi-Modal Fusion Spatio-Temporal Graph Convolutional Network (MM-ST-GCN), integrating visual skeletons, seat pressure and classroom interaction data, to achieve fine-grained and high-precision recognition of students' classroom behaviors. It models the classroom environment as a Partially Observable Markov Decision Process (POMDP) and uses the improved Soft Actor Critic (SAC) algorithm to generate the optimal intervention strategy that takes into account the learning benefits of students and the intervention costs of teachers. Experiments on the MMAct dataset show that the proposed behavior recognition model achieves 93.8% accuracy and a macro-average F1 score of 0.925. Simulation experiments indicate that the system has the potential to improve students' concentration and reduce distraction behavior. However, the above results are derived from the simulated environment and need to be further verified in the real classroom.

Jian Mu · 0 citations
#reinforcement learning Open access Aug 2026

Distribution Estimation Algorithm for Cloud Manufacturing Scheduling Optimization

To address the resource scheduling problem in complex production environments, this study proposes a production scheduling model based on Estimation of Distribution Algorithm.The model constructs a probability model using spatial distribution, evaluates the scheduling population based on high-quality individuals, introduces an archive mechanism to enhance solution diversity, and combines Deep Reinforcement Learning and Tabu Search algorithm for global optimization.It achieves adaptive production scheduling optimization under dynamically changing resources.In testing experiments, the model achieves an accuracy of 95.11 % in sample classification prediction tasks.The computational load and number of parameters for production data processing are 664.8FLOPs and 90.54 M, respectively.The scheduling delay rate and resource utilization are 4.39 % and 97.96 %, significantly outperforming comparison models.These results indicate that the model provides stable and efficient production scheduling optimization and multi-constraint conditions, offering reliable algorithm support for production scheduling in cloud-based networked manufacturing environments.

H. Wan, Y. Li · 0 citations
#reinforcement learning Review Open access Aug 2026

Learning by Watching: A Narrative Review of Imitation Learning from ALVINN to Generative Adversarial Imitation

Imitation learning---the learning of behavior from demonstrations instead of rewards---moved from Pomerleau's ALVINN driving network and Schaal's humanoid route through Ng and Russell's inverse reinforcement learning, Abbeel and Ng's apprenticeship learning, and Ziebart's maximum entropy to the robot learning from demonstration surveys, Ross's DAgger, Ho and Ermon's generative adversarial imitation, Finn's guided cost learning, and the algorithmic perspective's syntheses. This article presents a narrative review of that arc's canonical line: Pomerleau's 1989 ALVINN, Schaal's 1999 humanoid question, Ng and Russell's 2000 inverse RL, Abbeel and Ng's 2004 apprenticeship learning, Billard and colleagues's 2008 handbook chapter, Ziebart and colleagues's 2008 maximum entropy, Argall and colleagues's 2009 survey, Ross, Gordon, and Bagnell's 2011 DAgger, Ho and Ermon's 2016 GAIL, Finn and colleagues's 2016 guided cost learning, Hussein and colleagues's 2017 survey, and Osa and colleagues's 2018 algorithmic perspective. The review is organized around three themes: the foundations, in which the driving network's demonstrations, the humanoid's question, and the inverse reward's recovery defined the field's two programs; the demonstration's surveys, in which the robot programming's handbook and the LfD's survey systematized the practice; and the deep era, in which the DAgger's covariate correction, the adversarial's discrimination, and the algorithmic perspective's synthesis unified the field. It is concluded that imitation learning is the reward's workaround---and that its arc is the demonstrator's knowledge's transfer from the human's steering to the policy's distributions.

Zen Revista, 10 IA · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.