Skip to content

Category

reinforcement learning

155 papers

#reinforcement learning Open access Aug 2026

DIArc Foundational Note v0.1 — Minimum Claim Edition

Abstract The rapid development of artificial intelligence has significantly increased the availability of information, analytical capability, and machine-assisted reasoning. However, greater access to information does not necessarily produce better decisions. In many organizational contexts, the emerging bottleneck is no longer information acquisition, but the human and organizational capacity to determine what information is sufficient, when analysis should stop, when a decision should be made, and how outcomes should improve future judgment. This Foundational Note introduces Decision Intelligence Architecture (DIArc) as an architectural framework for Human–AI collaborative decision systems. DIArc is based on a central proposition: in the AI era, competitive advantage increasingly depends not on maximizing information, but on maximizing the rate at which high-quality decisions generate learning and improve judgment, under explicit constraints on information consumption and decision cycles. The architecture is organized into four theoretical layers. First, the Capability Inversion Hypothesis describes a structural shift in which information, knowledge, and analysis become increasingly abundant while judgment, commitment, execution, and learning become comparatively scarce capabilities. Second, Identity-driven Information Consumption (IDIC) describes a decision failure mechanism in which continued information consumption may serve identity reinforcement rather than decision improvement. Third, the Decision Constraint Architecture, comprising Decision Information Budget (DIB) and Decision Cycle Budget (DCB), introduces explicit constraints on information consumption and analytical iteration. Fourth, High-quality Decision Velocity (HQDV) describes the performance objective of accelerating completed high-quality decision loops, while Judgment Evolution Rate (JER) represents the longer-term evolutionary objective of improving judgment through outcome-based learning. This note constitutes the initial public disclosure of the DIArc architecture and establishes its theoretical baseline for subsequent research and branch concepts.

Lucas Xiaochun Xu · 0 citations
#reinforcement learning Open access Aug 2026

Evaluating a School Waste Bank Strategy for Strengthening Students’ Environmental Responsibility: A Case Study at SDN 4 Tanggungharjo

This study aims to evaluate the strategy for strengthening students’ environmental care character through the Waste Bank program at SDN 4 Tanggungharjo, Grobogan Regency. This study employed a qualitative approach with a case study design. Data were collected through observation, interviews, and documentation involving the principal, teachers, students, Waste Bank management team, and other relevant school stakeholders. Data analysis was conducted through data condensation, data display, and conclusion drawing, while data credibility was established through source and technique triangulation. The findings indicate that the evaluation of the Waste Bank strategy was conducted through four interconnected mechanisms: periodic evaluation and collective reflection, financial transparency and program accountability, adaptive responses to problems, and program sustainability. Periodic evaluation enabled the school to identify problems related to student participation, waste sorting, and program management and to formulate corrective actions collaboratively. Financial transparency was maintained through individual student savings records and the main Waste Bank financial records, strengthening accountability and trust. Adaptive responses transformed students’ mistakes in waste sorting into learning opportunities through additional explanation, guidance, and repeated practice. Program sustainability was supported by institutional planning and budgeting, adequate facilities, continued student participation, and cooperation with external waste-management partners. Overall, the evaluation process demonstrated that the Waste Bank had developed beyond a waste-collection activity into a school-based mechanism for strengthening environmental care character. The evaluation functioned as a feedback mechanism connecting reflection, corrective action, behavioral reinforcement, accountability, and institutional sustainability. The study concludes that the effectiveness of a school-based Waste Bank strategy depends not only on the implementation of environmental activities but also on the school’s capacity to continuously evaluate, adapt, and institutionalize the program to support students’ environmental responsibility.

Nur Solikin, Endang Wuryandini, Widya Kusumaningsih · 0 citations
#reinforcement learning Open access Aug 2026

Adaptive Controllable Emergence in Multi-Task Air and Space Defense Systems: A Framework for Mission Reconfiguration, Resilient Coordination, Intelligent Decision-Making, and Dynamic Resource Allocation

Emergent collective intelligence provides an important theoretical and computational perspective for understanding how locally interacting agents can generate coordinated global behaviors that cannot be explained by the behavior of individual agents alone. In large-scale air and space defense systems, this property is particularly relevant because heterogeneous sensing, decision-making, communication, and execution resources must operate under dynamic environments, incomplete information, changing mission requirements, limited resources, and potentially degraded communication conditions. However, conventional controllable-emergence models generally assume relatively stable task structures and predefined interaction rules, which limits their adaptability when multiple tasks arrive concurrently or when the network topology and available resources change over time.This study proposes an Adaptive Controllable Emergence (ACE) framework for multi-task air and space defense systems. The proposed framework extends graph-based multi-agent modeling and multi-agent reinforcement learning by introducing four coupled mechanisms: dynamic mission reconfiguration, resilient coordination, intelligent distributed decision-making, and dynamic resource allocation. The system is represented as a time-varying interaction graph in which sensing, decision, and execution agents dynamically modify their relationships according to mission requirements and resource availability. A decentralized partially observable Markov decision process is employed to formulate local decision-making under incomplete information. A multi-objective reward function jointly considers mission completion, coordination quality, resource utilization, network resilience, adaptation cost, and decision latency. Furthermore, a mission-reconfiguration mechanism is introduced to enable the system to modify task-agent assignments when the task set, network topology, or resource state changes. The resulting framework transforms controllable emergence from a static rule-design problem into an adaptive optimization process in which microscopic policies continuously modify macroscopic system behavior. The proposed mathematical formulation provides a basis for analyzing emergence quality, adaptation speed, coordination robustness, resource efficiency, and convergence. A simulation framework is also developed for evaluating the proposed architecture under static, dynamic, multi-task, and communication-degradation scenarios. The framework is intended as a general computational model for studying adaptive coordination in large-scale multi-agent systems rather than as a platform-specific operational defense procedure.

Nor Ahmed Gujar · 0 citations
#reinforcement learning Dataset Open access Aug 2026

A Twin Delayed Deep Deterministic-based control method for a full vehicle semi-active suspension system

Through this system, users can input parameters for a vehicle semi-active suspension using a magneto-rheological damper (sprung mass, sprung mass centroid position parameters, unsprung mass, magneto-rheological damper model parameters, suspension spring stiffness, wheel equivalent spring stiffness, balance bar torsion spring stiffness, and ground excitation). Through this program's calculations, a vehicle suspension deep reinforcement learning controller model can be obtained to optimize the shock absorption effect at the vehicle's center of mass. Development hardware environment: CPU Intel Core i5 11600k, Memory: 32GB, Hard disk space: 4TB; Runtime hardware environment: CPU Intel Core i5 10400, Memory: 8GB, Hard disk space: 500GB. Development software environment: Windows 10; Runtime software environment: Windows 10.

Yongjun Wang, Xiaoming Wang, Gang Zhi et al. · 0 citations
#reinforcement learning Open access Aug 2026

智能电动汽车品牌声音DNA的动态生成——基于深度学习的主动声音设计系统研究

(ASD) widespread adoption of intelligent electric vehicles (IEVs) has eliminated traditional engine noise, raising pedestrian safety concerns and accelerating product homogeneity, which severely weakens brand auditory identity. Active sound design (ASD) has thus become a core technology for reshaping brand sound DNA and enhancing in‑cabin immersion and interaction. However, existing ASD systems largely rely on static concatenation of audio samples or fixed rule‑based parameter mapping, struggling to cope with complex driving scenarios and failing to deliver dynamic evolution or personalized expression of brand sound DNA. To address this limitation, this paper presents a deep learning‑based system for dynamic brand sound DNA generation and active sound design. We construct a multidimensional acoustic feature corpus of brand sound DNA and quantitatively decode the deep mapping between acoustic parameters and brand emotional semantics (e.g., sense of technology, sportiness, and luxury). A sequential audio generation model is designed by fusing a conditional generative adversarial network (cGAN) with a long short‑term memory (LSTM) network. An online adaptation mechanism built on deep reinforcement learning is introduced, which collects real‑time physiological feedback and subjective evaluations to dynamically fine‑tune the sound generation policy, enabling personalized evolution of the brand sound. This work provides an innovative technical pathway for auditory interaction design in IEVs and advances automotive acoustic engineering from static presets toward dynamic intelligent generation.

瑞珅 张 · 0 citations
#reinforcement learning Open access Aug 2026

Implementasi Nilai Karakter Pada Cerita Rakyat Lampung dalam Pembelajaran Bahasa Lampung SDN 3 Tegineneng Kelas V

This study aims to describe the implementation of character values through Lampung folklore in Lampung language learning among fifth-grade students at SDN 3 Tegineneng, Pesawaran. A descriptive qualitative approach was employed, involving three purposively selected informants: one Lampung language teacher as the main informant, supported by one school principal and one fifth-grade student. Data were collected through non-participant observation, semi-structured interviews, and documentation and analyzed using data condensation, data display, and conclusion drawing and verification. The findings show that character values were integrated into the introductory, core, and closing stages of learning. Nine character values were identified and actualized: religiosity, honesty, discipline, responsibility, hard work, caring, cooperation, patriotism, and respect for diversity. The teacher strengthened these values through exemplary behavior, habituation, discussion, motivation, appreciation, and reflection. Implementation was supported by school commitment, teacher competence, learning resources, a conducive classroom environment, and student participation, while limited instructional time, differences in Lampung language proficiency, limited learning media, and students' confidence constituted constraints. The main contribution of this study is the conceptual pattern of value internalization–actualization–reinforcement, which explains how moral values represented in folklore are understood by students, practiced through classroom activities, and reinforced through teachers' pedagogical actions.

Mela Santika, Chairul Amriyah, Yudesta Erfayliana · 0 citations
#reinforcement learning Review Open access Aug 2026

Reinforcement Learning in Wearable Robotic Systems for Orthopedic Rehabilitation: An Elbow-Focused Narrative Review

Orthopedic rehabilitation after upper-limb trauma increasingly emphasizes protected early motion, quantitative monitoring, and patient-specific assistance. This narrative review examines how reinforcement learning (RL) may contribute to wearable robotic systems for orthopedic rehabilitation, with emphasis on elbow-centered applications and transferable evidence from upper-limb exoskeletons, prosthetic-control studies, and musculoskeletal simulation. Literature from clinical and engineering sources was synthesized across four themes: clinical rationale, device platforms, control architecture, and translational readiness. The reviewed evidence suggests that the most plausible near-term platform is an externally worn powered orthosis rather than an implanted robotic joint. Across studies, RL is most defensible as a supervisory or personalization layer that adapts assistance within hard constraints on torque, speed, and range of motion, rather than as an unconstrained end-to-end controller. Multimodal sensing, especially combinations of electromyography, kinematics, and interaction sensing, appears more robust than any single intent channel. Digital twins and musculoskeletal simulators provide a practical substrate for offline training and conservative policy transfer, but fracture-specific clinical validation remains limited. Key barriers include alignment, comfort, safety governance, and the persistent gap between simulation and bedside deployment. Overall, the literature supports a staged translational strategy centered on hierarchical control, conservative safety design, and clinically bounded personalization.

Yash Jayeshbhai Patel, Bikram Bhakat, Abhijeet Patel et al. · 0 citations
#reinforcement learning Review Open access Aug 2026

Reinforcement Learning in Wearable Robotic Systems for Orthopedic Rehabilitation: An Elbow-Focused Narrative Review

Orthopedic rehabilitation after upper-limb trauma increasingly emphasizes protected early motion, quantitative monitoring, and patient-specific assistance. This narrative review examines how reinforcement learning (RL) may contribute to wearable robotic systems for orthopedic rehabilitation, with emphasis on elbow-centered applications and transferable evidence from upper-limb exoskeletons, prosthetic-control studies, and musculoskeletal simulation. Literature from clinical and engineering sources was synthesized across four themes: clinical rationale, device platforms, control architecture, and translational readiness. The reviewed evidence suggests that the most plausible near-term platform is an externally worn powered orthosis rather than an implanted robotic joint. Across studies, RL is most defensible as a supervisory or personalization layer that adapts assistance within hard constraints on torque, speed, and range of motion, rather than as an unconstrained end-to-end controller. Multimodal sensing, especially combinations of electromyography, kinematics, and interaction sensing, appears more robust than any single intent channel. Digital twins and musculoskeletal simulators provide a practical substrate for offline training and conservative policy transfer, but fracture-specific clinical validation remains limited. Key barriers include alignment, comfort, safety governance, and the persistent gap between simulation and bedside deployment. Overall, the literature supports a staged translational strategy centered on hierarchical control, conservative safety design, and clinically bounded personalization.

Yash Jayeshbhai Patel, Bikram Bhakat, Abhijeet Patel et al. · 0 citations
#reinforcement learning Review Open access Aug 2026

Worlds Without Handcrafted Limits: A Narrative Review of Procedural Content Generation from Perlin Noise to Machine Learning

Procedural content generation (PCG)---the algorithmic creation of game levels, terrain, quests, and rules---has evolved from a memory-saving trick into one of game development's most active research areas. This article presents a narrative review of the field's canonical line: Perlin's 1985 image synthesizer, the search-based taxonomy of Togelius and colleagues, the ACM survey of Hendrikx and colleagues, the Springer volume of Shaker, Togelius, and Nelson, the AI-and-games synthesis of Yannakakis and Togelius, the machine-learning turn of Summerville and colleagues, and the reinforcement-learning frontier of Khalifa and colleagues. The synthesis is organized around three themes: foundations, in which noise functions, grammars, and search established the generative toolbox; design, in which PCG met authorship---level design as search space, evolution as game designer, mixed-initiative tools; and learning, in which generative models trained on human-authored corpora opened PCGML and its controllability problem. It is concluded that PCG's history is the progressive relocation of authorship---from the asset to the generator---and that controllability is the field's central open problem.

Zen Revista, 10 GAME · 0 citations
#reinforcement learning Open access Aug 2026

AgentCreditBench: A Conformance-Test Suite for Turn-Level Credit Estimators

AgentCreditBench is a CPU-first conformance-test suite for turn-level credit assignment in agentic reinforcement learning. It compares GRPO, RLOO, GAE, GiGPO, Monte Carlo, and custom estimator outputs with exact policy advantages on tiny finite-horizon Markov decision processes, and separately evaluates the induced policy-gradient signal.

Yi Yan Ng · 0 citations

From tech blogs

See all →
MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.