Code accompanying the manuscript "Deep Reinforcement Learning for Drought-Resilient Operation of Agricultural Reservoirs" (submitted to Water Resources Management).
Qiang Gao, Hyunho Yang· Zenodo (CERN European Organi...· 0 citations
Through this system, users can input parameters for a vehicle semi-active suspension using a magneto-rheological damper (sprung mass, sprung mass centroid position parameters, unsprung mass, magneto-rheological damper model parameters, suspension spring stiffness, wheel equivalent spring stiffness, balance bar torsion spring stiffness, and ground excitation). Through this program's calculations, a vehicle suspension deep reinforcement learning controller model can be obtained to optimize the shock absorption effect at the vehicle's center of mass. Development hardware environment: CPU Intel Core i5 11600k, Memory: 32GB, Hard disk space: 4TB; Runtime hardware environment: CPU Intel Core i5 10400, Memory: 8GB, Hard disk space: 500GB. Development software environment: Windows 10; Runtime software environment: Windows 10.
Yongjun Wang, Xiaoming Wang, Gang Zhi et al.· Zenodo (CERN European Organi...· 0 citations
(ASD) widespread adoption of intelligent electric vehicles (IEVs) has eliminated traditional engine noise, raising pedestrian safety concerns and accelerating product homogeneity, which severely weakens brand auditory identity. Active sound design (ASD) has thus become a core technology for reshaping brand sound DNA and enhancing in‑cabin immersion and interaction. However, existing ASD systems largely rely on static concatenation of audio samples or fixed rule‑based parameter mapping, struggling to cope with complex driving scenarios and failing to deliver dynamic evolution or personalized expression of brand sound DNA. To address this limitation, this paper presents a deep learning‑based system for dynamic brand sound DNA generation and active sound design. We construct a multidimensional acoustic feature corpus of brand sound DNA and quantitatively decode the deep mapping between acoustic parameters and brand emotional semantics (e.g., sense of technology, sportiness, and luxury). A sequential audio generation model is designed by fusing a conditional generative adversarial network (cGAN) with a long short‑term memory (LSTM) network. An online adaptation mechanism built on deep reinforcement learning is introduced, which collects real‑time physiological feedback and subjective evaluations to dynamically fine‑tune the sound generation policy, enabling personalized evolution of the brand sound. This work provides an innovative technical pathway for auditory interaction design in IEVs and advances automotive acoustic engineering from static presets toward dynamic intelligent generation.
This study aims to describe the implementation of character values through Lampung folklore in Lampung language learning among fifth-grade students at SDN 3 Tegineneng, Pesawaran. A descriptive qualitative approach was employed, involving three purposively selected informants: one Lampung language teacher as the main informant, supported by one school principal and one fifth-grade student. Data were collected through non-participant observation, semi-structured interviews, and documentation and analyzed using data condensation, data display, and conclusion drawing and verification. The findings show that character values were integrated into the introductory, core, and closing stages of learning. Nine character values were identified and actualized: religiosity, honesty, discipline, responsibility, hard work, caring, cooperation, patriotism, and respect for diversity. The teacher strengthened these values through exemplary behavior, habituation, discussion, motivation, appreciation, and reflection. Implementation was supported by school commitment, teacher competence, learning resources, a conducive classroom environment, and student participation, while limited instructional time, differences in Lampung language proficiency, limited learning media, and students' confidence constituted constraints. The main contribution of this study is the conceptual pattern of value internalization–actualization–reinforcement, which explains how moral values represented in folklore are understood by students, practiced through classroom activities, and reinforced through teachers' pedagogical actions.
Mela Santika, Chairul Amriyah, Yudesta Erfayliana· Jurnal Ilmu Pendidikan dan S...· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Orthopedic rehabilitation after upper-limb trauma increasingly emphasizes protected early motion, quantitative monitoring, and patient-specific assistance. This narrative review examines how reinforcement learning (RL) may contribute to wearable robotic systems for orthopedic rehabilitation, with emphasis on elbow-centered applications and transferable evidence from upper-limb exoskeletons, prosthetic-control studies, and musculoskeletal simulation. Literature from clinical and engineering sources was synthesized across four themes: clinical rationale, device platforms, control architecture, and translational readiness. The reviewed evidence suggests that the most plausible near-term platform is an externally worn powered orthosis rather than an implanted robotic joint. Across studies, RL is most defensible as a supervisory or personalization layer that adapts assistance within hard constraints on torque, speed, and range of motion, rather than as an unconstrained end-to-end controller. Multimodal sensing, especially combinations of electromyography, kinematics, and interaction sensing, appears more robust than any single intent channel. Digital twins and musculoskeletal simulators provide a practical substrate for offline training and conservative policy transfer, but fracture-specific clinical validation remains limited. Key barriers include alignment, comfort, safety governance, and the persistent gap between simulation and bedside deployment. Overall, the literature supports a staged translational strategy centered on hierarchical control, conservative safety design, and clinically bounded personalization.
Yash Jayeshbhai Patel, Bikram Bhakat, Abhijeet Patel et al.· Zenodo (CERN European Organi...· 0 citations
Orthopedic rehabilitation after upper-limb trauma increasingly emphasizes protected early motion, quantitative monitoring, and patient-specific assistance. This narrative review examines how reinforcement learning (RL) may contribute to wearable robotic systems for orthopedic rehabilitation, with emphasis on elbow-centered applications and transferable evidence from upper-limb exoskeletons, prosthetic-control studies, and musculoskeletal simulation. Literature from clinical and engineering sources was synthesized across four themes: clinical rationale, device platforms, control architecture, and translational readiness. The reviewed evidence suggests that the most plausible near-term platform is an externally worn powered orthosis rather than an implanted robotic joint. Across studies, RL is most defensible as a supervisory or personalization layer that adapts assistance within hard constraints on torque, speed, and range of motion, rather than as an unconstrained end-to-end controller. Multimodal sensing, especially combinations of electromyography, kinematics, and interaction sensing, appears more robust than any single intent channel. Digital twins and musculoskeletal simulators provide a practical substrate for offline training and conservative policy transfer, but fracture-specific clinical validation remains limited. Key barriers include alignment, comfort, safety governance, and the persistent gap between simulation and bedside deployment. Overall, the literature supports a staged translational strategy centered on hierarchical control, conservative safety design, and clinically bounded personalization.
Yash Jayeshbhai Patel, Bikram Bhakat, Abhijeet Patel et al.· Zenodo (CERN European Organi...· 0 citations
Procedural content generation (PCG)---the algorithmic creation of game levels, terrain, quests, and rules---has evolved from a memory-saving trick into one of game development's most active research areas. This article presents a narrative review of the field's canonical line: Perlin's 1985 image synthesizer, the search-based taxonomy of Togelius and colleagues, the ACM survey of Hendrikx and colleagues, the Springer volume of Shaker, Togelius, and Nelson, the AI-and-games synthesis of Yannakakis and Togelius, the machine-learning turn of Summerville and colleagues, and the reinforcement-learning frontier of Khalifa and colleagues. The synthesis is organized around three themes: foundations, in which noise functions, grammars, and search established the generative toolbox; design, in which PCG met authorship---level design as search space, evolution as game designer, mixed-initiative tools; and learning, in which generative models trained on human-authored corpora opened PCGML and its controllability problem. It is concluded that PCG's history is the progressive relocation of authorship---from the asset to the generator---and that controllability is the field's central open problem.
Zen Revista, 10 GAME· Zenodo (CERN European Organi...· 0 citations
Code accompanying the manuscript "Deep Reinforcement Learning for Drought-Resilient Operation of Agricultural Reservoirs" (submitted to Water Resources Management).
Qiang Gao, Hyunho Yang· Zenodo (CERN European Organi...· 0 citations
AgentCreditBench is a CPU-first conformance-test suite for turn-level credit assignment in agentic reinforcement learning. It compares GRPO, RLOO, GAE, GiGPO, Monte Carlo, and custom estimator outputs with exact policy advantages on tiny finite-horizon Markov decision processes, and separately evaluates the induced policy-gradient signal.
Yi Yan Ng· Zenodo (CERN European Organi...· 0 citations
In an era defined by extreme Volatility, Uncertainty, Complexity, and Ambiguity (VUCA), artificial intelligence (AI) governance must transcend passive compliance checklists to become an embedded, adaptive socio-technical architecture. This paper proposes a triadic synthesis of Reinforcement Learning (RL), Generative AI (GenAI), and Cybersecurity, organized within a Seven-Layer Integrated Architecture spanning perception, cognition, adaptation, generation, protection, embodiment, and governance. Central to the framework is a formal isomorphism between Predictive Processing (PP) and Reinforcement Learning, in which both systems minimize prediction error through Bayesian updating (Friston, 2010; Friston et al., 2009). This isomorphism is operationalized through a safety-constrained objective function that treats variational free energy as a regularizer, mitigating the class of failures known as “reward hacking” (Laidlaw et al., 2025; Shihab et al., 2025; Skalse et al., 2022). Illustrative comparison of the Asynchronous Advantage Actor-Critic (A3C) algorithm against legacy Q-Learning suggests materially faster and more stable policy convergence under the resource-constrained, high-packet-loss conditions typical of emerging economies. By integrating the sub-Saharan African relational philosophy of Ubuntu/Unhu with global AI4People principles (Floridi et al., 2018; Van Norren, 2023; Yilma, 2025), the framework embeds explicit digital forensics workflows and blockchain-anchored chain-of-custody protocols (Atlam et al., 2024; Patil et al., 2024). The framework is further extended and empirically grounded through a twentyproject, four-cluster Edge-AI case portfolio spanning domestic safety, environmental intelligence, sustainable energy and agriculture, and healthcare accessibility in the Indian context, demonstrating the triadic architecture’s applicability from enterprise-scale governance to grassroots micro, small, and medium enterprise (MSME) innovation. This synthesis serves as a blueprint for organizations in the Southern African Development Community (SADC) and India to assert digital sovereignty, ensuring that autonomous systems are antifragile, context-sensitive, and designed for communal flourishing rather than extractive optimization.
Gabriel Kabanda· Zenodo (CERN European Organi...· 0 citations
Data for "Auditing Single-Agent Reinforcement Learning for EV Charging Assignment: A Protocol-Amended Comparison of Trained, Untrained, and Heuristic Policies" Raw seed-level data and campaign manifests for a benchmark of five agents (Random, Adaptive Heuristic, Q-Learning, DQN, Double DQN) on EV charging-station assignment, simulated on real Rabat and Tangier (Morocco) road networks in SUMO. Includes: campaign manifests with SHA-256 provenance, raw per-seed CSVs for the confirmatory trained/untrained diagnostic (three scenarios) and the legacy 210-run benchmark, and the JSON summaries behind the manuscript's result tables. Integrity verifiable via the included SHA-256 manifest. Preliminary, data-only deposit. Simulation event logs and trained model weights are not included in this version; available from the corresponding author on request.
This survey develops a unified taxonomy that systematically integrates learning paradigms, agent architectures, coordination mechanisms, deployment models, application domains, and evaluation frameworks from a common analytical perspective and provides a structured foundation for future advances in intelligent agent systems.
Elias Dritsas, M. Trigka· Evolutionary Intelligence· 0 citations
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.
MIT News · Artificial Intelligence· news.mit.eduAug 24, 2026
A new method for surgically removing training examples from a model reveals that as datasets grow, the link between what a model learns and what it produces dissolves.