Category
reinforcement learning
461 papers
Antifragile Intelligence: A Triadic Framework for AI Governance, Digital Forensics, and Sovereignty in Emerging Economies
In an era defined by extreme Volatility, Uncertainty, Complexity, and Ambiguity (VUCA), artificial intelligence (AI) governance must transcend passive compliance checklists to become an embedded, adaptive socio-technical architecture. This paper proposes a triadic synthesis of Reinforcement Learning (RL), Generative AI (GenAI), and Cybersecurity, organized within a Seven-Layer Integrated Architecture spanning perception, cognition, adaptation, generation, protection, embodiment, and governance. Central to the framework is a formal isomorphism between Predictive Processing (PP) and Reinforcement Learning, in which both systems minimize prediction error through Bayesian updating (Friston, 2010; Friston et al., 2009). This isomorphism is operationalized through a safety-constrained objective function that treats variational free energy as a regularizer, mitigating the class of failures known as “reward hacking” (Laidlaw et al., 2025; Shihab et al., 2025; Skalse et al., 2022). Illustrative comparison of the Asynchronous Advantage Actor-Critic (A3C) algorithm against legacy Q-Learning suggests materially faster and more stable policy convergence under the resource-constrained, high-packet-loss conditions typical of emerging economies. By integrating the sub-Saharan African relational philosophy of Ubuntu/Unhu with global AI4People principles (Floridi et al., 2018; Van Norren, 2023; Yilma, 2025), the framework embeds explicit digital forensics workflows and blockchain-anchored chain-of-custody protocols (Atlam et al., 2024; Patil et al., 2024). The framework is further extended and empirically grounded through a twentyproject, four-cluster Edge-AI case portfolio spanning domestic safety, environmental intelligence, sustainable energy and agriculture, and healthcare accessibility in the Indian context, demonstrating the triadic architecture’s applicability from enterprise-scale governance to grassroots micro, small, and medium enterprise (MSME) innovation. This synthesis serves as a blueprint for organizations in the Southern African Development Community (SADC) and India to assert digital sovereignty, ensuring that autonomous systems are antifragile, context-sensitive, and designed for communal flourishing rather than extractive optimization.
Primary Field Cybernetics
Primary Field Cybernetics: From Spectral-Phase Flow to a Symformic Theory of Control This article develops Primary Field Cybernetics (PFC) as a cybernetic theory derived from the ontology of Symformism and the field architecture of Dynamical Informational Field Theory (DIFT). Symformism treats an enduring form not as a static object but as a relational organization that preserves its identity through continuous change, exchange, perturbation, and reorganization. DIFT provides a physical research architecture for this intuition through the complex spectral-phase Primary Organizational Field, organized phase current, structural memory, adaptive access geometry, mobility, impedance, and dynamostatic persistence. From these relations, the article derives a domain-neutral cybernetic architecture: Primary Field → organized flow → retained history → differential impedance → differential accessibility → viable future action → control. The central proposition is that control cannot be reduced to selecting an action from a fixed repertoire. A system’s history can alter the practical accessibility of its future responses. Consequently, an enduring actant participates in reorganizing the conditions under which its own later regulation remains possible. On this basis, PFC distinguishes state control, access control, reflexive control, and relational metacontrol, and introduces the concept of accessible variety: regulatory variety understood not merely as nominally available responses, but as responses that remain practically reachable within relevant constraints of cost, delay, compatibility, and viability. The paper situates this proposal in relation to classical and second-order cybernetics, including Maxwell, Wiener, Ashby, Beer, von Foerster, Maturana and Varela, and gives particular attention to Marian Mazur’s theory of autonomous systems. It also distinguishes the proposed architecture from active inference, allostasis, reinforcement learning, eligibility traces, synaptic plasticity, and meta-learning. A deliberately limited numerical model is included as an illustration of one consequence of the theory rather than as validation of PFC or DIFT. It examines whether different histories can produce different future response accessibility under an otherwise matched challenge. Causal ablation and a single-timescale trace equivalence control are used to delimit what this example does and does not establish. The article concludes with a falsification programme based on matched-state, divergent-history experiments, in which present observables are matched, accessibility is measured before a decisive response, identical perturbations are applied, and the proposed access mechanism is selectively manipulated. Candidate applications include physical, biological, neural, artificial, and institutional systems. The resulting formulation shifts the fundamental cybernetic question from “How does a system correct its present state?” toward “How does an enduring organization preserve and reorganize the conditions under which viable future action remains accessible?” Keywords: Primary Field Cybernetics; Symformism; DIFT; Primary Organizational Field; spectral-phase flow; phase current; dynamostasis; impedance; access geometry; accessible variety; cybernetics; control; memory; resilience.
The Cognitive Familiarity Supremacy Theory (CFST)
The Cognitive Familiarity Supremacy Theory (CFST) proposes that a substantial portion of human certainty, ideological attachment, collective identity formation, and perceived superiority emerges not primarily from objective rational evaluation, but from repeated familiarity encoding mechanisms operating within subconscious cognitive architectures.This framework argues that repeated environmental exposure, social conditioning, emotional reinforcement, identity fusion, symbolic repetition, and institutional amplification collectively construct familiarity-driven epistemic structures that are frequently mistaken for objective truth, rational certainty, or universal superiority. The theory synthesizes and mathematically formalizes principles from cognitive neuroscience, psychology, sociology, political theory, philosophy of mind, epistemology, systems theory, information theory, complexity science, behavioral economics, evolutionary biology, communication studies, artificial intelligence, anthropology, cybernetics, and cultural theory into a unified explanatory framework.CFST introduces a comprehensive causal chain model:Repeated Exposure \rightarrow Subconscious Encoding \rightarrow Identity Fusion \rightarrow Emotional Reinforcement \rightarrow Bias Formation \rightarrow Perceived Superiority.The theory proposes that human cognition operates through familiarity-weighted interpretive systems, where the subjective sensation of certainty often emerges from accumulated familiarity intensity rather than objective verification.The framework further integrates: Bayesian epistemology, predictive processing, Hebbian learning, social identity theory, information entropy, network propagation, algorithmic amplification, memetic evolution, cultural conditioning, political hegemony, and neurocognitive attractor-state dynamics.The theory also develops: formal mathematical models, belief topology equations, dynamic systems formulations, stochastic familiarity propagation systems, agent-based ideological simulations, network-theoretical belief diffusion structures, and computational cognitive equilibrium equations.At the civilizational level, CFST proposes that societies are partially constructed upon collectively reinforced familiarity architectures rather than purely objective truth systems. At the individual level, it explains ideological rigidity, nationalism, fanaticism, cultural supremacy perception, identity-protective cognition, and epistemic polarization.Finally, the theory proposes that genuine epistemic liberation requires conscious disruption of subconscious familiarity monopolies through critical reasoning, diversity exposure, meta-cognitive awareness, and reflective epistemological reconstruction.
Reach audiences
Advertise in front of researchers, engineers, and readers.
DIArc Foundational Note v0.1 — Minimum Claim Edition
Abstract The rapid development of artificial intelligence has significantly increased the availability of information, analytical capability, and machine-assisted reasoning. However, greater access to information does not necessarily produce better decisions. In many organizational contexts, the emerging bottleneck is no longer information acquisition, but the human and organizational capacity to determine what information is sufficient, when analysis should stop, when a decision should be made, and how outcomes should improve future judgment. This Foundational Note introduces Decision Intelligence Architecture (DIArc) as an architectural framework for Human–AI collaborative decision systems. DIArc is based on a central proposition: in the AI era, competitive advantage increasingly depends not on maximizing information, but on maximizing the rate at which high-quality decisions generate learning and improve judgment, under explicit constraints on information consumption and decision cycles. The architecture is organized into four theoretical layers. First, the Capability Inversion Hypothesis describes a structural shift in which information, knowledge, and analysis become increasingly abundant while judgment, commitment, execution, and learning become comparatively scarce capabilities. Second, Identity-driven Information Consumption (IDIC) describes a decision failure mechanism in which continued information consumption may serve identity reinforcement rather than decision improvement. Third, the Decision Constraint Architecture, comprising Decision Information Budget (DIB) and Decision Cycle Budget (DCB), introduces explicit constraints on information consumption and analytical iteration. Fourth, High-quality Decision Velocity (HQDV) describes the performance objective of accelerating completed high-quality decision loops, while Judgment Evolution Rate (JER) represents the longer-term evolutionary objective of improving judgment through outcome-based learning. This note constitutes the initial public disclosure of the DIArc architecture and establishes its theoretical baseline for subsequent research and branch concepts.
Learning by Consequence: A Narrative Review of Reinforcement Learning from Thorndike's Law of Effect to Deep Q-Networks and AlphaGo
Reinforcement learning---learning what to do from reward and punishment rather than from instruction---unifies animal psychology, optimal control, and machine learning into one computational program, and its deep-learning era delivered the field's most visible artificial intelligence achievements. This article presents a narrative review of the canonical line: Thorndike's 1911 law of effect, Bellman's 1957 dynamic programming, Samuel's 1959 checkers player, Sutton's 1988 temporal-difference learning, Watkins and Dayan's 1992 Q-learning, Tesauro's 1995 TD-Gammon, Sutton and Barto's 1998 synthesis, Mnih and colleagues' 2015 Deep Q-Network, Silver and colleagues' 2016 AlphaGo and 2017 AlphaGo Zero, Lillicrap and colleagues' continuous control with DDPG, and Schulman and colleagues' 2017 proximal policy optimization. The synthesis is organized around three themes: foundations, in which the credit-assignment problem received formal solutions in value functions and temporal difference; scaling, in which function approximation, experience replay, and self-play converted tabular theory into high-dimensional control; and algorithmic consolidation, in which actor-critic methods and policy gradients stabilized practice. It is concluded that reinforcement learning's contribution is a general grammar of goal-directed learning---and that its open problems, sample efficiency and reward specification, define the frontier between artificial and natural intelligence.
Student Behavior Recognition and Intervention Methods in Intelligent Classrooms Based on Deep Reinforcement Learning
In response to the problem that traditional classroom student behavior analysis relies on manual observation and is difficult to achieve real-time, precise and personalized intervention. This paper constructs an end-to-end intelligent classroom intervention system based on deep reinforcement learning. This system adopts the Multi-Modal Fusion Spatio-Temporal Graph Convolutional Network (MM-ST-GCN), integrating visual skeletons, seat pressure and classroom interaction data, to achieve fine-grained and high-precision recognition of students' classroom behaviors. It models the classroom environment as a Partially Observable Markov Decision Process (POMDP) and uses the improved Soft Actor Critic (SAC) algorithm to generate the optimal intervention strategy that takes into account the learning benefits of students and the intervention costs of teachers. Experiments on the MMAct dataset show that the proposed behavior recognition model achieves 93.8% accuracy and a macro-average F1 score of 0.925. Simulation experiments indicate that the system has the potential to improve students' concentration and reduce distraction behavior. However, the above results are derived from the simulated environment and need to be further verified in the real classroom.
Distribution Estimation Algorithm for Cloud Manufacturing Scheduling Optimization
To address the resource scheduling problem in complex production environments, this study proposes a production scheduling model based on Estimation of Distribution Algorithm.The model constructs a probability model using spatial distribution, evaluates the scheduling population based on high-quality individuals, introduces an archive mechanism to enhance solution diversity, and combines Deep Reinforcement Learning and Tabu Search algorithm for global optimization.It achieves adaptive production scheduling optimization under dynamically changing resources.In testing experiments, the model achieves an accuracy of 95.11 % in sample classification prediction tasks.The computational load and number of parameters for production data processing are 664.8FLOPs and 90.54 M, respectively.The scheduling delay rate and resource utilization are 4.39 % and 97.96 %, significantly outperforming comparison models.These results indicate that the model provides stable and efficient production scheduling optimization and multi-constraint conditions, offering reliable algorithm support for production scheduling in cloud-based networked manufacturing environments.
Learning by Watching: A Narrative Review of Imitation Learning from ALVINN to Generative Adversarial Imitation
Imitation learning---the learning of behavior from demonstrations instead of rewards---moved from Pomerleau's ALVINN driving network and Schaal's humanoid route through Ng and Russell's inverse reinforcement learning, Abbeel and Ng's apprenticeship learning, and Ziebart's maximum entropy to the robot learning from demonstration surveys, Ross's DAgger, Ho and Ermon's generative adversarial imitation, Finn's guided cost learning, and the algorithmic perspective's syntheses. This article presents a narrative review of that arc's canonical line: Pomerleau's 1989 ALVINN, Schaal's 1999 humanoid question, Ng and Russell's 2000 inverse RL, Abbeel and Ng's 2004 apprenticeship learning, Billard and colleagues's 2008 handbook chapter, Ziebart and colleagues's 2008 maximum entropy, Argall and colleagues's 2009 survey, Ross, Gordon, and Bagnell's 2011 DAgger, Ho and Ermon's 2016 GAIL, Finn and colleagues's 2016 guided cost learning, Hussein and colleagues's 2017 survey, and Osa and colleagues's 2018 algorithmic perspective. The review is organized around three themes: the foundations, in which the driving network's demonstrations, the humanoid's question, and the inverse reward's recovery defined the field's two programs; the demonstration's surveys, in which the robot programming's handbook and the LfD's survey systematized the practice; and the deep era, in which the DAgger's covariate correction, the adversarial's discrimination, and the algorithmic perspective's synthesis unified the field. It is concluded that imitation learning is the reward's workaround---and that its arc is the demonstrator's knowledge's transfer from the human's steering to the policy's distributions.
PBFT-CG-MARL
PBFT-CG-MAPPO is a consensus-conditioned multi-agent reinforcement learning framework that embeds a Practical Byzantine Fault Tolerance (PBFT) three-phase commit protocol inside the CTDE-MAPPO training loop as an additive consensus loss with zero initialization, complemented by an entropy floor that prevents premature policy collapse. Across three cooperative environments (MPE spread, SMAClite 5m_vs_6m, VMAS UAV coverage) with five seeds each, PBFT-CG-MAPPO reduces cross-seed return variance by 1.9–47.4× compared to MAPPO and produces zero catastrophic seeds. Under Byzantine injection (f = 1), the PBFT quorum maintains stable consensus rates against random and adversarial attacks, where disabling the consensus layer causes up to 93% return degradation—demonstrating that consensus provides critical protection under active attacks while remaining non-interfering in clean environments. The experimental results yield three design principles for safe consensus conditioning in MARL: replace only dissenting actions (never overwrite consenting agents), condition via additive loss with zero initialization, and enforce an entropy floor. The framework transfers from 4-agent cooperative navigation to 12-agent permafrost monitoring without modification, scaling the f < n/3 tolerance bound automatically. The codebase includes six algorithm baselines (MAPPO, MADDPG, QMIX, CommNet, TarMAC), three Byzantine attack types, ablation studies, cross-environment evaluation, and publication-quality figure generation scripts.
Learning Among Learners: A Narrative Review of Multi-Agent Reinforcement Learning from Markov Games to Deep Emergent Play
Multi-agent reinforcement learning---the learning of behavior when the environment's other agents learn too---moved from Tan's independent learners and Littman's Markov games framework through the cooperative dynamics' analyses and the surveys' question to the deep era's communication, actor-critics, value decompositions, and the large-scale emergent play of Capture the Flag. This article presents a narrative review of that arc's canonical line: Tan's 1993 independent versus cooperative agents, Littman's 1994 Markov games, Claus and Boutilier's 1998 cooperative dynamics, Hu and Wellman's 1998 framework, Shoham, Powers, and Grenager's 2007 question, Busoniu, Babuska, and De Schutter's 2008 survey, Foerster and colleagues' 2016 learning to communicate, Lowe and colleagues' 2017 multi-agent actor-critic, Sunehag and colleagues' 2018 value-decomposition networks, Rashid and colleagues' 2018 QMIX, Jaderberg and colleagues' 2019 3D multiplayer Capture the Flag, and Hernandez-Leal, Kartal, and Taylor's 2019 survey and critique. The review is organized around three themes: the foundational frames, in which the Markov game's formalization and the non-stationarity's, the coordination's, and the equilibrium's problems defined the field's difficulties; the theory's question, in which the surveys asked what learning among learners is for; and the deep era, in which the communications, the centralized critics, the monotonic factorizations, and the population-scale play made the multi-agent learning practical. It is concluded that multi-agent reinforcement learning is the non-stationarity's discipline---and that its deep era turned the other learners' obstruction into the curriculum's engine.
Integrasi Teori Perilaku Belajar dan Rekonstruksi Penguatan dalam Pembelajaran Pendidikan Agama Islam
In educational psychology, the theory of student learning behavior is one of the most important topics. The diversity of theories on this subject demonstrates the numerous ways teachers can easily understand students’ conditions and master how to teach the lessons they will teach. The purpose of this study is to focus on explaining the nature, definition, and principles of learning, the theory of learning behavior, the relationship between learning and teaching, and the application of learning principles in Islamic Religious Education (PAI) learning. The method used in this study was a literature review with a qualitative approach. The research findings indicate that behavioral theories, particularly the Contiguity theory, Connectionism theory, Classical Conditioning theory, Operant Conditioning theory, and Social Learning theory, can be easily applied to various aspects of Islamic Education learning by creating conducive associations and positive reinforcement. In PAI learning, the application of learning principles is crucial to increasing learning effectiveness, particularly in achieving one of the goals of Islamic religious education, namely shaping student behavior in accordance with religious teachings.
From tech blogs
See all →Transfer learning for genomic prediction in underrepresented populations
General Science
Mapping global methane emissions from space with deep learning
Climate & Sustainability
Looking beyond natural sequences
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.
Generating scenarios for extreme events, without extreme data
A new algorithm learns to anticipate the unprecedented scenarios that critical infrastructure and global supply chains are least prepared for.
When AI art has no author: Study finds generated images often can’t be traced to training data
A new method for surgically removing training examples from a model reveals that as datasets grow, the link between what a model learns and what it produces dissolves.
What We Learned by Reproducing 2,200 papers from ICML
We’re on a journey to advance and democratize artificial intelligence through open source and open science.