The results suggest that many psychopathology-relevant aspects may be interpreted as bounded cognitive systems operating under modern-ancestral environmental mismatch, positioning ERDM as a key computational cognitive tool that can be extended to other studies.
Abstract
This study introduces the evolutionarily recurrent decision model (ERDM), a computational reinforcement learning framework designed to examine how evolutionary mismatch, bounded rationality, and satisficing contribute to adaptive and maladaptive behavior. ERDM simulates agents across evolutionary recurrent environments, including threat, prey/goal-pursuits, and alliances. Agents learn through competing rewards abstracted from survival metrics. A validity study under varying adverse childhood experiences demonstrates that distinct adaptive and maladaptive strategies, such as learned helplessness, avoidance, healthy relationships, and aggression, emerge naturally without being hardwired. These results align with empirical literature, showcasing ecological validity. The results suggest that many psychopathology-relevant aspects may be interpreted as bounded cognitive systems operating under modern-ancestral environmental mismatch, positioning ERDM as a key computational cognitive tool that can be extended to other studies.
Decision-making in natural and artificial systems is often shaped by past interactions, as individuals adjust their behavior accordingly. A persistent challenge lies in identifying successful strategies and understanding how memory influences their performance, especially when individuals’ optimal actions conflict with collective outcomes. Most theoretical studies in evolutionary dynamics have focused on memory-1 strategies, and recent studies indicate that extending memory can improve strategies’ performance. Among these, all-or-none (AoNK) strategies perform well at short memory. Our analysis reveals, however, that their effectiveness deteriorates as memory extends, showing the inherent limitations of sole coordination. Building on this, we propose the adaptive coordination strategy, combining coordination and tolerance to enhance adaptability and stability. Our theoretical analysis precisely characterizes its behavior and key properties. This strategy outperforms classic memory-1 strategies and achieves optimal outcomes, providing a framework for understanding complex history-based behaviors beyond traditional memory-n models.
Feipeng Zhang, Te Wu, Long Wang· Science China Information Sc...· 0 citations
Real-world decision-making rarely occurs with perfect information. Instead, individuals must constantly weigh potential rewards against the probability of adverse outcomes.1 Failures of this process can lead to maladaptive decisions associated with reduced lifetime success, and numerous psychiatric disorders such as gambling addictions, bulimia nervosa, and substance use disorder.2,3 The neural computations that facilitate inference about the landscape of potential outcomes remain unclear, but are thought to occur in distributed frontotemporal circuits.4 Here we used deep reinforcement learning agents to predict distinct behavioral strategies and their underlying neural population dynamics during a risky decision-making task. Across a range of training conditions, deep reinforcement learning agents separated into strategies marked by either overly cautious exploration of the reward contingency space or a high-performing, risk-adaptive Bimodal strategy. The internal dynamics of high-performing Bimodal agents formed low-dimensional representations that segregated safe and risky states. In contrast, the cautious exploration agents were associated with more skewed and entangled neural representations. We found remarkably similar dynamical representations and their associated behavioral strategies in neuronal ensemble recordings from human epilepsy patients performing a similar risky decision-making task. These results reveal the structure of dynamical computations that underlie inferences about uncertain outcomes and their associated behavioral strategies.
T. Price, A. Liu, Rhiannon L. Cowan et al.· bioRxiv· 0 citations
This integrative review proposes that the reduction in uncertainty may function as a candidate primary reinforcer in humans. The proposed mechanism is positive prediction error coding by the mesolimbic dopamine system within nucleus accumbens and lateral habenula (NAc-LHb) opponent-process circuitry, gated and refined by later cerebellar and cortical additions. Because the dopaminergic signal responds to prediction resolution rather than to what is predicted, certainty operates as a reinforcer across sensory, motor, social, and propositional domains. The case is developed through three converging arguments: functional, mechanistic, and phylogenetic. The review first documents the conservation of NAc-LHb prediction error circuitry from lampreys to mammals and the matching law as a quantitative description of behavioral allocation. It then traces cerebellar emergence, forward modeling, and an ultrafast disynaptic cerebellum-to-NAc pathway that modulates reward valuation before cortical evaluation completes. It further traces cortical evolution into executive and language-supporting structures, integrating recent causal and human electroencephalogram (EEG) evidence for hierarchical cortical predictive coding as a layer above subcortical and cerebellar contributions. Finally, it addresses system integration and articulates a synthesis between free energy minimization and behavior-analytic motivating operations, proposing that motivating operations are the behavioral implementation of free energy gradients across domains, such that the list of primary reinforcers is principled rather than arbitrary. A falsifiable empirical test is proposed: pairing an arbitrary neutral cue with the resolution of a probabilistic prediction should produce conditioning curves quantitatively similar to those produced by pairing with food or water. On balance, uncertainty reduction is not yet established as a primary reinforcer in the same sense as food or water, but it is a strong candidate for a trans-domain motivational process that can recruit reward, aversion, precision-weighting, and belief-updating systems. Implications are noted for healthy belief revision, neurodevelopmental and psychiatric conditions, and contemporary digital information environments.
A. Lincoln· Academia Neuroscience and Br...· 0 citations
Large brains are metabolically costly, and associations with changing environments do not imply they evolved there, as the Cognitive Buffer Hypothesis (CBH) would suggest. They may instead evolve in stable conditions and later facilitate colonization of changing environments. Using neuro-evolution in an artificial seasonal foraging task, we compared agents evolving exclusively in changing environments to agents first evolved in static environments before transitioning. Results show that larger neural networks in dynamic environments arise mainly from prior static evolution, achieving superior performance under unpredictable changes. Our results challenge strict CBH predictions, provide agent-based (computational) support for a colonization-based account and highlight the role of evolutionary history in brain size evolution.
Sian Heesom-Green, Jonathan P. Shock, Geoff S. Nitschke· Proceedings of the Genetic a...· 0 citations