Skip to content

Author

Zen Revista

73 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#graph neural networks Review Open access Aug 2026

Learning on Edges: A Narrative Review of Graph Neural Networks from Recursive Networks to Geometric Deep Learning

Graph neural networks---learning over relational, irregular structure by passing messages between nodes---generalized deep learning's grids to the graph: molecules, social networks, knowledge bases, and the web. This article presents a narrative review of that arc's canonical line: Sperduti and Starita's 1997 structure classification, Gori, Monfardini, and Scarselli's 2005 graph-domain learning, Scarselli and colleagues' 2009 GNN model, Bruna and colleagues' 2014 spectral networks, Defferrard and colleagues' 2016 localized filtering, Kipf and Welling's 2017 graph convolutions, Gilmer and colleagues' 2017 message passing, Hamilton, Ying, and Leskovec's 2017 GraphSAGE, Velickovic and colleagues' 2018 attention, Ying and colleagues' 2019 GNNExplainer, Wu and colleagues' 2021 comprehensive survey, and Bronstein and colleagues' 2021 geometric deep learning. The synthesis is organized around three themes: recursion, in which state propagation over nodes founded learning on graphs; convolution, in which spectral theory and message passing gave the graph a deep architecture; and geometry, in which attention, pooling, explainability, and symmetry made the network general. It is concluded that the GNN is deep learning's relational settlement---convolution's invariance learned from graph geometry rather than grid regularity---and that its message-passing abstraction is one of machine learning's cleanest unifications.

Zen Revista, 10 IA · 0 citations
#reinforcement learning Review Open access Aug 2026

Worlds Without Handcrafted Limits: A Narrative Review of Procedural Content Generation from Perlin Noise to Machine Learning

Procedural content generation (PCG)---the algorithmic creation of game levels, terrain, quests, and rules---has evolved from a memory-saving trick into one of game development's most active research areas. This article presents a narrative review of the field's canonical line: Perlin's 1985 image synthesizer, the search-based taxonomy of Togelius and colleagues, the ACM survey of Hendrikx and colleagues, the Springer volume of Shaker, Togelius, and Nelson, the AI-and-games synthesis of Yannakakis and Togelius, the machine-learning turn of Summerville and colleagues, and the reinforcement-learning frontier of Khalifa and colleagues. The synthesis is organized around three themes: foundations, in which noise functions, grammars, and search established the generative toolbox; design, in which PCG met authorship---level design as search space, evolution as game designer, mixed-initiative tools; and learning, in which generative models trained on human-authored corpora opened PCGML and its controllability problem. It is concluded that PCG's history is the progressive relocation of authorship---from the asset to the generator---and that controllability is the field's central open problem.

Zen Revista, 10 GAME · 0 citations
#reinforcement learning Open access Aug 2026

Key Developments in Reinforcement Learning in Robotics and Practical Implications: A Comprehensive Survey

This narrative review examines the current state of knowledge regarding Reinforcement Learning in Robotics within the broader context of Artificial Intelligence. We survey the theoretical foundations, methodological approaches, and key findings that have shaped the field, identifying major themes and tracing the evolution of ideas over time. The review synthesizes evidence from multiple research traditions and highlights both established conclusions and areas of ongoing debate. Particular attention is given to recent advances that have opened new avenues for investigation and to the practical implications of theoretical developments. We conclude with a discussion of the most promising directions for future research, emphasizing the importance of interdisciplinary collaboration and methodological innovation.

Zen Revista, 10 IA · 0 citations
#reinforcement learning Review Open access Aug 2026

Worlds Without Handcrafted Limits: A Narrative Review of Procedural Content Generation from Perlin Noise to Machine Learning

Procedural content generation (PCG)---the algorithmic creation of game levels, terrain, quests, and rules---has evolved from a memory-saving trick into one of game development's most active research areas. This article presents a narrative review of the field's canonical line: Perlin's 1985 image synthesizer, the search-based taxonomy of Togelius and colleagues, the ACM survey of Hendrikx and colleagues, the Springer volume of Shaker, Togelius, and Nelson, the AI-and-games synthesis of Yannakakis and Togelius, the machine-learning turn of Summerville and colleagues, and the reinforcement-learning frontier of Khalifa and colleagues. The synthesis is organized around three themes: foundations, in which noise functions, grammars, and search established the generative toolbox; design, in which PCG met authorship---level design as search space, evolution as game designer, mixed-initiative tools; and learning, in which generative models trained on human-authored corpora opened PCGML and its controllability problem. It is concluded that PCG's history is the progressive relocation of authorship---from the asset to the generator---and that controllability is the field's central open problem.

Zen Revista, 10 GAME · 0 citations
#reinforcement learning Open access Aug 2026

Key Developments in Reinforcement Learning in Robotics and Practical Implications: A Comprehensive Survey

This narrative review examines the current state of knowledge regarding Reinforcement Learning in Robotics within the broader context of Artificial Intelligence. We survey the theoretical foundations, methodological approaches, and key findings that have shaped the field, identifying major themes and tracing the evolution of ideas over time. The review synthesizes evidence from multiple research traditions and highlights both established conclusions and areas of ongoing debate. Particular attention is given to recent advances that have opened new avenues for investigation and to the practical implications of theoretical developments. We conclude with a discussion of the most promising directions for future research, emphasizing the importance of interdisciplinary collaboration and methodological innovation.

Zen Revista, 10 IA · 0 citations
#reinforcement learning Review Open access Aug 2026

Learning by Consequence: A Narrative Review of Reinforcement Learning from Thorndike's Law of Effect to Deep Q-Networks and AlphaGo

Reinforcement learning---learning what to do from reward and punishment rather than from instruction---unifies animal psychology, optimal control, and machine learning into one computational program, and its deep-learning era delivered the field's most visible artificial intelligence achievements. This article presents a narrative review of the canonical line: Thorndike's 1911 law of effect, Bellman's 1957 dynamic programming, Samuel's 1959 checkers player, Sutton's 1988 temporal-difference learning, Watkins and Dayan's 1992 Q-learning, Tesauro's 1995 TD-Gammon, Sutton and Barto's 1998 synthesis, Mnih and colleagues' 2015 Deep Q-Network, Silver and colleagues' 2016 AlphaGo and 2017 AlphaGo Zero, Lillicrap and colleagues' continuous control with DDPG, and Schulman and colleagues' 2017 proximal policy optimization. The synthesis is organized around three themes: foundations, in which the credit-assignment problem received formal solutions in value functions and temporal difference; scaling, in which function approximation, experience replay, and self-play converted tabular theory into high-dimensional control; and algorithmic consolidation, in which actor-critic methods and policy gradients stabilized practice. It is concluded that reinforcement learning's contribution is a general grammar of goal-directed learning---and that its open problems, sample efficiency and reward specification, define the frontier between artificial and natural intelligence.

Zen Revista, 10 IA · 0 citations
#reinforcement learning Review Open access Aug 2026

Learning by Consequence: A Narrative Review of Reinforcement Learning from Thorndike's Law of Effect to Deep Q-Networks and AlphaGo

Reinforcement learning---learning what to do from reward and punishment rather than from instruction---unifies animal psychology, optimal control, and machine learning into one computational program, and its deep-learning era delivered the field's most visible artificial intelligence achievements. This article presents a narrative review of the canonical line: Thorndike's 1911 law of effect, Bellman's 1957 dynamic programming, Samuel's 1959 checkers player, Sutton's 1988 temporal-difference learning, Watkins and Dayan's 1992 Q-learning, Tesauro's 1995 TD-Gammon, Sutton and Barto's 1998 synthesis, Mnih and colleagues' 2015 Deep Q-Network, Silver and colleagues' 2016 AlphaGo and 2017 AlphaGo Zero, Lillicrap and colleagues' continuous control with DDPG, and Schulman and colleagues' 2017 proximal policy optimization. The synthesis is organized around three themes: foundations, in which the credit-assignment problem received formal solutions in value functions and temporal difference; scaling, in which function approximation, experience replay, and self-play converted tabular theory into high-dimensional control; and algorithmic consolidation, in which actor-critic methods and policy gradients stabilized practice. It is concluded that reinforcement learning's contribution is a general grammar of goal-directed learning---and that its open problems, sample efficiency and reward specification, define the frontier between artificial and natural intelligence.

Zen Revista, 10 IA · 0 citations
#reinforcement learning Review Open access Aug 2026

Learning by Watching: A Narrative Review of Imitation Learning from ALVINN to Generative Adversarial Imitation

Imitation learning---the learning of behavior from demonstrations instead of rewards---moved from Pomerleau's ALVINN driving network and Schaal's humanoid route through Ng and Russell's inverse reinforcement learning, Abbeel and Ng's apprenticeship learning, and Ziebart's maximum entropy to the robot learning from demonstration surveys, Ross's DAgger, Ho and Ermon's generative adversarial imitation, Finn's guided cost learning, and the algorithmic perspective's syntheses. This article presents a narrative review of that arc's canonical line: Pomerleau's 1989 ALVINN, Schaal's 1999 humanoid question, Ng and Russell's 2000 inverse RL, Abbeel and Ng's 2004 apprenticeship learning, Billard and colleagues's 2008 handbook chapter, Ziebart and colleagues's 2008 maximum entropy, Argall and colleagues's 2009 survey, Ross, Gordon, and Bagnell's 2011 DAgger, Ho and Ermon's 2016 GAIL, Finn and colleagues's 2016 guided cost learning, Hussein and colleagues's 2017 survey, and Osa and colleagues's 2018 algorithmic perspective. The review is organized around three themes: the foundations, in which the driving network's demonstrations, the humanoid's question, and the inverse reward's recovery defined the field's two programs; the demonstration's surveys, in which the robot programming's handbook and the LfD's survey systematized the practice; and the deep era, in which the DAgger's covariate correction, the adversarial's discrimination, and the algorithmic perspective's synthesis unified the field. It is concluded that imitation learning is the reward's workaround---and that its arc is the demonstrator's knowledge's transfer from the human's steering to the policy's distributions.

Zen Revista, 10 IA · 0 citations
#reinforcement learning Review Open access Aug 2026

Learning Among Learners: A Narrative Review of Multi-Agent Reinforcement Learning from Markov Games to Deep Emergent Play

Multi-agent reinforcement learning---the learning of behavior when the environment's other agents learn too---moved from Tan's independent learners and Littman's Markov games framework through the cooperative dynamics' analyses and the surveys' question to the deep era's communication, actor-critics, value decompositions, and the large-scale emergent play of Capture the Flag. This article presents a narrative review of that arc's canonical line: Tan's 1993 independent versus cooperative agents, Littman's 1994 Markov games, Claus and Boutilier's 1998 cooperative dynamics, Hu and Wellman's 1998 framework, Shoham, Powers, and Grenager's 2007 question, Busoniu, Babuska, and De Schutter's 2008 survey, Foerster and colleagues' 2016 learning to communicate, Lowe and colleagues' 2017 multi-agent actor-critic, Sunehag and colleagues' 2018 value-decomposition networks, Rashid and colleagues' 2018 QMIX, Jaderberg and colleagues' 2019 3D multiplayer Capture the Flag, and Hernandez-Leal, Kartal, and Taylor's 2019 survey and critique. The review is organized around three themes: the foundational frames, in which the Markov game's formalization and the non-stationarity's, the coordination's, and the equilibrium's problems defined the field's difficulties; the theory's question, in which the surveys asked what learning among learners is for; and the deep era, in which the communications, the centralized critics, the monotonic factorizations, and the population-scale play made the multi-agent learning practical. It is concluded that multi-agent reinforcement learning is the non-stationarity's discipline---and that its deep era turned the other learners' obstruction into the curriculum's engine.

Zen Revista, 10 IA · 0 citations
#reinforcement learning Review Open access Aug 2026

Learning by Watching: A Narrative Review of Imitation Learning from ALVINN to Generative Adversarial Imitation

Imitation learning---the learning of behavior from demonstrations instead of rewards---moved from Pomerleau's ALVINN driving network and Schaal's humanoid route through Ng and Russell's inverse reinforcement learning, Abbeel and Ng's apprenticeship learning, and Ziebart's maximum entropy to the robot learning from demonstration surveys, Ross's DAgger, Ho and Ermon's generative adversarial imitation, Finn's guided cost learning, and the algorithmic perspective's syntheses. This article presents a narrative review of that arc's canonical line: Pomerleau's 1989 ALVINN, Schaal's 1999 humanoid question, Ng and Russell's 2000 inverse RL, Abbeel and Ng's 2004 apprenticeship learning, Billard and colleagues's 2008 handbook chapter, Ziebart and colleagues's 2008 maximum entropy, Argall and colleagues's 2009 survey, Ross, Gordon, and Bagnell's 2011 DAgger, Ho and Ermon's 2016 GAIL, Finn and colleagues's 2016 guided cost learning, Hussein and colleagues's 2017 survey, and Osa and colleagues's 2018 algorithmic perspective. The review is organized around three themes: the foundations, in which the driving network's demonstrations, the humanoid's question, and the inverse reward's recovery defined the field's two programs; the demonstration's surveys, in which the robot programming's handbook and the LfD's survey systematized the practice; and the deep era, in which the DAgger's covariate correction, the adversarial's discrimination, and the algorithmic perspective's synthesis unified the field. It is concluded that imitation learning is the reward's workaround---and that its arc is the demonstrator's knowledge's transfer from the human's steering to the policy's distributions.

Zen Revista, 10 IA · 0 citations
#reinforcement learning Review Open access Aug 2026

Learning Among Learners: A Narrative Review of Multi-Agent Reinforcement Learning from Markov Games to Deep Emergent Play

Multi-agent reinforcement learning---the learning of behavior when the environment's other agents learn too---moved from Tan's independent learners and Littman's Markov games framework through the cooperative dynamics' analyses and the surveys' question to the deep era's communication, actor-critics, value decompositions, and the large-scale emergent play of Capture the Flag. This article presents a narrative review of that arc's canonical line: Tan's 1993 independent versus cooperative agents, Littman's 1994 Markov games, Claus and Boutilier's 1998 cooperative dynamics, Hu and Wellman's 1998 framework, Shoham, Powers, and Grenager's 2007 question, Busoniu, Babuska, and De Schutter's 2008 survey, Foerster and colleagues' 2016 learning to communicate, Lowe and colleagues' 2017 multi-agent actor-critic, Sunehag and colleagues' 2018 value-decomposition networks, Rashid and colleagues' 2018 QMIX, Jaderberg and colleagues' 2019 3D multiplayer Capture the Flag, and Hernandez-Leal, Kartal, and Taylor's 2019 survey and critique. The review is organized around three themes: the foundational frames, in which the Markov game's formalization and the non-stationarity's, the coordination's, and the equilibrium's problems defined the field's difficulties; the theory's question, in which the surveys asked what learning among learners is for; and the deep era, in which the communications, the centralized critics, the monotonic factorizations, and the population-scale play made the multi-agent learning practical. It is concluded that multi-agent reinforcement learning is the non-stationarity's discipline---and that its deep era turned the other learners' obstruction into the curriculum's engine.

Zen Revista, 10 IA · 0 citations
#large language models Review Open access Aug 2026

Attention Is Foundational: A Narrative Review of the Transformer Architecture from Sequence-to-Sequence to Large Language Models

The Transformer architecture---built on attention rather than recurrence---redrew the landscape of natural language processing and became the substrate of contemporary artificial intelligence. This article presents a narrative review of the architecture's canonical line: Sutskever and colleagues' 2014 sequence-to-sequence learning, Bahdanau and colleagues' 2015 attention alignment, Vaswani and colleagues' 2017 Attention Is All You Need, Devlin and colleagues' 2019 BERT pretraining, Radford and colleagues' 2019 GPT-2, Brown and colleagues' 2020 GPT-3 and few-shot learning, Raffel and colleagues' 2020 T5 transfer, Dosovitskiy and colleagues' 2021 Vision Transformer, Bommasani and colleagues' 2021 foundation-model framing, Hoffmann and colleagues' 2022 Chinchilla scaling laws, Ouyang and colleagues' 2022 InstructGPT alignment, and Touvron and colleagues' 2023 LLaMA openness. The synthesis is organized around three themes: architecture, in which self-attention's parallel sequence processing replaced recurrence and enabled scale; scaling, in which pretraining on text plus parameter growth yielded emergent few-shot capability and then compute-optimal correction; and alignment and access, in which instruction tuning, RL from feedback, and open weights reshaped capability's deployment. It is concluded that the Transformer is machine learning's most consequential architecture to date---its attention mechanism the field's new inductive bias---and that scaling's economics and governance now define its trajectory.

Zen Revista, 10 IA · 0 citations