Skip to content

Author

Zen Revista

33 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#reinforcement learning Open access Aug 2026

Key Developments in Reinforcement Learning in Robotics and Practical Implications: A Comprehensive Survey

This narrative review examines the current state of knowledge regarding Reinforcement Learning in Robotics within the broader context of Artificial Intelligence. We survey the theoretical foundations, methodological approaches, and key findings that have shaped the field, identifying major themes and tracing the evolution of ideas over time. The review synthesizes evidence from multiple research traditions and highlights both established conclusions and areas of ongoing debate. Particular attention is given to recent advances that have opened new avenues for investigation and to the practical implications of theoretical developments. We conclude with a discussion of the most promising directions for future research, emphasizing the importance of interdisciplinary collaboration and methodological innovation.

Zen Revista, 10 IA · 0 citations
#reinforcement learning Review Open access Aug 2026

Learning by Consequence: A Narrative Review of Reinforcement Learning from Thorndike's Law of Effect to Deep Q-Networks and AlphaGo

Reinforcement learning---learning what to do from reward and punishment rather than from instruction---unifies animal psychology, optimal control, and machine learning into one computational program, and its deep-learning era delivered the field's most visible artificial intelligence achievements. This article presents a narrative review of the canonical line: Thorndike's 1911 law of effect, Bellman's 1957 dynamic programming, Samuel's 1959 checkers player, Sutton's 1988 temporal-difference learning, Watkins and Dayan's 1992 Q-learning, Tesauro's 1995 TD-Gammon, Sutton and Barto's 1998 synthesis, Mnih and colleagues' 2015 Deep Q-Network, Silver and colleagues' 2016 AlphaGo and 2017 AlphaGo Zero, Lillicrap and colleagues' continuous control with DDPG, and Schulman and colleagues' 2017 proximal policy optimization. The synthesis is organized around three themes: foundations, in which the credit-assignment problem received formal solutions in value functions and temporal difference; scaling, in which function approximation, experience replay, and self-play converted tabular theory into high-dimensional control; and algorithmic consolidation, in which actor-critic methods and policy gradients stabilized practice. It is concluded that reinforcement learning's contribution is a general grammar of goal-directed learning---and that its open problems, sample efficiency and reward specification, define the frontier between artificial and natural intelligence.

Zen Revista, 10 IA · 0 citations
#reinforcement learning Review Open access Aug 2026

Learning by Consequence: A Narrative Review of Reinforcement Learning from Thorndike's Law of Effect to Deep Q-Networks and AlphaGo

Reinforcement learning---learning what to do from reward and punishment rather than from instruction---unifies animal psychology, optimal control, and machine learning into one computational program, and its deep-learning era delivered the field's most visible artificial intelligence achievements. This article presents a narrative review of the canonical line: Thorndike's 1911 law of effect, Bellman's 1957 dynamic programming, Samuel's 1959 checkers player, Sutton's 1988 temporal-difference learning, Watkins and Dayan's 1992 Q-learning, Tesauro's 1995 TD-Gammon, Sutton and Barto's 1998 synthesis, Mnih and colleagues' 2015 Deep Q-Network, Silver and colleagues' 2016 AlphaGo and 2017 AlphaGo Zero, Lillicrap and colleagues' continuous control with DDPG, and Schulman and colleagues' 2017 proximal policy optimization. The synthesis is organized around three themes: foundations, in which the credit-assignment problem received formal solutions in value functions and temporal difference; scaling, in which function approximation, experience replay, and self-play converted tabular theory into high-dimensional control; and algorithmic consolidation, in which actor-critic methods and policy gradients stabilized practice. It is concluded that reinforcement learning's contribution is a general grammar of goal-directed learning---and that its open problems, sample efficiency and reward specification, define the frontier between artificial and natural intelligence.

Zen Revista, 10 IA · 0 citations
#reinforcement learning Review Open access Aug 2026

Learning by Watching: A Narrative Review of Imitation Learning from ALVINN to Generative Adversarial Imitation

Imitation learning---the learning of behavior from demonstrations instead of rewards---moved from Pomerleau's ALVINN driving network and Schaal's humanoid route through Ng and Russell's inverse reinforcement learning, Abbeel and Ng's apprenticeship learning, and Ziebart's maximum entropy to the robot learning from demonstration surveys, Ross's DAgger, Ho and Ermon's generative adversarial imitation, Finn's guided cost learning, and the algorithmic perspective's syntheses. This article presents a narrative review of that arc's canonical line: Pomerleau's 1989 ALVINN, Schaal's 1999 humanoid question, Ng and Russell's 2000 inverse RL, Abbeel and Ng's 2004 apprenticeship learning, Billard and colleagues's 2008 handbook chapter, Ziebart and colleagues's 2008 maximum entropy, Argall and colleagues's 2009 survey, Ross, Gordon, and Bagnell's 2011 DAgger, Ho and Ermon's 2016 GAIL, Finn and colleagues's 2016 guided cost learning, Hussein and colleagues's 2017 survey, and Osa and colleagues's 2018 algorithmic perspective. The review is organized around three themes: the foundations, in which the driving network's demonstrations, the humanoid's question, and the inverse reward's recovery defined the field's two programs; the demonstration's surveys, in which the robot programming's handbook and the LfD's survey systematized the practice; and the deep era, in which the DAgger's covariate correction, the adversarial's discrimination, and the algorithmic perspective's synthesis unified the field. It is concluded that imitation learning is the reward's workaround---and that its arc is the demonstrator's knowledge's transfer from the human's steering to the policy's distributions.

Zen Revista, 10 IA · 0 citations
#reinforcement learning Review Open access Aug 2026

Learning Among Learners: A Narrative Review of Multi-Agent Reinforcement Learning from Markov Games to Deep Emergent Play

Multi-agent reinforcement learning---the learning of behavior when the environment's other agents learn too---moved from Tan's independent learners and Littman's Markov games framework through the cooperative dynamics' analyses and the surveys' question to the deep era's communication, actor-critics, value decompositions, and the large-scale emergent play of Capture the Flag. This article presents a narrative review of that arc's canonical line: Tan's 1993 independent versus cooperative agents, Littman's 1994 Markov games, Claus and Boutilier's 1998 cooperative dynamics, Hu and Wellman's 1998 framework, Shoham, Powers, and Grenager's 2007 question, Busoniu, Babuska, and De Schutter's 2008 survey, Foerster and colleagues' 2016 learning to communicate, Lowe and colleagues' 2017 multi-agent actor-critic, Sunehag and colleagues' 2018 value-decomposition networks, Rashid and colleagues' 2018 QMIX, Jaderberg and colleagues' 2019 3D multiplayer Capture the Flag, and Hernandez-Leal, Kartal, and Taylor's 2019 survey and critique. The review is organized around three themes: the foundational frames, in which the Markov game's formalization and the non-stationarity's, the coordination's, and the equilibrium's problems defined the field's difficulties; the theory's question, in which the surveys asked what learning among learners is for; and the deep era, in which the communications, the centralized critics, the monotonic factorizations, and the population-scale play made the multi-agent learning practical. It is concluded that multi-agent reinforcement learning is the non-stationarity's discipline---and that its deep era turned the other learners' obstruction into the curriculum's engine.

Zen Revista, 10 IA · 0 citations
#reinforcement learning Review Open access Aug 2026

Learning by Watching: A Narrative Review of Imitation Learning from ALVINN to Generative Adversarial Imitation

Imitation learning---the learning of behavior from demonstrations instead of rewards---moved from Pomerleau's ALVINN driving network and Schaal's humanoid route through Ng and Russell's inverse reinforcement learning, Abbeel and Ng's apprenticeship learning, and Ziebart's maximum entropy to the robot learning from demonstration surveys, Ross's DAgger, Ho and Ermon's generative adversarial imitation, Finn's guided cost learning, and the algorithmic perspective's syntheses. This article presents a narrative review of that arc's canonical line: Pomerleau's 1989 ALVINN, Schaal's 1999 humanoid question, Ng and Russell's 2000 inverse RL, Abbeel and Ng's 2004 apprenticeship learning, Billard and colleagues's 2008 handbook chapter, Ziebart and colleagues's 2008 maximum entropy, Argall and colleagues's 2009 survey, Ross, Gordon, and Bagnell's 2011 DAgger, Ho and Ermon's 2016 GAIL, Finn and colleagues's 2016 guided cost learning, Hussein and colleagues's 2017 survey, and Osa and colleagues's 2018 algorithmic perspective. The review is organized around three themes: the foundations, in which the driving network's demonstrations, the humanoid's question, and the inverse reward's recovery defined the field's two programs; the demonstration's surveys, in which the robot programming's handbook and the LfD's survey systematized the practice; and the deep era, in which the DAgger's covariate correction, the adversarial's discrimination, and the algorithmic perspective's synthesis unified the field. It is concluded that imitation learning is the reward's workaround---and that its arc is the demonstrator's knowledge's transfer from the human's steering to the policy's distributions.

Zen Revista, 10 IA · 0 citations
#reinforcement learning Review Open access Aug 2026

Learning Among Learners: A Narrative Review of Multi-Agent Reinforcement Learning from Markov Games to Deep Emergent Play

Multi-agent reinforcement learning---the learning of behavior when the environment's other agents learn too---moved from Tan's independent learners and Littman's Markov games framework through the cooperative dynamics' analyses and the surveys' question to the deep era's communication, actor-critics, value decompositions, and the large-scale emergent play of Capture the Flag. This article presents a narrative review of that arc's canonical line: Tan's 1993 independent versus cooperative agents, Littman's 1994 Markov games, Claus and Boutilier's 1998 cooperative dynamics, Hu and Wellman's 1998 framework, Shoham, Powers, and Grenager's 2007 question, Busoniu, Babuska, and De Schutter's 2008 survey, Foerster and colleagues' 2016 learning to communicate, Lowe and colleagues' 2017 multi-agent actor-critic, Sunehag and colleagues' 2018 value-decomposition networks, Rashid and colleagues' 2018 QMIX, Jaderberg and colleagues' 2019 3D multiplayer Capture the Flag, and Hernandez-Leal, Kartal, and Taylor's 2019 survey and critique. The review is organized around three themes: the foundational frames, in which the Markov game's formalization and the non-stationarity's, the coordination's, and the equilibrium's problems defined the field's difficulties; the theory's question, in which the surveys asked what learning among learners is for; and the deep era, in which the communications, the centralized critics, the monotonic factorizations, and the population-scale play made the multi-agent learning practical. It is concluded that multi-agent reinforcement learning is the non-stationarity's discipline---and that its deep era turned the other learners' obstruction into the curriculum's engine.

Zen Revista, 10 IA · 0 citations
#large language models Review Open access Aug 2026

Attention Is Foundational: A Narrative Review of the Transformer Architecture from Sequence-to-Sequence to Large Language Models

The Transformer architecture---built on attention rather than recurrence---redrew the landscape of natural language processing and became the substrate of contemporary artificial intelligence. This article presents a narrative review of the architecture's canonical line: Sutskever and colleagues' 2014 sequence-to-sequence learning, Bahdanau and colleagues' 2015 attention alignment, Vaswani and colleagues' 2017 Attention Is All You Need, Devlin and colleagues' 2019 BERT pretraining, Radford and colleagues' 2019 GPT-2, Brown and colleagues' 2020 GPT-3 and few-shot learning, Raffel and colleagues' 2020 T5 transfer, Dosovitskiy and colleagues' 2021 Vision Transformer, Bommasani and colleagues' 2021 foundation-model framing, Hoffmann and colleagues' 2022 Chinchilla scaling laws, Ouyang and colleagues' 2022 InstructGPT alignment, and Touvron and colleagues' 2023 LLaMA openness. The synthesis is organized around three themes: architecture, in which self-attention's parallel sequence processing replaced recurrence and enabled scale; scaling, in which pretraining on text plus parameter growth yielded emergent few-shot capability and then compute-optimal correction; and alignment and access, in which instruction tuning, RL from feedback, and open weights reshaped capability's deployment. It is concluded that the Transformer is machine learning's most consequential architecture to date---its attention mechanism the field's new inductive bias---and that scaling's economics and governance now define its trajectory.

Zen Revista, 10 IA · 0 citations
#large language models Review Open access Aug 2026

The Talk That Plays: A Narrative Review of Game Dialogue Systems from ELIZA to Personalized Agents

Dialogue systems---the machinery that lets games and agents converse---moved from Weizenbaum's 1966 ELIZA through the narrative engines of Si's Thespian, Mateas and Stern's Facade, and Cavazza's character-based storytelling, and the virtual humans of Swartout and Traum, to Rieser and Lemon's data-driven methodology, Serban's hierarchical networks, Vinyals and Le's neural conversational model, Zhang's personalized agents, Luger and Sellen's expectation gap, and Kiela's dodecathlon. This article presents a narrative review of that arc's canonical line: Weizenbaum's 1966 ELIZA, Cavazza, Charles, and Mead's 2002 storytelling, Si, Marsella, and Pynadath's 2005 Thespian, Mateas and Stern's 2005 Facade, Traum and colleagues's 2003 negotiation, Swartout and colleagues's 2006 virtual humans, Rieser and Lemon's 2011 methodology, Vinyals and Le's 2015 neural conversational model, Serban and colleagues's 2016 hierarchical networks, Luger and Sellen's 2016 expectation gap, Zhang and colleagues's 2018 personalization, and Kiela and colleagues's 2018 dodecathlon. The review is organized around three themes: the scripted origins, in which the pattern-matching's illusion, the drama managers's, and the virtual humans's architectures built the game's conversation; the statistical turn, in which the data-driven's methodology and the neural's hierarchies moved the dialogue from the rules to the corpora; and personalization and the language-model era, in which the personas's, the expectations's gap, and the benchmarks's carried the talk into the large models's future. It is concluded that dialogue systems are the game's social machinery---and that their arc is the conversation's engineering from the illusion's ELIZA to the persona's large models.

Zen Revista, 10 GAME · 0 citations
#large language models Review Open access Aug 2026

The Talk That Plays: A Narrative Review of Game Dialogue Systems from ELIZA to Personalized Agents

Dialogue systems---the machinery that lets games and agents converse---moved from Weizenbaum's 1966 ELIZA through the narrative engines of Si's Thespian, Mateas and Stern's Facade, and Cavazza's character-based storytelling, and the virtual humans of Swartout and Traum, to Rieser and Lemon's data-driven methodology, Serban's hierarchical networks, Vinyals and Le's neural conversational model, Zhang's personalized agents, Luger and Sellen's expectation gap, and Kiela's dodecathlon. This article presents a narrative review of that arc's canonical line: Weizenbaum's 1966 ELIZA, Cavazza, Charles, and Mead's 2002 storytelling, Si, Marsella, and Pynadath's 2005 Thespian, Mateas and Stern's 2005 Facade, Traum and colleagues's 2003 negotiation, Swartout and colleagues's 2006 virtual humans, Rieser and Lemon's 2011 methodology, Vinyals and Le's 2015 neural conversational model, Serban and colleagues's 2016 hierarchical networks, Luger and Sellen's 2016 expectation gap, Zhang and colleagues's 2018 personalization, and Kiela and colleagues's 2018 dodecathlon. The review is organized around three themes: the scripted origins, in which the pattern-matching's illusion, the drama managers's, and the virtual humans's architectures built the game's conversation; the statistical turn, in which the data-driven's methodology and the neural's hierarchies moved the dialogue from the rules to the corpora; and personalization and the language-model era, in which the personas's, the expectations's gap, and the benchmarks's carried the talk into the large models's future. It is concluded that dialogue systems are the game's social machinery---and that their arc is the conversation's engineering from the illusion's ELIZA to the persona's large models.

Zen Revista, 10 GAME · 0 citations
#small language model Review Open access Aug 2026

From Konigsberg's Bridges to Complex Networks: A Narrative Review of Graph Theory's Foundations, Landmark Theorems, and Applications

Graph theory began in 1736 as Euler's solution to the Konigsberg bridge problem and became mathematics' most versatile language for structure---in chemistry, sociology, computer science, and network science. This article presents a narrative review of the field's canonical line: Euler's 1736 paper, Kempe's 1879 four-color attempt, Konig's 1936 founding treatise, Erdos and Renyi's random graphs, Dirac's 1952 Hamiltonian theorem, the Appel--Haken four-color proof, Watts and Strogatz's small worlds, Barabási and Albert's scale-free networks, the Graph Minors program's completion by Robertson and Seymour, and the modern textbooks of Harary, Bondy and Murty, and Diestel. The synthesis is organized around three themes: foundations, in which graphs were formalized and their central problems---coloring, connectivity, traversability---defined; structure, in which random, small-world, and scale-free models quantified real networks; and depth, in which the Graph Minors program demonstrated the field's modern combinatorial power. It is concluded that graph theory's history is the refinement of a single idea---structure abstracted from substance---whose applications now feed back into the mathematics itself.

Zen Revista, 10 MATH · 0 citations
#small language model Review Open access Aug 2026

From Konigsberg's Bridges to Complex Networks: A Narrative Review of Graph Theory's Foundations, Landmark Theorems, and Applications

Graph theory began in 1736 as Euler's solution to the Konigsberg bridge problem and became mathematics' most versatile language for structure---in chemistry, sociology, computer science, and network science. This article presents a narrative review of the field's canonical line: Euler's 1736 paper, Kempe's 1879 four-color attempt, Konig's 1936 founding treatise, Erdos and Renyi's random graphs, Dirac's 1952 Hamiltonian theorem, the Appel--Haken four-color proof, Watts and Strogatz's small worlds, Barabási and Albert's scale-free networks, the Graph Minors program's completion by Robertson and Seymour, and the modern textbooks of Harary, Bondy and Murty, and Diestel. The synthesis is organized around three themes: foundations, in which graphs were formalized and their central problems---coloring, connectivity, traversability---defined; structure, in which random, small-world, and scale-free models quantified real networks; and depth, in which the Graph Minors program demonstrated the field's modern combinatorial power. It is concluded that graph theory's history is the refinement of a single idea---structure abstracted from substance---whose applications now feed back into the mathematics itself.

Zen Revista, 10 MATH · 0 citations