Skip to content

Author

Álvaro Serra-Gómez

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Sep 2026

Reinforcement Learning with Verifiable Rewards for Small Search Agents

Reinforcement Learning with Verifiable Rewards (RLVR) performs well on problems with clear rewards, such as mathematics and coding, but whether it also works where the reward is less clear remains open. The reason-over-search recipe applies RLVR to open-domain question answering, where retrieval grounds the answer and...

Gaurisankar Jayadas, A. Plaat, Álvaro Serra-Gómez et al. · 0 citations
#machine learning Preprint Sep 2026

GEM-MPC: Balancing Exploration and Exploitation through Expert-Guided Planning

Effective exploration in high-dimensional continuous control remains a central challenge in reinforcement learning. Planning-based methods address this by combining online planning with learned policies and value functions, but their components can become misaligned during training: learned sampling policies may diverg...

Álvaro Serra-Gómez, Thomas Moerland · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.