This position paper argues that, in such settings, the objective of offline RL should extend beyond immediate deployment and instead prioritize learning adaptive policy priors: policies that preserve the capacity to improve during subsequent interaction through memory, exploration, and self-correction.
Tian-Wei Ni, Vineet Jain, Akash Karthikeyan et al.· 1 citation
Building2Building (B2B), a large-scale suite of realistic HVAC control environments built on EnergyPlus, a state-of-the-art building simulator, is introduced, defining benchmark tasks targeting key open challenges in RL, including goal adaptation, dynamics adaptation, action-space shifts, and cross-domain transfer.
Vincent Taboga, Justine Veilleux, Doseok Jang et al.· arXiv.org· 0 citations
Tabular foundation models (TFMs) learn to fill in tables the way language models fill in text, and tables are arguably the format in which most physical measurement arrives. Did they learn any physics in the process? They are Bayesian by construction, so the question is what their prior contains. We probe it directly,...
Wassim Tenachi, Y. Hezaveh, L. P. Levasseur et al.· 0 citations
A novel unified operator is introduced that combines several regularized RL operators into a general framework that better targets peakier sampling distributions and is named trajectory general mellowmax (TGM), which is shown to identify higher quality, diverse candidates than baselines in both synthetic and real-world...
Marco Jiralerspong, Esther Derman, Danilo Vucetic et al.· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.