Building2Building (B2B), a large-scale suite of realistic HVAC control environments built on EnergyPlus, a state-of-the-art building simulator, is introduced, defining benchmark tasks targeting key open challenges in RL, including goal adaptation, dynamics adaptation, action-space shifts, and cross-domain transfer.
Abstract
Reinforcement learning (RL) has achieved strong results in control, yet learned policies remain brittle to changes in dynamics, action spaces, observation spaces, or goals, a critical limitation for real-world deployment. Existing benchmarks offer limited diversity and complexity, making it difficult to rigorously study transfer, multi-task learning, and meta-learning in RL. We introduce Building2Building (B2B), a large-scale suite of realistic Heating, Ventilation, and Air Conditioning (HVAC) control environments built on EnergyPlus, a state-of-the-art building simulator. B2B is fully compatible with the Gymnasium interface and features a parametric building generator, enabling the systematic generation of diverse building configurations with heterogeneous observation and action spaces. Based on this suite, we define benchmark tasks targeting key open challenges in RL, including goal adaptation, dynamics adaptation, action-space shifts, and cross-domain transfer. By providing a large-scale, diverse, and physically grounded testbed with standardized evaluation protocols, B2B enables systematic investigation of generalization and transfer in continuous control. Beyond advancing research on generalization in RL, this new benchmark also carries significant societal implications by enabling improved HVAC control at scale, one of the most energy-intensive systems in buildings.
This survey examines RL-based AD in modular and end-to-end pipelines and relates reported methods to task formulation and deployment evidence and examines deployment barriers, including safety, Sim2Real generalization, data efficiency, computation, embodied alignment, and evaluation readiness.
B. Shuai, Min Hua, Le-Tian Tao et al.· Communications in Transporta...· 0 citations
This work proposes QWM, a framework that leverages world models to perform test-time search over imagined trajectories on top of Q-learning to select high-value actions during both online rollouts and evaluation, and significantly outperforms strong prior state-of-the-art methods on both sample efficiency and performan...
Perry Dong, Yue-Ru Jia, Chelsea Finn et al.· 2 citations
This work instantiates Real-Time EXPO-FT, an RL framework for finetuning real-time VLA policies that meets the real-time control requirements of dynamic real-world manipulation, demonstrating rapid, sample-efficient adaptation to challenging real-world dynamics.
Perry Dong, Kuo-Han Hung, D. Sadigh et al.· 0 citations
Robust reinforcement learning (RRL) aims to develop a robust policy that maintains stable performance across diverse environments characterized by an uncertainty set. This set consists of perturbed environments derived from a nominal (training) environment that generates samples, thereby capturing potential discrepanci...
Ukjo Hwang, Songnam Hong· IEEE Transactions on Neural...· 0 citations
This work uses Sample-based Model Predictive Control entirely in simulation as an automated, rapidly tunable expert to generate massive offline datasets and validate the robustness of this sim-to-real framework by successfully deploying complex loco-manipulation skills across different morphologies.
Martin Schuck, Maks Sorokin, S. Manni et al.· 1 citation
WM-R1 is the first reinforcement learning framework that trains mobile GUI agents with world models instead of real environments, eliminating the need for real-environment interaction, supports massively parallelized and step-level granularized trajectory generation grounded in world models, and introduces a multi-dime...
Yu Han, Tianwen Qian· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.