This work presents RoboLDA, a Bayesian probabilistic model that decomposes VSR morphology generation into a four-level hierarchy:"task-robot-organ-voxel", and is trained via variational inference, which pioneers hierarchical generative modeling of robot morphology.
Abstract
Recent advances in robotics highlight hierarchical configurations of robot morphology, where multiple levels of functional substructures synergize to facilitate intelligent behaviors. This hierarchical perspective, while particularly advantageous for voxel-based soft robots (VSRs) to ease design and control complexities, is hindered by its heavy reliance on domain expertise. In this work, we address the following question: can we derive such hierarchical design principles solely from existing successful designs? We answer affirmatively by presenting RoboLDA, a Bayesian probabilistic model that decomposes VSR morphology generation into a four-level hierarchy:"task-robot-organ-voxel", and is trained via variational inference. Through extensive experiments on simulated VSRs, we verify the presence of consistent, intuitive hierarchical patterns underlying high-performing VSR designs and showcase RoboLDA's proficiency to extract and leverage these hierarchical priors for zero-shot robot design in unseen tasks. The generated designs, even without further optimization, achieve on average 106.4% of the optimized performance produced by evolutionary algorithms. Additionally, the organ structures inferred by RoboLDA serve as valid functional substructures, significantly enhancing synergistic motion control when integrated with modular control policies. Our work pioneers hierarchical generative modeling of robot morphology, offering a promising pathway towards more interpretable and generalizable development of embodied agents.
MISCO is developed, a novel evolutionary framework empowered by deep generative models to optimize VSR designs with theoretical guarantees that represents a step change towards more scalable and reliable soft robot development.
Jun-Ru Song, Huan Xiao, Yang Yang et al.· 0 citations
MorphIK is a flow-matching model that solves inverse kinematics for revolute-joint-based kinematic chains it has never seen during training, and allows learning and generalizing neural inverse kinematics for a multitude of known and unknown robots.
Lennart Clasmeier, Jan-Gerrit Habekost, C. Weber et al.· 0 citations
GIF, an agentic Generation framework for Interactive and Functional object compositions, recast this problem as disentangled reconstruction followed by relative pose recovery, revealing diversity scaling in both simulation and real-world deployment.
Long-Ji Xu, Zhi-Qi Zhang, Mi Yan et al.· 2 citations· ⚡2
This work introduces a physical intelligence framework in which distributed compliant interactions jointly reveal task-relevant information and organize manipulation behavior and demonstrates this principle through blind whole-arm grasping with a hybrid rigid-soft robotic arm that is equip with IMUs embedded directly w...
Chu-Han Zhang, E. Shahabi, K. Khomenko et al.· 0 citations
Wet-lab experimentation serves as the gold standard for hypothesis verification in scientific discovery; yet it is inherently labor-intensive, costly, and safety-critical. Embodied agents hold the promise of automating these tedious workflows, but their development is hindered by the scarcity of real-world training dat...
Chen-Xi Li, Hai-Yuan Wan, Rui Li et al.· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduJul 30, 2026
From feet to fingertips — we are teaching robots intelligent whole-body control, fine dexterity, and teamwork to complete a broad range of complex tasks.
With $2.1 million funding from Google.org, the open-source Public Transit Intelligence Hub will unify public transit monitoring, operations, and passenger communication.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.