NavGen: Visual Generative Models as a Scalable Data Engine for Embodied 3D Navigation
NavGen is introduced, a text-to-video data generation pipeline that produces diverse vision-language navigation episodes across indoor and outdoor scenes and a style-diversification method that scales up long-tail data that is difficult and costly to collect.
Xi-Jie Huang, Yong-Yang Wan, Cheng-Bin Dong et al.
· 0 citations