Real-world Earth observation (EO) agents must translate high-level scientific questions into executable workflows to acquire observations, prepare data, perform domain computations, and derive conclusions from runtime evidence. Existing EO agents typically start from supplied observations, while benchmarks typically pr...
Zhu-Tao Lv, Chen-Hao Dang, Yi-Cong Feng et al.· 0 citations
While diffusion base models such as GPT-Image-2 and Nano-Banana exhibit remarkable visual expressiveness, their end-to-end generation inherently yields flattened bitmaps with error-prone text, precluding layer-wise post-editing. Conversely, code-based visual generation via Coding Agents provides precise layout control...
Jun-Yan Ye, Wei Liu, Dongzhi Jiang et al.· 0 citations
This work introduces GeoMoE, a sparse mixture-of-experts dual encoder that decouples global multi-scale representation learning from local hierarchical search, and introduces VIGOR-M, a four-city benchmark with an explicit parent--child satellite hierarchy and held-out half-step galleries for single-resolution, cross-r...
Ruijie Fan, Jun-Yan Ye, Qi Zhu et al.· 0 citations
VideoCoCo, an agentic dual-engine framework in which executable Blender code serves as a process-level chain of thought, demonstrates that executable code provides an effective, controllable, and inspectable intermediate representation for physically consistent video generation.
Haodong Li, Tianfei Ren, Xiaoxiao Ma et al.· arXiv.org· 8 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.