This work constructs 11,911 text--staging pairs from 2,328 figurative paintings by reconstructing SMPL bodies, estimating low-frequency illumination, recovering camera parameters, and pairing each scene with ArtEmis descriptions, demonstrating the feasibility of generating editable, emotionally conditioned 3D staging references from text.
Abstract
Artists coordinate human pose, illumination, and camera placement to convey narrative and emotion, but existing generative methods typically model these elements independently. We introduce text-to-editable 3D staging, a task that jointly generates human poses, a dominant light, and a camera configuration from an affective description. We construct 11,911 text--staging pairs from 2,328 figurative paintings by reconstructing SMPL bodies, estimating low-frequency illumination, recovering camera parameters, and pairing each scene with ArtEmis descriptions. We train a flow-matching transformer that supports variable numbers of figures and produces multiple staging alternatives for each prompt. On held-out descriptions, the model achieves 32.2\% retrieval R@1, compared with 16.6\% for CLIP-based nearest-neighbor retrieval, while approximately preserving corpus-level diversity. These results demonstrate the feasibility of generating editable, emotionally conditioned 3D staging references from text.
Emotion-aware artistic image generation requires a model to satisfy semantic content, artistic style, and target emotion simultaneously. The key challenge is that artistic captions conflate these axes into underspecified free-form text, making fine-grained visual attributes such as brushwork, composition, and tonal atm...
Qianqian Tang, Jia-Yi Gao, Ting Lei et al.· 0 citations
TSLP, a text-driven and style-transferable pipeline for indoor furniture layout generation that first synthesizes an interior image from textual prompts via a diffusion-based model, followed by a style-transfer module for aesthetic customization, significantly enhancing both usability and adaptability.
Zhi-Guo Xu· Poster Volume 0007 The 2026...· 0 citations
Results demonstrate that immersive WebXR representations substantially strengthen user engagement and understanding, offering a scalable pathway for employer branding and digital recruitment, and point to interactivity as a key avenue for future work.
Louis Burk, Christoph Scharnagl, Uwe Wienkop· International Conference on...· 0 citations
LumiTokens is a framework that formulates 3D relighting as a direct transformation on latent scene tokens, without explicit 3D representations, rendering equations, or physics-based decomposition, and achieves comparable or superior relighting quality to other methods and supports progressive, composable lighting edits...
In recent years, text-to-image (T2I) generation models have made substantial progress, particularly in visual realism and the expression of prompt semantics. However, a key difficulty remains: how to evaluate generated results automatically in a way that is both comprehensive and interpretable, while still being practi...
A novel self-supervised framework that enables granular control over image generation through a visual abstraction set that provides a richer, more flexible paradigm for creative design compared to state-of-the-art baselines across diverse styles and compositions is introduced.
Amir Hertz, Noah Snavely, Google DeepMind· 0 citations
Known for his clear and elegant writing style, Bertsekas shaped fields from control and optimization to large-scale computation and artificial intelligence.