Skip to content

OmniStyle-INR: Universal and Multimodal Style Transfer for INRs

Jul 2026 · arXiv.org · Vol abs/2607.16362 · 0 citations · 57 references
Computer Science

TL;DR

OmniStyle-INR is introduced, a novel framework that leverages network-based continuous representations as a truly universal domain that performs high-quality style transfer across all visual modalities, guided seamlessly by both text prompts and visual exemplars.

Abstract

Style transfer remains a fundamental and highly important task across various data modalities, enabling creative manipulation conditioned by both reference images and textual descriptions. Recently, methods utilizing Gaussian Splatting have emerged as a unified representation for 2D images, video, 3D scenes, and 4D dynamics. However, representing videos and 2D images with Gaussian Splatting is structurally sub-optimal for dense continuous domains. The number of required Gaussians often approaches the total number of pixels, raising questions about the actual utility of such a representation for these specific modalities. In contrast, Implicit Neural Representations have established themselves as a much more popular and natural choice across all these data domains. Implicit Neural Representations naturally provide significant advantages, including data compression, inherent capabilities for super resolution, and seamless integration with deep generative models. To this end, we introduce OmniStyle-INR, a novel framework that leverages network-based continuous representations as a truly universal domain. Our approach successfully performs high-quality style transfer across all visual modalities, guided seamlessly by both text prompts and visual exemplars.

View source

Similar papers

Preprint Sep 2026

Isotropic Embedding Perturbations for Robust Vision Language Encoders

Aether is introduced, a simple plug-in method that applies diffusion-style random perturbations in the embedding space via controlled alpha-mixing, specifically designed to provide isotropic regularization that remains semantically consistent.

Hyesong Choi, Daeun Kim, Song Park et al. · 0 citations
Preprint Aug 2026

Exploring the Design Space of Representation Learning for Audio Transformations

This framework produces both a transformation embedding and a processed-audio embedding, and it finds that the two play complementary roles: distance-based tasks favor the former, while probe-based tasks favor the latter.

Sungho Lee, Marco A. Mart'inez-Ram'irez, Junghyun Koo et al. · 0 citations
Open access 2026

StyleSmith: A Signature Style Transfer Framework With Parametric and Temporal Attention

Artistic style transfer, which renders content images in the visual style of a reference artwork while preserving semantic structure, is a fundamental problem in computer vision and AIGC, with broad applications in digital art, commercial design, and creative media. Despite progress driven by diffusion-based approaches...

Yu Cheng, Anucha Pangkesorn, Jia-Rui Wu et al. · 0 citations

FlowLess: Controlling Abstract Image Generation

A novel self-supervised framework that enables granular control over image generation through a visual abstraction set that provides a richer, more flexible paradigm for creative design compared to state-of-the-art baselines across diverse styles and compositions is introduced.

Amir Hertz, Noah Snavely, Google DeepMind · 0 citations
Preprint Aug 2026

PoseAdapter: Dual-Stream 2.5D Controllable Image Generation for Complex Multi-Object Scenes

PoseAdapter, a lightweight framework for high-fidelity 2.5D controllable image generation, and a Context-Aware Dual-Stream Representation, to resolve the generative trade-off between strict instance isolation and global coherence.

Yu-Feng Chi, Hui-Min Ma, Fan Gao et al. · 0 citations
#small language model Preprint Oct 2026

Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation

This work presents Omni-Embed-Mini, a 0.9B-parameter model that maps text, speech, audio, images, video, and visually-rich documents into a single shared cosine space without updating any text-side parameter, and is competitive with the closed gemini-embedding-2, edging ahead of it on the overall-modality average.

Mohammed Irfan Kurpath, Jaseel Muhammad Kaithakkodan, Sahal Shaji Mullappilly et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.