Skip to content
Preprint

S$^2$GS: Structured Sparse Gaussian Streaming for Efficient Free-Viewpoint Video Reconstruction on Edge-IoT Devices

Aug 2026 · 0 citations · 51 references
Computer Science

TL;DR

Structured Sparse Gaussian Streaming (S$^2$GS), an FVV reconstruction framework that exploits structure-aware temporal sparsity to selectively update Gaussian residuals, enabling efficient streaming without compromising visual fidelity is proposed.

Abstract

Streaming reconstruction of Free-Viewpoint Videos (FVVs) supports immersive Internet of Things (IoT) services, such as telepresence and digital twin visualization. Existing methods suffer from high per-frame optimization time and large storage footprints, limiting deployment on resource-constrained Edge-IoT devices. To address these challenges, we propose Structured Sparse Gaussian Streaming (S$^2$GS), an FVV reconstruction framework that exploits structure-aware temporal sparsity to selectively update Gaussian residuals, enabling efficient streaming without compromising visual fidelity. In the spatial domain, a streaming octree hierarchically organizes Gaussian residuals, capturing spatial correlations that guide residual updates. In the temporal domain, a structured gating mechanism, comprising hierarchical feature propagation (HFP) and Gumbel-Sigmoid sampling, converts hierarchical dynamic cues into sparse residual update decisions under differentiable optimization. A multi-level discrete scheme is further adopted to provide fine-grained control over residual updates while preserving intricate dynamic details. Extensive experiments across consumer GPUs, industrial edge IoT devices, and a physical telepresence testbed demonstrate that S$^2$GS consistently reduces per-frame optimization time and storage footprint while maintaining competitive visual quality. Compared with QUEEN, S$^2$GS reduces per-frame optimization time by 59% and storage costs by 85% on an RTX 4090 GPU. On the Jetson AGX Orin, S$^2$GS delivers the highest rendering throughput (60+ FPS) and the lowest energy consumption among the evaluated methods, demonstrating its potential for deployment in resource-constrained systems.

View source

Similar papers

Preprint Sep 2026

DecoGS: Adaptive Static-Dynamic Decoupling of 3D Gaussians for Free-Viewpoint Video Streaming

Streaming 3D reconstruction demands both speed and temporal fidelity, goals that existing methods undermine by updating every Gaussian every frame, even in static regions. We present DecoGS, a method for efficient online training of 3D Gaussians from streaming videos. Unlike prior methods that update the entire scene i...

Idil Sulo, Alexey Supikov, Ilke Demir et al. · 0 citations
Preprint Aug 2026

QuARC-GS: Quantized Anchored Residual Coding for Compact Dynamic Scene Streaming with Gaussian Splatting

Quantized Anchored Residual Coding Gaussian Streaming (QuARC-GS), a quantization-aware 4D scene optimization framework for online dynamic scene reconstruction that achieves ultra-high compression while maintaining reconstruction speed and quality, is proposed.

V. Nguyen, Yu-Chen Wang, Kyung Chul Lee et al. · 0 citations
Aug 2026

S $^{2}$ Q-VDiT$^+$: Accurate Quantized Video Diffusion Transformer with Multi-Resolution Sampling and Structural Distillation.

Large-scale video diffusion models (V-DMs) have achieved remarkable text-to-video generation quality, yet their massive computational complexity makes deployment costly. Post-Training Quantization (PTQ) offers an appealing route to accelerate inference without retraining, but existing diffusion PTQ methods remain fragi...

Wei-Lun Feng, Chuan-Guang Yang, Haotong Qin et al. · 4 citations
Preprint Sep 2026

Spatial-Temporal Multi-scale Network for Screen Content Video Quality Enhancement

Different from natural videos, Screen Content Videos (SCVs) are characterized by abrupt motion, scene switches, and high-frequency details such as text and graphics. Conventional video enhancement methods, which rely heavily on temporal continuity, often suffer from performance degradation when processing SCVs due to t...

Zi-Yin Huang, Sik-Ho Tsang, Xin Qin et al. · 0 citations
Preprint Aug 2026

GS$^{2}$CI: Robust Gaussian Splatting For Snapshot Compressive Imaging via Large Vision Model Priors

This work proposes a novel framework that reconstructs high-quality 3D scenes from a single SCI measurement by leveraging 3D Gaussian Splatting and the powerful priors of large-scale vision foundation models (VFMs) and introduces Opacity-Guided Splitting and Growth Regulation (OSGR), an SCI-specific densification strat...

Yan-Ming Yang, Chen-Xi Song, Ping Wang et al. · 0 citations
Open access 2026

MSP-Edge: Multi-Stage Pruning With Neuron Reconstruction for Privacy-Aware Edge Vision-Based Crowd Sensing in IoT-Enabled Rail Transit Systems

Internet-of-Things (IoT)-enabled surveillance systems increasingly utilize edge intelligence to support real-time analytics while reducing network dependency and limiting the transmission of raw visual data, thereby reducing potential data exposure. In rail transit environments, crowd monitoring requires low-latency on...

Xiao-Fei Li, Chuan-He Wang, Bin Hu · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.