Digital content infringement has increased, with geometric cropping, photometric filtering, and stochastic noise significantly affecting redistributed media and destroying automated detection systems. Traditional systems achieve limited adversarial robustness because they use limited annotated data and computationally inefficient, high-dimensional, multimodal descriptors for large-scale retrieval. This research utilizes contrastive self-supervised representation learning to increase adversarial resilience and advanced multimodal feature compression to reduce retrieval complexity while retaining embedding discrimination. This research proposes the MoCo-OSGF framework (momentum contrast with orthogonal subspace and global feature learning), an integrated framework for high-fidelity multimodal infringement analysis. The momentum contrast encoder employs contrastive learning on a 65 K-negative queue and over 1 million unlabeled multimodal samples to derive adversarially robust and semantically consistent feature representations. Orthogonality-based regularization, subspace alignment, product-quantization-style feature partitioning, and multi-scale aggregation compress features from 2048 to 256D while preserving over 92% of their discriminative ability in the orthogonal subspace multi granularity compression module. Spatial orthogonality divergence, multi-head attention, and saliency-based spatial modeling in the Global Spatial Feature Learning module preserve fine-grained infringement cues even under high adversary distortions. Key findings indicate 18-22% resilience against adversarial changes and 35% large-scale retrieval accuracy improvement. Results show a 27.6% decrease in retrieval delay and a nearly 48.2% decrease in computational overhead. Finally, MoCo-OSGF is scalable and resilient for next-generation multimodal digital content infringement detection.
Haibo Wu, Jie Zhao, Feng Zhou et al.· Scientific Reports· 0 citations
Animating objects in a static 3D Gaussian scene requires an explicit object-level dynamic state and a controllable model of object motion. Existing dynamic Gaussian methods primarily reconstruct time-varying scenes or simulate deformation, rather than provide compact object states for direct control. To address this gap, we present NewtonGS, a physics-structured framework for object-level state rollout and Gaussian scene animation. NewtonGS represents each object with a 22-dimensional state covering pose, linear and angular velocity, anisotropic scale and its rate, mass, and contact properties. Its Gaussian Neural Newtonian Dynamics (Gaussian-NND) model combines analytic translation, quaternion kinematics, gravity, damping, and scale-restoration dynamics with learned continuous and contact residuals. A discrete event map handles floor contact. Predicted poses and scales define a shared affine transformation that updates the means and covariances of all Gaussians associated with each object. We construct two procedurally generated datasets: State-32 for state-rollout evaluation and Gaussian-32 for state-to-Gaussian transformation. On both the in-distribution and velocity-range-shift splits of State-32, NewtonGS achieves lower trajectory RMSE, final displacement error, and velocity RMSE than five analytic baselines. Experiments on Gaussian-32 further demonstrate effective conversion from predicted states to animated Gaussian objects.
Lianlei Shan, Feiyang Ye, Yan Chen et al.· 0 citations