Skip to content
Preprint

HybridSim: A Physics-Learning Hybrid Digital Twin for mmWave Human Sensing

Jul 2026 · 0 citations · 38 references
Computer Science

Abstract

High-fidelity simulation of mmWave radar signals for dynamic human motion is valuable for developing radar-based human sensing models; yet collecting accurately labeled measurements for a specific deployment site remains expensive. We present HybridSim, a physics-learning hybrid simulator that synthesizes mmWave radar signals from dynamic human meshes under a fixed indoor room configuration, explicitly decoupling propagation into two components. To parameterize the human subject, we use a tri-plane representation to extract human features and a Graph Convolutional Network to stabilize optimization and mitigate gradient instability. The direct signal path is modeled via an inverse-rendering formulation with a microfacet BRDF to capture primary surface reflections. In parallel, the indirect path is approximated by combining 3D Gaussian Splatting with a virtual-receiver geometry to fit and reproduce site-specific multipath interference patterns, achieving substantially lower computational cost than explicit full ray tracing. Experiments in a fixed-room setting show improved agreement with a physically based reference and consistent gains on downstream radar-based human sensing tasks when using HybridSim for site-specific data augmentation.

View source

Similar papers

Open access Jul 2026

Lightweight mmWave radar human action recognition via knowledge distillation and physically enhanced representation

Human activity recognition (HAR) based on millimeter-wave (mmWave) radar has attracted significant attention due to its advantages in non-contact sensing and privacy preservation. However, extracting robust fine-grained features from sparse and nonuniform mmWave point clouds remains challenging when relying solely on raw 3D coordinates, often leading to measurement uncertainty in complex dynamic scenarios. Furthermore, existing high-performance models incur high computational overhead, limiting their deployment on resource-constrained edge sensing devices. To address these challenges, this paper proposes a lightweight mmWave radar HAR framework based on knowledge distillation. First, we introduce a physically enhanced input representation module that alleviates the representation limitations of sparse mmWave radar point clouds by explicitly incorporating inter-frame centroid differences and normalized timestamps as auxiliary motion and temporal cues. Second, we employ a high-accuracy dual-branch teacher model to guide a student model that maintains macro-architectural consistency but utilizes simplified operators. A multi-granularity feature joint distillation strategy transfers representation capabilities to the student without incurring the teacher’s computational burden. Experimental results demonstrate that the student model achieves average recognition accuracies of 98.38% and 98.41% on the RadHAR and Pantomime datasets, respectively. Notably, the model requires only 0.27 M parameters and 0.029 GFLOPS.

Fengdai Liu, Yinan Wang, Zuoheng Liu et al. · 0 citations
Preprint Jul 2026

mmSimPrior: Learning Simulation Priors for Data-Efficient and Generalizable Real-World Radar-based Human Motion Reconstruction

Millimeter-wave (mmWave) radar enables privacy-preserving and illumination-robust human motion reconstruction, but training generalizable models typically requires costly paired radar-motion recordings. Simulation can scale such supervision, yet even physics-based simulators cannot fully reproduce real-world multipath, clutter, hardware-specific response statistics, or distance-dependent resolution degradation, leaving a sim-to-real gap. We present mmSimPrior, a simulation-pretrained framework that factorizes transferable knowledge into signal, motion, and radar-to-motion mapping priors. To learn transferable signal and motion priors, we pretrain a multimodal radar encoder with a physics-informed domain-randomization curriculum designed to mitigate the sim-to-real gap by approximating real-world propagation- and acquisition-level variations, while a joint-temporal tokenizer learns a discrete prior over plausible human motion. A dual-mode mapping module predicts either motion-code distributions for structurally constrained zero-shot reconstruction or continuous motion parameters for flexible adaptation from limited real data. We further construct a 4.2M-frame, 31K-sequence dataset suite and introduce a No-Overlap Setting that prevents any exact subject-environment-location-motion tuple from appearing in both the adaptation and test sets. Experiments on mmSimPrior-Real and RT-Pose demonstrate consistent gains: with only 24 paired real sequences, mmSimPrior-Reg reduces MPJPE by 24.7-39.0% over the strongest baseline across the three environments, while mmSimPrior-Cls reduces zero-shot MPJPE by 8.5% without fine-tuning.

Cheng Guo, Qiming Cao, Shengkai Xu et al. · 0 citations
Open access Jul 2025

Photon Splatting: A Physics-Guided Neural Surrogate for Real-Time Channel Modeling in Quasi-Static Wireless Environments

We present Photon Splatting, a physics-guided neural surrogate for wireless channel modeling in complex quasi-static environments. The work represents wave-environment interactions using surface-attached virtual sources, or photons, which carry directional wave signatures informed by the scene geometry and transmitter (Tx) configuration. At runtime, channel impulse responses (CIRs) are predicted by splatting photon contributions onto the angular domain of the receiver (Rx) using a geodesic rasterizer and aggregating them in delay. The model is trained to learn a physically grounded representation that maps Tx–Rx configurations to full channel responses. After a one-time offline training phase, the surrogate generalizes to unseen Tx locations, Rx positions, and antenna beam patterns without retraining. We demonstrate the framework on Sionna ray-tracing data for canonical 3-D scenes and a complex indoor café with 1000 Rxs. The results show millisecond-level inference latency and accurate CIR predictions across a range of configurations.

Ge Cao, G. Gradoni, Zhen Peng · 2 citations
#artificial intelligence Preprint Aug 2026

Physics-Unrolled Neural Operator for Wireless Field Modeling

This work proposes Physics-Unrolled Hybrid Neural Operator (PU-HNO), a three-stage cascade that predicts high-fidelity indoor radio maps from low-fidelity ray-tracing outputs and scene priors by progressively capturing reflection, diffraction, and scattering effects, rather than treating radio maps as generic images.

Rafid Umayer Murshed, Saif Ur Rahman, Mingyue Tang et al. · 0 citations
Open access 2026

Enhancing D-Band FMCW Radar Tracking for In-Air Writing Through Weakly Supervised Deep Association

Reconstructing precise character trajectories in multipath environments represents a primary challenge for high-fidelity reconstruction for radar-based in-air writing. Existing tracking methods typically rely on dominant target assumptions, often failing to account for the interference of multipath reflections. Conversely, end-to-end deep learning approaches generally lack the physical interpretability necessary for accurate geometric drawing. In this work, we propose a novel tracking framework that combines high-resolution D-band sensing with a data-driven probabilistic data association (PDA) architecture. We leverage a 56 GHz bandwidth frequency-modulated continuous-wave (FMCW) radar setup to capture fine-grained kinematic signatures, supported by a rigorous sensor placement and calibration strategy. To address multipath interference without relying on rigid parametric assumptions, we propose the Temporal Deep PDA (TD-PDA). This architecture replaces heuristic hard assignments and classical parametric clutter models with a covariance-aware causal temporal convolutional network (TCN) fused directly into the statistical covariance spread equations of an extended Kalman filter (EKF). To train the network without heuristic bias, we introduce a track quality indicator (TQI) ensemble to automatically extract high-fidelity pseudo-labels. The framework is rigorously validated via simulated ground-truth data to quantify the geometric impacts of hardware constraints, and experimentally proven on a 1000-gesture dataset using tenfold leave-one-Subject-Out (LOSO) cross-validation. The proposed TD-PDA generalizes to unseen users with an ultralow inference latency, successfully reconstructing legible trajectories even in the presence of strong multipath interference. The model demonstrates improvements over prior deep-association architectures and achieves stability comparable to a well-tuned classical PDA filter via a purely data-driven design.

Salah Abouzaid, Leander Nothelle, Nils Pohl · 0 citations
Open access Aug 2026

Transformer-Based Physics Prior-Enhanced Residual Learning for OFDM Channel Estimation: The PERL Architecture

Accurate channel estimation in millimeter-wave MIMO-OFDM systems is hindered by limited pilot resources and the high sensitivity of physical priors to propagation conditions. This paper proposes PERL, a physics-enhanced residual learning framework that integrates ray-tracing (RT) priors with deep neural networks to refine coarse RT-based estimates. Unlike direct channel reconstruction, PERL learns a residual correction atop the RT baseline, with the correction magnitude adaptively gated according to noise level and RT reliability, thereby relying more on physical priors under low SNR and exploiting pilot observations for refinement at high SNR. The framework leverages a Transformer-based multimodal attention mechanism to deeply fuse sparse pilot observations, path-level propagation features (delay, angle, power, and phase), and scene-level statistical descriptors, enabling physical constraints and data-driven refinement to interact effectively. Experiments on a 28 GHz urban macro-cellular MIMO-OFDM scenario demonstrate that PERL achieves an overall NMSE of −28.82 dB, outperforming the RT baseline by 5.30 dB and the MMSE estimator by over 20 dB, with notably larger gains under non-line-of-sight conditions where the RT prior is less accurate. Link-level evaluations further confirm improved error vector magnitude and maintained bit/block error rates relative to the RT baseline, validating that the proposed physical-data collaborative paradigm not only enhances estimation accuracy but also preserves communication reliability, while offering a promising foundation for future detection-aware and integrated-sensing-and-communication optimizations.

Xiaotian Shi, Yuanjian Liu · 0 citations