Reconstruction of gene regulatory networks (GRNs) is essential for uncovering regulatory relationships between transcription factors (TFs) and target genes. With advances in single-cell RNA sequencing (scRNA-seq), cell-type-specific GRN inference has become an important direction in systems biology; however, existing deep learning methods still struggle with noisy expression, symmetric link scoring that cannot capture regulatory direction, and difficulty separating co-expression and co-regulation from direct regulation. Built upon the variational autoencoder (VAE)-graph attention network (GAT) framework of GRANet [1], we propose DMVD-GRN with enhanced VAE multi-view denoising, skew-symmetric directed decoding, and structure-aware joint decoding, integrating expression correlation, second-order adjacency, and Jaccard co-regulation priors. On 14 tasks of the STRING benchmark, DMVD-GRN achieves superior AUROC and AUPRC, with more pronounced AUPRC gains, thereby improving identification of true regulatory edges under extreme class imbalance.
Joint-embedding predictive architectures (JEPAs) learn world models that predict in a compact latent space rather than in pixels, reducing the pressure to model nuisance appearance. Yet this provides no guarantee against visual perturbations: they can still alter the encoded representation and affect subsequent action-conditioned predictions. Bisimulation captures this requirement precisely: two observations should be treated as the same state only when their action-conditioned consequences agree. Guided by this criterion, we introduce Action-Conditioned Predictive Consistency (ACPC), a diagnostic that measures how far a clean history and a visually perturbed view of it diverge after being rolled forward under the same action sequence. We prove that this divergence bounds the perturbation-induced change in multi-step prediction error and planner cost. Building on pairwise ACPC, we define two complementary measures: the Invariance Radius (IR) summarizes clean-perturbed rollout spread, while the Separation Rate (SR) checks whether different states remain distinguishable after rollout. Experiments on four visual control tasks show that pairwise ACPC predicts perturbation-induced prediction and cost changes. On LeWM, the IR-SR screen transfers across tasks, and the joint diagnostic remains informative under blur and resize. PLDM exhibits similar diagnostic trends under a different architecture.
Guo An, Zijing Wu, Hongzhuang Dong et al.· 0 citations