Experiments show that H-FedSN reduces communication costs by up to 477 times compared to baseline methods while maintaining high accuracy, making it well-suited for hierarchical FL in IoT deployments.
Jiechao Gao, Yuangang Li, Jie Wang et al.· arXiv.org· 4 citations
This work presents MEGA-GNN, a model-agnostic message-passing framework for edge-attributed multigraphs, and shows that MEGA-GNN is permutation equivariant and has the same asymptotic complexity as standard GNNs with edge updates.
H. Bilgi, Kubilay Atasu, Aggregation Aggregation· 2 citations
The teacher, an auxiliary behavior model, is trained to sample high-loss regions of the student and can generalize across unexplored modes, thereby enhancing mode coverage by providing an efficient training curriculum.
Minsu Kim, Sanghyeok Choi, Taeyoung Yun et al.· International Conference on...· 27 citations· ⚡6
It is established that gradient-based training can induce an implicit regularization towards low rank for several neural network architectures, and it is demonstrated empirically that this phenomenon may facilitate an explanation of generalization over natural data.
Noam Razin· arXiv.org· 1 citation
Reach audiences
Advertise in front of researchers, engineers, and readers.
It is established that wide residual networks (ResNets) with constant scaling factors become asymptotically unlearnable as depth increases, and the generalization capability of wide ResNets can be approximated by kernel regression associated with the Neural Tangent Kernel (NTK).
This study unveils the capability of attackers to generate adversarial policies even when restricted to partial observations of the victims in multi-agent competitive environments, and proposes a novel black-box attack (SUB-PLAY) that incorporates the concept of constructing multiple subgames to mitigate the impact of partial observability.
Oubo Ma, Yuwen Pu, L. Du et al.· Conference on Computer and C...· 16 citations
This work interprets chain-of-thought reasoning as a latent variable modeling problem and demonstrates that this distribution-matching paradigm of LLM fine-tuning can serve as an effective alternative to maximum-likelihood training and reward-maximizing policy optimization.
Edward J. Hu, Moksh Jain, Eric Elmoznino et al.· International Conference on...· 110 citations· ⚡19
This work shows that EFlowNets outperform other GFlowNet formulations in stochastic tasks such as protein design and extends the concept of EflowNets to adversarial environments, proposing adversarial flow networks (A FlowNets) for two-player zero-sum games.
Marco Jiralerspong, Bilun Sun, Danilo Vucetic et al.· International Conference on...· 11 citations· ⚡1
A new algorithm for amortized inference in sparse probabilistic graphical models (PGMs) is presented that enables off-policy training but avoids the need to instantiate all the random variables for each parameter update, thus speeding up training considerably.
J. Falet, Haebeom Lee, Esmeralda S. Whitammer et al.· International Conference on...· 9 citations
This paper builds bridges between two families of probabilistic algorithms: (hierarchical) variational inference (VI), which is typically used to model distributions over continuous spaces, and generative flow networks (GFlowNets), which have been used for distributions over discrete structures such as graphs. We demonstrate that, in certain cases, VI algorithms are equivalent to special cases of GFlowNets in the sense of equality of expected gradients of their learning objectives. We then point out the differences between the two families and show how these differences emerge experimentally. Notably, GFlowNets, which borrow ideas from reinforcement learning, are more amenable than VI to off-policy training without the cost of high gradient variance induced by importance sampling. We argue that this property of GFlowNets can provide advantages for capturing diversity in multimodal target distributions.
Esmeralda S. Whitammer, S. Lahlou, T. Deleu et al.· International Conference on...· 120 citations· ⚡9
A general framework for implementing NMMs using Template Model Builder (TMB), which automatically integrates out random effects and evaluates the marginal objective function alongside its exact gradients, eliminating the need for manual derivations or ad hoc approximations.
Nan Zheng, H. Cheung, Vibhu Sharma et al.· 0 citations
The results indicate that a converged moment-matching loss is not a reliable measure of generalization, and that train-classical, deploy-quantum workflows will need approaches that target generalization directly, leaving open whether better training objectives suffice or whether the model architectures themselves must change.
S. Raj, Natansh Mathur, A. Perdomo-Ortiz· 0 citations
A weeklong summer workshop brought higher education faculty to campus to explore how AI and machine learning materials can be adapted for their classrooms.
Adaptive AI agents can help make BIM data more machine-readable by navigating IFC models, interpreting inconsistent information, and mapping it to defined standards. In this blog, Alok Rawat shares findings from a real-world pilot in construction workflows. The post Adaptive AI Agents in Construction Workflows appeared first on GPT-Lab.