The Sparse-Activation-ReLU (SAR) layer is proposed, a single-step alternative that promotes activation sparsity without surrogate-gradient training while remaining compatible with event-based computing and is a step towards energy-efficient virtual sensing.
The proposed Wireless GPU Computing Infrastructure (WiCi) can reduce time to first token by up to 90%, improve the token rate by approximately 39x compared to local inference on mobile devices for the same model, and support much larger models.
Yibin Shen, Wei Li, Kaiqiang Xu et al.· 0 citations
A mesh-free discretization in which a single neural network represents the displacement and phase fields and is trained by minimizing the incremental energy directly is proposed.
Han Zhang, M. Alamdari, B. Shahbodagh et al.· 0 citations
MPR image quality is typically inferior to axial images and deteriorates further when derived from thicker axial slices, therefore, appropriate selection of the reconstruction kernel, pixel size and slice thickness are essential to maintain diagnostic image quality.
P. Monnin, A. Viry, F. Becce et al.· Journal of Applied Clinical...· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
The feasibility of integrating Edge AI with microfluidic biosensing concepts for low-latency and energy-aware monitoring of critical biomarkers is demonstrated and full analytical and device-level validation remains future work.
Salman Khan, Sunny Barua, Ahsan Zahid Satti et al.· Microfluidics and Nanofluidi...· 0 citations
This paper proposes a heterogeneous multi-agent proximal policy optimization (MAPPO)-based framework where both user devices and UAVs act as heterogeneous agents and utilizes a centralized training and decentralized execution (CTDE) paradigm to enable collaborative strategies between computing requesters and providers.
Ming Cheng, Canlin Zhu, Jiang-Hang Tang et al.· Journal of King Saud Univers...· 0 citations
Three-dimensional time-harmonic Maxwell simulations generate massive complex indefinite systems whose mesh coarsening is strictly limited by phase accuracy. Although matrix-free finite element kernels utilize GPU throughput efficiently, standard multilevel solvers are ultimately bottlenecked by the memory and communication costs of exact coarse-grid factorizations. We present a fully matrix-free, factorization-free three-grid preconditioner for curl-conforming N{\'e}delec discretizations with perfectly matched layers (PML) and optimally blended quadrature. The method employs an outer FGMRES to solve the unshifted fine-grid equation, while an intermediate-grid correction is computed by a fixed-work FGMRES preconditioned with a complex-shifted $2h$--$4h$ cycle. This strategically confines the complex shift to an auxiliary preconditioner, preserving the physical Maxwell operator. A local Fourier analysis derives the blended Maxwell branches and compatible edge transfers, identifying robust shift and Jacobi damping parameters. Validated against the analytical Maxwell Green tensor, our approach demonstrates extreme scalability: using a single solver configuration, both homogeneous and highly heterogeneous systems with approximately 10.89 billion complex edge unknowns are solved in 42.0--72.0 seconds on just 64 NVIDIA A100 GPUs.
An inference-aware AirMoE error metric is constructed to quantify aggregation distortion effects on end-to-end (E2E) inference accuracy via perturbation-based layer-sensitivity calibration, and an activation- and channel-aware expert placement strategy is developed that assigns more important experts to devices with lower channel-power cost.
This work introduces Misanthrope, a novel privacy-preserving keypoint detector trained through self-distillation to avoid detecting keypoints on people, thus mitigating inversion attacks at the source rather than through post-hoc obfuscation.
F. Vultaggio, Predrag Djindjic, Markus Gerke et al.· 0 citations
A hierarchical joint optimization algorithm is developed within a multi-agent deep reinforcement learning (MADRL) framework to coordinate UAVs and MTs in a distributed manner and outperforms other benchmarks under varying network scales and capabilities by jointly optimizing UAV operations and resource utilization.
Tiankui Zhang, Wenlong Xu, Tianyi Shi et al.· IEEE Internet of Things Jour...· 0 citations
A compact and configurable event-driven autoencoder that efficiently compresses neuromorphic data while preserving essential spatiotemporal structure for downstream inference and demonstrates the potential of compact event-driven models to advance environmentally conscious, low-power AI systems for high-speed perception in autonomous, mobile, and embedded computing environments.
Riadul Islam, Joey Mulé, Dhandeep Challagundla et al.· 0 citations
PRISM, a prediction-guided runtime framework that jointly selects model variants and CPU allocations for containerized edge microservices, and adapts each pipeline stage in place and minimizes predicted CPU-package energy under deadline, resource, and offline model-level Quality of Result constraints is presented.
Uwe Gropengießer, Thomas Reuter, Dominik Schön et al.· 0 citations
What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research.
MIT News · Artificial Intelligence· news.mit.eduAug 31, 2026
With millions of users across the world, Julia has been used to conduct cutting-edge research and to design new drugs, jet engines, heat pumps, and more.
MIT News · Artificial Intelligence· news.mit.eduAug 27, 2026
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.