This work shows that one can extract LLM assets during inference, namely embeddings, attention, and quantized MLP weights, activations, and other inference states, from localized memories and compute subcircuits from localized memories and compute subcircuits by deploying laser voltage imaging.
Dev M. Mehta, Lily Dukette, William Folan et al.· 0 citations
A real-time status monitoring and traceability system based on the edge nodes of the Internet of Things is constructed, providing a replicable technical paradigm for the digital protection of intangible cultural heritage processes.
MHA-CSP achieves robust structured reasoning via synthetic distance rectification---powered by Mahalanobis-based attention---and efficient information bypass inherited from the CSP backbone, highlighting the effectiveness of complex-valued state propagation with collaborative multi-head rectification in capturing symbolic structures.
The notion of a pseudo e-net is introduced, which decomposes the surface into e-thin cylinders together with a Delaunay triangulation over an e-net of the remaining thick part of the hyperbolic surface.
V. Delecroix, Vincent Despré, Camille Lanuel et al.· 3 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
The design is ~82x faster than MedSOM, the CUDA implementation behind the authors' earlier MEDLINE atlases and, at 128x128, 621x faster than the best available multicore-CPU library.
The Sparse-Activation-ReLU (SAR) layer is proposed, a single-step alternative that promotes activation sparsity without surrogate-gradient training while remaining compatible with event-based computing and is a step towards energy-efficient virtual sensing.
The proposed Wireless GPU Computing Infrastructure (WiCi) can reduce time to first token by up to 90%, improve the token rate by approximately 39x compared to local inference on mobile devices for the same model, and support much larger models.
Yibin Shen, Wei Li, Kaiqiang Xu et al.· 0 citations
A mesh-free discretization in which a single neural network represents the displacement and phase fields and is trained by minimizing the incremental energy directly is proposed.
Han Zhang, M. Alamdari, B. Shahbodagh et al.· 0 citations
MPR image quality is typically inferior to axial images and deteriorates further when derived from thicker axial slices, therefore, appropriate selection of the reconstruction kernel, pixel size and slice thickness are essential to maintain diagnostic image quality.
P. Monnin, A. Viry, F. Becce et al.· Journal of Applied Clinical...· 0 citations
The feasibility of integrating Edge AI with microfluidic biosensing concepts for low-latency and energy-aware monitoring of critical biomarkers is demonstrated and full analytical and device-level validation remains future work.
Salman Khan, Sunny Barua, Ahsan Zahid Satti et al.· Microfluidics and Nanofluidi...· 0 citations
This paper proposes a heterogeneous multi-agent proximal policy optimization (MAPPO)-based framework where both user devices and UAVs act as heterogeneous agents and utilizes a centralized training and decentralized execution (CTDE) paradigm to enable collaborative strategies between computing requesters and providers.
Ming Cheng, Canlin Zhu, Jiang-Hang Tang et al.· Journal of King Saud Univers...· 0 citations
Three-dimensional time-harmonic Maxwell simulations generate massive complex indefinite systems whose mesh coarsening is strictly limited by phase accuracy. Although matrix-free finite element kernels utilize GPU throughput efficiently, standard multilevel solvers are ultimately bottlenecked by the memory and communication costs of exact coarse-grid factorizations. We present a fully matrix-free, factorization-free three-grid preconditioner for curl-conforming N{\'e}delec discretizations with perfectly matched layers (PML) and optimally blended quadrature. The method employs an outer FGMRES to solve the unshifted fine-grid equation, while an intermediate-grid correction is computed by a fixed-work FGMRES preconditioned with a complex-shifted $2h$--$4h$ cycle. This strategically confines the complex shift to an auxiliary preconditioner, preserving the physical Maxwell operator. A local Fourier analysis derives the blended Maxwell branches and compatible edge transfers, identifying robust shift and Jacobi damping parameters. Validated against the analytical Maxwell Green tensor, our approach demonstrates extreme scalability: using a single solver configuration, both homogeneous and highly heterogeneous systems with approximately 10.89 billion complex edge unknowns are solved in 42.0--72.0 seconds on just 64 NVIDIA A100 GPUs.
What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research.
MIT News · Artificial Intelligence· news.mit.eduAug 31, 2026
With millions of users across the world, Julia has been used to conduct cutting-edge research and to design new drugs, jet engines, heat pumps, and more.
MIT News · Artificial Intelligence· news.mit.eduAug 27, 2026
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.