Aug 2026· Computer graphics forum (Print)· Vol 45· 0 citations· 29 references
Computer Science
TL;DR
This paper presents Mobile Radiance Caching (MobileRC), an online trainable radiance caching approach based on a plenoxel representation, designed to accelerate path tracing on mobile devices, and exploits the localized interaction between learnable plenoxel weights and training samples, designing a mobile‐friendly training method.
Abstract
Real‐time path tracing for global illumination has recently become feasible on high‐performance desktop GPUs, but achieving similar performance on mobile platforms remains a significant challenge due to computational limitations. As mobile devices begin to integrate ray tracing capabilities, new methods are required to bridge the performance gap and enable advanced rendering techniques on constrained hardware. In this paper, we present Mobile Radiance Caching (MobileRC), an online trainable radiance caching approach based on a plenoxel representation, designed to accelerate path tracing on mobile devices. Unlike neural network‐based radiance caching methods, which rely on matrix multiplication accelerators unavailable on current mobile GPUs, MobileRC uses a voxel‐based representation where each voxel stores spherical harmonic coefficients to represent angular dependencies, making it more suitable for mobile hardware. Specifically, we exploit the localized interaction between learnable plenoxel weights and training samples, designing a mobile‐friendly training method. While our approach incurs a mild loss in cache quality compared to neural methods optimized for high‐end GPUs, it significantly improves image quality while reducing rendering time, achieving interactive frame rates for full HD images in room‐sized scenes on mobile hardware.
Recent GPU generations include special-purpose ray tracing (RT) cores for graphics applications. While RT cores are primarily used for rendering, recent works show they can be leveraged for general-purpose tasks, including similarity searches. However, existing approaches do not support datasets exceeding three dimensi...
A hierarchical search-space planning framework for GPU kernel optimization that delivers stronger overall implementation validity, sample efficiency, and optimization performance than existing training-free methods, while remaining competitive with the training-based CUDA-L1 without additional model training is propose...
Jing-Hao Wang, Qiqi Gu, Chenpeng Wu et al.· 0 citations
Bloom and glow post-processing effects are perceptually significant components of modern real-time rendering pipelines, yet systematic evaluations of competing bloom implementations on mobile Graphics Processing Unit (GPU) hardware remain scarce. This paper presents a comparative benchmarking study of five bloom techni...
Jos Timanta Tarigan· Engineering, Technology &...· 0 citations
The core of LEO lies in introducing hierarchical caching to construct a hybrid-grained software pipeline, enabling efficient computation-communication overlap both within and across GPU kernels.
Jia-Qi Si, De-Zun Dong· ACM Transactions on Architec...· 0 citations
The proposed Wireless GPU Computing Infrastructure (WiCi) can reduce time to first token by up to 90%, improve the token rate by approximately 39x compared to local inference on mobile devices for the same model, and support much larger models.
Yibin Shen, Wei Li, Kai-Qiang Xu et al.· 0 citations
Atlas is presented, an on device city scale 3DGS rendering framework that enables scalable rendering without runtime Internet access and proposes temporal aware LoD search and stereo rasterization to avoid redundant computation in VR.
He Zhu, Zheng Liu, Xingyang Li et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.