Skip to content

Author

Shao-Huai Shi

We have 6 of 24 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Sep 2026

Laplacian Frequency Hierarchies for Efficient 3D Gaussian Splatting Training

A key bottleneck in 3D Gaussian Splatting training is the continual growth of Gaussian primitives, which increases optimization cost and slows convergence, especially at high resolutions. We propose Laplacian Frequency Hierarchies, a simple yet efficient 3DGS scheme that combines Laplacian image decomposition with coar...

Yixiong Yang, Sirius Z. Zhang, Q. Yan et al. · 0 citations
Book Open access Sep 2026

Low-bit and Sparsified Gradient Communication for Accelerating Distributed Deep Learning with Convergence Guarantees

Communication poses a dominant bottleneck in distributed data parallel training with synchronous stochastic gradient descent, whereas the traditional AllReduce collective used for gradient synchronization limits the efficient utilization of communication compression strategies. In this paper, we propose a low-bit and s...

Jia-Qi Li, Shao-Huai Shi, Jing Peng et al. · 0 citations
#artificial intelligence Preprint Sep 2026

LeanGRPO: Eliminating Redundant Recomputation in Diffusion RL

LeanGRPO is presented by restructuring the data-parallel layout and introducing two recompute-free training schedules for trajectory-logprob diffusion RL, which achieves up to 1.83x end-to-end speedup while preserving the original optimization objective.

Si-Jie Wang, Zhi-Qiang Tan, Xin-Rui Yang et al. · 0 citations
Preprint Jul 2026

Bidirectional Resource Scheduling for Disaggregated and Asynchronous RL Post-Training

BiDiRL, a hybrid time-space multiplexing architecture for asynchronous, disaggregated RL designed to reduce resource idleness, is presented, including a hot-switch runtime that enables rapid switching between rollout and training resources with negligible overhead and a static, scheduling-aware planner based on time-pe...

Zhiqiang Tan, Maoxin Wang, Sijie Wang et al. · 1 citation
Jul 2026

Xema: Efficient Diffusion Serving through Fine-Grained Memory Management and Auto-Configuration

Xema is presented, a memory-efficient diffusion serving system that exploits predictable tensor lifetimes for trace-guided memory optimization and introduces an offline planner that jointly selects parallelism, concurrency, and memory control under GPU memory and SLO constraints.

Xueze Kang, Guangyu Xiang, Suyi Li et al. · 1 citation
Jun 2026

KernelFlume: Elastic Core-Attention Scaling for Agentic Long-Context Decoding

KernelFlume is presented, a decode-centric architecture that disaggregates the stable projection/FFN path from core-attention computation: weight nodes execute dense projection/FFN kernels, while weightless attention nodes store token-range KV partitions and scale with request-state demand.

Guangyu Xiang, Xueze Kang, Lin Zhang et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.