Skip to content

Author

Chengru Song

We have 6 of 27 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

TEMPO: Makespan-Aware Expert-Parallel Load Balancing Across Memory- and Compute-Bound Regimes

This work formalizes per-batch dispatch as a fixed-charge makespan problem---NP-hard on two fully replicated GPUs, polynomial in degenerate limits---and presents TEMPO, a makespan-aware dispatcher solving it in milliseconds off the critical path; its SGLang integration runs out-of-process and fuses dispatch with count...

Jie Li, Chen-Xin Jia, Jinliang Shen et al. · 0 citations
Jul 2026

JAGG: Jacobian-Aggregated Group Gradient for Efficient GRPO Training of Diffusion Models

JAGG approximates intermediate-step Jacobians via $t$-weighted interpolation of the endpoint Jacobians, then aggregates per-step upstream signals into two composite gradients applied through a single joint backward pass, and it is proved this interpolation is exact when the velocity is linear in $(z,t)$, and a cosine-s...

Rui-Ying Ding, Jie Li, He Kang et al. · 0 citations
Preprint Aug 2026

Latent Reward Registers for Diffusion Preference Alignment

This work proposes Latent Reward Registers, a mechanism that estimates terminal preference directly from intermediate noisy latents, and achieves significant reward improvement with a favorable reward-quality balance against training-free baselines.

Yuan-Shen Guan, Zipeng Feng, Cheng-Ru Song et al. · 0 citations
Jul 2026

X-Stage: An Overlooked Pipeline Stage for Communication-Computation Overlap in DiT Inference

A lightweight Burst-Gap model parameterized by backpressure-free issue time, effective drain rate, and outstanding capacity predicts issue overhead, recovery between bursts, and the onset of backpressure, and two communication-computation fused kernels are redesigned.

Jianwen Xian, Zhiyuan Xu, Yucheng Li et al. · 0 citations
Jul 2026

HeadCast: Casting Attention Heads for Efficient Autoregressive Video Generation

HeadCast is proposed, a training-free, plug-and-play acceleration framework built on the observation that a pre-trained AR model's attention heads exhibit stable, heterogeneous behaviors that accelerates inference by up to 1.62x at 720P and 1.95x at 1080P, while keeping VBench quality on par with full attention and lar...

Jinliang Shen, Li Su, Zheming Li et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.