Skip to content

Author

Gongming Zhao

We have 6 of 97 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Sep 2026

HOCCL: Offloading Collective Communication from GPU Cores to Accelerate Distributed Training

Large language model training involves massive computation on GPU streaming multiprocessors (SMs), the primary compute units of GPUs. Since SMs host specialized accelerators such as Tensor Cores, their efficient utilization is critical to training efficiency. Unfortunately, existing collective communication systems com...

Yao Fei, Gong-Ming Zhao, Hong-Li Xu et al. · 0 citations
Preprint Sep 2026

VarioPath: Workload-Aware All-to-All Communication for PCIe GPU Clusters

AlltoAllv communication is a critical primitive in distributed large-model inference, particularly for mixture-of-experts (MoE) models. The growing adoption of PCIe GPU systems for cost-efficient inference makes AlltoAllv performance on these systems increasingly important. Without a dedicated scale-up interconnect (e....

Yao Fei, Jin Fang, Si-Ze Zheng et al. · 0 citations
#machine learning Preprint Sep 2026

TopoEP: Topology-Aware Load Balancing for Expert-Parallel MoE Training

A GPU-native, topology-aware load-balancing system for large-scale MoE training that converts the current routing result into hot-expert replication and token-rerouting decisions and executes the resulting plan without data-dependent host synchronization, reducing critical-path overhead.

Jia-Cheng Zhu, Xie Zhao, Gong-Ming Zhao et al. · 0 citations
Conference Mar 2026

Utility-Guided Orchestration for Cost-Efficient Tool-Augmented LLM Services

A lightweight utility-guided orchestration framework that formulates agent control as a costaware sequential decision problem over a compact action space, intended as an inspectable control layer for practical LLM services rather than a universally dominant accuracy optimizer.

Bo-Yang Liu, Gongming Zhao, Hong-Liu Xu et al. · 4 citations
Conference Jul 2026

Utility-Guided Orchestration for Cost-Efficient Tool-Augmented LLM Services

Tool-augmented large language model (LLM) services can solve complex tasks through retrieval and external tools, but current execution paradigms often trade adaptability for efficiency. Fixed workflows are predictable but rigid, while freeform reasoning loops such as ReAct may over-execute and issue redundant tool call...

Bo-Yang Liu, Gongming Zhao, Hongliu Xu et al. · 0 citations
Book Open access Aug 2026

Rethinking Cloud Optimization: Volatility-Driven for Better Outcomes

Hestia is proposed, a framework that achieves long-term stable oversubscription through workload aggregation through a smoothing-based method to classify workloads suitable for aggregation according to their periodicity, and an aggregation algorithm to minimize the overall MCV.

Baoqing Wang, Gongming Zhao, Hongli Xu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.