Skip to content

Author

Hongyang Du

We have 5 of 36 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Sep 2026

NebulaSD: Many-for-Many Speculative Decoding

Speculative decoding accelerates Large Language Model (LLM) inference by using a lightweight draft model to propose candidate tokens for parallel verification by a target model. Drafting and verification, however, exhibit different service characteristics and favor different batch configurations, making fixed draft-tar...

Jun-Hao He, Hong-Yang Du · 0 citations
#artificial intelligence Preprint Sep 2026

Designer-RSI: Evolving Procedural Memory from User Traffic for Agentic Graphic Design

Professional graphic design is a long-horizon agentic task in which structured, editable artifacts emerge from many interdependent actions, yet outcomes admit no reliable programmatic oracle. We introduce a continual adaptation framework in which a frozen frontier model operates professional design software through mor...

Hong-Yang Du, Lan Yan, Christian Flores et al. · 0 citations
#natural language process... Preprint Aug 2026

CROP: Task Relevance via Counterfactuals for Selective On-Policy Distillation

On-policy distillation (OPD) supervises a student language model on trajectories sampled from its current policy, but assigns equal credit to response tokens with unequal supervision value. Selective OPD addresses this limitation by allocating supervision non-uniformly across response tokens according to their estimate...

Enhan Li, Jun-Hao He, Hong-Yang Du · 4 citations
Jul 2026

HACO: Hedged Agent Computing for Reliable LLM Systems

HACO is proposed, a runtime control scheme that treats each role request as a reliability-constrained selection problem over candidate agent instances, each coupling a role type, an LLM, and a concrete execution environment.

Enhan Li, Hongyang Du · 0 citations
Preprint Jul 2026

MORES: Mobile Reasoning-as-a-Service via Distributed LLM Inference-Time Scaling

A Mobile Reasoning-as-aService (MORES) framework that treats reasoning as a computational service accessible to edge devices over wireless networks, and focuses on implicit reasoning, which achieves an approximately 18% improvement in system throughput over the baseline Soft Actor-Critic (SAC) algorithm.

Guanchen Liu, Hongyang Du, Kaibin Huang · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.