Skip to content

Author

Hai-Peng Yao

We have 2 of 67 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Book Open access Aug 2026

HyNA: Taming Tail Latency in MoE Training with Hybrid Switch Silicon

The transition to trillion-parameter models, particularly Mixture-of-Experts (MoE), shifts the bottleneck of distributed training from computation to communication. However, existing Parameter Server (PS) architectures succumb to incast congestion, while state-of-the-art In-Network Aggregation (INA) solutions like ATP fail to handle the sparse, bursty traffic of MoE workloads. These solutions suffer from severe tail latency amplification due to their reliance on slow, host-based fallbacks for collisions and overflows. To dismantle this communication wall, we propose HyNA, a fully serverless aggregation system that eliminates dedicated parameter-server nodes by leveraging a novel hardware-software co-designed switch architecture. HyNA couples wire-speed Reconfigurable Match Tables (RMT) with embedded RISC-V cores. By adhering to a strict on-chip closure principle, the system processes all traffic anomalies—including hash collisions and floating-point variances—entirely within the switch ASIC, converting unpredictable network RTT into deterministic on-chip latency. We validate our design through a 100 Gbps FPGA prototype and a 7nm ASIC synthesis analysis. Results demonstrate that HyNA incurs less than 3% silicon area overhead while improving aggregation throughput by 7.35X over BytePS and 1.4X over ATP. Crucially, in the MoE gradient synchronization phase, the system eliminates the fallback penalty and reduces synchronization time by up to 1.6X compared to dynamic INA baselines, without compromising bit-level model accuracy.

Yang Liu, Tianxiang Liu, Hai-Peng Yao · 0 citations
2026

A2ProSFC: Agentic AI-Enabled Proactive SFC Orchestration in Embodied Edge Intelligence Networks

Embodied Artificial Intelligence (AI) integrates multimodal large models into Embodied Agents (EAs), driving the evolution of Embodied Edge Intelligence Networks (EEINs) to handle the heterogeneous requests generated by EAs. To guarantee service performance for heterogeneous requests, Service Function Chain (SFC) orchestration has emerged as a critical solution, involving the sequential deployment of Network Functions (NFs) to satisfy customized service requirements. However, realizing SFC orchestration in EEINs presents several challenges, including limited forwarding performance, dynamic environment evolution, and high-dimensional decision spaces. To tackle these issues, we present A2ProSFC, an agentic AI-enabled SFC orchestration system that facilitates perception–reasoning–action loops by leveraging programmable switches. Specifically, we formulate a long-term SFC orchestration problem aimed at maximizing served SFC throughput while ensuring system load balancing. Subsequently, we employ Lyapunov optimization to decouple the long-term orchestration into a sequence of online optimization subproblems and design DiffOrch, a diffusion-based SFC orchestration algorithm. By leveraging In-band Network Telemetry (INT), DiffOrch perceives network state information and adaptively generates orchestration decisions. Furthermore, we design a pipeline integrating INT perception and SFC orchestration to validate the system’s effectiveness. Experimental results demonstrate that A2ProSFC improves throughput by 40.53% and enhances load balancing efficiency by 36.97% compared to existing baselines.

Tianhao Ouyang, Yichi Zhang, Xiaoxu Ren et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.