Skip to content

Author

Leyuan Fang

3 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Aug 2026

Multiview Clustering via Enhanced Multiorder Bipartite Graph Learning.

Although existing bipartite graph-based multiview clustering (MVC) methods effectively exploit the structural relationships within multiview data, they exhibit three major limitations: 1) they primarily focus on direct similarities between data points and anchors, neglecting underlying neighborhood structures; 2) most existing methods fail to capture high-order correlations across bipartite graphs from different views; and 3) they overlook the relationships among anchor points, limiting the discriminative power of the learned graph. To address these challenges, we propose a unified framework, termed enhanced multiorder bipartite graph learning (EMOBGL) for MVC. The proposed EMOBGL method first constructs a second-order bipartite graph (SOBG) to capture both local and neighboring structural relationships between data points and anchors through first-order similarity (FOS) and second-order similarity (SOS). Then, the tensor Schatten- $p$ regularizer is incorporated to construct a multiorder bipartite graph (MOBG) to capture third-order similarity (TOS) across views. Meanwhile, the anchor structure regularization (ASR) is introduced to model anchor-anchor interactions, further enhancing the structural expressiveness and discriminability of the bipartite graph. The resulting EMOBGL model effectively integrates multiorder and multiview relationships within a unified framework, achieving robust and discriminative clustering performance. An efficient alternating direction method of multipliers (ADMMs) is developed to optimize the model, and we theoretically prove that the solution converges to a Karush-Kuhn-Tucker (KKT) stationary point. Extensive comparative experiments on 13 benchmark datasets demonstrate that the proposed EMOBGL consistently outperforms 13 state-of-the-art methods in both clustering accuracy and robustness. The source code is available at https://github.com/DongHuangTaiYi871/EMOBGL.

Yangjun Deng, Wenhao Deng, Longfei Ren et al. · 0 citations
Book Open access Aug 2026

Towards Efficient Embodied Reasoning: Mixture-of-Depth Compute Allocation for Vision-Language-Action Model

Vision-Language-Action (VLA) model plays a crucial role in embodied decision making. While practical deployment requires fast inference under limited onboard computation, a full forward pass through the vision-language model makes such deployment challenging. To address this issue, existing methods typically employ lightweight techniques to compress the backbone. However, these information-lossy methods degrade spatial representations for action generation. In contrast, rate-distortion principles aim to reduce computation while retaining control-sufficient information. Inspired by this insight, we introduce Effective-Edge Flow, an action-aligned attribution measure that quantifies the marginal contribution of token interactions across network depth. This analysis reveals a consistent depth asymmetry, with visual evidence dominating early layers and linguistic reasoning sustaining task-relevant influence into deeper layers. Building on this structure, we propose MoDeVLA, the first rate-distortion driven efficient VLA model that performs token-wise depth allocation via Mixture-of-Depth Conditioning and integrates shallow visual-spatial with deep textual-logical features for action conditioning. Extensive real-robot evaluations across 20 tasks and multiple embodiments demonstrate that MoDeVLA preserves task performance while reducing latency by about 38% and FLOPs by 86% on edge device NVIDIA Jetson Orin, highlighting its strong ability for embodied systems deployment.

Weiying Xie, Qingcheng Zeng, Zihan Meng et al. · 0 citations
Book Open access Aug 2026

Towards Efficient Embodied Reasoning: Mixture-of-Depth Compute Allocation for Vision-Language-Action Model

Vision-Language-Action (VLA) model plays a crucial role in embodied decision making. While practical deployment requires fast inference under limited onboard computation, a full forward pass through the vision-language model makes such deployment challenging. To address this issue, existing methods typically employ lightweight techniques to compress the backbone. However, these information-lossy methods degrade spatial representations for action generation. In contrast, rate-distortion principles aim to reduce computation while retaining control-sufficient information. Inspired by this insight, we introduce Effective-Edge Flow, an action-aligned attribution measure that quantifies the marginal contribution of token interactions across network depth. This analysis reveals a consistent depth asymmetry, with visual evidence dominating early layers and linguistic reasoning sustaining task-relevant influence into deeper layers. Building on this structure, we propose MoDeVLA, the first rate-distortion driven efficient VLA model that performs token-wise depth allocation via Mixture-of-Depth Conditioning and integrates shallow visual-spatial with deep textual-logical features for action conditioning. Extensive real-robot evaluations across 20 tasks and multiple embodiments demonstrate that MoDeVLA preserves task performance while reducing latency by about 38% and FLOPs by 86% on edge device NVIDIA Jetson Orin, highlighting its strong ability for embodied systems deployment.

Weiying Xie, Qingcheng Zeng, Zihan Meng et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.