Hybrid attention models have emerged as a crucial architecture for Large Language Models (LLMs) (e.g., the Qwen3.5 and Kimi series). Their memory and computational efficiency make them highly attractive for on-device inference, forming a promising synergy with edge Neural Processing Units (NPUs). However, naive executi...
Yin-Yuan Zhang, Da-Liang Xu, Xiao-Long Huang et al.· 0 citations
PhyAI, a Physical AI inference engine with a single runtime that keeps architecture-specific conditioning, solver, cache, and output logic in model adapters while sharing graph execution, kernels, memory management, and parallel services, is built.
Cheng-Hua Wang, Daliang Xu, Dongqi Cai et al.· 1 citation
A systems vision for AI infrastructure in space is developed as the systems layer that manages AI capabilities across spacecraft, orbital networks, ground stations, and cloud backends, while treating orbital and physical state as part of the resource model.
Qing Li, Qiyang Zhang, Daliang Xu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.