The rapid growth of large-scale deep neural networks has pushed inference workloads toward increasingly diverse shapes. In practical inference, these models are often invoked with tensors of varying shapes rather than a single fixed shape. Such shape variability reshapes operator loops and intermediate tensors, leading...
Hon-Gen Shao, Cen Chen, Xin Fang et al.· Journal of systems architect...· 0 citations
Diffusion Transformers (DiTs) have become a dominant architecture for video generation, but their efficiency is limited by the quadratic complexity of full attention. Sparse attention reduces this cost by retrieving important blocks and computing attention only within them, but inaccurate retrieval can either degrade g...
Yun-Wei Dai, Jia-Rui Wen, Hui-Ping Zhuang et al.· 0 citations
An analytic imbalance rectifier (AIR) algorithm for real-world CL is proposed that addresses class imbalance with an analytic reweighting module (ARM) that calculates a reweighting factor for each class in the loss function to equalize total sample weights across classes.
Di Fang, Yinan Zhu, Run-Ze Fang et al.· arXiv.org· 12 citations· ⚡1
X-SG$^2$S is the first framework to unify 1D-to-3D watermarking and enable simultaneous multi-modal watermark embedding in 3DGS, achieving this with minimal rendering interference and zero modifications to parameters or pipelines.
EgoSafe-Bench is introduced, a benchmark specifically designed to probe forensic reasoning in egocentric safety scenarios, generated by pairing each of the 3,000 video clips with a QA chain governed by the proposed Hierarchical Reasoning Evaluation (HRE) protocol.