We propose monotonic consistent caching (MCC), a cache scheme for applications that demand transactional guarantees. MCC warrants that a transaction-like request always sees a consistent view of the backend database and that observed writes over the cache will not be lost, even if it operates on conventional cache syst...
Shuai An, Yang Cao, Jia Li et al.· ACM Transactions on Database...· 0 citations
Formal feature explanations strictly maintain perfect conformity but are intractable to compute, while heuristic methods are much faster but can lead to problematic explanations due to lack of conformity guarantees. We propose relative keys that have the best of both worlds. Relative keys associate feature explanations...
Shuai An, Yang Cao, Jia Li et al.· The VLDB journal· 1 citation
This work presents a cross-layer analysis approach for single-GPU LLM inference that jointly characterizes latency and memory behavior, and systematically characterize representative dense and Mixture-of-Experts models under diverse workloads on a single A100 GPU.
Zong-Xing Zhao, Xia-Qing Li, Ze-Kai Meng et al.· Proceedings of the Internati...· 0 citations
The performance of single-GPU LLM inference is characterized by strong cross-layer interactions spanning model architecture, runtime scheduling, operator execution, and GPU microarchitecture. Unfortunately, a unified understanding of single-GPU LLM inference bottlenecks is still lacking due to two limitations: the lack...
Zong-Xing Zhao, Xiaqing Li, Ze-Kai Meng et al.· Proceedings of the Internati...· 0 citations
We propose monotonic consistent caching (MCC), a cache scheme for applications that demand transactional guarantees. MCC warrants that a transaction-like request always sees a consistent view of the backend database and that observed writes over the cache will not be lost, even if it operates on conventional cache syst...
Shuai An, Yang Cao, Jia Li et al.· ACM Transactions on Database...· 0 citations
Federated learning (FL) is a distributed machine learning (ML) paradigm that has been widely used to train ML models on massive amounts of data in edge computing (EC) environments. However, FL faces significant challenges from device heterogeneity, edge dynamics, and limited communication resources. To address these ch...
Junyi Deng, Jiahua Liu, Yanheng Liu et al.· IEEE Internet of Things Jour...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.