A low-latency MLIR compilation backend based on TPDE that by-passes the LLVM lowering and compilation pipeline is introduced and techniques that reduce MLIR’s intrinsic overhead are proposed.
eBPF allows user-defined programs to safely extend Linux kernel functionality at runtime, but its final machine code comes from a compilation pipeline that differs from native targets, and how efficient that pipeline is has no clear reference point. Our work constructs one: using the standard LLVM x86 backend as an app...
Hoang Duong, Hao Sun, Zhendong Su· Proceedings of the 4th Works...· 0 citations
Training and serving large language models (LLMs) has become a core business for AI providers. To ensure a high-quality user experience while optimizing infrastructure costs, providers need to closely monitor the performance of LLM executions in production. However, existing performance profiling tools fall short in th...
Wei Liu, Yong-Chao He, Bo-Han Zhao et al.· Proceedings of the ACM SIGOP...· 0 citations
This thesis presents a performance evaluation of three widely used open-source inference engines: vLLM, SGLang, and llama.cpp, and summarizes the experimental findings into selection recommendations for practical deployment scenarios, providing a reference for developers and researchers in choosing an appropriate infer...
This paper systematically evaluates the performance across three different native vectorized execution engines, including C++ based Velox and ClickHouse, and Rust based DataFusion, for the widely used big data engine Apache Spark, and proposes a cost model driven operator placement strategy that adaptively schedules op...
Torchy is presented, a tracing JIT compiler for PyTorch, one of the mainstream eager-mode frameworks, that achieves similar performance as data-flow frameworks, while providing the same semantics of straight-away execution.
Nuno P. Lopes· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.