Skip to content

Similar papers

Book Open access Sep 2026

Quantifying the Code-Size Overhead of eBPF JIT Compilation

eBPF allows user-defined programs to safely extend Linux kernel functionality at runtime, but its final machine code comes from a compilation pipeline that differs from native targets, and how efficient that pipeline is has no clear reference point. Our work constructs one: using the standard LLVM x86 backend as an app...

Hoang Duong, Hao Sun, Zhendong Su · 0 citations
Book Open access Sep 2026

LMTracer: Fine-Grained and Real-Time Performance Profiling for Production LLM Systems

Training and serving large language models (LLMs) has become a core business for AI providers. To ensure a high-quality user experience while optimizing infrastructure costs, providers need to closely monitor the performance of LLM executions in production. However, existing performance profiling tools fall short in th...

Wei Liu, Yong-Chao He, Bo-Han Zhao et al. · 0 citations

Performance Evaluation of LLM Inference Engines

This thesis presents a performance evaluation of three widely used open-source inference engines: vLLM, SGLang, and llama.cpp, and summarizes the experimental findings into selection recommendations for practical deployment scenarios, providing a reference for developers and researchers in choosing an appropriate infer...

Unknown authors · 0 citations

[Experiments & Analysis] Understanding the Performance of Native Execution in Big Data Engines: The Good, the Bad, and How to Fix It

This paper systematically evaluates the performance across three different native vectorized execution engines, including C++ based Velox and ClickHouse, and Rust based DataFusion, for the widely used big data engine Apache Spark, and proposes a cost model driven operator placement strategy that adaptively schedules op...

Hai-Kai Zhao · 0 citations

Torchy: A Tracing JIT Compiler for PyTorch (Extended Version)

Torchy is presented, a tracing JIT compiler for PyTorch, one of the mainstream eager-mode frameworks, that achieves similar performance as data-flow frameworks, while providing the same semantics of straight-away execution.

Nuno P. Lopes · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.