AgentReplay: Token-Wise Trace Replay Is Essential for Fair Serving System Performance Benchmarking
LLM-based agents execute multi-turn workflows with interleaved model inference and tool calls, making efficient serving increasingly important. However, evaluating serving optimizations is challenging because identical tasks can produce different execution trajectories. Changes in generated tokens can alter subsequent...