Preprint
Jul 2026
FlashBEV: Fast and Memory-Efficient Exact BEV Transformation with IO-Awareness
FlashBEV is proposed, a fully fused and IO-aware execution strategy mathematically equivalent to Tensorized Sampling-VT (same operator output) while substantially reducing global memory traffic and kernel-launch overhead, and achieves more than an order of magnitude lower peak GPU memory and significant inference-latency speedups.
Shunsuke Yokokawa, Hironori Kasahara
· 0 citations