Upper-Expectile Multi-Step Q-Learning for Off-Policy Reinforcement Learning
Using a single expectile level $\tau=0.8$ and a fixed backup horizon across 27 manipulation and navigation task instances, ENQ is competitive with LQL on aggregate, achieves higher measured training-step throughput in the authors' profiling study, and benefits more from a ten-critic ensemble in a controlled scaling experiment.