Skip to content

Author

Amritansh Mishra

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

Efficient Reinforcement Learning for Long-Horizon Tool-Use Agentic Tasks

SINKFLEX-RL, a modular training system for RL in dual-control tool-use environments that combines a Gymnasium-compatible environment wrapper, a VERL-style rollout dataflow, group-relative policy optimization without a separate value model, and a sink-aware FlexAttention path designed to preserve model-specific sink scaling under causal and sliding-window masks is presented.

Zelei Cheng, Amritansh Mishra, Sambit Sahu et al. · 0 citations