Skip to content

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective

ConSPO is a Contrastive Sequence-level Policy Optimization method that uses length-normalized sequence log-probabilities as rollout scores and contrasts verified positive rollouts against negative distractors within the same group, and outperforms strong baselines on challenging reasoning benchmarks.

Feng Zhang, Xin-Hong Ma, Zi-Qiang Dong et al. · 1 citation

HISA: Efficient Hierarchical Indexing for Fine-Grained Sparse Attention

This work proposes HISA (Hierarchical Indexed Sparse Attention), a plug-and-play replacement for the indexer that rewrites the search path from a flat token scan into a two-stage hierarchical procedure: a block-level coarse filtering stage that scores pooled block representations to discard irrelevant regions, followed...

Yufei Xu, Fan-Xu Meng, Fan Jiang et al. · 10 citations · ⚡1

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.