Skip to content

Author

Zhiquan Hu

We have 3 of 3 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Jul 2026

EviBack: Search-Agent Reinforcement Learning via Evidence-Constrained Teacher Backoff

EviBack is presented, an evidence- constrained Teacher backoff that supplies auxiliary super- vision to such groups while preserving verifiable Actor re- wards, and separates evidence assessment from answer refine- ment, preventing reference answers from overriding evidence- insufficiency judgments.

Xiao Ma, Zhiquan Hu, Yi Wei et al. · 0 citations
#natural language process... Preprint Aug 2026

JPO: Juris Policy Optimization for Structured Legal Reasoning in Criminal Judgment Prediction

Juris Policy Optimization (JPO), a post-training framework for structured legal reasoning in Chinese criminal judgment prediction, is proposed and experiments show that JPO consistently improves both judgment prediction and reasoning quality over supervised fine-tuning and reinforcement learning baselines.

Zhao-Lu Kang, Yan-Tao Liu, Tailong Luo et al. · 1 citation
#natural language process... Preprint Aug 2026

Modality Fault Lines: Structural Corruptions Reveal Fragile Omni-Modal Reasoning

SCEval (Structure-Corruption Evaluation) a diagnostic evaluation protocol that keeps the question, answer space, and modality channels fixed while applying controlled structural corruptions to text, vision, and audio individually and jointly is introduced.

Zhaolu Kang, Mei-Xin Wu, Yu Xue et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.