Skip to content

Author

Sheng-Jun Xiong

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Aug 2026

PartHackBench: Certified Equal-Progress Stress Tests for Partial-Credit Tool-Agent Evaluation

PartHackBench provides a certified control for testing whether evaluator credit changes while all benchmark-defined task-relevant progress remains fixed, and provides a certified control for testing whether evaluator credit changes while all benchmark-defined task-relevant progress remains fixed.

Hong-Ye Yang, Zhi Xie, Sheng-Jun Xiong · 0 citations
#artificial intelligence Preprint Sep 2026

DepthBenchCAD: When Does Deeper Auditing Yield More Reliable Conclusions?

Generative CAD models are expected to remain behaviorally correct after parameter edits, so increasing the number of edit checks is often treated as a direct route to more reliable evaluation. Under a fixed budget, however, auditing each program more thoroughly reduces the number of tasks and independent generations th...

Hong-Ye Yang, Zhi-Hao Xie, Sheng-Jun Xiong et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.