Skip to content

Author

Shiu-hong Kao

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Sep 2026

Aligning Thoughts with Answers: Probability Rewards to Tame Thinking Drift

This paper studies \textbf{thinking--answer consistency} in vision-language models. We focus on Visual Intention Grounding, where a model infers a target object based on a human intention query and predicts a bounding box. We reveal that previous IoU-based reinforcement learning (RL) frameworks suffer from ``thinking d...

Peng-Zhan Sun, Shiu-hong Kao, Shi-Jie Li et al. · 1 citation
#artificial intelligence Preprint Oct 2026

Rethinking Probability-Based Reinforcement Learning From Posterior Concentration

Verifier-free reinforcement learning with probability-based rewards offers a promising way to train LLMs on general reasoning tasks where external verifiers are unavailable. Yet the reliability of these rewards, especially in long-horizon reasoning, remains underexplored. This work identifies a length-dependent failure...

Shiu-hong Kao, Yu-Bo Zhao, Zhen Tian et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.