Skip to content

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

OdysSim: Building Foundation Models for Human Behavior Simulation

It is shown that LLM-as-judge RL induces reward-hacking patterns, and that LLM-as-judge RL detectors can mitigate them during post-training, suggesting that behavioral foundation models require rethinking the LLM training paradigm.

Xuhui Zhou, Weiwei Sun, Weihua Du et al. · 8 citations · ⚡1
#artificial intelligence Preprint Sep 2026

Efficient Test-Time Adaptation through Human-AI Interaction

This work proposes test-time adaptation through human-agent interaction (TAHI), which integrates these signals into agent context and weights, and crystallizes each user's training and evaluation criteria via an evolving rubric module.

Zora Zhiruo Wang, Apurva Gandhi, Rulin Shao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.